BeamModel

Deploying Beam

The weights are not public yet. The commands below follow each engine’s standard usage and will be tested and updated on release day. Replace the placeholder repo id with the official one.

Before you start

At FP8, the weights alone take about 500 GB, so plan for a multi-GPU node such as 8× H200 or 8× B200. For quantized GGUF builds on Apple Silicon or CPU, see the VRAM calculator.

1. Download the weights

pip install -U "huggingface_hub[cli]"
huggingface-cli download <beam-repo-id> --local-dir ./beam

2. vLLM (recommended for serving)

pip install -U vllm
vllm serve ./beam \
  --tensor-parallel-size 8 \
  --max-model-len 131072 \
  --port 8000

3. SGLang

pip install -U "sglang[all]"
python -m sglang.launch_server \
  --model-path ./beam \
  --tp 8 \
  --port 8000

4. llama.cpp (GGUF, consumer hardware)

Community GGUF quantizations usually appear within days of a release. Large models are split into several files, so point -m at the first shard.

llama-server \
  -m ./beam-gguf/<first-shard>.gguf \
  -c 32768 \
  -ngl 99 \
  --port 8000

5. Test the OpenAI-compatible endpoint

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "./beam", "messages": [{"role": "user", "content": "Hello"}]}'