GLM-5.2-Vision (FP8)
GLM-5.2 with sight. A vision-language model that bolts the MoonViT vision encoder from Kimi-K2.6 onto GLM-5.2 through a trained PatchMerger projector.
GLM-5.2 is a strong open reasoning model with no vision input. This checkpoint adds it, without touching a single GLM weight: the text backbone and the vision tower are both frozen and byte-identical to their upstream releases. The only newly-trained parameters are the 49.5M-parameter projector that maps MoonViT's 1152-dim patch embeddings into GLM's 6144-dim token space.
| Component | Detail |
|---|---|
| Text backbone | GLM-5.2 (744B total / A40B active, MoE + MLA + DSA sparse attention) — frozen |
| Vision tower | MoonViT-3d from Kimi-K2.6, 27 layers, 1152-dim — frozen |
| Projector | PatchMerger MLP (pre_norm → linear_1 → GELU → linear_2), 1152→4608→6144 — trained |
| Text weights | block-FP8, from zai-org/GLM-5.2-FP8 |
| Size | ~757 GB |
| Hardware | 8×B200 or 8×H200 |
| Image tokens | up to 4096 per image (16384 MoonViT patches, 2×2 merge) |
| Max context | 1048576 (1M tokens) |
The vision tower and projector are bf16 — only the GLM text Linears are quantized. This is the most thoroughly validated build; if you are unsure which variant to use, use this one.
At ~757 GB the weights do not fit on four GPUs. For a 4-GPU deployment use
baseten/GLM-5.2-Vision-NVFP4.
Quickstart
SGLang needs a small out-of-tree plugin because Glm5vForConditionalGeneration is not yet an
upstream architecture. It ships inside this repo, so there is nothing else to clone:
uvx --from huggingface-hub hf download baseten/GLM-5.2-Vision-FP8 \
--include 'plugins/*' --local-dir ./glm5v
uv pip install ./glm5v/plugins
SGLang
export SGLANG_EXTERNAL_MODEL_PACKAGE=sglang_glm5v
export SGLANG_EXTERNAL_MM_PROCESSOR_PACKAGE=sglang_glm5v
export SGLANG_EXTERNAL_MM_MODEL_ARCH=Glm5vForConditionalGeneration
python -m sglang_glm5v.patch
8×B200 / 8×H200 — full 1M context
python -m sglang.launch_server \
--model-path baseten/GLM-5.2-Vision-FP8 --trust-remote-code \
--tp-size 8 \
--attention-backend dsa --mm-attention-backend sdpa \
--kv-cache-dtype fp8_e4m3 --page-size 64 \
--mem-fraction-static 0.85 \
--context-length 1048576 \
--reasoning-parser glm45 --tool-call-parser glm47 \
--served-model-name glm-5.2-vision \
--port 30000
Query it
Standard OpenAI multimodal messages deliver the image as image_url:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="none")
r = client.chat.completions.create(
model="glm-5.2-vision",
messages=[{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "https://ultralytics.com/images/bus.jpg"}},
{"type": "text", "text": "Describe this image in detail."},
]}],
temperature=1.0, top_p=0.95, max_tokens=512,
)
print(r.choices[0].message.content)
GLM-5.2 is a reasoning model: with --reasoning-parser glm45, the chain of thought arrives in
message.reasoning_content and the answer in message.content.
Deploy on Baseten
The repository includes ready-to-push Truss configs. The only credential you need is an API key for your own Baseten account; no Hugging Face token or pre-created Baseten secret is required.
- Install
uvand create a Baseten API key. - Export the key, download the small Truss directory, and deploy the model:
export BASETEN_API_KEY="your-baseten-api-key"
uvx truss login --api-key "$BASETEN_API_KEY" --remote baseten --non-interactive
uvx --from huggingface-hub hf download baseten/GLM-5.2-Vision-FP8 \
--include 'truss/*' --local-dir ./glm5v
cd glm5v/truss
# FP8 requires 8×B200 and provides the full 1M-token context.
uvx truss push --remote baseten --config config.yaml --wait --output json
The command creates a new model and published deployment in your Baseten account and prints
JSON containing model_id, model_version_id, predict_url, and logs_url. It does not
promote the deployment to production.
Set PREDICT_URL to the returned predict_url, then query the model:
export PREDICT_URL="https://model-...api.baseten.co/deployment/.../predict"
curl -fsS "$PREDICT_URL" \
-H "Authorization: Api-Key $BASETEN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.2-vision",
"messages": [{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "https://ultralytics.com/images/bus.jpg"}},
{"type": "text", "text": "Describe this image in detail."}
]
}],
"max_tokens": 512,
"temperature": 1.0,
"top_p": 0.95
}'
The first deployment downloads about 757 GB of weights and initializes SGLang, so startup can take several minutes.
License
MIT, following both parents: GLM-5.2 (MIT) and Kimi-K2.6 (Modified MIT). The projector weights are released under MIT. Redistributed upstream weights remain under their original terms.
Acknowledgements
Built on Z.ai's GLM-5.2 and Moonshot AI's Kimi-K2.6. Neither team was involved in this work; please do not direct issues with this checkpoint to them.
- Downloads last month
- 24
Model tree for baseten/GLM-5.2-Vision-FP8
Base model
zai-org/GLM-5.2