OpenThai-SystemOne — mlx-nvfp4

OpenThai-SystemOne is an open Thai + English System One decision model: one forward pass answers typed questions (choice over up to 255 options, ordinal score, yes/no noul) about a text / JSON state with calibrated probabilities, no text generation. It is a Qwen3.5-0.8B text tower (Thai continued pre-training) plus a 256-slot decision head. This repo is a quantization of v0.3 (commit f3709948).

What is quantized: the tower including the token embeddings (MLX quantizes the embedding table too). The 256-slot decision head and the per-type temperatures stay in fp32 (head.safetensors). Quantization therefore only perturbs the hidden state the head reads.

Format: MLX NVFP4 for Apple Silicon (mlx-lm). mlx-lm runs the tower; the included client applies the decision head on the final hidden states. Size: 424 MB.

Measured on a MacBook Pro M3 Max: 4-bit ≈ 19 ms per 3-question Thai decision (the PyTorch model on MPS: ~150 ms).

Usage

pip install mlx-lm torch transformers safetensors pydantic
huggingface-cli download iapp/OpenThai-SystemOne-MLX-nvfp4 --local-dir openthai-mlx
import sys; sys.path.insert(0, "openthai-mlx")
from openthai_systemone.mlx_client import MLXSystemOneClient
c = MLXSystemOneClient("openthai-mlx")
r = c.system_one("ร้านนี้อาหารอร่อยมาก แต่รอนานเกือบชั่วโมง", {"sentiment": {"type": "choice", "instructions": "ความรู้สึก",
                 "criteria": {"บวก": None, "ลบ": None, "กลาง": None}}})
print(r.answers["sentiment"].choice, r.answers["sentiment"].probabilities)

Accuracy vs the bf16 original (same records, single option order, first 800 per set)

Macro: public 73.2 (original 74.3), Thai 79.4 (original 80.1).

subset bf16 original this Δ
public 13-subset bench
aegis2 (noul) 83.2 80.8 -2.4
boolq (noul) 79.7 78.7 -1.0
civil_comments (noul) 79.0 81.3 +2.3
helpsteer2 (score) 41.6 42.8 +1.2
massive-de-DE (choice) 88.3 84.0 -4.3
massive-en-US (choice) 88.3 87.7 -0.6
multinli (choice) 89.0 85.6 -3.3
paws (noul) 94.0 92.8 -1.2
pubmedqa (choice) 64.0 62.4 -1.6
squad2 (noul) 89.3 86.3 -3.0
summeval-consistency (score) 75.0 79.2 +4.2
summeval-relevance (score) 21.7 20.8 -0.8
vitaminc-dev (choice) 72.5 69.8 -2.7
macro, public 13-subset bench 74.3 73.2 -1.0
Thai held-out / eval sets
banking77 (choice) 59.1 60.6 +1.5
contrastive_th (choice) 80.7 80.7 +0.0
contrastive_th (noul) 83.5 81.0 -2.4
contrastive_th (score) 78.6 75.0 -3.6
massive_th (choice) 90.6 89.8 -0.9
prachathai (choice) 98.3 98.8 +0.5
prachathai (noul) 93.4 93.9 +0.4
sib200_th (choice) 77.9 76.5 -1.5
wisesight (choice) 48.9 48.6 -0.2
wongnai (score) 64.5 64.6 +0.1
xlam_tools (choice) 99.4 99.2 -0.1
xnli_th (choice) 79.8 77.0 -2.8
xnli_th (noul) 86.8 86.1 -0.6
macro, Thai held-out / eval sets 80.1 79.4 -0.7

Notes

  • Scores are single-option-order accuracy on the first 800 records of each set (scripts/06_eval.py --limit 800), the same records for the original and the quantization. score subsets report exact level accuracy.
  • Base model, data, training and the full benchmark tables: iapp/OpenThai-SystemOne.
  • License Apache-2.0 (same as the base). Built by iApp Technology / OpenThaiGPT.
Downloads last month
-
Safetensors
Model size
0.8B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for iapp/OpenThai-SystemOne-MLX-nvfp4

Quantized
(14)
this model