Qolda-AVL
Qolda-AVL is a 5B audio-vision-language model designed to operate in Kazakh, Russian, and English. The model extends Qwen3-VL with an audio branch built on a fine-tuned Whisper encoder and a dedicated audio projection module. All three modalities are adapted to Kazakh through a staged training pipeline, with the audio branch covering speech recognition, speech translation, audio classification, and environmental sound captioning.
To improve audio feature injection into the language backbone, we apply the DeepStack mechanism to the audio branch, mirroring the vision processing pipeline of Qwen3-VL 💜
The model is our step towards omni-modal systems for the Kazakh language.
The name "Qolda" reflects both its design and purpose in Kazakh: "in hand" (қолда) for its compact accessibility, and "to support" (қолдау) for its assistive nature.
Evaluation
The benchmark suites and Qolda-AVL model collection are available here:
- Text benchmarks: https://huggingface.co/collections/issai/qolda-language-benchmarks
- Vision benchmarks: https://huggingface.co/collections/issai/qolda-vision-benchmarks
- Audio benchmarks: https://huggingface.co/collections/issai/qolda-avl-audio-benchmarks
- Qolda model: https://huggingface.co/issai/Qolda
The following tables report benchmark results across the Qolda-AVL family together with Qolda baselines. Higher is better unless otherwise noted, and the best result within each row is shown in bold.
5B / 9B / 34B = Qolda-AVL variants · Qolda-NT = Qolda No-Think · Qolda-T = Qolda Think
Click to for Text benchmark results
| Benchmark | Language | Qolda-AVL-5B | Qolda-AVL-9B | Qolda-AVL-34B | Qolda-NT | Qolda-T |
|---|---|---|---|---|---|---|
| MMLU | Kazakh | 73.07 | 73.89 | 81.94 | 58.39 | 69.28 |
| English | 79.09 | 80.96 | 86.71 | 70.03 | 76.46 | |
| MMLU-Pro | Kazakh | 62.21 | 64.28 | 73.67 | 40.62 | 57.68 |
| English | 71.26 | 72.61 | 79.00 | 58.09 | 66.31 | |
| Russian | 66.79 | 68.29 | 76.34 | 47.42 | 62.45 | |
| GPQA | Kazakh | 47.91 | 46.97 | 58.77 | 31.89 | 38.60 |
| English | 52.68 | 53.36 | 63.81 | 39.46 | 45.62 | |
| Russian | 48.78 | 51.34 | 59.36 | 32.60 | 40.19 | |
| ARC | Kazakh | 93.59 | 94.27 | 96.76 | 86.14 | 92.30 |
| English | 96.90 | 97.29 | 98.17 | 94.22 | 96.11 | |
| Russian | 96.11 | 96.62 | 97.63 | 91.17 | 94.17 | |
| GSM8K | Kazakh | 85.75 | 83.17 | 90.52 | 73.01 | 83.00 |
| English | 95.22 | 95.83 | 96.44 | 62.85 | 90.22 | |
| Russian | 90.90 | 92.04 | 94.31 | 84.99 | 83.98 | |
| MMLU-Redux | Kazakh | 76.24 | 76.92 | 84.39 | 60.06 | 72.38 |
| English | 82.65 | 84.56 | 88.11 | 72.91 | 79.40 | |
| KazCulture | Kazakh | 44.75 | 56.39 | 62.37 | 53.00 | 47.45 |
| KazMMLU | Kazakh | 69.27 | 73.04 | 78.98 | 58.11 | 66.14 |
| KazBench | Kazakh | 61.83 | 64.61 | 70.05 | 64.23 | 61.12 |
| Belebele | Kazakh | 81.70 | 84.76 | 88.78 | 81.07 | 82.91 |
| PIQA | Kazakh | 81.00 | 78.00 | 85.00 | 63.00 | 70.00 |
| INCLUDE | Kazakh | 53.80 | 57.20 | 61.80 | 45.20 | 46.00 |
| Russian | 58.15 | 64.49 | 70.83 | 59.17 | 56.52 | |
| KKCOPA | Kazakh | 76.60 | 78.11 | 79.60 | 70.00 | 73.79 |
| NIS Math | Kazakh | 94.00 | 93.00 | 98.00 | 66.00 | 87.88 |
| KazQAD | Kazakh | 42.28 | 65.76 | 70.36 | 70.99 | 67.40 |
| RAGBench | Kazakh | 52.12 | 62.91 | 69.98 | 54.95 | 66.81 |
Click to for Vision benchmark results
| Benchmark | Language | Qolda-AVL-5B | Qolda-AVL-9B | Qolda-AVL-34B | Qolda-NT | Qolda-T |
|---|---|---|---|---|---|---|
| RealWorldQA | Kazakh | 52.81 | 50.07 | 55.42 | 53.86 | 48.10 |
| English | 67.84 | 70.07 | 71.76 | 61.57 | 61.70 | |
| Russian | 59.35 | 60.39 | 66.41 | 56.08 | 57.12 | |
| MMStar | Kazakh | 67.45 | 68.89 | 72.50 | 53.08 | 59.60 |
| English | 70.27 | 72.93 | 75.93 | 58.48 | 65.04 | |
| Russian | 66.47 | 69.45 | 73.93 | 55.48 | 59.84 | |
| AI2D | Kazakh | 72.18 | 73.34 | 79.05 | 63.48 | 66.26 |
| English | 79.40 | 81.96 | 84.55 | 73.99 | 75.61 | |
| MathVista | Kazakh | 70.17 | 72.34 | 74.65 | 58.32 | 66.33 |
| English | 76.75 | 80.94 | 82.57 | 63.14 | 71.04 | |
| MathVision | Kazakh | 52.07 | 55.75 | 62.06 | 35.41 | 44.38 |
| English | 54.93 | 58.24 | 63.90 | 42.00 | 48.05 | |
| MMBench | Kazakh | 87.69 | 89.19 | 90.18 | 79.97 | 83.85 |
| English | 87.89 | 89.01 | 90.43 | 83.05 | 84.40 | |
| OCRBench | Kazakh | 30.61 | 51.25 | 53.97 | 49.89 | 46.49 |
| English | 73.20 | 77.20 | 79.20 | 69.90 | 68.70 |
Click to for Audio benchmark results
| Benchmark | Language | Qolda-AVL-5B | Qolda-AVL-9B | Qolda-AVL-34B |
|---|---|---|---|---|
| SAKURA · Animal · Multi | Kazakh | 48.58 | 69.54 | 70.00 |
| English | 64.60 | 82.20 | 83.00 | |
| Russian | 50.40 | 81.76 | 83.00 | |
| SAKURA · Animal · Single | Kazakh | 55.31 | 87.80 | 86.00 |
| English | 52.20 | 87.80 | 83.80 | |
| Russian | 53.23 | 88.40 | 88.20 | |
| SAKURA · Emotion · Multi | Kazakh | 33.87 | 36.95 | 39.20 |
| English | 35.80 | 37.40 | 43.20 | |
| Russian | 38.20 | 37.20 | 40.20 | |
| SAKURA · Emotion · Single | Kazakh | 39.03 | 47.27 | 45.58 |
| English | 42.60 | 44.40 | 47.40 | |
| Russian | 40.28 | 44.80 | 47.60 | |
| SAKURA · Gender · Multi | Kazakh | 67.94 | 75.40 | 81.60 |
| English | 72.00 | 83.20 | 82.80 | |
| Russian | 70.42 | 77.60 | 83.20 | |
| SAKURA · Gender · Single | Kazakh | 70.88 | 82.80 | 84.20 |
| English | 80.00 | 85.00 | 87.80 | |
| Russian | 76.92 | 82.60 | 87.40 | |
| SAKURA · Language · Multi | Kazakh | 87.40 | 88.80 | 92.60 |
| English | 92.60 | 92.00 | 94.40 | |
| Russian | 87.80 | 90.80 | 93.60 | |
| SAKURA · Language · Single | Kazakh | 96.00 | 97.60 | 97.40 |
| English | 96.40 | 97.40 | 97.40 | |
| Russian | 97.19 | 97.60 | 97.60 | |
| SpokenMQA · Long Digit | Kazakh | 81.40 | 88.95 | 91.28 |
| English | 88.37 | 93.02 | 94.19 | |
| SpokenMQA · Multi-step Reasoning | Kazakh | 74.57 | 77.66 | 87.28 |
| English | 87.42 | 90.23 | 92.53 | |
| SpokenMQA · Short Digit | Kazakh | 84.00 | 92.00 | 95.00 |
| English | 88.00 | 94.00 | 92.00 | |
| SpokenMQA · Single-step Reasoning | Kazakh | 90.20 | 92.74 | 93.92 |
| English | 93.92 | 95.10 | 94.76 | |
| ASR · WER ↓ | Kazakh | 0.1707 | 0.1452 | 0.1688 |
| English | 0.0662 | 0.0634 | 0.0601 | |
| Russian | 0.1136 | 0.1095 | 0.1077 | |
| WavCaps | Kazakh | 0.64 | 6.24 | 6.53 |
| English | 3.82 | 14.28 | 15.03 | |
| Russian | 1.27 | 8.38 | 8.32 | |
| WavCaps-QA | Kazakh | 8.55 | 25.33 | 25.33 |
| English | 24.34 | 35.86 | 34.21 | |
| Russian | 19.08 | 32.24 | 31.25 |
Model Usage
1. Transformers inference
To run the inference with transformers, complete the preliminary setup:
uv venv venv
source venv/bin/activate
uv pip install torch accelerate transformers
Then initialize the model and processor:
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
model = AutoModelForCausalLM.from_pretrained(
"issai/Qolda-AVL-5B",
trust_remote_code=True,
dtype=torch.bfloat16,
device_map="auto"
)
processor = AutoProcessor.from_pretrained("issai/Qolda-AVL-5B", trust_remote_code=True)
Depending on the required modalities, define the messages list:
Language:
messages = [
{
"role": "user",
"content": [
{"type": "text", "text": "y = (lnx)^2 функциясының туындысын тап. JSON форматында жауап бер: {'answer': '...'}"},
],
}
]
Vision-Language:
messages = [
{
"role": "user",
"content": [
{"type": "image", "image": "assets/sample_image.jpg"}, # or provide link to the image
{"type": "text", "text": "Суретті егжей-тегжейлі сипаттап бер. Неше жылқы көріп тұрсың және олардың түстері қандай?"},
],
}
]
Audio-Language:
Note: The model was not trained to answer questions posed directly in the audio. Provide a detailed text instruction alongside the audio describing the task you want performed on it.
prompt = """Математикалық есепті шеш.
Respond ONLY with this JSON format: {"explanation": "<your step-by-step reasoning>", "answer": <integer or float number>}
The answer must be a number (integer or float). No text, no units, just the number.
"""
messages = [
{
"role": "user",
"content": [
{"type": "audio", "audio": "assets/sample_audio.wav"}, # or provide link to the audio
{"type": "text", "text": prompt}
],
}
]
Audio-Vision-Language:
messages = [
{
"role": "user",
"content": [
{"type": "audio", "audio": "assets/question_audio.wav"},
{"type": "image", "image": "assets/sample_image.jpg"},
{"type": "text", "text": "Answer the question"},
],
}
]
Finally, pass the messages to the model for inference:
inputs = processor.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True,
return_dict=True, return_tensors="pt",
).to(model.device)
generated_ids = model.generate(
**inputs,
max_new_tokens=4096,
temperature=0.7,
top_p=0.95,
top_k=20,
do_sample=True,
repetition_penalty=1.0,
)
generated_ids_trimmed = [
out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(generated_ids_trimmed, skip_special_tokens=True)
print(processor.batch_decode(generated_ids_trimmed, skip_special_tokens=True)[0])
2. vLLM inference
Alternatively, you can run the model via a vLLM server. Note that we use a custom vLLM package. First, complete the preliminary setup:
uv venv venv
source venv/bin/activate
# Install this fork (precompiled binaries)
git clone https://github.com/IS2AI/vLLM-Qolda-AVL.git
cd vLLM-Qolda-AVL
VLLM_USE_PRECOMPILED=1 uv pip install -e .
Then start the OpenAI-compatible server (adjust parameters to your settings):
vllm serve issai/Qolda-AVL-5B \
--served-model-name qolda-avl \
--trust-remote-code \
--tensor-parallel-size 4 \
--dtype bfloat16 \
--max-model-len 16384 \
--limit-mm-per-prompt '{"audio": 1, "image": 1}'
To run inference, you can use the following code:
import base64
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY"
)
def encode_audio_base64(path: str | Path) -> str:
with open(path, "rb") as f:
return base64.b64encode(f.read()).decode("utf-8")
def encode_image_base64(path: str | Path) -> str:
with open(path, "rb") as f:
return base64.b64encode(f.read()).decode("utf-8")
audio_path = "assets/sample_audio.wav"
audio_b64 = encode_audio_base64(audio_path)
stream = client.chat.completions.create(
model=client.models.list().data[0].id,
messages=[
{
"role": "user",
"content": [
{
"type": "input_audio",
"input_audio": {
"data": audio_b64,
"format": "wav",
},
},
{
"type": "text",
"text": (
"Analyze the voice in the audio and identify the speaker's "
"gender (male or female). Also transcribe what is said. "
"Return your answer as JSON in the following format: "
'{"answer": "<male or female>",'
'"transcription": "<transcription>"}'
),
},
],
}
],
max_tokens=4096,
temperature=0.7,
top_p=0.8,
stream=True,
stream_options={"include_usage": True},
)
text = ""
usage = None
for chunk in stream:
if chunk.usage:
usage = chunk.usage
if chunk.choices and chunk.choices[0].delta.content:
token = chunk.choices[0].delta.content
print(token, end="", flush=True)
text += token
License
Apache License 2.0
Citation
@article{qolda-avl-bdcc,
AUTHOR = {Arystanbekov, Batyr and Maxutov, Akylbek and Nurimanov, Aspandiyar and Varol, Huseyin Atakan},
TITLE = {Extending a Vision–Language Model with Audio Understanding: Introducing Qolda-AVL for the Kazakh Language},
JOURNAL = {Big Data and Cognitive Computing},
VOLUME = {10},
YEAR = {2026},
NUMBER = {6},
ARTICLE-NUMBER = {192},
URL = {https://www.mdpi.com/2504-2289/10/6/192},
ISSN = {2504-2289},
DOI = {10.3390/bdcc10060192}
}
- Downloads last month
- 272