Qolda-AVL

Qolda-AVL is a 5B audio-vision-language model designed to operate in Kazakh, Russian, and English. The model extends Qwen3-VL with an audio branch built on a fine-tuned Whisper encoder and a dedicated audio projection module. All three modalities are adapted to Kazakh through a staged training pipeline, with the audio branch covering speech recognition, speech translation, audio classification, and environmental sound captioning.

To improve audio feature injection into the language backbone, we apply the DeepStack mechanism to the audio branch, mirroring the vision processing pipeline of Qwen3-VL 💜

Qolda-AVL architecture

The model is our step towards omni-modal systems for the Kazakh language.

The name "Qolda" reflects both its design and purpose in Kazakh: "in hand" (қолда) for its compact accessibility, and "to support" (қолдау) for its assistive nature.

Evaluation

The benchmark suites and Qolda-AVL model collection are available here:

The following tables report benchmark results across the Qolda-AVL family together with Qolda baselines. Higher is better unless otherwise noted, and the best result within each row is shown in bold.

5B / 9B / 34B = Qolda-AVL variants · Qolda-NT = Qolda No-Think · Qolda-T = Qolda Think

Click to for Text benchmark results
Benchmark Language Qolda-AVL-5B Qolda-AVL-9B Qolda-AVL-34B Qolda-NT Qolda-T
MMLU Kazakh 73.07 73.89 81.94 58.39 69.28
English 79.09 80.96 86.71 70.03 76.46
MMLU-Pro Kazakh 62.21 64.28 73.67 40.62 57.68
English 71.26 72.61 79.00 58.09 66.31
Russian 66.79 68.29 76.34 47.42 62.45
GPQA Kazakh 47.91 46.97 58.77 31.89 38.60
English 52.68 53.36 63.81 39.46 45.62
Russian 48.78 51.34 59.36 32.60 40.19
ARC Kazakh 93.59 94.27 96.76 86.14 92.30
English 96.90 97.29 98.17 94.22 96.11
Russian 96.11 96.62 97.63 91.17 94.17
GSM8K Kazakh 85.75 83.17 90.52 73.01 83.00
English 95.22 95.83 96.44 62.85 90.22
Russian 90.90 92.04 94.31 84.99 83.98
MMLU-Redux Kazakh 76.24 76.92 84.39 60.06 72.38
English 82.65 84.56 88.11 72.91 79.40
KazCulture Kazakh 44.75 56.39 62.37 53.00 47.45
KazMMLU Kazakh 69.27 73.04 78.98 58.11 66.14
KazBench Kazakh 61.83 64.61 70.05 64.23 61.12
Belebele Kazakh 81.70 84.76 88.78 81.07 82.91
PIQA Kazakh 81.00 78.00 85.00 63.00 70.00
INCLUDE Kazakh 53.80 57.20 61.80 45.20 46.00
Russian 58.15 64.49 70.83 59.17 56.52
KKCOPA Kazakh 76.60 78.11 79.60 70.00 73.79
NIS Math Kazakh 94.00 93.00 98.00 66.00 87.88
KazQAD Kazakh 42.28 65.76 70.36 70.99 67.40
RAGBench Kazakh 52.12 62.91 69.98 54.95 66.81
Click to for Vision benchmark results
Benchmark Language Qolda-AVL-5B Qolda-AVL-9B Qolda-AVL-34B Qolda-NT Qolda-T
RealWorldQA Kazakh 52.81 50.07 55.42 53.86 48.10
English 67.84 70.07 71.76 61.57 61.70
Russian 59.35 60.39 66.41 56.08 57.12
MMStar Kazakh 67.45 68.89 72.50 53.08 59.60
English 70.27 72.93 75.93 58.48 65.04
Russian 66.47 69.45 73.93 55.48 59.84
AI2D Kazakh 72.18 73.34 79.05 63.48 66.26
English 79.40 81.96 84.55 73.99 75.61
MathVista Kazakh 70.17 72.34 74.65 58.32 66.33
English 76.75 80.94 82.57 63.14 71.04
MathVision Kazakh 52.07 55.75 62.06 35.41 44.38
English 54.93 58.24 63.90 42.00 48.05
MMBench Kazakh 87.69 89.19 90.18 79.97 83.85
English 87.89 89.01 90.43 83.05 84.40
OCRBench Kazakh 30.61 51.25 53.97 49.89 46.49
English 73.20 77.20 79.20 69.90 68.70
Click to for Audio benchmark results
Benchmark Language Qolda-AVL-5B Qolda-AVL-9B Qolda-AVL-34B
SAKURA · Animal · Multi Kazakh 48.58 69.54 70.00
English 64.60 82.20 83.00
Russian 50.40 81.76 83.00
SAKURA · Animal · Single Kazakh 55.31 87.80 86.00
English 52.20 87.80 83.80
Russian 53.23 88.40 88.20
SAKURA · Emotion · Multi Kazakh 33.87 36.95 39.20
English 35.80 37.40 43.20
Russian 38.20 37.20 40.20
SAKURA · Emotion · Single Kazakh 39.03 47.27 45.58
English 42.60 44.40 47.40
Russian 40.28 44.80 47.60
SAKURA · Gender · Multi Kazakh 67.94 75.40 81.60
English 72.00 83.20 82.80
Russian 70.42 77.60 83.20
SAKURA · Gender · Single Kazakh 70.88 82.80 84.20
English 80.00 85.00 87.80
Russian 76.92 82.60 87.40
SAKURA · Language · Multi Kazakh 87.40 88.80 92.60
English 92.60 92.00 94.40
Russian 87.80 90.80 93.60
SAKURA · Language · Single Kazakh 96.00 97.60 97.40
English 96.40 97.40 97.40
Russian 97.19 97.60 97.60
SpokenMQA · Long Digit Kazakh 81.40 88.95 91.28
English 88.37 93.02 94.19
SpokenMQA · Multi-step Reasoning Kazakh 74.57 77.66 87.28
English 87.42 90.23 92.53
SpokenMQA · Short Digit Kazakh 84.00 92.00 95.00
English 88.00 94.00 92.00
SpokenMQA · Single-step Reasoning Kazakh 90.20 92.74 93.92
English 93.92 95.10 94.76
ASR · WER ↓ Kazakh 0.1707 0.1452 0.1688
English 0.0662 0.0634 0.0601
Russian 0.1136 0.1095 0.1077
WavCaps Kazakh 0.64 6.24 6.53
English 3.82 14.28 15.03
Russian 1.27 8.38 8.32
WavCaps-QA Kazakh 8.55 25.33 25.33
English 24.34 35.86 34.21
Russian 19.08 32.24 31.25

Model Usage

1. Transformers inference

To run the inference with transformers, complete the preliminary setup:

uv venv venv
source venv/bin/activate
uv pip install torch accelerate transformers

Then initialize the model and processor:

import torch
from transformers import AutoModelForCausalLM, AutoProcessor

model = AutoModelForCausalLM.from_pretrained(
    "issai/Qolda-AVL-5B",
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="auto"
)

processor = AutoProcessor.from_pretrained("issai/Qolda-AVL-5B", trust_remote_code=True)

Depending on the required modalities, define the messages list:

Language:

messages = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "y = (lnx)^2 функциясының туындысын тап. JSON форматында жауап бер: {'answer': '...'}"},
        ],
    }
]

Vision-Language:

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "assets/sample_image.jpg"}, # or provide link to the image
            {"type": "text", "text": "Суретті егжей-тегжейлі сипаттап бер. Неше жылқы көріп тұрсың және олардың түстері қандай?"},
        ],
    }
]

Audio-Language:

Note: The model was not trained to answer questions posed directly in the audio. Provide a detailed text instruction alongside the audio describing the task you want performed on it.

prompt = """Математикалық есепті шеш.
Respond ONLY with this JSON format: {"explanation": "<your step-by-step reasoning>", "answer": <integer or float number>}
The answer must be a number (integer or float). No text, no units, just the number.
"""

messages = [
    {
        "role": "user",
        "content": [
            {"type": "audio", "audio": "assets/sample_audio.wav"}, # or provide link to the audio
            {"type": "text", "text": prompt}
        ],
    }
]

Audio-Vision-Language:

messages = [
    {
        "role": "user",
        "content": [
            {"type": "audio", "audio": "assets/question_audio.wav"},
            {"type": "image", "image": "assets/sample_image.jpg"},
            {"type": "text", "text": "Answer the question"},
        ],
    }
]

Finally, pass the messages to the model for inference:

inputs = processor.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    return_dict=True, return_tensors="pt",
).to(model.device)

generated_ids = model.generate(
    **inputs,
    max_new_tokens=4096,
    temperature=0.7,
    top_p=0.95,
    top_k=20,
    do_sample=True,
    repetition_penalty=1.0,
)
generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(generated_ids_trimmed, skip_special_tokens=True)

print(processor.batch_decode(generated_ids_trimmed, skip_special_tokens=True)[0])

2. vLLM inference

Alternatively, you can run the model via a vLLM server. Note that we use a custom vLLM package. First, complete the preliminary setup:

uv venv venv
source venv/bin/activate

# Install this fork (precompiled binaries)
git clone https://github.com/IS2AI/vLLM-Qolda-AVL.git
cd vLLM-Qolda-AVL
VLLM_USE_PRECOMPILED=1 uv pip install -e .

Then start the OpenAI-compatible server (adjust parameters to your settings):

vllm serve issai/Qolda-AVL-5B \
    --served-model-name qolda-avl \
    --trust-remote-code \
    --tensor-parallel-size 4 \
    --dtype bfloat16 \
    --max-model-len 16384 \
    --limit-mm-per-prompt '{"audio": 1, "image": 1}'

To run inference, you can use the following code:

import base64
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1", 
    api_key="EMPTY"
)

def encode_audio_base64(path: str | Path) -> str:
    with open(path, "rb") as f:
        return base64.b64encode(f.read()).decode("utf-8")

def encode_image_base64(path: str | Path) -> str:
    with open(path, "rb") as f:
        return base64.b64encode(f.read()).decode("utf-8")

audio_path = "assets/sample_audio.wav"
audio_b64 = encode_audio_base64(audio_path)

stream = client.chat.completions.create(
    model=client.models.list().data[0].id,
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "input_audio",
                    "input_audio": {
                        "data": audio_b64,
                        "format": "wav",
                    },
                },
                {
                    "type": "text",
                    "text": (
                        "Analyze the voice in the audio and identify the speaker's "
                        "gender (male or female). Also transcribe what is said. "
                        "Return your answer as JSON in the following format: "
                        '{"answer": "<male or female>",'
                        '"transcription": "<transcription>"}'
                    ),
                },
            ],
        }
    ],
    max_tokens=4096,
    temperature=0.7,
    top_p=0.8,
    stream=True,
    stream_options={"include_usage": True},
)

text = ""
usage = None
for chunk in stream:
    if chunk.usage:
        usage = chunk.usage
    if chunk.choices and chunk.choices[0].delta.content:
        token = chunk.choices[0].delta.content
        print(token, end="", flush=True)
        text += token

License

Apache License 2.0

Citation


@article{qolda-avl-bdcc,
AUTHOR = {Arystanbekov, Batyr and Maxutov, Akylbek and Nurimanov, Aspandiyar and Varol, Huseyin Atakan},
TITLE = {Extending a Vision–Language Model with Audio Understanding: Introducing Qolda-AVL for the Kazakh Language},
JOURNAL = {Big Data and Cognitive Computing},
VOLUME = {10},
YEAR = {2026},
NUMBER = {6},
ARTICLE-NUMBER = {192},
URL = {https://www.mdpi.com/2504-2289/10/6/192},
ISSN = {2504-2289},
DOI = {10.3390/bdcc10060192}
}
Downloads last month
272
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for issai/Qolda-AVL-5B

Finetuned
(28)
this model
Quantizations
1 model

Space using issai/Qolda-AVL-5B 1

Collection including issai/Qolda-AVL-5B