issai/Qolda-AVL-5B

Model

2

stars

6

commits

1

repos using this model

1

linked in READMEs

Aug 18, 2026

updated

custom_code
qwen3_vl
safetensors

README

Qolda-AVL

Qolda-AVL is a 5B audio-vision-language model designed to operate in Kazakh, Russian, and English. The model extends Qwen3-VL with an audio branch built on a fine-tuned Whisper encoder and a dedicated audio projection module. All three modalities are adapted to Kazakh through a staged training pipeline, with the audio branch covering speech recognition, speech translation, audio classification, and environmental sound captioning.

To improve audio feature injection into the language backbone, we apply the DeepStack mechanism to the audio branch, mirroring the vision processing pipeline of Qwen3-VL 💜

Qolda-AVL architecture

The model is our step towards omni-modal systems for the Kazakh language.

The name "Qolda" reflects both its design and purpose in Kazakh: "in hand" (қолда) for its compact accessibility, and "to support" (қолдау) for its assistive nature.

Evaluation

The benchmark suites and Qolda-AVL model collection are available here:

The following tables report benchmark results across the Qolda-AVL family together with Qolda baselines. Higher is better unless otherwise noted, and the best result within each row is shown in bold.

5B / 9B / 34B = Qolda-AVL variants · Qolda-NT = Qolda No-Think · Qolda-T = Qolda Think

Click to for Text benchmark results
BenchmarkLanguageQolda-AVL-5BQolda-AVL-9BQolda-AVL-34BQolda-NTQolda-T
MMLUKazakh73.0773.8981.9458.3969.28
English79.0980.9686.7170.0376.46
MMLU-ProKazakh62.2164.2873.6740.6257.68
English71.2672.6179.0058.0966.31
Russian66.7968.2976.3447.4262.45
GPQAKazakh47.9146.9758.7731.8938.60
English52.6853.3663.8139.4645.62
Russian48.7851.3459.3632.6040.19
ARCKazakh93.5994.2796.7686.1492.30
English96.9097.2998.1794.2296.11
Russian96.1196.6297.6391.1794.17
GSM8KKazakh85.7583.1790.5273.0183.00
English95.2295.8396.4462.8590.22
Russian90.9092.0494.3184.9983.98
MMLU-ReduxKazakh76.2476.9284.3960.0672.38
English82.6584.5688.1172.9179.40
KazCultureKazakh44.7556.3962.3753.0047.45
KazMMLUKazakh69.2773.0478.9858.1166.14
KazBenchKazakh61.8364.6170.0564.2361.12
BelebeleKazakh81.7084.7688.7881.0782.91
PIQAKazakh81.0078.0085.0063.0070.00
INCLUDEKazakh53.8057.2061.8045.2046.00
Russian58.1564.4970.8359.1756.52
KKCOPAKazakh76.6078.1179.6070.0073.79
NIS MathKazakh94.0093.0098.0066.0087.88
KazQADKazakh42.2865.7670.3670.9967.40
RAGBenchKazakh52.1262.9169.9854.9566.81
Click to for Vision benchmark results
BenchmarkLanguageQolda-AVL-5BQolda-AVL-9BQolda-AVL-34BQolda-NTQolda-T
RealWorldQAKazakh52.8150.0755.4253.8648.10
English67.8470.0771.7661.5761.70
Russian59.3560.3966.4156.0857.12
MMStarKazakh67.4568.8972.5053.0859.60
English70.2772.9375.9358.4865.04
Russian66.4769.4573.9355.4859.84
AI2DKazakh72.1873.3479.0563.4866.26
English79.4081.9684.5573.9975.61
MathVistaKazakh70.1772.3474.6558.3266.33
English76.7580.9482.5763.1471.04
MathVisionKazakh52.0755.7562.0635.4144.38
English54.9358.2463.9042.0048.05
MMBenchKazakh87.6989.1990.1879.9783.85
English87.8989.0190.4383.0584.40
OCRBenchKazakh30.6151.2553.9749.8946.49
English73.2077.2079.2069.9068.70
Click to for Audio benchmark results
BenchmarkLanguageQolda-AVL-5BQolda-AVL-9BQolda-AVL-34B
SAKURA · Animal · MultiKazakh48.5869.5470.00
English64.6082.2083.00
Russian50.4081.7683.00
SAKURA · Animal · SingleKazakh55.3187.8086.00
English52.2087.8083.80
Russian53.2388.4088.20
SAKURA · Emotion · MultiKazakh33.8736.9539.20
English35.8037.4043.20
Russian38.2037.2040.20
SAKURA · Emotion · SingleKazakh39.0347.2745.58
English42.6044.4047.40
Russian40.2844.8047.60
SAKURA · Gender · MultiKazakh67.9475.4081.60
English72.0083.2082.80
Russian70.4277.6083.20
SAKURA · Gender · SingleKazakh70.8882.8084.20
English80.0085.0087.80
Russian76.9282.6087.40
SAKURA · Language · MultiKazakh87.4088.8092.60
English92.6092.0094.40
Russian87.8090.8093.60
SAKURA · Language · SingleKazakh96.0097.6097.40
English96.4097.4097.40
Russian97.1997.6097.60
SpokenMQA · Long DigitKazakh81.4088.9591.28
English88.3793.0294.19
SpokenMQA · Multi-step ReasoningKazakh74.5777.6687.28
English87.4290.2392.53
SpokenMQA · Short DigitKazakh84.0092.0095.00
English88.0094.0092.00
SpokenMQA · Single-step ReasoningKazakh90.2092.7493.92
English93.9295.1094.76
ASR · WER ↓Kazakh0.17070.14520.1688
English0.06620.06340.0601
Russian0.11360.10950.1077
WavCapsKazakh0.646.246.53
English3.8214.2815.03
Russian1.278.388.32
WavCaps-QAKazakh8.5525.3325.33
English24.3435.8634.21
Russian19.0832.2431.25

Model Usage

1. Transformers inference

To run the inference with transformers, complete the preliminary setup:

uv venv venv
source venv/bin/activate
uv pip install torch accelerate transformers

Then initialize the model and processor:

import torch
from transformers import AutoModelForCausalLM, AutoProcessor

model = AutoModelForCausalLM.from_pretrained(
    "issai/Qolda-AVL-5B",
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="auto"
)

processor = AutoProcessor.from_pretrained("issai/Qolda-AVL-5B", trust_remote_code=True)

Depending on the required modalities, define the messages list:

Language:

messages = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "y = (lnx)^2 функциясының туындысын тап. JSON форматында жауап бер: {'answer': '...'}"},
        ],
    }
]

Vision-Language:

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "assets/sample_image.jpg"}, # or provide link to the image
            {"type": "text", "text": "Суретті егжей-тегжейлі сипаттап бер. Неше жылқы көріп тұрсың және олардың түстері қандай?"},
        ],
    }
]

Audio-Language:

Note: The model was not trained to answer questions posed directly in the audio. Provide a detailed text instruction alongside the audio describing the task you want performed on it.

prompt = """Математикалық есепті шеш.
Respond ONLY with this JSON format: {"explanation": "<your step-by-step reasoning>", "answer": <integer or float number>}
The answer must be a number (integer or float). No text, no units, just the number.
"""

messages = [
    {
        "role": "user",
        "content": [
            {"type": "audio", "audio": "assets/sample_audio.wav"}, # or provide link to the audio
            {"type": "text", "text": prompt}
        ],
    }
]

Audio-Vision-Language:

messages = [
    {
        "role": "user",
        "content": [
            {"type": "audio", "audio": "assets/question_audio.wav"},
            {"type": "image", "image": "assets/sample_image.jpg"},
            {"type": "text", "text": "Answer the question"},
        ],
    }
]

Finally, pass the messages to the model for inference:

inputs = processor.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    return_dict=True, return_tensors="pt",
).to(model.device)

generated_ids = model.generate(
    **inputs,
    max_new_tokens=4096,
    temperature=0.7,
    top_p=0.95,
    top_k=20,
    do_sample=True,
    repetition_penalty=1.0,
)
generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(generated_ids_trimmed, skip_special_tokens=True)

print(processor.batch_decode(generated_ids_trimmed, skip_special_tokens=True)[0])

2. vLLM inference

Alternatively, you can run the model via a vLLM server. Note that we use a custom vLLM package. First, complete the preliminary setup:

uv venv venv
source venv/bin/activate

# Install this fork (precompiled binaries)
git clone https://github.com/IS2AI/vLLM-Qolda-AVL.git
cd vLLM-Qolda-AVL
VLLM_USE_PRECOMPILED=1 uv pip install -e .

Then start the OpenAI-compatible server (adjust parameters to your settings):

vllm serve issai/Qolda-AVL-5B \
    --served-model-name qolda-avl \
    --trust-remote-code \
    --tensor-parallel-size 4 \
    --dtype bfloat16 \
    --max-model-len 16384 \
    --limit-mm-per-prompt '{"audio": 1, "image": 1}'

To run inference, you can use the following code:

import base64
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1", 
    api_key="EMPTY"
)

def encode_audio_base64(path: str | Path) -> str:
    with open(path, "rb") as f:
        return base64.b64encode(f.read()).decode("utf-8")

def encode_image_base64(path: str | Path) -> str:
    with open(path, "rb") as f:
        return base64.b64encode(f.read()).decode("utf-8")

audio_path = "assets/sample_audio.wav"
audio_b64 = encode_audio_base64(audio_path)

stream = client.chat.completions.create(
    model=client.models.list().data[0].id,
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "input_audio",
                    "input_audio": {
                        "data": audio_b64,
                        "format": "wav",
                    },
                },
                {
                    "type": "text",
                    "text": (
                        "Analyze the voice in the audio and identify the speaker's "
                        "gender (male or female). Also transcribe what is said. "
                        "Return your answer as JSON in the following format: "
                        '{"answer": "<male or female>",'
                        '"transcription": "<transcription>"}'
                    ),
                },
            ],
        }
    ],
    max_tokens=4096,
    temperature=0.7,
    top_p=0.8,
    stream=True,
    stream_options={"include_usage": True},
)

text = ""
usage = None
for chunk in stream:
    if chunk.usage:
        usage = chunk.usage
    if chunk.choices and chunk.choices[0].delta.content:
        token = chunk.choices[0].delta.content
        print(token, end="", flush=True)
        text += token

License

Apache License 2.0

Citation


@article{qolda-avl-bdcc,
AUTHOR = {Arystanbekov, Batyr and Maxutov, Akylbek and Nurimanov, Aspandiyar and Varol, Huseyin Atakan},
TITLE = {Extending a Vision–Language Model with Audio Understanding: Introducing Qolda-AVL for the Kazakh Language},
JOURNAL = {Big Data and Cognitive Computing},
VOLUME = {10},
YEAR = {2026},
NUMBER = {6},
ARTICLE-NUMBER = {192},
URL = {https://www.mdpi.com/2504-2289/10/6/192},
ISSN = {2504-2289},
DOI = {10.3390/bdcc10060192}
}

Contributors

batyrme

3 commits

issai/Qolda-AVL-5B

Model

2

stars

6

commits

1

repos using this model

1

linked in READMEs

Aug 18, 2026

updated

custom_code
qwen3_vl
safetensors

README

Qolda-AVL

Qolda-AVL is a 5B audio-vision-language model designed to operate in Kazakh, Russian, and English. The model extends Qwen3-VL with an audio branch built on a fine-tuned Whisper encoder and a dedicated audio projection module. All three modalities are adapted to Kazakh through a staged training pipeline, with the audio branch covering speech recognition, speech translation, audio classification, and environmental sound captioning.

To improve audio feature injection into the language backbone, we apply the DeepStack mechanism to the audio branch, mirroring the vision processing pipeline of Qwen3-VL 💜

Qolda-AVL architecture

The model is our step towards omni-modal systems for the Kazakh language.

The name "Qolda" reflects both its design and purpose in Kazakh: "in hand" (қолда) for its compact accessibility, and "to support" (қолдау) for its assistive nature.

Evaluation

The benchmark suites and Qolda-AVL model collection are available here:

The following tables report benchmark results across the Qolda-AVL family together with Qolda baselines. Higher is better unless otherwise noted, and the best result within each row is shown in bold.

5B / 9B / 34B = Qolda-AVL variants · Qolda-NT = Qolda No-Think · Qolda-T = Qolda Think

Click to for Text benchmark results
BenchmarkLanguageQolda-AVL-5BQolda-AVL-9BQolda-AVL-34BQolda-NTQolda-T
MMLUKazakh73.0773.8981.9458.3969.28
English79.0980.9686.7170.0376.46
MMLU-ProKazakh62.2164.2873.6740.6257.68
English71.2672.6179.0058.0966.31
Russian66.7968.2976.3447.4262.45
GPQAKazakh47.9146.9758.7731.8938.60
English52.6853.3663.8139.4645.62
Russian48.7851.3459.3632.6040.19
ARCKazakh93.5994.2796.7686.1492.30
English96.9097.2998.1794.2296.11
Russian96.1196.6297.6391.1794.17
GSM8KKazakh85.7583.1790.5273.0183.00
English95.2295.8396.4462.8590.22
Russian90.9092.0494.3184.9983.98
MMLU-ReduxKazakh76.2476.9284.3960.0672.38
English82.6584.5688.1172.9179.40
KazCultureKazakh44.7556.3962.3753.0047.45
KazMMLUKazakh69.2773.0478.9858.1166.14
KazBenchKazakh61.8364.6170.0564.2361.12
BelebeleKazakh81.7084.7688.7881.0782.91
PIQAKazakh81.0078.0085.0063.0070.00
INCLUDEKazakh53.8057.2061.8045.2046.00
Russian58.1564.4970.8359.1756.52
KKCOPAKazakh76.6078.1179.6070.0073.79
NIS MathKazakh94.0093.0098.0066.0087.88
KazQADKazakh42.2865.7670.3670.9967.40
RAGBenchKazakh52.1262.9169.9854.9566.81
Click to for Vision benchmark results
BenchmarkLanguageQolda-AVL-5BQolda-AVL-9BQolda-AVL-34BQolda-NTQolda-T
RealWorldQAKazakh52.8150.0755.4253.8648.10
English67.8470.0771.7661.5761.70
Russian59.3560.3966.4156.0857.12
MMStarKazakh67.4568.8972.5053.0859.60
English70.2772.9375.9358.4865.04
Russian66.4769.4573.9355.4859.84
AI2DKazakh72.1873.3479.0563.4866.26
English79.4081.9684.5573.9975.61
MathVistaKazakh70.1772.3474.6558.3266.33
English76.7580.9482.5763.1471.04
MathVisionKazakh52.0755.7562.0635.4144.38
English54.9358.2463.9042.0048.05
MMBenchKazakh87.6989.1990.1879.9783.85
English87.8989.0190.4383.0584.40
OCRBenchKazakh30.6151.2553.9749.8946.49
English73.2077.2079.2069.9068.70
Click to for Audio benchmark results
BenchmarkLanguageQolda-AVL-5BQolda-AVL-9BQolda-AVL-34B
SAKURA · Animal · MultiKazakh48.5869.5470.00
English64.6082.2083.00
Russian50.4081.7683.00
SAKURA · Animal · SingleKazakh55.3187.8086.00
English52.2087.8083.80
Russian53.2388.4088.20
SAKURA · Emotion · MultiKazakh33.8736.9539.20
English35.8037.4043.20
Russian38.2037.2040.20
SAKURA · Emotion · SingleKazakh39.0347.2745.58
English42.6044.4047.40
Russian40.2844.8047.60
SAKURA · Gender · MultiKazakh67.9475.4081.60
English72.0083.2082.80
Russian70.4277.6083.20
SAKURA · Gender · SingleKazakh70.8882.8084.20
English80.0085.0087.80
Russian76.9282.6087.40
SAKURA · Language · MultiKazakh87.4088.8092.60
English92.6092.0094.40
Russian87.8090.8093.60
SAKURA · Language · SingleKazakh96.0097.6097.40
English96.4097.4097.40
Russian97.1997.6097.60
SpokenMQA · Long DigitKazakh81.4088.9591.28
English88.3793.0294.19
SpokenMQA · Multi-step ReasoningKazakh74.5777.6687.28
English87.4290.2392.53
SpokenMQA · Short DigitKazakh84.0092.0095.00
English88.0094.0092.00
SpokenMQA · Single-step ReasoningKazakh90.2092.7493.92
English93.9295.1094.76
ASR · WER ↓Kazakh0.17070.14520.1688
English0.06620.06340.0601
Russian0.11360.10950.1077
WavCapsKazakh0.646.246.53
English3.8214.2815.03
Russian1.278.388.32
WavCaps-QAKazakh8.5525.3325.33
English24.3435.8634.21
Russian19.0832.2431.25

Model Usage

1. Transformers inference

To run the inference with transformers, complete the preliminary setup:

uv venv venv
source venv/bin/activate
uv pip install torch accelerate transformers

Then initialize the model and processor:

import torch
from transformers import AutoModelForCausalLM, AutoProcessor

model = AutoModelForCausalLM.from_pretrained(
    "issai/Qolda-AVL-5B",
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="auto"
)

processor = AutoProcessor.from_pretrained("issai/Qolda-AVL-5B", trust_remote_code=True)

Depending on the required modalities, define the messages list:

Language:

messages = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "y = (lnx)^2 функциясының туындысын тап. JSON форматында жауап бер: {'answer': '...'}"},
        ],
    }
]

Vision-Language:

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "assets/sample_image.jpg"}, # or provide link to the image
            {"type": "text", "text": "Суретті егжей-тегжейлі сипаттап бер. Неше жылқы көріп тұрсың және олардың түстері қандай?"},
        ],
    }
]

Audio-Language:

Note: The model was not trained to answer questions posed directly in the audio. Provide a detailed text instruction alongside the audio describing the task you want performed on it.

prompt = """Математикалық есепті шеш.
Respond ONLY with this JSON format: {"explanation": "<your step-by-step reasoning>", "answer": <integer or float number>}
The answer must be a number (integer or float). No text, no units, just the number.
"""

messages = [
    {
        "role": "user",
        "content": [
            {"type": "audio", "audio": "assets/sample_audio.wav"}, # or provide link to the audio
            {"type": "text", "text": prompt}
        ],
    }
]

Audio-Vision-Language:

messages = [
    {
        "role": "user",
        "content": [
            {"type": "audio", "audio": "assets/question_audio.wav"},
            {"type": "image", "image": "assets/sample_image.jpg"},
            {"type": "text", "text": "Answer the question"},
        ],
    }
]

Finally, pass the messages to the model for inference:

inputs = processor.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    return_dict=True, return_tensors="pt",
).to(model.device)

generated_ids = model.generate(
    **inputs,
    max_new_tokens=4096,
    temperature=0.7,
    top_p=0.95,
    top_k=20,
    do_sample=True,
    repetition_penalty=1.0,
)
generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(generated_ids_trimmed, skip_special_tokens=True)

print(processor.batch_decode(generated_ids_trimmed, skip_special_tokens=True)[0])

2. vLLM inference

Alternatively, you can run the model via a vLLM server. Note that we use a custom vLLM package. First, complete the preliminary setup:

uv venv venv
source venv/bin/activate

# Install this fork (precompiled binaries)
git clone https://github.com/IS2AI/vLLM-Qolda-AVL.git
cd vLLM-Qolda-AVL
VLLM_USE_PRECOMPILED=1 uv pip install -e .

Then start the OpenAI-compatible server (adjust parameters to your settings):

vllm serve issai/Qolda-AVL-5B \
    --served-model-name qolda-avl \
    --trust-remote-code \
    --tensor-parallel-size 4 \
    --dtype bfloat16 \
    --max-model-len 16384 \
    --limit-mm-per-prompt '{"audio": 1, "image": 1}'

To run inference, you can use the following code:

import base64
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1", 
    api_key="EMPTY"
)

def encode_audio_base64(path: str | Path) -> str:
    with open(path, "rb") as f:
        return base64.b64encode(f.read()).decode("utf-8")

def encode_image_base64(path: str | Path) -> str:
    with open(path, "rb") as f:
        return base64.b64encode(f.read()).decode("utf-8")

audio_path = "assets/sample_audio.wav"
audio_b64 = encode_audio_base64(audio_path)

stream = client.chat.completions.create(
    model=client.models.list().data[0].id,
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "input_audio",
                    "input_audio": {
                        "data": audio_b64,
                        "format": "wav",
                    },
                },
                {
                    "type": "text",
                    "text": (
                        "Analyze the voice in the audio and identify the speaker's "
                        "gender (male or female). Also transcribe what is said. "
                        "Return your answer as JSON in the following format: "
                        '{"answer": "<male or female>",'
                        '"transcription": "<transcription>"}'
                    ),
                },
            ],
        }
    ],
    max_tokens=4096,
    temperature=0.7,
    top_p=0.8,
    stream=True,
    stream_options={"include_usage": True},
)

text = ""
usage = None
for chunk in stream:
    if chunk.usage:
        usage = chunk.usage
    if chunk.choices and chunk.choices[0].delta.content:
        token = chunk.choices[0].delta.content
        print(token, end="", flush=True)
        text += token

License

Apache License 2.0

Citation


@article{qolda-avl-bdcc,
AUTHOR = {Arystanbekov, Batyr and Maxutov, Akylbek and Nurimanov, Aspandiyar and Varol, Huseyin Atakan},
TITLE = {Extending a Vision–Language Model with Audio Understanding: Introducing Qolda-AVL for the Kazakh Language},
JOURNAL = {Big Data and Cognitive Computing},
VOLUME = {10},
YEAR = {2026},
NUMBER = {6},
ARTICLE-NUMBER = {192},
URL = {https://www.mdpi.com/2504-2289/10/6/192},
ISSN = {2504-2289},
DOI = {10.3390/bdcc10060192}
}

Contributors

batyrme

3 commits