Papers
arxiv:2607.10387

GigaChat Audio: Time-aware Large Audio Language Model

Published on Jul 11
· Submitted by
Aleksandr Kutsakov
on Jul 21
Authors:
,
,
,
,

Abstract

Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.

Community

Paper author Paper submitter

GigaChat Audio 10B (A1.8B) — an open-source, MIT-licensed audio-native LLM for speech and long-form audio understanding. It combines the GigaAM encoder with a 10B MoE decoder using 1.8B active parameters. The model supports ASR, translation, audio QA, emotion recognition, and temporal grounding for recordings up to two hours. With inter-timings—periodic temporal anchors embedded into the audio stream—it reaches 48.3 mIoU on temporal localization, compared with approximately 0.0 for other open-source models.

Sign up or log in to comment

Models citing this paper 3

Datasets citing this paper 1

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.