Abstract
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.
Community
GigaChat Audio 10B (A1.8B) — an open-source, MIT-licensed audio-native LLM for speech and long-form audio understanding. It combines the GigaAM encoder with a 10B MoE decoder using 1.8B active parameters. The model supports ASR, translation, audio QA, emotion recognition, and temporal grounding for recordings up to two hours. With inter-timings—periodic temporal anchors embedded into the audio stream—it reaches 48.3 mIoU on temporal localization, compared with approximately 0.0 for other open-source models.
Models citing this paper 3
ai-babai/gigachat-audio-mlx-q8-bf16
Datasets citing this paper 1
ai-sage/TimeGround-1M
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper