SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding Paper • 2604.13023 • Published Apr 14 • 2
LTX-2.5 Collection LTX-2.5 base models, quantized models and accompanying LoRAs and IC-LoRAs • 4 items • Updated 3 days ago • 38
Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors Paper • 2606.19325 • Published Jun 17 • 2
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 14 days ago • 38
view article Article IDEOGRAM-4 for inpainting with Modular Diffusers and Differential Diffusion OzzyGT • 9 days ago • 5
QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment Paper • 2607.24598 • Published 18 days ago • 7
StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation Paper • 2607.26754 • Published 16 days ago • 18
view article Article Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident +2 hlarcher, XciD, raphael-gl, chris-rannou • 18 days ago • 468
Laguna S 2.1 Collection Our most capable model to date, designed for long-horizon work. • 13 items • Updated 11 days ago • 43
view article Article Experimenting with the proposed Cross-Origin Storage API in Transformers.js tomayac • Jun 23 • 8
view article Article Be Ready Before the Attack: A Practical Guide to Self-Hosting an Open Model for Cyber Defense jeffboudier • 25 days ago • 20
view article Article NEO-unify: Building Native Multimodal Unified Models End to End sensenova • Mar 5 • 175