Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published 3 days ago • 159
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published 25 days ago • 115
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? Paper • 2605.16679 • Published May 15 • 56
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering Paper • 2511.13998 • Published Nov 17, 2025 • 3
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering Paper • 2511.13998 • Published Nov 17, 2025 • 3 • 2
LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Paper • 2509.09614 • Published Sep 11, 2025 • 7
UserBench: An Interactive Gym Environment for User-Centric Agents Paper • 2507.22034 • Published Jul 29, 2025 • 31
Align and Attend: Multimodal Summarization with Dual Contrastive Losses Paper • 2303.07284 • Published Mar 13, 2023
MultiSum: A Dataset for Multimodal Summarization and Thumbnail Generation of Videos Paper • 2306.04216 • Published Jun 7, 2023
MHMS: Multimodal Hierarchical Multimedia Summarization Paper • 2204.03734 • Published Apr 7, 2022 • 1
Entity6K: A Large Open-Domain Evaluation Dataset for Real-World Entity Recognition Paper • 2403.12339 • Published Mar 19, 2024
Semantics-Consistent Cross-domain Summarization via Optimal Transport Alignment Paper • 2210.04722 • Published Oct 10, 2022 • 1
LiveSeg: Unsupervised Multimodal Temporal Segmentation of Long Livestream Videos Paper • 2210.05840 • Published Oct 12, 2022 • 1
Transfer Knowledge from Natural Language to Electrocardiography: Can We Detect Cardiovascular Disease Through Language Models? Paper • 2301.09017 • Published Jan 21, 2023