Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild Paper • 2605.24213 • Published May 22 • 15
From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality Paper • 2607.13196 • Published 9 days ago • 26
From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality Paper • 2607.13196 • Published 9 days ago • 26
Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality Paper • 2509.10402 • Published Sep 12, 2025 • 6
Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality Paper • 2509.10402 • Published Sep 12, 2025 • 6 • 2
Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality Paper • 2509.10402 • Published Sep 12, 2025 • 6