Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents Paper • 2609.37236 • Published 3 days ago • 33
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness Paper • 2502.19412 • Published Feb 17
Reinforcement Learning with Large Action Spaces for Neural Machine Translation Paper • 2210.03053 • Published Oct 6, 2022
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results Paper • 2606.14516 • Published Jun 12 • 7
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Paper • 2606.09809 • Published Jun 8 • 5
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration Paper • 2604.12843 • Published Apr 15 • 1
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks Paper • 2605.28556 • Published May 27 • 75
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents Paper • 2605.22608 • Published May 21 • 6
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Paper • 2605.19196 • Published May 18
Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization Paper • 2510.05038 • Published Oct 6, 2025 • 1
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents Paper • 2605.22608 • Published May 21 • 6
AlephBERT:A Hebrew Large Pre-Trained Language Model to Start-off your Hebrew NLP Application With Paper • 2104.04052 • Published Apr 8, 2021
Lexical Generalization Improves with Larger Models and Longer Training Paper • 2210.12673 • Published Oct 23, 2022