One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows Paper • 2608.19741 • Published 5 days ago • 7
GameXpert-Bench: How Far Are Coding Agents from Expert Game Development? Paper • 2608.21833 • Published 3 days ago • 10
HPSD: Hybrid-Policy Self-Distillation for Text-Image-to-Video Diffusion Models Paper • 2608.13205 • Published 12 days ago • 3
Intern-S2-Preview: Scientific Agentic Foundation Model Paper • 2608.13505 • Published 12 days ago • 70
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design Paper • 2608.10299 • Published 15 days ago • 135
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published about 1 month ago • 125
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published Jul 23 • 26
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Paper • 2607.13705 • Published Jul 15 • 45
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published Jul 16 • 106
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published Jul 14 • 235
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Paper • 2607.05382 • Published Jul 9 • 88
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published Jul 11 • 71
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published Jul 13 • 85
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Paper • 2607.11849 • Published Jul 13 • 33
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 149
WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence Paper • 2607.06838 • Published Jul 7 • 14
PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception Paper • 2606.28322 • Published Jun 26 • 43
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning Paper • 2607.01191 • Published Jul 1 • 19
Dockerless: Environment-Free Program Verifier for Coding Agents Paper • 2606.28436 • Published Jun 26 • 116