Improving Test-Time Scaling with Adaptive Looped Transformers Paper • 2609.35748 • Published 10 days ago • 59
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure Paper • 2609.26550 • Published 16 days ago • 42
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Paper • 2609.24972 • Published 17 days ago • 223
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation Paper • 2608.21500 • Published Aug 21 • 41
Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning Paper • 2511.16043 • Published Nov 20, 2025 • 110
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models Paper • 2511.08577 • Published Nov 11, 2025 • 110
When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents? Paper • 2510.17862 • Published Oct 15, 2025 • 8
Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs? Paper • 2510.01161 • Published Oct 1, 2025 • 14