PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents Paper • 2608.26530 • Published 6 days ago • 30
Agentic Transaction: Towards ACID-Compliant Agent Systems Paper • 2608.13900 • Published 19 days ago • 27
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks Paper • 2608.14905 • Published 19 days ago • 30
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published 16 days ago • 150
view article Article Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers +1 tomaarsen, NohTow, raphaelsty • 15 days ago • 100