GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture Paper • 2608.15875 • Published 12 days ago • 98
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 4 days ago • 197
ANEForge: Python for direct computation on the Apple Neural Engine Paper • 2606.17090 • Published Jun 12 • 3
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL Paper • 2608.17253 • Published 9 days ago • 94
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe Paper • 2608.13667 • Published 15 days ago • 16
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution Paper • 2608.08311 • Published 20 days ago • 91
Next-Latent Prediction Transformers Learn Compact World Models Paper • 2511.05963 • Published Nov 8, 2025 • 4
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents Paper • 2608.05013 • Published 24 days ago • 37
view article Article Training a coding agent using the OpenCode harness in remote HF sandboxes with TRL and OpenEnv sergiopaniego • 23 days ago • 26
Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency Paper • 2604.04847 • Published Apr 6 • 3
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents Paper • 2607.22798 • Published Jul 24 • 62
IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation Paper • 2607.22375 • Published Jul 24 • 9
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Paper • 2607.20465 • Published May 19 • 56
3rd Place at CVPR 2026 CASTLE Challenge: Agentic Multi-View Long-Context Video Understanding via Hierarchical Knowledge Graph Retrieval Paper • 2606.01933 • Published Jun 1 • 1
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published Jul 23 • 26