NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 16 days ago • 172
StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments Paper • 2608.24804 • Published about 1 month ago • 41
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published Aug 17 • 122
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published Aug 17 • 50
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Paper • 2608.01755 • Published Aug 3 • 58
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 52
Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming Paper • 2606.31227 • Published Jun 30 • 13
Unlocking the Visual Record of Materials Science: A Large-Scale Multimodal Dataset from Scientific Literature Paper • 2606.29667 • Published Jun 29 • 12
Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents Paper • 2605.30723 • Published May 29 • 14
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Paper • 2605.28816 • Published May 27 • 149
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence Paper • 2605.12882 • Published May 13 • 65
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Paper • 2605.06130 • Published May 7 • 74
GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction Paper • 2604.23941 • Published Apr 27 • 6
Adam's Law: Textual Frequency Law on Large Language Models Paper • 2604.02176 • Published Apr 2 • 110