UCSC-VLAA/openvision2-vit-so400m-patch14-384-vision-only Image-to-Text • Updated 16 days ago • 15 • 1
VisualClaw: A Real-Time, Personalized Agent for the Physical World Paper • 2606.16295 • Published Jun 15 • 29
Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents Paper • 2605.30621 • Published May 28 • 26
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Paper • 2605.20177 • Published May 19 • 10
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning Paper • 2605.20176 • Published May 19 • 12
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning Paper • 2605.20176 • Published May 19 • 12
m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning with Large Language Models Paper • 2504.00869 • Published Apr 1, 2025 • 11
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Paper • 2605.20177 • Published May 19 • 10