view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 2 days ago • 46
view article Article Training a coding model to paint watercolours with TRL and OpenEnv sergiopaniego • 2 days ago • 45
Running on CPU Upgrade 64 H3 Acceleration Arena 🥇 64 Blind A/B ranking of MiniMax-H3 acceleration variants
Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion Paper • 2608.26794 • Published 9 days ago • 16
Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models Paper • 2608.23478 • Published 12 days ago • 30
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published 9 days ago • 36