Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models Paper • 2608.13760 • Published 29 days ago • 3
Multimodal Model Diffing for Feature Discovery and Control Paper • 2608.09928 • Published Aug 10 • 11
Multimodal Model Diffing for Feature Discovery and Control Paper • 2608.09928 • Published Aug 10 • 11
Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis Paper • 2603.06507 • Published Mar 6 • 7
ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels Paper • 2608.03507 • Published Aug 4 • 4
ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels Paper • 2608.03507 • Published Aug 4 • 4
MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following Paper • 2605.03858 • Published May 5 • 1
Towards Understanding Multimodal Fine-Tuning: Spatial Features Paper • 2602.08713 • Published Feb 6 • 1
EVCL: Elastic Variational Continual Learning with Weight Consolidation Paper • 2406.15972 • Published Jun 23, 2024 • 1
Reasoning Fine-Tuning Induces Persistent Latent Policy States Paper • 2607.18532 • Published Jul 20 • 1
Constitutional Midtraining: Content Presence Drives Alignment Gains Paper • 2607.26654 • Published Jul 29 • 7
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks Paper • 2605.01417 • Published May 2 • 2
Measuring what Matters: Construct Validity in Large Language Model Benchmarks Paper • 2511.04703 • Published Nov 3, 2025 • 8
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought Paper • 2403.05518 • Published Mar 8, 2024 • 3
Constitutional Midtraining: Content Presence Drives Alignment Gains Paper • 2607.26654 • Published Jul 29 • 7