SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published 16 days ago • 157
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Paper • 2607.27616 • Published Jul 30 • 39
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms Paper • 2607.26497 • Published Jul 30 • 52
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Paper • 2607.27816 • Published Jul 30 • 34
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory Paper • 2607.24368 • Published Jul 27 • 33