AI summaries of arXiv papers
Paste any arXiv URL or paper ID. Or just swap arxiv.org → arxivtldr.org in your address bar.
Add to Chrome — Free ExtensionToday's Papers
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
VBVR-Pro is a new testbed for native visual reasoning, offering scalable tasks, verifiable rewards, and tools for studying generative models.
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
Zero-WAM enables robots to generalize to unseen manipulation tasks by using human videos as in-context guidance, achieving strong zero-shot performance.
RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing
RefVideo-6M is a new large-scale dataset with 6M visual references and reliable supervision for training advanced, controllable video editing models.
A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training
A novel VDA framework enables Multimodal Unsupervised Continual Post-Training (MU-CPT) by leveraging token-level visual dependence to prevent forgetting and boost new-task learning.
MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching
MyoMechanix introduces a multimodal dataset and AI system for biomechanically-grounded, fine-grained understanding and coaching of skilled physical activities.
Agentic Autoresearch for Cell-Edge Power Control: Radically Redefining the Researcher's Role
An AI agent autonomously designs and optimizes ML algorithms for complex wireless power control, achieving near-optimal performance with significantly reduced inference cost.
PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans
PlanSightRAG is a visual-first multimodal RAG system that automates civil plan compliance checking, achieving high accuracy by reasoning directly over plan imagery.
Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders
This paper applies sparse autoencoders to a neutrino foundation model to find interpretable physical concepts and significantly improve angular resolution.
From Producing to Validating: How AI Is Deskilling Freelancers
Generative AI is deskilling freelancers by shifting their roles from production to validation, posing risks to skill development and job security.
Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings
Planetary Prediction Engine (PPE) autonomously generates high-fidelity geospatial predictions by intelligently selecting data and using foundation model embeddings.
TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development
TraceML analyzes human-agent ML development, revealing agents' narrow planning loops compared to human experts, and that instructions only partially close the gap.
ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing
ICON decomposition provides multivariate concept-level explanations for deep representations, accurately quantifying concept importance by accounting for correlations.
Browse by Category
Browse by author →Browse by Topic
All topics →📬 Weekly AI Paper Digest
Get the top 10 AI/ML arXiv papers from the week — summarized, scored, and delivered to your inbox every Monday.
The URL swap trick
Reading a paper on arxiv.org? Just change the domain to arxivtldr.org in your address bar:
Paste a link
Enter any arXiv URL or paper ID
AI summarizes
Get a TLDR, key bullets, and why it matters
Read & share
Get key insights in seconds, share with colleagues