ArXiv TLDR
21,482 papers summarized

AI summaries of arXiv papers

Paste any arXiv URL or paper ID. Or just swap arxiv.org arxivtldr.org in your address bar.

Add to Chrome — Free Extension
Computer Vision

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

VBVR-Pro is a new testbed for native visual reasoning, offering scalable tasks, verifiable rewards, and tools for studying generative models.

2608.26105
Robotics

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

Zero-WAM enables robots to generalize to unseen manipulation tasks by using human videos as in-context guidance, achieving strong zero-shot performance.

2608.26103
Computer Vision

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

RefVideo-6M is a new large-scale dataset with 6M visual references and reliable supervision for training advanced, controllable video editing models.

2608.26101
Computer Vision

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

A novel VDA framework enables Multimodal Unsupervised Continual Post-Training (MU-CPT) by leveraging token-level visual dependence to prevent forgetting and boost new-task learning.

2608.26095
Computer Vision

MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching

MyoMechanix introduces a multimodal dataset and AI system for biomechanically-grounded, fine-grained understanding and coaching of skilled physical activities.

2608.26094
Machine Learning

Agentic Autoresearch for Cell-Edge Power Control: Radically Redefining the Researcher's Role

An AI agent autonomously designs and optimizes ML algorithms for complex wireless power control, achieving near-optimal performance with significantly reduced inference cost.

2608.26093
Information Retrieval

PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans

PlanSightRAG is a visual-first multimodal RAG system that automates civil plan compliance checking, achieving high accuracy by reasoning directly over plan imagery.

2608.26091

Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders

This paper applies sparse autoencoders to a neutrino foundation model to find interpretable physical concepts and significantly improve angular resolution.

2608.26090
Human-Computer Interaction

From Producing to Validating: How AI Is Deskilling Freelancers

Generative AI is deskilling freelancers by shifting their roles from production to validation, posing risks to skill development and job security.

2608.26089
Artificial Intelligence

Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings

Planetary Prediction Engine (PPE) autonomously generates high-fidelity geospatial predictions by intelligently selecting data and using foundation model embeddings.

2608.26088
Machine Learning

TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

TraceML analyzes human-agent ML development, revealing agents' narrow planning loops compared to human experts, and that instructions only partially close the gap.

2608.26086
Machine Learning

ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing

ICON decomposition provides multivariate concept-level explanations for deep representations, accurately quantifying concept importance by accounting for correlations.

2608.26083

📬 Weekly AI Paper Digest

Get the top 10 AI/ML arXiv papers from the week — summarized, scored, and delivered to your inbox every Monday.

The URL swap trick

Reading a paper on arxiv.org? Just change the domain to arxivtldr.org in your address bar:

arxiv.org/abs/2401.12345
arxivtldr.org/abs/2401.12345
1

Paste a link

Enter any arXiv URL or paper ID

2

AI summarizes

Get a TLDR, key bullets, and why it matters

3

Read & share

Get key insights in seconds, share with colleagues