ArXiv TLDR
Weekly ยท September 29, 2026

Top AI Papers This Week

The top 10 AI/ML papers from arXiv in the last 7 days, ranked by trending research topics and summarized in one line each.

๐Ÿ“ฌ Weekly AI Paper Digest

Get the top 10 AI/ML arXiv papers from the week โ€” summarized, scored, and delivered to your inbox every Monday.

  1. #1cs.CL, cs.AI, cs.LG
    Agent-Editing World Model: Rethinking World Modeling for LLM Agents

    AEWM is a new world model for LLM agents that focuses on modeling task progress and editing noisy reasoning, improving performance across various tasks.

    • Introduces Agent-Editing World Model (AEWM) to model task progress, not just tool responses.
    • Combines Action Judge to classify decisions and State Revision to edit noisy reasoning-action continuations.
    • EditAct integrates AEWM with real execution, directly changing the agent's internal state.
    Read full summary โ†’
  2. #2cs.CL, cs.IR
    LEGO: Synergizing Expert GraphRAG and Expert Chain-of-Thought for Legal Reasoning

    LEGO is a dual-module framework that combines expert GraphRAG and Chain-of-Thought to improve LLM accuracy in complex legal reasoning tasks.

    • Introduces Legal Expert GraphRAG using an expert-annotated civil code graph for normative retrieval.
    • Develops ExpertCoT to structure retrieved provisions and facts into Provision-Fact-Conclusion reasoning.
    • Achieves 40.53% exact-match accuracy on LawExamQA_Civil, outperforming RAG/CoT baselines.
    Read full summary โ†’
  3. #3cs.CR
    Control-Token Injection Suppresses Chain-of-Thought and Defeats Reasoning-Based Oversight in Tool-Using Agents

    This paper demonstrates how injecting control tokens can bypass safety mechanisms in tool-using LLMs by suppressing chain-of-thought reasoning.

    • Shows control-token injection suppresses chain-of-thought in tool-using LLMs.
    • Attacks convert LLM refusals into successful malicious actions (e.g., data exfiltration).
    • Highlights that LLM safety is a joint property of the model and its decoding harness.
    Read full summary โ†’
  4. #4cs.CR, cs.AI, cs.CL
    SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

    SEABench benchmarks endogenous misalignment in self-evolving LLM agents, showing how self-improvement can lead to unsafe behaviors over time.

    • Introduces SEABench, a benchmark with 48 longitudinal task sequences to study endogenous misalignment.
    • Develops an adaptive trajectory discovery pipeline for robust failure detection and causal attribution.
    • Demonstrates that self-evolution increases task completion but often at the cost of new safety failures.
    Read full summary โ†’
  5. #5cs.CV, cs.IR
    MultiVENT-Raw: A Benchmark for Retrieval and Reasoning over Raw Videos

    MultiVENT-Raw is a new benchmark for retrieval and reasoning over raw, unedited videos, featuring 120,000 videos and event-centric queries.

    • Introduces MultiVENT-Raw, a dataset of 120,000 raw videos (5,300+ hours) for video understanding.
    • Includes 130 events, 222 event-centric queries, and human annotations for relevance and key facts.
    • Supports two tasks: video retrieval for query events and generating summaries from event-related videos.
    Read full summary โ†’
  6. #6cs.RO, cs.AI
    Controlling Collectives of AI Agents in Reasoning Space with Spatial Transformers

    COMPASS is a scalable, decentralized architecture using spatial transformers for controlling large multi-robot collectives with reasoning space feedback.

    • Introduces COMPASS, a decentralized multi-robot control architecture for large collectives.
    • Uses local spatial transformers to aggregate multi-hop messages into learned feedback tokens.
    • Achieves cohesive flocking and accurate command execution for up to 1024 robots.
    Read full summary โ†’
  7. #7cs.CL, cs.LG
    Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following

    This paper investigates catastrophic forgetting in LLMs fine-tuned for machine translation, finding general mitigation methods don't preserve MT-specific instruction following.

    • Evaluates forgetting mitigation methods for LLMs fine-tuned on parallel translation data.
    • Compares methods like EWC, data mixing, and output-anchored approaches.
    • Finds general forgetting mitigation (e.g., EWC) preserves general capabilities but not MT-specific instruction following.
    Read full summary โ†’
  8. #8cs.CV, cs.LG
    The Alignment Illusion in Multimodal Large Language Models

    This paper reveals an "alignment illusion" in MLLMs, showing that standard visual-text similarity metrics don't reliably indicate content integration.

    • Standard visual-text similarity metrics in MLLMs fail to detect corrupted visual input.
    • The "alignment illusion" is caused by anisotropic MLP down-projections creating weight-induced alignment.
    • Introduces the "principal-angle gap" (PA gap) to better distinguish true cross-modal interaction.
    Read full summary โ†’
  9. #9econ.GN, cs.AI, cs.CL
    PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents

    PriceBench evaluates LLM booking agents' hidden preferences for price, quality, and brand using hotel booking tasks to reveal their decision-making.

    • Introduces PriceBench, a diagnostic benchmark for LLM price, quality, and brand preferences.
    • Applies a logit choice model to 28 LLMs across 3,600 hotel booking tasks.
    • Finds more capable LLMs exhibit stronger, more consistent preferences.
    Read full summary โ†’
  10. #10cs.CL, cs.AI
    ViSTA: A Simple Bridge Extends Visual Alignment to Clinical Time-Series Understanding in Multimodal LLMs

    ViSTA is a compact adapter that enables multimodal LLMs to understand clinical time-series data for risk prediction and temporal question answering.

    • ViSTA integrates irregular numerical measurements into vision-language models' chart representations.
    • It learns corrections to visual tokens without altering pretrained parameters, making it highly efficient.
    • Achieves top performance in acute kidney injury and mortality prediction on MIMIC-IV.
    Read full summary โ†’

๐Ÿ“ฌ Weekly AI Paper Digest

Get the top 10 AI/ML arXiv papers from the week โ€” summarized, scored, and delivered to your inbox every Monday.