Agentic RL Daily

VOL. 027

8 research signals

DAILY EDITION / SAVED SNAPSHOT

Daily Signals: One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.

EDITOR'S VIEW

Three judgments

  1. The edition is based only on primary sources or official project releases.
  2. Older signals remain visible as continuing observations, not as rewritten news.
  3. Headline claims are constrained by the evidence included in this dated snapshot.

FULL EDITION

All signals in this edition

Archived / 2026-08-13

01

PapersPapers

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM pol

Training AlgorithmsAgent Capabilities安全与对齐评估体系
arXiv ->

02

PapersPapers

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely bottlenecked by the slow, manual process of dense reward shaping. To bypass this limitation, we leverage Sample-based Model Predictive Control (SMPC)

Training AlgorithmsReward and Credit AssignmentAgent Capabilities
arXiv ->

03

PapersPapers

DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation

DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation

Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observability. Although recent VLA models offer a promising perception-to-action paradigm, adapting them t

Agent CapabilitiesSystems Engineering记忆与自进化评估体系
arXiv ->

04

PapersPapers

Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents

Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents

LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two sequential control points to untrusted publishers: a static skill may steer an otherwise correct task onto an unne

Agent CapabilitiesData Loops安全与对齐评估体系
arXiv ->

05

PapersPapers

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stochasticity causes minor variations in textual prompts or hyperpa

Training AlgorithmsAgent Capabilities评估体系
arXiv ->

06

PapersPapers

Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation

Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation

As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation problems such as the Prisoner's Dilemma are resolva

Training AlgorithmsAgent Capabilities
arXiv ->

07

PapersPapers

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap betw

Training AlgorithmsAgent Capabilities安全与对齐
arXiv ->

08

PapersPapers

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K}nowledge \textbf{R}etrieval \textbf{A}gents), a benchmark of over $

Agent Capabilities评估体系
arXiv ->