Agentic RL Daily

VOL. 006

8 research signals

DAILY EDITION / SAVED SNAPSHOT

Daily Signals: CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.

EDITOR'S VIEW

Three judgments

  1. The edition is based only on primary sources or official project releases.
  2. Older signals remain visible as continuing observations, not as rewritten news.
  3. Headline claims are constrained by the evidence included in this dated snapshot.

FULL EDITION

All signals in this edition

Archived / 2026-07-22

01

PapersPapers

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer.

Training AlgorithmsAgent Capabilities评估体系
arXiv ->

02

PapersPapers

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted.

Training AlgorithmsAgent CapabilitiesData Loops安全与对齐
arXiv ->

03

PapersPapers

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics.

Training AlgorithmsAgent CapabilitiesSystems Engineering记忆与自进化评估体系
arXiv ->

04

Open SourceOpen Source

alibaba/ROLL v0.3.0

alibaba/ROLL v0.3.0

ROLL发布了v0.3.0版本,新增Video RLVR、AgentRunner 2.0、MTP训练、Router Replay、Multi-Teacher OPD等重要特性。

Training AlgorithmsReward and Credit AssignmentAgent CapabilitiesData LoopsSystems Engineering
GitHub ->

05

PapersPapers

Supra Cognitive Modes: A Routed Architecture for Agent Memory

Supra Cognitive Modes: A Routed Architecture for Agent Memory

Agent-memory workloads mix direct factual lookup, relation-chain and current-state reasoning, and broad synthesis over long histories.

Training AlgorithmsAgent Capabilities记忆与自进化评估体系
arXiv ->

06

PapersPapers

Agents in the Wild: Where Research Meets Deployment

Agents in the Wild: Where Research Meets Deployment

Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and more.

Agent Capabilities安全与对齐评估体系
arXiv ->

07

PapersPapers

They'll Verify. They Just Won't Act. How Authority Framing and Laundered Code Turn a Trusted Agentic CI/CD Pipeline Into an Attack Surface

They'll Verify. They Just Won't Act. How Authority Framing and Laundered Code Turn a Trusted Agentic CI/CD Pipeline Into an Attack Surface

We study a five-agent CI/CD pipeline (triage -> developer -> security-scan -> review -> approve/deploy), built from five distinct production LLMs across three providers, behind an LLM firewall in shadow mode.

Agent CapabilitiesSystems Engineering
arXiv ->

08

PapersPapers

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters, and assemble actors, objects, cameras, poses, and trajectories.

Agent Capabilities安全与对齐评估体系
arXiv ->