Agentic RL Daily

VOL. 004

8 research signals

DAILY EDITION / SAVED SNAPSHOT

Agentic RL Daily

A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.

EDITOR'S VIEW

Three judgments

  1. The edition is based only on primary sources or official project releases.
  2. Older signals remain visible as continuing observations, not as rewritten news.
  3. Headline claims are constrained by the evidence included in this dated snapshot.

FULL EDITION

All signals in this edition

Archived / 2026-07-20

01

PapersPapers

When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis

When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis

Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting this substitution is essentially never tested against the baseline it claims to replace: methods merge independently released agents precisely because a joint model is unavailable. We b

Training AlgorithmsAgent Capabilities安全与对齐评估体系
arXiv ->

02

PapersPapers

Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving

Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving

Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires. Existing closed-loop agents hide this gap by invoking the model on alternate simulation ticks and replaying the previous command

Training AlgorithmsAgent CapabilitiesData Loops评估体系
arXiv ->

03

PapersPapers

Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents

Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents

Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic components such as reflection and memory lead to observable and controllable improvements over fixed LLM workflows. We study this question through conference-paper

Agent Capabilities记忆与自进化评估体系
arXiv ->

04

IndustryIndustry

alibaba/ROLL v0.3.0

alibaba/ROLL v0.3.0

A verified industry signal preserved from the original daily snapshot. The primary source is linked for full context, while the archive keeps the original publication date and source attribution intact.

Training AlgorithmsReward and Credit AssignmentAgent CapabilitiesData LoopsSystems Engineering
GitHub ->

05

PapersPapers

SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery

SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery

Scientific work increasingly spans heterogeneous artifacts -- papers, code, datasets, scientific file formats, model outputs, figures, manuscripts, and team decisions -- yet general-purpose AI assistants rarely preserve these objects as a coherent, auditable research state. We present SciForge, a mu

Training AlgorithmsAgent Capabilities
arXiv ->

06

PapersPapers

Perceived AGI: Believability as Dimensional Completeness, Not Capability

Perceived AGI: Believability as Dimensional Completeness, Not Capability

Large language models are broadly capable, yet in sustained one-to-one conversation they still read as flat: competent, responsive, and somehow not quite the presence of a mind. We hypothesize that a central missing ingredient is not more capability but dimensional completeness. We propose that the

Agent Capabilities评估体系
arXiv ->

07

PapersPapers

AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets

AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets

Geospatial datasets support applications from urban planning to climate modeling, yet consistent assessment of FAIR compliance is difficult. Existing evaluators use different rubrics and evidence sources and may fail on JavaScript-rendered pages or repository-specific identifiers. For 50 datasets fr

Training AlgorithmsAgent Capabilities安全与对齐评估体系
arXiv ->

08

PapersPapers

Knowledge-Centric Agents for Workflow Generation

Knowledge-Centric Agents for Workflow Generation

Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions. Existing large language model (LLM) approaches often treat this as a direct text-to-JSON generation task, struggling with structural brittlene

Agent Capabilities评估体系
arXiv ->