Agentic RL Daily

VOL. 038

8 research signals

DAILY EDITION / SAVED SNAPSHOT

Daily Signals: CoAnchor: Robust Collaborative Perception under Spatio-Temporal Misalignment via Object-Level Anchors

A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.

EDITOR'S VIEW

Three judgments

  1. The edition is based only on primary sources or official project releases.
  2. Older signals remain visible as continuing observations, not as rewritten news.
  3. Headline claims are constrained by the evidence included in this dated snapshot.

FULL EDITION

All signals in this edition

Archived / 2026-08-24

02

Open SourceOpen Source

OpenRLHF/OpenRLHF Release v0.11.0

OpenRLHF/OpenRLHF Release v0.11.0

Fix two latent bugs: dr_grpo n=1 guard and masked_normalize broadcast by @hijkzzz in https://github.com/OpenRLHF/OpenRLHF/pull/1250

Training AlgorithmsAgent Capabilities
GitHub ->

03

Open SourceOpen Source

verl-project/verl v0.9.0

verl-project/verl v0.9.0

Training with Megatron Lite (`m

Training AlgorithmsSystems Engineering记忆与自进化评估体系
GitHub ->

05

PapersPapers

AID-Guard: Stateful Authorization for Delegated Agent Effects

AID-Guard: Stateful Authorization for Delegated Agent Effects

Tool-using AI agents turn delegated tasks into provider effects, yet authorization often ends at admission while provider state, delivery, retry, and recovery evolve.

Training AlgorithmsAgent CapabilitiesData Loops
arXiv ->

06

PapersPapers

Large Language Models at the Intersection of Software Engineering and Software Security:An Evidence-Centered Structured Survey and Research Agenda

Large Language Models at the Intersection of Software Engineering and Software Security:An Evidence-Centered Structured Survey and Research Agenda

Large Language Models (LLMs) are moving from code completion toward repository-scale agents that retrieve context, edit files, execute tools, and participate in security-sensitive workflows.

Training AlgorithmsAgent Capabilities评估体系
arXiv ->