Agentic RL Daily

VOL. 029

8 research signals

DAILY EDITION / SAVED SNAPSHOT

Daily Signals: OpenRLHF/OpenRLHF Release v0.11.0

A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.

EDITOR'S VIEW

Three judgments

  1. OpenRLHF/OpenRLHF Release v0.11.0
  2. verl-project/verl v0.9.0
  3. Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks

FULL EDITION

All signals in this edition

Archived / 2026-08-15

01

releaserelease

OpenRLHF/OpenRLHF Release v0.11.0

OpenRLHF/OpenRLHF Release v0.11.0

## What's Changed * Fix two latent bugs: dr_grpo n=1 guard and masked_normalize broadcast by @hijkzzz in https://github.com/OpenRLHF/OpenRLHF/pull/1250 * fix: allow eval_dataset with MultiTurnAgentExecutor (#1242) by @codewithyug06 in https://github.com/OpenRLHF/OpenRLHF/pull/1251 * Fix Qwen3.5 ZeRO

Training Algorithms
GitHub ->

02

releaserelease

verl-project/verl v0.9.0

verl-project/verl v0.9.0

# v0.9.0 ## Highlights ### Training #### Megatron - DeepSeek-V4 GRPO end-to-end with Megatron-Bridge actor/ref, vLLM rollout and FP8/MXFP4 weight transfer (#6473), plus a contiguous context-parallel layout (#7221) and CP fixes that make long-context DeepSeek-V4 runnable (#7297). - [Megatron Lite (`m

Training AlgorithmsSystems Engineering记忆与自进化评估体系
GitHub ->

03

researchresearch

Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks

Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks

6G networks will not be serving as communication infrastructures only; rather, they are expected to evolve into intelligent systems, where thousands of autonomous artificial intelligence (AI) agents are interconnected. The agents are deployed across a wide range of platforms including low Earth orbi

Agent CapabilitiesSystems Engineering安全与对齐评估体系
arXiv ->

04

researchresearch

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents

Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a single reward signal. On-policy distillation provides dense per-tok

Training AlgorithmsReward and Credit AssignmentAgent CapabilitiesData Loops安全与对齐
arXiv ->

05

researchresearch

Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test

Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test

Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared state. We model this failure with a finite \emph{capability sheaf}: stalks encode typed behavior signatures, restriction maps retain shared fields, and accepted r

Training AlgorithmsAgent Capabilities
arXiv ->

06

researchresearch

Vero: Can AI Agents Build Formally Verified Software Repositories?

Vero: Can AI Agents Build Formally Verified Software Repositories?

AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code generation, in which an agent produces both an implementation and a machine-checked proof of its specification, offers a stronger path toward trustworthy AI-generated

Training AlgorithmsAgent Capabilities评估体系
arXiv ->

07

researchresearch

MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with e

Training AlgorithmsAgent Capabilities评估体系
arXiv ->

08

researchresearch

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empiric

Training AlgorithmsAgent Capabilities评估体系
arXiv ->