Agentic RL Daily

VOL. 040

8 research signals

DAILY EDITION / SAVED SNAPSHOT

Daily Signals: OpenRLHF/OpenRLHF Release v0.11.0

A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.

EDITOR'S VIEW

Three judgments

  1. The edition is based only on primary sources or official project releases.
  2. Older signals remain visible as continuing observations, not as rewritten news.
  3. Headline claims are constrained by the evidence included in this dated snapshot.

FULL EDITION

All signals in this edition

Archived / 2026-08-26

01

releaserelease

OpenRLHF/OpenRLHF Release v0.11.0

OpenRLHF/OpenRLHF Release v0.11.0

Fix two latent bugs: dr_grpo n=1 guard and masked_normalize broadcast by @hijkzzz in https://github.com/OpenRLHF/OpenRLHF/pull/1250

Training Algorithms
GitHub ->

03

releaserelease

verl-project/verl v0.9.0

verl-project/verl v0.9.0

Megatron - DeepSeek-V4 GRPO end-to-end with Megatron-Bridge actor/ref, vLLM rollout and FP8/MXFP4 weight transfer (#6473), plus a contiguous context-parallel layout (#7221) and CP fixes that make long-context DeepSeek-V4 runnable (#7297).

Training Algorithms
GitHub ->

07

researchresearch

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those errors compound.

Training Algorithms
arXiv ->

08

Open SourceOpen Source

Confident at the moment of action: belief miscalibration in LLM play under hidden information

Confident at the moment of action: belief miscalibration in LLM play under hidden information

官方来源补充信号:Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidence tracks correctness at the moment of acting. We test this in a hidden-information chess variant where royal status can be secretly, repeatedly relocated between pieces, and where an agent's stated probability distribution over the opponent's hidden royal piece -- elicited every turn, separately from the move it chooses -- is scored against ground truth recoverable after the game. Across two independent batches, captures made at high stated confidence ($\geq 0.5$) about the hidden piece's locat

Training AlgorithmsAgent Capabilities评估体系
arXiv ->