01
OpenRLHF/OpenRLHF Release v0.11.0
OpenRLHF/OpenRLHF Release v0.11.0
Fix two latent bugs: dr_grpo n=1 guard and masked_normalize broadcast by @hijkzzz in https://github.com/OpenRLHF/OpenRLHF/pull/1250
VOL. 040
8 research signalsDAILY EDITION / SAVED SNAPSHOT
A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.
EDITOR'S VIEW
FULL EDITION
Archived / 2026-08-26
01
OpenRLHF/OpenRLHF Release v0.11.0
Fix two latent bugs: dr_grpo n=1 guard and masked_normalize broadcast by @hijkzzz in https://github.com/OpenRLHF/OpenRLHF/pull/1250
02
Simthesizer: An Agent-Driven Simulation Framework for LLM Serving Systems
System-level simulation is an essential tool for exploring the rapidly expanding design space of LLM serving systems
03
verl-project/verl v0.9.0
Megatron - DeepSeek-V4 GRPO end-to-end with Megatron-Bridge actor/ref, vLLM rollout and FP8/MXFP4 weight transfer (#6473), plus a contiguous context-parallel layout (#7221) and CP fixes that make long-context DeepSeek-V4 runnable (#7297).
04
SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL
Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and variable tool-use trajectories.
05
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions.
06
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation.
07
CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those errors compound.
08
Confident at the moment of action: belief miscalibration in LLM play under hidden information
官方来源补充信号:Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidence tracks correctness at the moment of acting. We test this in a hidden-information chess variant where royal status can be secretly, repeatedly relocated between pieces, and where an agent's stated probability distribution over the opponent's hidden royal piece -- elicited every turn, separately from the move it chooses -- is scored against ground truth recoverable after the game. Across two independent batches, captures made at high stated confidence ($\geq 0.5$) about the hidden piece's locat