Agentic RL Daily

VOL. 028

8 research signals

DAILY EDITION / SAVED SNAPSHOT

Daily Signals: OpenRLHF/OpenRLHF Release v0.11.0

A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.

EDITOR'S VIEW

Three judgments

  1. The edition is based only on primary sources or official project releases.
  2. Older signals remain visible as continuing observations, not as rewritten news.
  3. Headline claims are constrained by the evidence included in this dated snapshot.

FULL EDITION

All signals in this edition

Archived / 2026-08-14

01

PapersPapers

OpenRLHF/OpenRLHF Release v0.11.0

OpenRLHF/OpenRLHF Release v0.11.0

Fix two latent bugs: dr_grpo n=1 guard and masked_normalize broadcast by @hijkzzz in https://github.com/OpenRLHF/OpenRLHF/pull/1250

Agent CapabilitiesTraining Algorithms
GitHub ->

04

PapersPapers

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows.

Agent CapabilitiesTraining Algorithms
arXiv ->

05

PapersPapers

Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation

Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation

As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes.

Agent CapabilitiesTraining Algorithms
arXiv ->

08

Open SourceOpen Source

VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

官方来源补充信号:Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the full range of vulnerable software versions. Existing vulnerability datasets suffer from limited programming language coverage, restricted patch complexity, and narrow project scope. Through our dual annotation by human experts and an agentic workflow, we create a benchmark - VICBench - of 100 verified VICs for 100 CVEs across 88 projects in Python, Java, and C++, coverin

Agent Capabilities评估体系
arXiv ->