Agentic RL Daily

VOL. 033

8 research signals

DAILY EDITION / SAVED SNAPSHOT

Daily Signals: GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities

A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.

EDITOR'S VIEW

Three judgments

  1. The edition is based only on primary sources or official project releases.
  2. Older signals remain visible as continuing observations, not as rewritten news.
  3. Headline claims are constrained by the evidence included in this dated snapshot.

FULL EDITION

All signals in this edition

Archived / 2026-08-19

01

PapersPapers

GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities

GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities

LLM-driven agents can autonomously exchange opinions on online platforms and form communities. Such agent-operated social platforms raise a new security concern: attackers may manipulate agents to induce group polarization. Existing methods manipulate agent prompts or construct echo chambers, both o

Agent Capabilities
arXiv ->

02

Open SourceOpen Source

OpenRLHF/OpenRLHF Release v0.11.0

OpenRLHF/OpenRLHF Release v0.11.0

## What's Changed * Fix two latent bugs: dr_grpo n=1 guard and masked_normalize broadcast by @hijkzzz in https://github.com/OpenRLHF/OpenRLHF/pull/1250 * fix: allow eval_dataset with MultiTurnAgentExecutor (#1242) by @codewithyug06 in https://github.com/OpenRLHF/OpenRLHF/pull/1251 * Fix Qwen3.5 ZeRO

Training AlgorithmsAgent Capabilities
GitHub ->

03

Open SourceOpen Source

verl-project/verl v0.9.0

verl-project/verl v0.9.0

# v0.9.0 ## Highlights ### Training #### Megatron - DeepSeek-V4 GRPO end-to-end with Megatron-Bridge actor/ref, vLLM rollout and FP8/MXFP4 weight transfer (#6473), plus a contiguous context-parallel layout (#7221) and CP fixes that make long-context DeepSeek-V4 runnable (#7297). - [Megatron Lite (`m

Training AlgorithmsSystems Engineering记忆与自进化评估体系
GitHub ->

04

PapersPapers

Benchmarking Automated Security Patch Backporting: How Far Are We?

Benchmarking Automated Security Patch Backporting: How Far Are We?

Automated security patch backporting is critical for mitigating N-day vulnerabilities. Recent tools report success rates above 80% on their respective datasets. However, these evaluations are often confined to homogeneous environments, such as one repository or specific project versions. Consequentl

Agent Capabilities
arXiv ->

05

PapersPapers

Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating

Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating

Autonomous LLM agents that converse on a user's behalf are an emerging design pattern in matching platforms, yet their viability depends on a condition rarely examined: users must accept not only delegating conversation to an agent, but also receiving agent-mediated communication from others. We stu

Agent Capabilities
arXiv ->

06

PapersPapers

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually dema

Agent Capabilities评估体系
arXiv ->

07

PapersPapers

EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection

EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection

Financial time series exhibit non-stationary and heterogeneous statistical properties, making change-point detection challenging because no single unsupervised algorithm performs consistently across assets and market regimes. Conventional workflows consequently depend heavily on expert-driven model

Training AlgorithmsAgent CapabilitiesData Loops记忆与自进化评估体系
arXiv ->

08

PapersPapers

D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory

D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory

Memory is a key capability of LLM agents. Persistent memory extends this across sessions---enabling recall, revision, and personalization. Yet its multi-stage pipeline (ingestion, retrieval, filtering, generation) makes failures difficult to localize: end-to-end evaluation reveals that an error occu

Agent CapabilitiesData LoopsSystems Engineering记忆与自进化评估体系
arXiv ->