Agentic RL Daily

VOL. 035

8 research signals

DAILY EDITION / SAVED SNAPSHOT

Daily Signals: Growth Without Us: Machine Consumers, Corporate Circularity, and the Decoupling of GDP from Humanity after AGI

A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.

EDITOR'S VIEW

Three judgments

  1. The edition is based only on primary sources or official project releases.
  2. Older signals remain visible as continuing observations, not as rewritten news.
  3. Headline claims are constrained by the evidence included in this dated snapshot.

FULL EDITION

All signals in this edition

Archived / 2026-08-21

01

PapersPapers

Growth Without Us: Machine Consumers, Corporate Circularity, and the Decoupling of GDP from Humanity after AGI

Growth Without Us: Machine Consumers, Corporate Circularity, and the Decoupling of GDP from Humanity after AGI

The standard objection to full automation is demand-side: if humans earn nothing, who buys the output? This confuses an accounting role with a biological species. We model a post-AGI economy in which corporations own populations of AI and robotic agents that are both producers and consumers of energy.

Agent CapabilitiesSystems Engineering安全与对齐
arXiv ->

02

Open SourceOpen Source

OpenRLHF/OpenRLHF Release v0.11.0

OpenRLHF/OpenRLHF Release v0.11.0

## What's Changed * Fix two latent bugs: dr_grpo n=1 guard and masked_normalize broadcast by @hijkzzz in https://github.com/OpenRLHF/OpenRLHF/pull/1250 * fix: allow eval_dataset with MultiTurnAgentExecutor (#1242) by @codewithyug06 in https://github.com/OpenRLHF/OpenRLHF/pull/1251 * Fix Qwen3.5 ZeRO

Training AlgorithmsAgent Capabilities
GitHub ->

03

Open SourceOpen Source

verl-project/verl v0.9.0

verl-project/verl v0.9.0

# v0.9.0 ## Highlights ### Training #### Megatron - DeepSeek-V4 GRPO end-to-end with Megatron-Bridge actor/ref, vLLM rollout and FP8/MXFP4 weight transfer (#6473), plus a contiguous context-parallel layout (#7221) and CP fixes that make long-context DeepSeek-V4 runnable (#7297). - [Megatron Lite (`m

Training AlgorithmsSystems Engineering记忆与自进化评估体系
GitHub ->

04

PapersPapers

From Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation

From Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation

Technical documentation is written for human developers, but an increasing share of software changes is now authored by autonomous coding agents. Which documents they consult, when, and what follows remain unknown. We conduct a behaviour-grounded study of agent-documentation interaction across two p

Training AlgorithmsAgent Capabilities
arXiv ->

05

PapersPapers

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

We present a novel approach to efficient LLM agent harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the underlying model weights. Existi

Training AlgorithmsAgent Capabilities评估体系
arXiv ->

06

PapersPapers

Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees

Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees

Loading reusable skill documents into a bounded context window is now the primary way large language model (LLM) agents acquire task-specific capabilities, which makes skill selection a first-order determinant of task performance and token cost. Yet current agents score skills independently by seman

Training AlgorithmsAgent Capabilities评估体系
arXiv ->

07

PapersPapers

Repo0: Design-Driven Zero-to-All Code Generation

Repo0: Design-Driven Zero-to-All Code Generation

Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly from natural-language

Agent Capabilities安全与对齐
arXiv ->

08

PapersPapers

A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries

A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries

Patients often submit short, underspecified queries to healthcare chatbots that lack the patient-specific information needed to determine an appropriate response. Although these queries may be linguistically clear, they can support multiple plausible answers depending on undisclosed factors such as

Training AlgorithmsAgent Capabilities安全与对齐评估体系
arXiv ->