Agentic RL Daily

VOL. 021

8 research signals

DAILY EDITION / SAVED SNAPSHOT

Daily Signals: Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.

EDITOR'S VIEW

Three judgments

  1. The edition is based only on primary sources or official project releases.
  2. Older signals remain visible as continuing observations, not as rewritten news.
  3. Headline claims are constrained by the evidence included in this dated snapshot.

FULL EDITION

All signals in this edition

Archived / 2026-08-07

01

PapersPapers

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the principle that governance should control an AI agent through resource allocation so as to make authorization self enforcing via compute budgets. The mechanism see

Agent CapabilitiesTraining Algorithms安全与对齐
arXiv ->

02

PapersPapers

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajectory that is responsible for the final failure. However, progress face

Agent CapabilitiesTraining AlgorithmsData Loops评估体系
arXiv ->

03

PapersPapers

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization -- the iterative and evaluation-guid

Agent CapabilitiesTraining Algorithms记忆与自进化评估体系
arXiv ->

04

PapersPapers

From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over Networks

From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over Networks

Despite advances in artificial intelligence (AI) across multiple sectors, today's AI tools, including deep learning and generative AI, still fail when embedded into physical systems, such as robots and vehicles operating under real-world physical laws. This stems from their inability to maintain rel

Agent CapabilitiesTraining AlgorithmsSystems Engineering评估体系
arXiv ->

05

PapersPapers

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the result is settled or stop before the agents can be

Reward and Credit AssignmentAgent Capabilities评估体系
arXiv ->

06

PapersPapers

Learning Globally Reusable Skills for Coding Agents

Learning Globally Reusable Skills for Coding Agents

Automated skill evolution enables Large Language Model (LLM) agents to continuously improve without expensive retraining. However, existing approaches typically treat skill evolution as a sequence of local updates, overlooking relationships among skills and often producing overfitted skill updates t

Agent CapabilitiesData Loops
arXiv ->

07

PapersPapers

Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture

Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture

AI agents performing cryptographic operations (signing Git commits, authenticating API calls, issuing certificates) currently store private keys in software-accessible locations: plaintext files, environment variables, or container memory. Any process with sufficient read privileges can extract the

Agent Capabilities记忆与自进化评估体系
arXiv ->

08

PapersPapers

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults and requires integrating fragmented EHR data wi

Agent CapabilitiesTraining Algorithms评估体系
arXiv ->