Agentic RL Daily

VOL. 004

8 research signals

DAILY EDITION / SAVED SNAPSHOT

Agentic RL Daily

A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.

EDITOR'S VIEW

Three judgments

  1. The edition is based only on primary sources or official project releases.
  2. Older signals remain visible as continuing observations, not as rewritten news.
  3. Headline claims are constrained by the evidence included in this dated snapshot.

FULL EDITION

All signals in this edition

Archived / 2026-07-15

01

PapersPapers

Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing

Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing

The rapid emergence of LLM-powered autonomous and semi-autonomous agents is reshaping software systems from static, request-response components into goal-directed, adaptive, and tool-using computational actors. As these agents move from isolated cognitive prototypes into complex distributed workflow

Agent Capabilities
arXiv ->

02

PapersPapers

Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration

Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration

The paper introduces the Internet of Agentic Things (IoAT), an architectural framework that integrates agentic AI, IoT, cyber-physical systems, Physical AI, edge computing, and digital twins into a unified closed-loop orchestration framework. The proposed architecture consists of cloud, edge/fog, an

Agent Capabilities
arXiv ->

03

PapersPapers

A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study

A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study

Clinical notes contain many of the signs and symptoms that bring patients to care, yet this information rarely reaches structured fields. Existing extraction approaches either rely on context-insensitive rules that generate false positives or on supervised models that require substantial fine-tuning

Training Algorithms
arXiv ->

04

Open SourceOpen Source

OpenRLHF/OpenRLHF Release v0.10.4

OpenRLHF/OpenRLHF Release v0.10.4

## What's Changed * fix: only pass min_lr_rate to schedulers that accept it by @matteolippi in https://github.com/OpenRLHF/OpenRLHF/pull/1238 * Upgrade vLLM to 0.22.1 and DeepSpeed to 0.19.1 by @hijkzzz in https://github.com/OpenRLHF/OpenRLHF/pull/1248 * Fix token-level loss (global token-mean acros

Systems Engineering
GitHub ->

05

PapersPapers

TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains. We present TerraZero, a

Agent Capabilities
arXiv ->

06

PapersPapers

Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determine parameters with one-shot estimators, which makes

Agent Capabilities
arXiv ->

07

PapersPapers

Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs

Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs

Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted deductions need not hold up under formal scrutiny. We present EG-VAR (Evidence-Grounded Verified Agentic Reasoning), a Lean 4-based tool-calling architecture in

Agent Capabilities
arXiv ->

08

PapersPapers

Bulkhead: Automated Semantic Detection and Remediation of Container Escape Vulnerabilities

Bulkhead: Automated Semantic Detection and Remediation of Container Escape Vulnerabilities

Filesystem isolation in container ecosystems is often weakened by cross-boundary path misresolution, causing path traversal (PaTra) vulnerabilities. These vulnerabilities stem from insecure host-container interactions and have become increasingly pervasive as cloud systems mount shared resources, su

Agent Capabilities
arXiv ->