Agentic RL Daily

VOL. 012

8 research signals

DAILY EDITION / SAVED SNAPSHOT

Daily Signals: Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL

A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.

EDITOR'S VIEW

Three judgments

  1. The edition is based only on primary sources or official project releases.
  2. Older signals remain visible as continuing observations, not as rewritten news.
  3. Headline claims are constrained by the evidence included in this dated snapshot.

FULL EDITION

All signals in this edition

Archived / 2026-07-29

01

PapersPapers

Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL

Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL

Large language model agents often spend substantial wall-clock time waiting for tool call results. Tool-call speculation can hide this latency by predicting and pre-executing an agent's next tool call if the prediction matches the agent's eventual tool call, but existing speculators are typically se

Training AlgorithmsAgent Capabilities
arXiv ->

02

PapersPapers

Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects a small task-relevant subset from a library of thousands of tools before the agent acts, has therefore become a critical component of LLM agent pipelines. Howeve

Training AlgorithmsAgent Capabilities
arXiv ->

03

PapersPapers

Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories

Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories

Agent skills package reusable procedures that improve downstream performance. Their lightweight, portable form enables marketplace monetization and private deployment behind cloud-hosted agent interfaces, giving providers incentives to keep high-value skills proprietary. Yet hiding the artifacts doe

Agent Capabilities评估体系
arXiv ->

04

Open SourceOpen Source

alibaba/ROLL v0.3.0

alibaba/ROLL v0.3.0

A verified open source signal preserved from the original daily snapshot. The primary source is linked for full context, while the archive keeps the original publication date and source attribution intact.

Training AlgorithmsReward and Credit AssignmentAgent CapabilitiesData LoopsSystems Engineering
GitHub ->

05

PapersPapers

Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines

Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines

Quattrociocchi and colleagues warn that the fluent outputs of large language models may allow linguistic plausibility to substitute for epistemic evaluation, producing the condition they call *Epistemia*: the experience of possessing knowledge without undertaking the practices through which judgment

Agent Capabilities评估体系
arXiv ->

06

PapersPapers

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for r

Agent Capabilities评估体系
arXiv ->

07

PapersPapers

MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents

MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents

Recently, memory management has become a key infrastructure for LLM-based agents, as it directly affects long-horizon reasoning, personalized responses, and knowledge reuse. However, existing LLM memory systems typically adopt a coarse-grained (utility-agnostic) manner that treats heterogeneous user

Agent Capabilities评估体系
arXiv ->

08

PapersPapers

Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks

Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks

Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management standards lack a standardized mechanism for cross-vendor trust visibility. When a tool from Vendor B is compromised, agents from Vendor A continue invoking it --

Agent Capabilities评估体系
arXiv ->