Agentic RL Daily

VOL. 020

8 research signals

DAILY EDITION / SAVED SNAPSHOT

Daily Signals: ContextWeave: A Real-World Workflow Benchmark

A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.

EDITOR'S VIEW

Three judgments

  1. ContextWeave: A Real-World Workflow Benchmark
  2. CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs
  3. NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment

FULL EDITION

All signals in this edition

Archived / 2026-08-06

01

PapersPapers

ContextWeave: A Real-World Workflow Benchmark

ContextWeave: A Real-World Workflow Benchmark

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering.

Training Algorithms
arXiv ->

06

PapersPapers

EviGraph: Evidence-Guided Autonomous Research Agents

EviGraph: Evidence-Guided Autonomous Research Agents

Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported claims and inconsistencies between research questions, experiments, results, and conclusions.

评估体系
arXiv ->