Daily Signals: SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
A dated Agentic RL Daily snapshot with 8 verified primary-source signals across papers, official releases, deployment evidence, and safety or alignment findings.
EDITOR'S VIEW
Three judgments
The edition is based only on primary sources or official project releases.
Older signals remain visible as continuing observations, not as rewritten news.
Headline claims are constrained by the evidence included in this dated snapshot.
SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control.
Towards Expert-level Medical AI for Real-time Video Consultations
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues.
CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environments remains a challenge.
## What's Changed * fix: only pass min_lr_rate to schedulers that accept it by @matteolippi in https://github.com/OpenRLHF/OpenRLHF/pull/1238 * Upgrade vLLM to 0.22.1 and DeepSpeed to 0.19.1 by @hijkzzz in https://github.com/OpenRLHF/OpenRLHF/pull/1248 * Fix token-level loss (global token-mean acros
DSLE: A Learning Environment for Dark Souls Boss Encounters
官方来源补充信号:We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounters of Dark Souls: Remastered as game-playing agent benchmarks through a Gymnasium-style interface. DSLE combines real-time combat, high-dimensional visual input, and sparse terminal rewards, with each environment step being a real action executed against the running game. To support controlled comparison, we define DSLE-5, a representative five-boss subset, spanning a melee fight, a spatially constrained arena, an environmental-hazard fight, a multi-target fight, and a fast final
Training AlgorithmsReward and Credit AssignmentAgent Capabilities评估体系