← Back to explorer

Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

Type
paper
Venue
arXiv
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Diagnoses which multi-turn tool-use call states are actually trainable with RL. Uses nested sampling to separate action-dependent reward variation from continuation noise, then trains the selected states as contextual bandits rather than applying RL to every step of a tool-use trajectory.

Keywords

reinforcement-learning · tool-use · multi-turn-agents · state-selection · bfcl · bandits

Topics

reinforcement-learning, tool-use, multi-turn-agents, state-selection, bfcl

Research notes

  • Discovery: Posted in #random-papers on 2026-09-28 with the DAIR academy URL, which now returns 404. Canonical paper located via arXiv search: 2609.24985v1.
  • Method: Nested sampling estimator that decomposes reward variation into an action-dependent component vs. continuation (luck/noise) component; states with high trainability signal are selected and optimized as contextual bandits.
  • Key findings: On BFCL v4, training on the selected states improved missing-function performance by about 14 percentage points; training on alternative (unselected) states was flat or worse, suggesting much multi-turn tool-use RL trains on untrainable noise.
  • Limitations: Code and dataset URLs not verified from available sources. Results reported on BFCL v4 only in the available summary; generality to other agentic benchmarks unverified.
  • Discovery URL (academy.dair.ai papers page) now 404s; canonical arXiv ID 2609.24985v1 used instead.