When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents
Researchers introduce ToolMaze, a benchmark testing how AI language models handle real-world tool failures and recovery scenarios, revealing that implicit semantic failures cause performance drops of ~37% and that fault-tolerance improves significantly slower than basic task performance as models scale.