Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use
Researchers introduced PhysTool-Bench, a benchmark testing how well multimodal large language models (MLLMs) can recognize and use physical tools in real-world scenarios. Testing 13 leading models revealed significant limitations: even the best performer (Gemini-3.1-Pro) identified only 58.7% of tools in scenes and completed just 21% of end-to-end tasks, exposing critical gaps in perception and functional reasoning for embodied AI applications.