y0news
AnalyticsDigestsSourcesTopicsRSSAICrypto

#procedural-reasoning News & Analysis

3 articles tagged with #procedural-reasoning. AI-curated summaries with sentiment analysis and key takeaways from 50+ sources.

3 articles
AINeutralarXiv – CS AI · Jun 257/10
🧠

InvestPhilBench: A Multi-Layer Dynamic Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

Researchers introduce InvestPhilBench, a comprehensive benchmark for testing large language models' ability to reconstruct and apply expert investment decision frameworks. The v0.6 release reveals that while state-of-the-art models achieve high composite scores (0.932), they exhibit significant procedural reasoning deficits (GRA scores of 0.57-0.77), indicating that fluent prose masks deeper gaps in step-by-step investment logic.

🧠 Claude
AINeutralarXiv – CS AI · Jun 236/10
🧠

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models

Researchers present Answer Engineering, a runtime technique that improves large language model compliance with procedural protocols by editing reasoning trajectories during generation. Testing on clinical decision-making shows the method increased protocol adherence from 25-54% to 78-84% without retraining models, addressing a critical safety gap in high-stakes domains.

AINeutralarXiv – CS AI · Jun 126/10
🧠

Constructing Evaluation Datasets for Procedural Reasoning: Balancing Naturalness, Grounding, and Multi-Hop Coverage

Researchers present a framework for evaluating procedural reasoning datasets in AI-supported learning systems by comparing three question-generation strategies based on Task-Method-Knowledge (TMK) models. The study demonstrates that strict TMK generation produces the most grounded and usable datasets (96.5% grounded), while transcript-based approaches sacrifice representational alignment for naturalness, highlighting the trade-off between learner-like phrasing and formal grounding in evaluation dataset construction.