y0news
AnalyticsDigestsSourcesTopicsRSSAICrypto

#benchmark-evaluation News & Analysis

53 articles tagged with #benchmark-evaluation. AI-curated summaries with sentiment analysis and key takeaways from 50+ sources.

53 articles
AINeutralarXiv – CS AI · May 126/10
🧠

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces

Researchers introduce OPT-BENCH, a benchmark evaluating whether large language models can self-improve through iterative feedback in complex problem spaces. Testing 19 LLMs across machine learning and NP-hard problems reveals that while stronger models adapt better, even the most advanced systems remain constrained by their base capabilities and fall short of human expert performance.

AIBullisharXiv – CS AI · Apr 156/10
🧠

Long-Horizon Plan Execution in Large Tool Spaces through Entropy-Guided Branching

Researchers introduce SLATE, a large-scale benchmark for evaluating AI agents using APIs, and propose Entropy-Guided Branching (EGB), a search algorithm that improves task success rates and computational efficiency. The work addresses critical limitations in deploying language models within complex tool environments by establishing rigorous evaluation frameworks and reducing the computational burden of exploring massive decision spaces.

AINeutralarXiv – CS AI · Apr 156/10
🧠

Cooperative Memory Paging with Keyword Bookmarks for Long-Horizon LLM Conversations

Researchers propose cooperative paging, a method for managing long LLM conversations by replacing evicted context with compact keyword bookmarks and providing a recall tool for on-demand retrieval. The technique outperforms existing solutions on the LoCoMo benchmark across multiple models, though bookmark discrimination remains a critical limitation.

🧠 GPT-4🧠 Claude
← PrevPage 3 of 3