y0news
AnalyticsDigestsSourcesTopicsRSSAICrypto
🤖All94,305🧠AI22,940⛓️Crypto17,363💎DeFi1,799🤖AI × Crypto1,480📰General50,723

AI × Crypto News Feed

Real-time AI-curated news from 94,318+ articles across 50+ sources. Sentiment analysis, importance scoring, and key takeaways — updated every 15 minutes.

94318 articles
AINeutralarXiv – CS AI · Jun 26/10
🧠

Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories

Researchers introduce TELBench, a benchmark for identifying errors in deep-research AI agent trajectories, and propose DRIFT, a claim-centric auditing framework that improves error localization accuracy by up to 30 percentage points. The work addresses a critical gap in AI evaluation by moving beyond final-answer assessment to analyze intermediate steps in agent reasoning.

AINeutralarXiv – CS AI · Jun 26/10
🧠

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

Merkle has developed BADGER, a unified evaluation framework that combines text-to-SQL assessment with agentic behavior evaluation for enterprise AI systems. The framework achieves substantial agreement with human expert judgment (Cohen's kappa=0.717) and outperforms six competing evaluation approaches, addressing a critical gap in production-grade AI system assessment.

AIBullisharXiv – CS AI · Jun 26/10
🧠

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning

Researchers propose EAPO, a reinforcement learning framework that teaches AI agents to use external tools selectively rather than excessively. The method improves accuracy while reducing redundant tool calls by 18-25% across multiple language models, demonstrating that agents can learn optimal tool-use patterns without compromising reasoning capabilities.

🧠 Llama
AIBullisharXiv – CS AI · Jun 26/10
🧠

S3TS: Stochastic Scenario-Structured Tree Search for Advanced Planning Under Uncertainty

Researchers introduce S3TS, a novel algorithm combining Monte Carlo Tree Search with stochastic optimization to handle both non-linear complexity and uncertainty in energy grid scheduling. The approach demonstrates near-optimal performance in linear settings and significantly outperforms existing methods in non-linear scenarios, achieving up to 51% cost reductions compared to baseline algorithms.

AINeutralarXiv – CS AI · Jun 25/10
🧠

An Abstract Worlds Semantic Framework for Belief Change Operators

Researchers propose Abstract Worlds Semantics (AWS), a set-theoretic framework for modeling belief change operators without assuming logical syntax. The framework unifies classical and non-prioritized belief change constructions, providing a homogeneous account of AGM, KM, and Multiple Change models in propositional logic.

AIBullisharXiv – CS AI · Jun 26/10
🧠

From Capability Models to Automated Planning: An AAS-Native Approach for Automatic PDDL Generation

Researchers have developed an automated method to generate PDDL planning problems directly from Asset Administration Shell (AAS) capability models using Industry 4.0 standards, eliminating the need for specialized planning expertise. This approach enables production engineers to design and verify manufacturing system layouts without requiring knowledge of formal planning languages, significantly reducing barriers to adopting automated planning in industrial settings.

AINeutralarXiv – CS AI · Jun 26/10
🧠

CEON: Circular Economy Ontology Network

Researchers have developed CEON (Circular Economy Ontology Network), a semantic framework designed to improve information sharing across industries to promote circular economy practices. The ontology addresses the challenge of enabling cross-sector communication along product life cycles in construction, electronics, and textiles, facilitating resource reuse and recycling strategies.

AINeutralarXiv – CS AI · Jun 26/10
🧠

Repair Before Veto: Repair-Augmented Constraint Learning for Contextual Decisions

Researchers introduce Repair-Augmented Constraint Learning (RACL), a machine learning framework that decides whether to repair constraint violations before rejecting candidates, rather than applying hard vetoes immediately. The method achieves significantly lower false-veto rates (0.25%) compared to baseline approaches (26.4%) on real-world airline data, with applications to automated decision systems.

AIBullisharXiv – CS AI · Jun 26/10
🧠

Forget Attention: Importance-Aware Attention Is All You Need

Researchers propose SISA (SSM-Informed Softmax Attention), a hybrid architecture that integrates state space model importance signals directly into transformer attention mechanisms at the score level. The approach achieves superior performance on language modeling benchmarks, particularly excelling at long-context retrieval tasks while maintaining computational efficiency through standard operations.

AINeutralarXiv – CS AI · Jun 26/10
🧠

Coordination Graphs for Constrained Multi-Agent Reinforcement Learning

Researchers introduce CG-CMARL, a framework combining coordination graphs with Lagrangian duality to solve constrained multi-agent reinforcement learning problems. The approach decomposes complex joint action spaces into manageable pairwise regions, enabling scalability to larger agent teams while maintaining convergence guarantees and allowing dynamic Pareto front tracing without retraining.

AIBullisharXiv – CS AI · Jun 26/10
🧠

SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training

Researchers introduce SIRI, a three-phase reinforcement learning framework that enables LLM agents to autonomously discover, validate, and internalize reusable skills without external skill generators or inference-time skill banks. Testing on ALFWorld and WebShop benchmarks shows meaningful performance improvements over baseline methods while reducing deployment complexity and latency.

AIBullisharXiv – CS AI · Jun 26/10
🧠

MOC: Multi-Order Communication in LLM-based Multi-Agent Systems

Researchers propose Multi-Order Communication (MOC), a new framework for improving how large language model-based multi-agent systems exchange information. The scheme addresses limitations in current message-passing approaches by capturing multi-hop dependencies and consolidating messages efficiently, demonstrating consistent performance improvements across multiple datasets while reducing communication costs.

AIBullisharXiv – CS AI · Jun 26/10
🧠

Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses

Researchers introduce Harness-1, a 20B parameter search agent that separates semantic decision-making from state management by externalizing working memory to a stateful harness environment. The system achieves 73% average curated recall across eight retrieval benchmarks, outperforming comparable open-source searchers by 11.4 points while generalizing well to held-out transfer tasks.

AINeutralarXiv – CS AI · Jun 26/10
🧠

Spatial Representation Learning Beyond Pixels: Unifying Raster Data and Vector Semantics for Human-Centric Geospatial Foundation Models

Researchers propose a paradigm shift in Earth Observation Foundation Models by integrating raster satellite imagery with vector data (like OpenStreetMap) into unified embedding spaces. This multimodal approach aims to create more semantically grounded geospatial AI systems that combine continuous physical patterns from imagery with discrete human-centric geographic entities and their relationships.

AINeutralarXiv – CS AI · Jun 25/10
🧠

A Mathematical Conflict Framework for Contextual Data Modulation

Researchers present a mathematical framework that treats data conflict as an explicit, operator-based phenomenon rather than an implicit optimization byproduct. The generalized approach models structural discrepancies between raw and contextual data as local, directional quantities, offering a unified abstraction applicable across problem classes without dependency on specific algorithms.

AINeutralarXiv – CS AI · Jun 26/10
🧠

Bridging the Sim-to-Real Gap in Semiconductor Visual Program Synthesis via Input Binarization

Researchers propose a visual program synthesis framework using Vision-Language Models to convert semiconductor inspection images into editable code, addressing the costly challenge of obtaining real training data for circuit metrology. By applying input binarization to strip texture noise from real Scanning Electron Microscope images, the approach bridges the gap between synthetic training data and real-world application, improving geometric accuracy detection by 19.6%.

AINeutralarXiv – CS AI · Jun 26/10
🧠

LLM-Evolved Pattern Generators for Optimal Classical Planning

Researchers have developed a novel method using large language models and evolutionary algorithms to automatically generate admissible heuristics for optimal classical planning problems. Unlike existing learned heuristics that improve search speed but cannot guarantee optimal solutions, this approach preserves A* optimality guarantees while matching or exceeding the performance of traditional domain-independent methods.

AINeutralarXiv – CS AI · Jun 26/10
🧠

HLL: Can Agents Cross Humanity's Last Line of Verification?

Researchers introduced HLL (Humanity's Last Line of Verification), a benchmark testing whether multimodal AI agents can bypass CAPTCHA protections designed to verify human users. Testing eight frontier models revealed significant brittleness: agent performance varies sharply across CAPTCHA types, degrades under realistic conditions, and fails when solutions must be supported by valid action traces, exposing gaps in localization, action calibration, and process consistency.

AINeutralarXiv – CS AI · Jun 26/10
🧠

AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents

Researchers introduce AgentCL, an evaluation framework for assessing continual learning in language agents, along with MemProbe, a memory design method that helps agents accumulate and reuse knowledge across tasks while avoiding interference. The framework uses controlled task streams to rigorously measure how well agents learn and transfer knowledge over time, revealing that current memory designs struggle to balance learning plasticity with stable knowledge reuse.

AINeutralarXiv – CS AI · Jun 26/10
🧠

MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation

Researchers introduced MCP-Persona, a new benchmark for evaluating how well AI agents handle personalized tools and applications through the Model Context Protocol (MCP). The benchmark tests agent performance on real-world personal applications like Reddit, Slack, and Lark, revealing significant gaps in current AI systems' ability to work with individualized, account-specific tools.

AINeutralarXiv – CS AI · Jun 26/10
🧠

Iteris: Agentic Research Loops for Computational Mathematics

Researchers have developed Iteris, an agentic AI system designed to tackle open problems in computational mathematics by combining language models with numerical experimentation and algorithm design. Applied to two unsolved problems from a Simons Workshop, Iteris generated verified results including a phase diagram for optimization algorithms and a counterexample about QR factorization, demonstrating that AI agents can contribute meaningfully to mathematical research when paired with human expertise.

AINeutralarXiv – CS AI · Jun 26/10
🧠

RASER: Recoverability-Aware Selective Escalation Router for Multi-Hop Question Answering

Researchers introduce RASER, a cost-efficient routing system for multi-hop question-answering that reduces token consumption by 51-59% compared to always-escalating methods while maintaining competitive accuracy. The system leverages six features from one-shot retrieval to intelligently decide whether additional retrieval rounds are necessary, eliminating wasteful LLM calls.

AINeutralarXiv – CS AI · Jun 26/10
🧠

Bridging the Last Mile of Time Series Forecasting with LLM Agents

Researchers present an LLM-agent framework that enhances time series forecasting by incorporating business context and expert judgment into statistical predictions. The system bridges the gap between raw forecasts and decision-ready outputs through structured reasoning, contextual evidence retrieval, and auditable revision mechanisms.

AINeutralarXiv – CS AI · Jun 26/10
🧠

Tracking the Behavioral Trajectories of Adapting Agents

Researchers present a methodology for measuring and tracking behavioral changes in AI agents by analyzing edits to their configuration files through embedding-space trait vectors. The approach achieves 91.2% accuracy in detecting specific behavioral traits like propensity to seek sensitive data, with potential applications in agent-to-agent trust protocols.

AIBullisharXiv – CS AI · Jun 26/10
🧠

A Novel Data Augmentation Strategy for Robust Deep Learning Classification of Biomedical Time-Series Data: Application to ECG and EEG Analysis

Researchers propose a unified deep learning framework combining ResNet-based CNNs with attention mechanisms and novel data augmentation techniques for analyzing biomedical time-series signals like ECG and EEG. The approach achieves near-perfect accuracy (99.78-100%) on benchmark datasets while remaining lightweight enough for wearable deployment, addressing critical gaps in multi-signal analysis and class imbalance handling.

← PrevPage 1354 of 3773Next →
Filters
Sentiment
Importance
Sort
Stay Updated
Everything combined