y0news
AnalyticsDigestsSourcesTopicsRSSAICrypto
🤖All95,870🧠AI22,940⛓️Crypto17,363💎DeFi1,799🤖AI × Crypto1,480📰General52,288

AI × Crypto News Feed

Real-time AI-curated news from 95,870+ articles across 50+ sources. Sentiment analysis, importance scoring, and key takeaways — updated every 15 minutes.

95870 articles
AIBullisharXiv – CS AI · May 296/10
🧠

Harnessing non-adversarial robustness in large language models

Researchers propose a debiasing fine-tuning method to improve Large Language Model robustness against semantically-neutral prompt variations without expensive full retraining. The approach identifies perturbation-induced bias in neural network outputs and demonstrates theoretical and experimental evidence that targeted debiasing can enhance model resilience to prompt alterations.

AIBullisharXiv – CS AI · May 296/10
🧠

OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation

OptSkills, a new AI system, advances automated optimization problem-solving by clustering problems by underlying mathematical archetypes rather than surface narratives, achieving 68.27% accuracy on diverse benchmarks and outperforming DeepSeek-V3.2-Thinking on large-scale problems. The system uses skill distillation and trajectory learning to improve generalization across both known and novel problem types.

AINeutralarXiv – CS AI · May 296/10
🧠

OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

Researchers introduced OmniMatBench, a comprehensive multimodal reasoning benchmark containing 3,171 expert-curated problems across 19 materials science subfields. Evaluation of 13 major language models revealed significant gaps in AI reasoning capabilities, with the best model achieving only 37.2% accuracy, highlighting the need for improved scientific AI systems.

AINeutralarXiv – CS AI · May 296/10
🧠

Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories

Researchers introduce RedundancyBench, a new benchmark for detecting redundant steps in LLM-based agent trajectories, revealing that current methods struggle significantly with this task—the best approach achieves only 24.88% accuracy. This work highlights a critical gap in agent evaluation: while task success is commonly measured, execution efficiency and resource optimization remain largely unmeasured, suggesting AI agents require substantial improvements in reasoning efficiency.

AINeutralarXiv – CS AI · May 296/10
🧠

On the Geometry of Games and their Solvers

Researchers propose a novel framework for understanding equilibrium computation in games by mapping the geometric structure of game spaces to solver effectiveness. Rather than studying algorithms in isolation, they develop a learned representation that identifies which solver mechanisms work best across different game regimes, revealing continuous regions of algorithmic validity and suggesting that solvability is governed by underlying structural properties.

AINeutralarXiv – CS AI · May 296/10
🧠

Toward AI Systems That Understand Self and Others: A Multi-Phase Inference Framework for Human Cognitive Diversity and World-Model Alignment

Researchers propose a Multi-Phase Inference Mechanism (MIM) framework that models how AI systems can understand diverse human cognition and world-models without forcing consensus. The framework formalizes how different agents form different representations and predictions from identical observations, offering a constructive approach to AI alignment and human-AI understanding.

AINeutralarXiv – CS AI · May 296/10
🧠

It`s All About Speed: AI`s Impact on Workflow in Music Production

An ethnographic study examines how AI and automated tools reshape music production workflows among professional engineers, mixers, and producers. The research identifies key tensions between automation benefits (speed and efficiency) and creative concerns (controllability and artistic agency), offering insights into how tool design can better balance these competing demands.

AINeutralarXiv – CS AI · May 296/10
🧠

Make LLM Learn to Synthesize from Streaming Experiences through Feedback

Researchers introduce StreamSynth, a new framework enabling large language models to learn and improve synthetic data generation across sequential tasks by accumulating experience and transferring knowledge between related synthesis problems. The SynLearner framework demonstrates that LLMs can leverage historical task insights to enhance future data generation quality, establishing synthetic data creation as an experience-driven process rather than isolated operations.

AINeutralarXiv – CS AI · May 296/10
🧠

MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization

Researchers introduce MuPHI, a dataset and training framework for detecting implicit multimodal harm in image-text pairs where danger emerges from context-dependent reasoning rather than surface features. The proposed MuPHIRM framework uses reward optimization to improve vision-language models' ability to reason about compositional harm while demonstrating stronger generalization to out-of-distribution scenarios.

AINeutralarXiv – CS AI · May 296/10
🧠

Meta-Programming for Linear-time Temporal Answer Set Programming

Researchers propose a meta-programming framework that enables flexible implementation of temporal logic extensions for Answer Set Programming (ASP) through a unified declarative system. The work introduces metasp, a tool that allows rapid exploration of different temporal logics—including linear-time (TEL), metric (MEL), and dynamic (DEL) variants—without modifying core ASP system code.

AINeutralarXiv – CS AI · May 296/10
🧠

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation

Researchers introduce Cookie-Bench, a comprehensive 1,000-query web development benchmark, and Cookie-Frame, an autonomous evaluation framework that assesses LLM-generated web applications through static perception, agent-driven interaction, and dynamic scoring. The approach eliminates reliance on reference implementations while aligning closely with human expert ratings, revealing significant performance gaps across 13 frontier LLMs.

AIBullisharXiv – CS AI · May 296/10
🧠

KairosAgent: Agentic Time Series Forecasting with Fused Semantic Reasoning

Researchers introduce KairosAgent, an agentic framework combining large language models with time series foundation models to improve multimodal forecasting across domains. The system uses semantic reasoning from LLMs fused with numerical forecasting capabilities, achieving superior zero-shot performance through reinforcement learning and structured tool integration.

AINeutralarXiv – CS AI · May 296/10
🧠

From GPS Points to Travel Patterns: Flexible and Semantic Trajectory Generation with LLMs

Researchers propose HTP, an LLM-based framework that generates realistic urban trajectories by first synthesizing travel patterns and then GPS points, addressing privacy concerns in smart city applications. The method outperforms existing approaches by 29.78% and can generate variable-length trajectories under multiple conditions, advancing synthetic data generation for urban analytics.

AINeutralarXiv – CS AI · May 296/10
🧠

RAISE: RAG Design as an Architecture Search Problem

Researchers introduce RAISE, a comprehensive framework for optimizing retrieval-augmented generation (RAG) systems by treating architecture design as a hyperparameter search problem. The study evaluates 13 optimization algorithms across seven datasets, revealing that RAG performance is highly task-dependent and no single optimization strategy universally outperforms others, highlighting the need for systematic rather than heuristic-based configuration approaches.

🏢 Meta
AIBullisharXiv – CS AI · May 296/10
🧠

Teaching Values to Machines: Simulating Human-Like Behavior in LLMs

Researchers successfully induced human-like values in Large Language Models using psychological theory and tested them against 5+ million questions, finding strong alignment between value-prompted LLMs and human behavior patterns. This work demonstrates that LLMs can simulate coherent value structures comparable to humans, opening possibilities for more realistic behavioral modeling.

AINeutralarXiv – CS AI · May 296/10
🧠

Learning to Choose: An Empowerment-Guided Multi-Agent System with semantic communication for Adaptive Method Selection

Researchers introduce a multi-agent framework that combines contextual bandits with semantic checkpoints to prevent 'semantic drift' in automated scientific computing workflows. The system ensures that computational strategies selected by AI agents are faithfully executed and remain causally attributable throughout multi-agent pipelines, improving convergence and robustness in adaptive decision-making.

AINeutralarXiv – CS AI · May 296/10
🧠

Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers

Researchers introduce SafeDIG, a safety steering framework designed to make text-to-image diffusion transformers like FLUX.1 and Stable Diffusion 3.5 resistant to generating harmful content. The method uses sparse autoencoders and adaptive decoding to maintain safety controls across different risk domains while preserving image quality.

🧠 Stable Diffusion
AINeutralarXiv – CS AI · May 296/10
🧠

Conformal Certification of Reasoning Trace Prefixes

Researchers introduce CROP, a statistical certification method for language model reasoning traces that identifies the longest reliable prefix before errors occur. The technique enables safer deployment of AI systems by providing rigorous guarantees about which intermediate reasoning steps can be trusted, while routing uncertain portions for human review or automated repair.

AINeutralarXiv – CS AI · May 296/10
🧠

Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison

Researchers introduce a benchmark for evaluating how AI systems handle conflicting information across multiple memory sources, addressing a critical gap in testing personal AI agents. The study compares various approaches including fusion methods and LLMs, revealing that trained fusion models outperform prompt-based LLMs by 10+ percentage points on accuracy, with selective abstention improving performance further.

AINeutralarXiv – CS AI · May 296/10
🧠

VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

Researchers introduce VLA-Trace, a diagnostic framework for analyzing Vision-Language-Action models that reveals how these AI systems transform multimodal inputs into physical control actions. The study identifies that popular VLA models like π₀.₅ and OpenVLA exhibit distinct adaptation patterns, rely on different routing strategies during decision-making, but struggle with fine-grained semantic understanding despite excelling at visual grounding.

AIBullisharXiv – CS AI · May 296/10
🧠

Enhancing Multi-Agent Communication through Attention Steering with Context Relevance

Researchers introduce Agent-Radar, a training-free context management method that improves multi-agent LLM systems by dynamically filtering irrelevant information from long conversation histories. The technique uses temporal and spatial decay mechanisms to maintain focus on relevant context, achieving up to 7.64% performance improvements across five benchmarks.

AINeutralarXiv – CS AI · May 296/10
🧠

AgentSchool: An LLM-Powered Multi-Agent Simulation for Education

Researchers introduce AgentSchool, an LLM-powered multi-agent simulator that models student learning through state transitions rather than simple role-play, featuring cognitively growable student agents with knowledge graphs and adaptive teachers operating within the Zone of Proximal Development. The system addresses the challenge of validating educational AI interventions in real classrooms by creating a configurable simulation environment that reproduces plausible learning outcomes and social dynamics without requiring institutional constraints or ethical compromises of live trials.

AINeutralarXiv – CS AI · May 296/10
🧠

Anchorless Diversification for Parallel LLM Ideation

Researchers present methods for improving how large language models generate diverse pools of creative ideas during parallel inference without relying on seed examples. Their findings show that semantic direction stratification—organizing generations across different semantic directions with a single planning call—outperforms anchor-dependent baselines while maintaining quality and computational efficiency.

AINeutralarXiv – CS AI · May 296/10
🧠

Temporal Stability and Few-Shot Prompting in Math Task Assessment

A longitudinal study examined how AI models (Gemini and Coteach) perform on mathematics task classification using the Task Analysis Guide, testing stability across model versions and responsiveness to few-shot prompting. Results showed newer model versions produced mixed effects, but few-shot prompting consistently improved both models' accuracy, suggesting prompt engineering is more reliable than passive model updates for specialized educational tasks.

🧠 Gemini
AINeutralarXiv – CS AI · May 296/10
🧠

Modularizing Educational LLM-Agency for Fostering Responsible Learning Assistance

Researchers propose a modular architecture for educational AI chatbots designed to enforce pedagogical principles and prevent negative learning outcomes. The approach addresses structural limitations in current monolithic LLM solutions by incorporating targeted modules at different exercise-solving stages, enabling more transparent and controlled student guidance.

← PrevPage 1417 of 3835Next →
Filters
Sentiment
Importance
Sort
Stay Updated
Everything combined