Real-time AI-curated news from 100,397+ articles across 50+ sources. Sentiment analysis, importance scoring, and key takeaways — updated every 15 minutes.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers propose a method to improve RLHF (Reinforcement Learning from Human Feedback) by treating the rationality parameter as context-dependent rather than fixed, using an LLM-as-judge to detect cognitive biases in human annotations and downweight unreliable comparisons. This approach enables training more robust AI models even when human feedback contains systematic biases.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers introduce MOCI (Multi-Objective Constraint Inference), a novel framework that uses inverse reinforcement learning to extract safety constraints and individual preferences from diverse expert demonstrations where multiple experts have different objectives. The approach addresses limitations in existing methods that assume homogeneous expert behavior and offers improved computational efficiency.
AINeutralarXiv – CS AI · May 115/10
🧠Researchers present a solution for selecting cost-effective experiments to narrow uncertainty bounds on partially identifiable causal effects from observational data. They formalize this as an NP-hard optimization problem and develop pruning algorithms that eliminate 50-88% of candidate experiments without exhaustive computation, demonstrated on real epidemiological datasets.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers introduce an adaptive auditing framework for AI systems that maintains statistical rigor while evaluating generative AI failure modes with limited observations. Using Safe Anytime-Valid Inference, the method enables auditors to draw reliable conclusions from as few as 20 test cases through sequential hypothesis testing, addressing a critical bottleneck in AI safety evaluation.
AIBullisharXiv – CS AI · May 116/10
🧠Researchers present a 2.5-D decomposition method that improves LLM-based spatial reasoning for autonomous construction tasks by constraining language models to 2D horizontal planning while deterministic systems handle vertical placement. The approach achieves 94.6% structural accuracy on benchmark tests, significantly outperforming existing methods and demonstrating practical deployment on edge hardware.
🏢 Nvidia🧠 GPT-4
AINeutralarXiv – CS AI · May 116/10
🧠TeamBench is a new benchmark evaluating multi-agent AI coordination under enforced role separation, revealing that prompt-only instructions fail to prevent role violations and that agent teams often underperform single agents on well-solved tasks. The study demonstrates that passing rates can mask coordination failures and misaligned team dynamics.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers introduce the Online Shared Supply Allocation (OSSA) problem, a theoretical framework for allocating limited resources across multiple locations before demand is known, common in humanitarian logistics and vaccine distribution. The proposed GPA algorithm achieves a 4/3-approximation ratio to optimal offline solutions, with proven tight bounds and a learning-augmented variant that incorporates forecasts.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers introduce ARMOR, an agentic framework that improves chemical reaction feasibility prediction by intelligently combining multiple AI tools rather than relying on single models. The system uses hierarchical tool organization and memory-augmented reasoning to resolve conflicting predictions, demonstrating significant performance gains especially when different tools disagree on outcomes.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers introduce AdaTKG, a novel machine learning approach for temporal knowledge graph reasoning that maintains adaptive per-entity memory updated with each interaction, enabling better predictions on evolving relational data and improved handling of unseen entities compared to existing static representation methods.
AINeutralarXiv – CS AI · May 116/10
🧠SREGym is a new open-source benchmark platform that enables realistic evaluation of AI agents designed to diagnose and fix failures in production systems. The framework simulates high-fidelity failure scenarios across cloud-native stacks and currently includes 90 SRE problems, revealing significant performance variations among frontier AI models.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers introduce Repeated Deceptive Path Planning (RDPP), a framework addressing how agents can conceal destinations from learning adversaries who adapt over time. The proposed Deceptive Meta Planning (DeMP) algorithm uses two-level optimization to sustain deception against evolving observers, outperforming existing static-observer approaches while maintaining reasonable path costs.
AINeutralarXiv – CS AI · May 115/10
🧠Researchers propose a Three-in-One world-model architecture using Deep Boltzmann Machines to unify marketing decision-making by simultaneously capturing consumer heterogeneity, predicting outcomes, and enabling counterfactual reasoning about interventions. The approach outperforms existing causal inference baselines in recovering treatment effects, particularly for confounded price-promotion scenarios.
AIBullisharXiv – CS AI · May 116/10
🧠Researchers introduce AIDA, an autonomous agent framework designed to transform complex enterprise data into actionable business insights by combining large language models with a domain-specific language and reinforcement learning. The system outperforms traditional workflow-based approaches in analyzing multi-dimensional retail data, demonstrating the potential for AI-driven autonomous intelligence in enterprise business intelligence systems.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers introduce HMACE, a multi-agent AI framework that uses specialized language model agents to design heuristics for combinatorial optimization problems. The system achieves competitive results on benchmark problems while using significantly fewer computational tokens than existing methods, demonstrating improved efficiency in automated algorithm design.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers introduce MemoRepair, a system that addresses cascade failures in agentic memory by preventing stale or invalidated information from corrupting downstream AI agent decisions. Using a barrier-first approach and graph-based optimization, the system reduces invalid memory exposure from 69-94% to 0% while maintaining 91-94% of valid successor states with significantly lower repair costs.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers introduce EnvSimBench, a benchmark for evaluating how well large language models can simulate interactive environments for AI agent training. The study reveals a critical flaw: LLMs achieve near-perfect accuracy when environment state remains static but fail catastrophically when multiple simultaneous state changes occur, exposing a fundamental capability gap in LLM-based simulation.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers introduce ChemCost, a benchmark for evaluating LLM agents on chemical cost estimation from reaction descriptions. The study reveals that even frontier LLMs achieve only 50.6% accuracy on clean inputs and degrade significantly with realistic noise, exposing brittleness in parsing, evidence integration, and tool use despite access to domain-specific tools.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers introduce Structured Role-Aware Policy Optimization (SRPO), a reinforcement learning method that improves multimodal AI reasoning by assigning credit to different token types based on their functional roles. The approach enhances vision-language models' ability to ground answers in visual evidence without requiring external reward models, advancing more reliable multimodal reasoning systems.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers present a signal-reshaping framework for GRPO (Group Relative Policy Optimization) that improves code-agent reinforcement learning under weak feedback conditions. The approach combines layered rewards, process-level credit assignment, and execution-aware rollout governance to increase strict compile-and-semantic accuracy from 38.5% to 53.5% on agentic code repair tasks.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers propose Structured Opponent Modeling (SOM), a two-stage framework using Structural Causal Models to improve how LLM-based agents predict and adapt to opponent behavior in multi-agent environments. The approach separates opponent model construction from prediction, enabling more accurate strategic decision-making in game-theoretic scenarios.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers present a scale-conditioned evaluation protocol for AI agent memory systems that tests whether stored evidence remains usable as irrelevant data accumulates. Testing across multiple memory architectures and language models reveals that reliability degrades unpredictably with scale, with some models exceeding computational budgets while others maintain performance, suggesting memory scalability claims must be conditioned on specific agent-interface-scale combinations.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers introduce DoLQ, a new method that combines large language models with symbolic regression to discover ordinary differential equations from observational data. The approach integrates both qualitative physical reasoning and quantitative metrics through a multi-agent architecture, demonstrating superior performance over existing methods in recovering accurate symbolic equations.
AIBullisharXiv – CS AI · May 116/10
🧠GraphReAct introduces a new reasoning-acting framework that enhances large language models for multi-step inference over graph-structured data by combining topological and semantic retrieval actions with context refinement. The framework demonstrates consistent improvements over existing methods across six benchmark datasets, advancing how AI systems can reason about interconnected, structured information.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers propose Posterior Sampling-based Policy Optimization (PSPO), a novel approach to offline reinforcement learning that addresses the critical challenge of balancing model generalization with robustness against exploitation errors. By formulating dynamics modeling as Bayesian inference, PSPO enables safer learning from out-of-distribution data while maintaining theoretical convergence guarantees.
AINeutralarXiv – CS AI · May 116/10
🧠Researchers extend bounded fitting—a machine learning paradigm for logical formula discovery—to more expressive description logics beyond ALC, maintaining PAC-style guarantees while implementing practical solutions via SAT solvers. The work demonstrates that this approach scales to complex logical systems with inverse roles and qualified restrictions, achieving competitive results against existing concept learners.