y0news
AnalyticsDigestsSourcesTopicsRSSAICrypto
🤖All95,119🧠AI22,940⛓️Crypto17,363💎DeFi1,799🤖AI × Crypto1,480📰General51,537

AI × Crypto News Feed

Real-time AI-curated news from 95,119+ articles across 50+ sources. Sentiment analysis, importance scoring, and key takeaways — updated every 15 minutes.

95119 articles
AIBullisharXiv – CS AI · Jun 87/10
🧠

OpenSkill: Open-World Self-Evolution for LLM Agents

OpenSkill introduces a framework enabling LLM agents to self-evolve in open-world environments without task-specific supervision, bootstrapping both skills and verification signals from public documentation and web resources. The approach demonstrates superior performance across benchmarks while maintaining transferability across different models, addressing a critical gap in autonomous agent deployment.

AINeutralarXiv – CS AI · Jun 87/10
🧠

Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics

A position paper argues that AI research must shift from analyzing finished models to studying the training dynamics that produce model behaviors. The authors propose that a rigorous science of AI requires understanding how data, objectives, and optimization shape model properties—enabling prediction and intervention during training rather than post-hoc fixes.

AIBearisharXiv – CS AI · Jun 87/10
🧠

Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models

Researchers measured how well frontier AI models perform complex reasoning without explicit chain-of-thought (CoT) tokens, finding that no-CoT task-completion time horizons have doubled yearly over six years. GPT-5.5 now reaches over 3 minutes of reasoning complexity, with projections suggesting frontier models could exceed 7 minutes by 2028 and 25 minutes by 2030, raising concerns about the effectiveness of current AI safety monitoring approaches.

🧠 GPT-5
AIBullisharXiv – CS AI · Jun 87/10
🧠

Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory

Lean4Agent introduces a formal verification framework using Lean4, a dependent-type language, to model and verify LLM agent workflows. The system demonstrates 11.94% performance improvement for verification-passing workflows and 7.47% additional gains through LeanEvolve optimization, establishing a new approach to ensuring AI agent reliability.

AIBullisharXiv – CS AI · Jun 87/10
🧠

The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective

Researchers propose formalizing the evaluation of foundation model agents through a classical sim-to-real framework based on Markov Decision Processes, addressing the gap between simulated training and real-world deployment. The work advocates adopting established robotics solutions like domain randomization and establishing standardized benchmarks to build more reliable AI agents for production applications.

AIBullisharXiv – CS AI · Jun 87/10
🧠

How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope

A study of Perplexity's autonomous AI agents reveals they perform 26 minutes of productive work per session versus 33 seconds for traditional search, reducing task completion time by 87% while improving quality and expanding the scope of work users attempt. This research demonstrates how AI agents are transitioning from conversational tools to end-to-end task executors that fundamentally reshape knowledge work.

🏢 Perplexity
AIBullisharXiv – CS AI · Jun 87/10
🧠

FIGMA: Towards FIne-Grained Music retrievAl

Researchers introduce FIGMA, a new multi-view contrastive learning architecture that significantly improves music retrieval based on fine-grained musical attributes like tempo, key, and chord progression. The work addresses a fundamental limitation in existing CLAP-based models that fail to process detailed musical descriptions, achieving up to 73.3% relative improvement and contributing a new 380K music-caption dataset (FGMCaps) to the field.

AIBullisharXiv – CS AI · Jun 87/10
🧠

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding

Researchers introduce LyraV, a streaming video-language model that maintains real-time synchronization between video perception and language generation without pausing. The system uses a hierarchical control framework with two key components—a Frame-Driven Transition Controller and Streaming Token Pacer—to interleave video frames with generated tokens at 3.89 FPS with 98.29% synchrony.

AINeutralarXiv – CS AI · Jun 87/10
🧠

Auditing Training Data in Domain-adapted LLMs: LoRA-MINT

Researchers introduce LoRA-MINT, a methodology for detecting whether specific data samples were used to train fine-tuned large language models, achieving 77-92% precision. This auditing tool addresses growing concerns about intellectual property protection and sensitive data exposure in adapted AI models, with implications for responsible AI deployment.

🏢 Perplexity
AIBullisharXiv – CS AI · Jun 87/10
🧠

OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios

Researchers introduce OpenHalDet, an open-source benchmark framework that standardizes hallucination detection evaluation across diverse LLM scenarios. The unified framework addresses reproducibility challenges by providing consistent evaluation pipelines and supporting multiple detector types (black-box, gray-box, white-box), enabling more reliable comparison of hallucination detection methods.

AIBullisharXiv – CS AI · Jun 87/10
🧠

STREAM: Stochastic Riemannian Flow Matching with Anisotropic Decoder for Digital Histopathology Image Generation

Researchers introduce STREAM, a novel framework applying Riemannian flow matching to synthetic histopathology image generation. The approach leverages pretrained Vision Foundation Models as latent space rather than conditioning signals, addressing the "conditioning collapse" problem and achieving state-of-the-art results for medical image synthesis.

AIBearisharXiv – CS AI · Jun 87/10
🧠

Hearing the Unspoken: Language Model Priors for Acoustic Adversarial Attacks

Researchers demonstrate a new adversarial attack called Semantic Gambit that exploits Large Language Models to significantly compromise real-time Automatic Speech Recognition systems. By leveraging predictive context from LLMs, the attack achieves a 35.6% Word Error Rate—three times higher than previously documented attacks—revealing a critical vulnerability in ASR pipelines that operate under temporal constraints.

AIBearisharXiv – CS AI · Jun 87/10
🧠

Does Topic Sentiment Cause Perceived Ideology? Comparing Human and LLM Annotations in Political News Articles

A research study compares how human annotators and large language models (GPT-4o-mini, Llama-3.3-70B) assign political ideology labels to news articles, finding that fine-tuned GPT-4o-mini models develop spurious correlations between sentiment and ideology that don't exist in human judgment. This reveals a critical vulnerability in using LLM annotations as training data for downstream tasks.

🧠 GPT-4🧠 Llama
AIBullisharXiv – CS AI · Jun 87/10
🧠

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating

Researchers introduce SlimSearcher, a framework that trains AI web agents to perform complex information-seeking tasks with 17-58% fewer tool calls while maintaining or improving accuracy. The approach combines efficient trajectory filtering during supervised fine-tuning with adaptive reward gating during reinforcement learning to eliminate wasteful search behaviors.

AIBullisharXiv – CS AI · Jun 87/10
🧠

dots.tts Technical Report

Researchers have developed dots.tts, a 2-billion parameter text-to-speech model that achieves state-of-the-art performance through innovations in continuous speech modeling, full-history conditioning, and self-corrective training. The model demonstrates exceptional multilingual capabilities and enables low-latency speech generation, with code and weights released open-source under Apache 2.0 license.

AIBullisharXiv – CS AI · Jun 87/10
🧠

DaX: Learning General Pathology Representations Across Scales

Researchers present DaX, a pathology vision foundation model that adapts self-supervised learning to whole-slide histopathology imaging. The model demonstrates strong performance across a standardized benchmark of 161 clinical tasks, establishing a reproducible evaluation framework for computational pathology applications.

AIBearisharXiv – CS AI · Jun 87/10
🧠

How reliable are LLMs when it comes to playing dice?

A comprehensive study of 8 state-of-the-art language models reveals significant limitations in probabilistic reasoning, with accuracy dropping from 96% on standard problems to 59% on counterintuitive ones. The research demonstrates that LLMs are vulnerable to token bias and prompt manipulation, suggesting they lack genuine probability reasoning despite excelling at other mathematical tasks.

AIBullisharXiv – CS AI · Jun 87/10
🧠

MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism

Researchers introduce MemDreamer, a framework that enables Vision-Language Models to process hours-long videos by decoupling perception from reasoning through hierarchical graph memory and agentic retrieval. The approach achieves state-of-the-art results while reducing computational context requirements to 2% of full video ingestion, establishing a new paradigm for long-form multimodal understanding.

AIBullisharXiv – CS AI · Jun 87/10
🧠

ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning

ThinkBooster is a unified framework that standardizes test-time compute scaling for large language models, providing a modular library, benchmarking suite, and production-ready API for improving LLM reasoning efficiency during inference. The framework enables developers to evaluate and deploy adaptive reasoning strategies with transparent performance-compute trade-offs across mathematical and coding tasks.

🏢 OpenAI
AINeutralarXiv – CS AI · Jun 87/10
🧠

The Three-Ring Architecture: Governing Agents in the Era of On-Platform Organisations

A research paper proposes the Three-Ring Architecture as a governance framework for enterprise AI deployment, arguing that organizations deploying agentic AI systems lack adequate control infrastructure. The framework separates deterministic, strategies-based agents (Ring 2) from non-deterministic LLM-based agents (Ring 3), positioning Ring 2 as essential operating system-level governance to prevent the 95% project failure rates seen in previous AI deployment waves.

AIBullisharXiv – CS AI · Jun 87/10
🧠

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

A comprehensive survey examines latent space as an emerging computational substrate for language models, arguing that continuous latent representations are more efficient than explicit token-level generation for critical internal processes. The research identifies four mechanistic developments (architecture, representation, computation, optimization) and seven capability areas (reasoning, planning, modeling, perception, memory, collaboration, embodiment) that latent space enables.

AIBearisharXiv – CS AI · Jun 87/10
🧠

CULTURESCORE: Evaluating Cultural Faithfulness in Video Generation Models

Researchers introduce CultureScore, a new evaluation framework for assessing cultural faithfulness in video generation models, revealing that leading AI systems like Veo 3.1 and LTX-2 fail to accurately represent diverse global cultures. Testing across 10 countries shows the best model achieves only 56.8% cultural accuracy, with human evaluators valuing cultural representation over visual quality metrics.

← PrevPage 236 of 3805Next →
Filters
Sentiment
Importance
Sort
Stay Updated
Everything combined