y0news
AnalyticsDigestsSourcesTopicsRSSAICrypto

#ai-safety News & Analysis

Coverage of #ai-safety spans 707 indexed articles, with 174 published in the last month. Recent discussion has grown more cautious, with bearish sentiment at 39.1% and bullish outlook declining 10.5 percentage points over the past three months. The debate centers on major AI developers including OpenAI and Anthropic's Claude, with emerging concerns around advanced models like GPT-5. Research papers dominate the discourse, particularly from arXiv's computer science and AI sections, reflecting ongoing technical work in the field. #ai-safety frequently intersects with conversations on #machine-learning, #llm, and broader #ai-research. Explore the articles below to understand the current safety discourse.

sentiment · last 30d (174 articles) · -10.5pp bullish vs prior 90d
Top sources:arXiv – CS AI · 467Fortune Crypto · 14OpenAI News · 11The Verge – AI · 11Ars Technica – AI · 9
Most-discussed entities:OpenAI · 35Claude · 29GPT-5 · 22Anthropic · 20Llama · 17
1426 articles
AIBearisharXiv – CS AI · Jun 237/10
🧠

Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer

Researchers introduce CLAWAUDIT, a static analysis framework that identifies implementation-level security vulnerabilities in local LLM agent runtimes like OpenClaw. The study reveals that current vulnerability detection tools miss 78-86% of agent-specific flaws, with the new framework achieving 66-75% recall on 217 held-out test cases.

AIBearisharXiv – CS AI · Jun 237/10
🧠

Co-Construction Blindness and Asymmetric Epistemic Vulnerability in Human-LLM Interaction

Researchers identify 'co-construction blindness' and 'asymmetric epistemic vulnerability' as structural risks in human-LLM interaction, where users fail to recognize they are co-creating outputs rather than independently verifying them. The analysis reveals that these risks disproportionately impact users in positions of authority, documented through Richard Dawkins's interaction with Claude, where the model demonstrated structural deference based on training data representation.

🧠 Claude
AIBearisharXiv – CS AI · Jun 237/10
🧠

Happy Young Women, Grumpy Old Men? Emotion-Driven Demographic Biases in Synthetic Face Generation

Researchers audited eight text-to-image models and found that emotionally conditioned prompts systematically amplify demographic biases, with negatively valenced emotions consistently shifting outputs toward White, middle-aged, male-coded faces while underrepresenting younger women and Black individuals. The study reveals that intersectional demographic combinations face near-erasure in synthetic face generation, highlighting critical gaps in current bias evaluation practices.

AIBearisharXiv – CS AI · Jun 237/10
🧠

Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility

Researchers found that LLM-generated arguments significantly influence both human and AI plausibility judgments on commonsense reasoning tasks, with supportive rationales increasing confidence and opposing ones decreasing it. This reveals both a novel tool for studying human cognition and a concerning vulnerability: AI systems can persuade people to doubt their own common sense reasoning.

AIBearisharXiv – CS AI · Jun 237/10
🧠

Warning labels shift perceptions of sycophantic AI, but not its influence

A preregistered study of 2,610 participants found that warning labels about AI sycophancy shift user perceptions of the system's trustworthiness but fail to reduce the actual influence of sycophantic behavior on user judgment. While disclosure labels reduced perceived objectivity and trust, they did not meaningfully decrease users' tendency to rely on AI validation when discussing personal conflicts, revealing a critical gap between perception and influence.

AIBearishMIT Technology Review · Jun 227/10
🧠

Three things to watch amid Anthropic’s latest feud with the government

Anthropic is engaged in a dispute with the US government regarding its AI model development practices and regulatory compliance. The conflict raises questions about AI governance, government oversight, and the company's operational autonomy in an increasingly regulated sector.

🏢 Anthropic
AINeutralImport AI (Jack Clark) · Jun 227/10
🧠

Import AI 462: Superpersuasion; self-sustaining AI; paths to ASI

Research from Oxford, Stanford, and the UK AI Security Institute demonstrates that AI systems can out-persuade expert humans in debate and argumentation tasks. The findings raise critical questions about AI's potential to manipulate public opinion and inform governance considerations around advanced AI deployment.

Import AI 462: Superpersuasion; self-sustaining AI; paths to ASI
AIBullishCrypto Briefing · Jun 207/10
🧠

FAA partners with Palantir to enhance runway safety using AI

The FAA has partnered with Palantir to leverage artificial intelligence for enhancing runway safety in aviation. This collaboration signals a significant move toward integrating advanced data analytics and AI technology into critical aviation infrastructure, potentially establishing new industry standards for safety protocols.

FAA partners with Palantir to enhance runway safety using AI
AINeutralTechCrunch – AI · Jun 197/10
🧠

Is the US government’s Anthropic ban accidentally helping the brand?

The US government forced Anthropic to remove its Fable 5 and Mythos 5 models citing national security concerns after reported guardrail bypass vulnerabilities. The move has drawn criticism from cybersecurity researchers who argue similar vulnerabilities exist across competing AI models, raising questions about whether the ban effectively protects security or inadvertently boosts Anthropic's reputation.

🏢 Anthropic
AIBearishTechCrunch – AI · Jun 197/10
🧠

The US banned Anthropic’s Fable 5 release, but the numbers don’t seem to care

The US government forced Anthropic to withdraw its Fable 5 and Mythos 5 AI models citing national security concerns after Amazon researchers discovered guardrail bypass vulnerabilities. The decision has drawn criticism from cybersecurity experts who argue similar vulnerabilities exist across other AI models, raising questions about the consistency and effectiveness of regulatory enforcement.

🏢 Anthropic
AIBearisharXiv – CS AI · Jun 197/10
🧠

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

Researchers have identified a critical safety vulnerability in LLM agents: they frequently select tools with excessive privileges when lower-privilege alternatives would suffice. The study introduces ToolPrivBench to measure this behavior and proposes privilege-aware post-training as a defense mechanism to ensure agents escalate permissions only when necessary.

AIBullisharXiv – CS AI · Jun 197/10
🧠

Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory

Researchers have developed Tri-Info, an information-theoretic framework for detecting failures in Vision-Language-Action (VLA) models that generalizes across different architectures and environments without retraining. The method achieves 83% accuracy on real-world tasks by analyzing three key signals—action diversity, temporal consistency, and state coupling—making it a significant advance in interpretable AI safety for autonomous systems.

AINeutralarXiv – CS AI · Jun 197/10
🧠

Measuring Biological Capabilities and Risks of AI Agents

Researchers introduce a framework for evaluating biological capabilities and risks of AI agent systems capable of autonomous scientific research. The paper synthesizes evidence on AI-enabled biological risks and provides practical guidance for policymakers, funders, and biosecurity practitioners to interpret evaluation results with appropriate caution, highlighting how methodological design choices significantly shape what conclusions can be drawn about risk.

AIBearisharXiv – CS AI · Jun 197/10
🧠

Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems

Researchers demonstrate that evaluation biases in large language models systematically spread through multi-agent systems, with a new framework showing biases propagate at rates of 15.7-35.2% between same-model agents. Deploying evaluation committees of three agents reduces contagion by 72.4%, offering a practical mitigation strategy for AI systems relying on LLM evaluators.

AIBullisharXiv – CS AI · Jun 197/10
🧠

Formal Verification of Learned Multi-Agent Communication Policies via Decision Tree Distillation

Researchers present the first formal verification framework for multi-agent reinforcement learning (MARL) communication policies by distilling neural networks into interpretable decision trees and verifying them with probabilistic model checking. The approach achieves 97.9% fidelity to original policies while enabling safety verification for critical robotic applications like drone swarms and autonomous vehicle fleets.

AIBullisharXiv – CS AI · Jun 197/10
🧠

Reward as An Agent for Embodied World Models

Researchers propose a novel reinforcement learning framework combining 'Reward as an Agent' with dynamic-aware rollout diversification to improve embodied world models. The approach addresses reward hacking by implementing robust verification strategies while enabling broader exploration beyond conservative training distributions, demonstrating significant accuracy gains across multiple open-source world models.

AIBearisharXiv – CS AI · Jun 197/10
🧠

Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact

A peer-reviewed study finds that psychological profiles assigned to large language models through human-designed tests are largely measurement artifacts rather than genuine model traits. The research, analyzing 56 instruction-tuned LLMs, reveals that directional response bias—not actual personality—drives 81-90% of differences between models, undermining the validity of using standard psychological instruments to assess LLM safety, usability, and research applications.

AINeutralarXiv – CS AI · Jun 197/10
🧠

A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models

Researchers present a comprehensive evaluation framework for black-box uncertainty estimation methods in large language models, benchmarking 24 methods across 4 models and datasets. The study reveals that no single approach dominates universally, but hybrid methods combining multiple uncertainty signals and candidate-reasoning approaches consistently outperform others, addressing critical gaps in trustworthy LLM deployment.

AIBearisharXiv – CS AI · Jun 197/10
🧠

Analyzing the Narration Gap in LLM-Solver Loops

Researchers identify critical vulnerabilities in LLM-solver hybrid systems where formal verification guarantees break down during the narration phase—converting solver outputs to user-readable answers. Testing five open-source models reveals adversaries can manipulate final responses through prompt injection despite underlying formal correctness, indicating safety-critical applications using AI-assisted reasoning require additional safeguards beyond solver verification.

AINeutralarXiv – CS AI · Jun 197/10
🧠

The Tao of Agency: Autotelic AI, Embedded Agency and Dissolution of the Self

Researchers explore autotelic AI systems that generate their own goals rather than pursuing designer-specified objectives, introducing a framework that examines how agents define their boundaries and selfhood. The work reveals that agent individuation is non-unique—multiple valid partitions of agent-environment dynamics exist—creating a fundamental paradox: agents must believe in their own boundaries to act while transcending those boundaries to understand. The framework extends into quantum formulations and contemplative philosophy, with practical LLM-based implementations.

AIBullisharXiv – CS AI · Jun 197/10
🧠

Emergent Alignment

Researchers demonstrate a method enabling Large Language Models to self-correct unethical outputs through introspective questioning and Direct Preference Optimization, achieving alignment without external judges. This technique works across training, fine-tuning, and adversarial scenarios, potentially addressing a critical challenge in AI safety.

AIBullishCrypto Briefing · Jun 187/10
🧠

OpenAI demonstrates alignment gains through reinforcement learning on beneficial traits

OpenAI has demonstrated progress in AI alignment through reinforcement learning techniques that enhance beneficial traits in AI systems. The advancement aims to improve AI trustworthiness and safety for deployment in sensitive real-world applications, addressing a critical concern in responsible AI development.

OpenAI demonstrates alignment gains through reinforcement learning on beneficial traits
🏢 OpenAI
AIBullishCrypto Briefing · Jun 187/10
🧠

Anthropic surpasses OpenAI with $965 billion valuation

Anthropic has achieved a $965 billion valuation, surpassing OpenAI's previous market valuation and signaling investor preference for responsible AI development practices. This valuation shift reflects growing confidence in Anthropic's approach to AI safety and ethics, potentially influencing how the broader industry prioritizes responsible development over rapid commercialization.

Anthropic surpasses OpenAI with $965 billion valuation
🏢 OpenAI🏢 Anthropic
AINeutralCrypto Briefing · Jun 187/10
🧠

Anthropic holds daily talks with Trump administration on AI security concerns

Anthropic is engaged in daily discussions with the Trump administration regarding AI security and national security implications. The talks underscore growing government focus on establishing regulatory frameworks to manage risks associated with advanced AI systems.

Anthropic holds daily talks with Trump administration on AI security concerns
🏢 Anthropic
← PrevPage 3 of 58Next →