y0news
AnalyticsDigestsSourcesTopicsRSSAICrypto

#software-engineering News & Analysis

105 articles tagged with #software-engineering. AI-curated summaries with sentiment analysis and key takeaways from 50+ sources.

105 articles
AINeutralOpenAI News · Feb 186/106
🧠

Introducing the SWE-Lancer benchmark

A new benchmark called SWE-Lancer has been introduced to evaluate whether frontier large language models can earn $1 million through real-world freelance software engineering work. This benchmark tests AI capabilities in practical, revenue-generating programming tasks rather than traditional academic assessments.

AIBullishOpenAI News · Aug 135/105
🧠

Introducing SWE-bench Verified

SWE-bench Verified is being released as a human-validated subset of the original SWE-bench benchmark. This new version aims to provide more reliable evaluation of AI models' capabilities in solving real-world software engineering problems.

AINeutralarXiv – CS AI · Apr 75/10
🧠

Measuring LLM Trust Allocation Across Conflicting Software Artifacts

Researchers developed TRACE, a framework to evaluate how LLMs allocate trust between conflicting software artifacts like code, documentation, and tests. The study found that current LLMs are better at identifying natural-language specification issues than detecting subtle code-level problems, with models showing systematic blind spots when implementations drift while documentation remains plausible.

AINeutralarXiv – CS AI · Mar 175/10
🧠

Describing Agentic AI Systems with C4: Lessons from Industry Projects

Researchers propose a new C4-based documentation framework specifically designed for agentic AI systems, which operate through specialized agents collaborating via artifact exchange and tool invocation. The approach provides structured modeling vocabulary and hierarchical description techniques to capture the unique architectural patterns of these systems for industrial applications.

AINeutralOpenAI News · Feb 114/106
🧠

Harness engineering: leveraging Codex in an agent-first world

This appears to be a technical article by Ryan Lopopolo discussing engineering approaches for leveraging Codex (OpenAI's code generation model) in agent-first development environments. The article focuses on practical implementation strategies for integrating AI code generation tools into modern software development workflows.

← PrevPage 5 of 5