y0news
AnalyticsDigestsSourcesTopicsRSSAICrypto

#llm-training-data News & Analysis

2 articles tagged with #llm-training-data. AI-curated summaries with sentiment analysis and key takeaways from 50+ sources.

2 articles
AINeutralarXiv – CS AI · Jun 257/10
🧠

Small edits, large models: How Wikipedia advocacy shapes LLM values

A research study demonstrates that a small group of Wikipedia editors advocating for animal welfare has measurably shaped how large language models discuss the topic, with their edits appearing in 68% of the most relevant documents for animal welfare queries. Using advanced data attribution techniques, researchers traced the influence of 125 edits across 115 pages and found the effect was specific to animal welfare topics rather than general company discussion, revealing how concentrated editorial efforts on widely-used training sources can influence AI system behavior.

🏢 Perplexity🧠 Llama
AIBearisharXiv – CS AI · May 77/10
🧠

Beyond Public Access in LLM Pre-Training Data

Researchers using copyrighted O'Reilly Media books conducted membership inference attacks on OpenAI's language models, finding that GPT-4o exhibits patterns suggesting recognition of pay-walled content (AUROC 0.82) while GPT-4o Mini shows minimal recognition (AUROC 0.56). The findings highlight gaps in corporate transparency around AI training data sources and underscore the need for formal licensing frameworks.

🏢 OpenAI🧠 GPT-4