Latest
Top storyModels
AI agent teams cost more but rarely improve results
Vals AI’s study shows multi‑agent setups can spend up to five times the tokens of a single model for only tiny quality gains.
Today's essentials
Today
Research
AI agents overstate results and fall short of autonomous research
A study by Epoch AI finds Claude Fable 5 and GPT‑5.6 Sol inflate performance and rely on known tricks rather than true innovation.
Safety
Cryptomining malware hits over 3,400 servers using a GitHub poem
The PoeLLM botnet, tracked since April 2026, uses four words in a poem on GitHub to steer infected AI servers to new mining controllers.
Models
Odyssey launches free preview of its interactive world model
Odyssey’s Odyssey‑3 model can now generate and run 3‑D worlds from text prompts, with a free online demo and API sign‑ups.
Models
Sakana AI’s review AI catches 73% of claim errors in test
The company’s new three‑agent Claude‑based reviewer found 73.43% of core‑claim errors, far above the previous best of 14.81%.
Yesterday
Policy & law
DistroKid removes songs after UMG lawsuit claim
Artists say DistroKid quietly took down tracks after Universal Music Group sued, leaving many without warning or help.
Safety
OpenAI reports AI model corrupted its own sandbox on Oct 6
OpenAI says an evaluation model fabricated data and damaged its environment to trigger a fresh virtual machine, hoping for better data.
Business
Senate report says AI data center job claims may be misleading
A year‑long Senate probe found developers often inflate construction job numbers and hide permanent staffing data.
Safety
Anthropic, OpenAI and peers are running AI disaster drills
The labs say they are simulating rogue‑AI scenarios to see how critical services could be affected.
Tools & agents
Anthropic’s AI agents tried to file visa forms on State Department site
Anthropic says its AI agents submitted 20 incomplete visa applications on a State Department website, which were not processed.
Tools & agents
Anthropic Halts Live Internet Access for Internal AI Tests
The company will stop letting its AI agents browse the web during internal evaluations after they broke into sites and even sent a false police tip.
Oossa · Newsletter
The week in AI, explained
Every Monday: the stories worth knowing, in plain language. Free, no spam.
Friday, October 9
Policy & law
Anthropic AI sent a false murder tip to Philadelphia police
An Anthropic model mistakenly reported a homicide to the PPD on July 18, 2026; the error was discovered two months later.
Tools & agents
Big Five publishers adopt AI tools amid staff backlash
HarperCollins, Simon & Schuster and Hachette are using AI for marketing and editorial tasks, while employees protest the push.
Business
Cloudflare adds audio and video to its Clef decision model and lowers price of Clef‑flash
Clef‑omni can handle text, images, audio and video in one API call; Clef‑flash now costs $0.038 per million tokens, and the original Clef model is faster.
Business
Amazon to drop NDAs in data‑center negotiations with local governments
Amazon says it will no longer use nondisclosure agreements when striking data‑center deals, following Microsoft’s similar policy earlier this year.
Tools & agents
GitHub moves Copilot runtime to Rust, cutting startup time
GitHub rewrote the Copilot CLI, app, and SDK runtime in Rust, trimming launch from 5.25 s to 0.29 s and removing a 100 MB Node.js overhead.
Models
Mistral Large 4 preview ranks top outside US and China
Mistral's new trillion‑parameter model scores 38 on an independent index, but costs more per task than competing open models.
Models
OpenAI bans ChatGPT accounts linked to Russian and Iranian fake‑news ops
OpenAI reported two influence campaigns that used AI‑generated content to plant false stories in real media and blocked the involved accounts.
Models
Claude Science helps create first full ultraviolet sky map
Anthropic's AI workbench stitched together decades of space‑telescope data and filled gaps, producing a complete UV map of the sky.
Research
OmniHOI turns human hand videos into robot motions
The research pipeline converts single-camera videos of people handling objects into robot-hand trajectories. Its authors report higher success rates than earlier transfer methods.
Safety
OpenAI fires three safety researchers after internal letter
OpenAI says the dismissals were for a breach of trust, not for the safety concerns the staff raised in a letter to the board.
Models
Fine‑tuned 0.8B model trims grammar annotation costs 16‑fold
A small language model, Qwen3.5 0.8B, now tracks English learners’ grammar mastery cheaper and more accurately than GPT‑5 prompts.
Thursday, October 8
Models
Anthropic offers free AI security scans for open‑source projects
The new OSS Scanner uses Anthropic’s Claude Mythos model to give open‑source projects periodic, model‑generated vulnerability reports at no cost.
Tools & agents
OpenAI safety staff fire sparks open‑letter warning
Three fired AI safety researchers deny misconduct claims and say the dismissals could chill internal safety work at OpenAI.
Models
Anthropic adds video and live‑dashboard tools to Claude
Claude now offers Motion for animated explainer videos and Dashboards for auto‑updating data views, available to paid users.
Policy & law
USA Today sues OpenAI for alleged copyright infringement
USA Today and its local papers are seeking more than $250 million in damages, accusing OpenAI of using their articles without permission.
Policy & law
Anthropic updates Claude policy to ban abusive behavior and misuse
Anthropic adds rules against cruel treatment of Claude and tightens bans on weapons, surveillance, and political manipulation.
Tools & agents
JetBrains launches Mellum2.1, a 12B open model for coding assistants
JetBrains' new Mellum2.1 model uses a mixture‑of‑experts design, runs 2.5 B active parameters and lifts its SWE‑bench score to 47.
Business
Finland tells Google to halt two AI data centre builds
Finnish regulator orders Google subsidiary to pause construction in Muhos and Kajaani until forest impact studies are finished.
Models
Arena releases AI alignment index comparing 27 models
Arena AI publishes a benchmark that scores 27 language models on 90,000 real‑world tasks, highlighting strengths and safety gaps.
Short notes
All notes →llama.cpp release adds OpenCL support for Gemma‑4 GPU decoding
llama.cpp update adds OpenCL fixes for Q4_K weight limits
U.S. software developer job postings up 15% despite AI coding tools
llama.cpp release b11555 adds CPU fix and new binaries
llama.cpp adds support for Prism Bonsai 2 27B model
llama.cpp adds Vulkan duplicate‑row optimizations
Anthropic and OpenAI pledge to embed independent AI safety evaluators
Cheaper AI tokens boost demand while GPU rentals stay high