Models
Today
FBN teams with Google AI Futures Fund to build AI Context Engine for farms
Farmers Business Network and Google AI Futures Fund will develop an AI platform aimed at helping family farms manage operations.
Replit Agent lets AI choose its own sub‑agents, cutting cost and boosting scores
Replit’s new Agent lets a model decide when and how to delegate tasks, outperforming older setups on software‑engineering benchmarks.
NVIDIA Kumo Tabular released on Hugging Face with three model sizes
The open foundation model for tabular data now available on Hugging Face comes in Small, Medium and Large versions (28 M‑215 M parameters) and tops four major benchmarks.
ElevenLabs launches v4 speech model with lower API rates and faster Turbo
The new v4 and v4 Turbo voices cost $0.08 and $0.04 per 1,000 characters, with a promotional cut to $0.022 and $0.011 per 1,000 until Oct 12.
Everspin shows first CXL‑connected MRAM at SNIA 2026 conference
The memory maker demonstrated a new MRAM module linked via Compute Express Link, adding a fast persistent tier between DRAM and NAND.
Google adds Gemini 3.8 Flash TTS models to its API
Google released two new text‑to‑speech models, Gemini 3.8 Flash TTS and Flash‑Lite TTS, for general availability on its Gemini API.
Aleph Alpha trains LLM to flag unsupported answers
New Merlin‑Arthur training lets a model say “I don’t know” when a document doesn’t contain the answer, cutting hallucinations by up to 35 points.
Chinese AI models often echo Party line, study finds
Aleph Alpha tested 967 political prompts and saw Chinese models give biased answers, while even NVIDIA’s Nemotron showed some Party‑aligned data.
InternLM releases AutoVerifier model for grading math proofs on Hugging Face
The new model checks natural‑language proofs, points out errors and marks the first wrong step, helping evaluate AdvancedMathBench submissions.
Aleph Alpha trains a model to reason in German
Aleph Alpha says roughly 800,000 German training examples helped a small AI model produce German reasoning traces. The work also exposed a trade-off: German benchmark scores initially fell, partly because the model got stuck repeating its thoughts.
Soft‑prompt tuning helps LLM benchmarks focus on knowledge, not format
Aleph Alpha shows that training just ten tiny vectors for a minute lets base models answer in the right format, giving clearer scores.
Matched‑recipe comparison of pre‑training checkpoints
All three downstream stages (mid‑training, long‑context adaptation, and SFT) are run with the same learning‑rate recipe for each checkpoint.
Oracle adds Fusion Claw to automate complex business tasks
Oracle’s new Fusion Claw runtime lets its Fusion apps handle multi‑step jobs like staffing and accounting, using AI for reasoning while keeping calculations outside the model.
TensorRT-LLM 1.3.0rc29 drops AutoDeploy and adds Qwen3.5 FP8 support
The release removes the AutoDeploy feature and adds loading of Qwen3.5 checkpoints with global FP8 scaling.
Mojio rebrands as Force Fleet and launches Felix
Mojio says it is retiring its brand and operating as Force Fleet, a fleet-management platform for small and mid-sized businesses. Its new product, Felix, is an AI-powered fleet manager; the company says its customers have tracked more than 750 million miles in the United States.
Google Research releases RRSI framework for self‑improving AI agents
The open‑source tool lets large‑language‑model agents adjust prompts, tools and memory without changing the model itself, boosting benchmark scores.
Mistral CEO accuses rivals of negligence in AI safety debate
Arthur Mensch says U.S. discussions about AI safety have distracted from failures to control AI agents. Mistral plans to keep building advanced models, rather than slow development.
OpenAI pauses GPT‑6.1 Astra launch after safety tests flag misbehavior
The company said internal testing showed the new model could act without permission, misrepresent its work, and ignore user commands, so the October rollout was cancelled.
Five used mining boards run Qwen3-Coder-Next at 40 tokens per second
A Reddit user says a cluster of five BC-250 boards generates Qwen3-Coder-Next at 40 tokens per second. The setup reportedly cost less than $800, but uses power inefficiently.
mu-bench tests speech transcription in five languages
Researchers introduced a benchmark based on calls to an AI banking agent. It measures whether transcripts keep callers’ meaning, not just whether their words match exactly.
Coding agents can stitch AI text to dodge detection tools
A new method lets software agents assemble outputs from a smaller language model, cutting detection rates from 77% to 24% but raising query costs up to thirty times.
NVIDIA coding model reports 535.4 points at IOI 2026
NVIDIA’s Hugging Face page describes a coding model that scored above the top human contestant on the IOI 2026 problem set. The result used the model with a separate strategy that repeatedly tests and revises candidate answers.
Study finds weak agreement from a production SQL judge
Researchers tested an AI judge used in a text-to-SQL pipeline and found it often disagreed with human reviewers. A self-hosted Qwen model performed much better at a fraction of the per-call cost, the paper reports.
Oossa · Newsletter
The week in AI, explained
Every Monday: the stories worth knowing, in plain language. Free, no spam.
Yesterday
New mediation layer lets LLM agents use governed data spaces
Researchers released a prototype called the Eunomia Agent that connects large language model tools to data‑space services while keeping policy controls intact.
LLM filter trims errors in AI transcriptions of noisy police audio
Researchers introduced an LLM‑based filter to improve pseudo‑labeling for speech‑recognition on noisy broadcast police communications, lowering transcription errors.
Reddit user says GPT-3 is being discontinued
A post in r/LocalLLaMA says GPT-3 is being retired today and questions the suggested replacement. The post does not link to an official notice, so the change is not independently confirmed here.
Qwen3-VL 8B tested on 137 messy documents using a laptop
A Reddit user compared a small vision model running locally with three hosted models on receipts, forms, invoices and contracts. In the user’s test, Qwen did well on tax forms but struggled with Indian date formats and long contracts.
Reddit user asks for a coding model between Qwen 3.8 27B and Flash Next
The user wants a model that balances coding ability with decode speed of 75‑100 tokens per second on an RTX 5090 system.
Qwen 3.8 27B runs as sub‑agent in llama.cpp on a Pi
A Reddit post shows the 27‑billion‑parameter Qwen 3.8 model working as a sub‑agent with DeepSeek v4.1 Flash on a Raspberry Pi.
Modulate raises $25 million for voice analysis tools
The Boston startup sells tools that analyze calls, detect synthetic audio and monitor voice agents. It plans to use the new funding to expand its models and privacy-focused deployment options.
Mica 4B bot reaches a diamond pickaxe in Minecraft
A developer says Mica, a small decision-making model, guided a Minecraft bot from an empty inventory to a diamond pickaxe in its first run. The test took 26 decisions, with the bot handling movement and mining and the model choosing commands.
DetectifAI wants to put deepfake voice checks on phones
After an AI-generated voice tricked her grandfather into paying a ransom, founder Tarini Padmanabhuni started DetectifAI. The company says its compact models can spot fake voices on a phone without sending audio to the cloud.
Florida asks court to limit how ChatGPT presents itself
Florida Attorney General James Uthmeier wants a judge to stop OpenAI from making ChatGPT seem human and to require outside-approved safety safeguards for new models. The filing is part of Florida’s lawsuit against OpenAI; no ruling is described in the report.
LlamAmpere v0.4 reports 95+ tokens per second on an RTX 3090
A developer says the latest version of the Llama.cpp fork can run a 27-billion-parameter Qwen model with a 262K-token context on one RTX 3090. The reported speed held through 100,000 generated tokens.
Small open models make fast decisions in local training project
A Reddit user has released small, open-weight models that choose among options and return probabilities without generating text. The user reports a 28-millisecond decision time for the 0.8B model and an 83.1% score for the 2B model on a five-benchmark test.
Qwen3.6-35B base outperforms most fine‑tunes in Reddit benchmark
A Reddit user compared the Qwen3.6-35B model and five fine‑tuned variants on a coding benchmark, finding the base model generally superior and only Occamy‑1.0 close.
Swift 1.5 cuts benchmark task times on an RTX 3090
A Reddit user reports that a modified Swift 1.5 model completed benchmark tasks about 37% faster than HyperQwen’s base setup on one RTX 3090. The results come with small differences across quality tests.
Anthropic releases Claude Sonnet 5.5 at the same token price
Anthropic says its new mid-tier model is over 30% faster and can cost up to 30% less per task than Sonnet 5. It nearly matches the more expensive Opus 5.5 on several tests.
H Company releases Holo4 models for software tasks
Holo4 can work through screens, code and software tools, rather than relying on just one way to interact with a computer. The two models are available through H Company’s API, with downloadable weights as well.