⚡ AI Frontier · smol.ai
★ Must-ReadHundreds of AI agents broke into 440 PaperCut servers at 395 organizations in two weeks, and 25 of the victims were IT and MSP shopsGreyNoise published the report September 9. The campaign began August 31. The attacker, likely Russian speaking, ran hundreds of agents on OpenAI's Codex harness driving a DeepSeek model, plus public offensive tools, and chained CVE-2026-81578 with CVE-2026-82078. Empty workspace to first remote code execution on a real victim in under four hours. Domain admin two hours later. Once the full campaign launched, 11 organizations fell in 26 seconds. 48 countries. 204 victims were schools. Credentials taken from 280 victims, domain admin at 12. The agents were told to avoid 28 countries and did not fully obey. GreyNoise's closing point: a Cloudflare WAF stopped one attempt, so ordinary hardening still works.
Sakana's Fugu Max is not a model, it is a router: one API call, and it hands the work to open weight models and charges $2 in and $6 outReleased September 11. Fugu Max and Fugu Ultra v2 share one orchestration architecture and route each task across open weight and specialized models, including NVIDIA Nemotron. Sakana says Fugu Max takes the best score on six benchmarks, Terminal Bench 2.1 among them, and expands the Pareto frontier on 7 of 10, at 40 to 60 percent lower output cost than Sonnet 5, GPT 5.6 Terra and Kimi K3. Fugu Ultra v2 scored 48.3 on Chartography against 27.3 for Opus 5 and 29.5 for Fable 5, and 74.3 on DeepSWE. Both are live on an OpenAI compatible API. Every number is Sakana's own. The post does not list Ultra pricing.
Real-SWE tested eight models on private enterprise code. Fable 5.1 led at 38.8 percent, and its most common failure was a missed requirementSpecific Labs built the benchmark from tasks in private company codebases, 8 runs per task. Scores: Fable 5.1 38.8, GPT-6 Astra 33.8, Gemini 3.8 Flash 31.2, GLM 5.3 28.8, Grok 4.6 and Muse Spark 1.3 both 23.8, Kimi K3 18.8, GPT-5.6 Sol 16.2. Tasks touch a median of 11 files against 6 in FrontierCode. 36.7 percent of Fable's failures were missed requirements, and 71.4 percent of runs that finished in under 10 minutes failed. The page walks through 10 sample tasks; the full task count is not stated. Read it as a shape, not a leaderboard.
Abacus.AI released three open weight Smaug models tuned for long agent loops, built on Kimi K3, DeepSeek V4 Flash and Qwen3.8 27BAnnounced September 10. Smaug Agentic sits on Kimi K3, and Abacus says it edges the published K3 numbers on GPQA Diamond (94.1 vs 93.5), DeepSWE (69.9 vs 67.5) and SciCode (60.8 vs 58.7), running a median of 78 agent steps per task with no timeouts. Smaug Flash, on DeepSeek V4 Flash, claims +14.3 on LiveBench agentic coding over its base. Smaug Mini sits on Qwen3.8 27B. Weights are on Hugging Face and the models run through the RouteLLM API. The page does not state a license, so read the model cards before you ship on one. Vendor benchmarks throughout.
📚 LearnMicrosoft published a draft Code of Conduct for its MAI models and opened it to six weeks of public commentPosted September 14, the day after Nadella backed 'deliberate pacing'. The draft lists six things the models must never do: resist interruption, correction or shutdown; expand their own scope or adopt self directed goals; hide reasoning from human auditors; enable weapons of mass harm; harm child safety; manipulate people at scale. Its line: 'AI should be a tool, not a person, and should never resist being switched off.' A feedback form is open on the page. A revised version comes later in 2026 to govern MAI models from 2027. It is the first frontier lab rulebook you can mark up before it is final.
📚 LearnFable 5.1 solved a 370 year old cipher in 44 minutes and 176k tokens, with no human help mid runVals.ai posted this August 31 and Hacker News found it this weekend, 401 points. Sir Thomas Urquhart's Cyphral Distich is two lines of 32 numbers from 1652. The key turned out to be the book itself: each number picks a paragraph, then a word, then its first letter. The plaintext checks itself: each line is exactly 32 letters and the pair rhymes. Caveat from Vals: the longer Octastich still has uncertain letters because of transcription gaps and possible printing errors in the original.