The Morning Wire

AI NEWS REPORT

Latest AI news weekdays · New edition by 9 AM PT
FRONTIER AI PRICE WAR: OPUS 5.5 AND GPT-6 SOL LAND 90 MINUTES APART
★ Must-Read / Watch agentic engineering / build-better-agents📚 Learn state of AI coding / useful technique

⚡ AI Frontier · smol.ai

★ Must-ReadGPT-6 Sol and Luna cut OpenAI's prices in half, but Artificial Analysis says the intelligence barely moved: Sol goes from 47 to 48
The Decoder, September 22. Sol drops from $4 in and $20 out to $2 and $10. Luna drops from 20 cents in and $1.20 out to 10 cents and 50 cents, with a 90 percent discount on cached input. OpenAI credits 'improvements in caching and inference.' The Decoder's read of Artificial Analysis: per-task cost halves, but Sol only moves from 47 to 48 on the Intelligence Index and Luna stays at 37. On GDPval-AA, a test of office work, Sol loses about 100 Elo points against its predecessor. The Decoder also calls OpenAI's benchmark picks 'cherry-picked,' since the charts compare against Opus 5 and skip Terminal-Bench 4.0. OpenAI says Sol makes about half as many mistakes as GPT-5.6 Sol. Live now in ChatGPT Work, Codex and the API as gpt-6-sol and gpt-6-luna. Why it matters: if you already run GPT-5.6 Sol, you just got a 50 percent discount for doing nothing but changing a model name.
★ Must-ReadOpus 5.5 sends most cybersecurity work to the older Opus 4.8, and does it without refusing, so you may not notice
September 22. Opus 5.5 is the first Opus model to ship with the same kind of safeguards as Claude Fable 5.1 for cybersecurity, biology and distillation. Anthropic's launch page says finding and fixing bugs in your own code as part of normal development stays on Opus 5.5, 'but most cybersecurity tasks will be re-routed to Opus 4.8.' The fallback is transparent, meaning you get an answer from the older model instead of a refusal. Biology work falls back the same way. Distillation attempts are blocked outright with no fallback. The way out is Anthropic's Cyber Verification Program, which will add Opus 5.5 in the coming weeks with three tiers of access, and a new Life Sciences Verification Program for labs and drug companies. Why it matters: an MSP or security team that benchmarks Opus 5.5 on pentest or triage work may be benchmarking Opus 4.8 without knowing it.
📚 LearnClaude Code only reads AGENTS.md when telemetry is on. Turn telemetry off and the file is skipped with no warning
Posted September 23, near the top of Hacker News within an hour. Claude Code 2.1.277 added support for AGENTS.md, the shared instruction file that Codex, Cursor and other agents read. The author tested 2.1.280 with a canary word and found the loader sits behind a remote feature flag. If DISABLE_TELEMETRY=1 or CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 is set, the flag cannot be fetched, it falls back to off, and your AGENTS.md is silently ignored. Setting those variables back to 0 does not fix it. The workaround is one line: create a CLAUDE.md that contains @AGENTS.md, which uses the normal import path and skips the flag. No Anthropic response yet. Why it matters: privacy-minded teams are exactly the ones who turn telemetry off, and they are the ones losing their instructions.
Unreal Agent: an open-source harness that runs tool calls in the background, claiming up to 40 percent lower cost than Codex
Released September 22 by Unreal Labs, a Sequoia and First Round backed team. 217 points on Hacker News. The idea: the harness handles tool calls asynchronously, so the model does not have to spend tokens on waits, polls and heartbeats, and the agent can keep taking your input while tools run. Their published runs: Terminal-Bench 4.0 at 57.9 percent for $1,428 total, SWE-Atlas at 65.8 percent for $936, DeepSWE 1.1 at 72.4 percent for $1,367. They claim up to 40 percent savings against Codex and up to 20 percent against Pi. Why it matters: on a day when model prices dropped, this is the other lever, which is spending fewer tokens per task. What to watch: these are the company's own numbers.
Alibaba's Qwen Audio 3.1: five new speech models and price cuts of up to 95 percent on speech-to-text
September 23, on Qwen Cloud. Two speech recognition models, one of which can tell speakers apart and pick up emotion and background sound. Two text-to-speech models with voice transfer. And a real-time model that can listen and talk at the same time and take interruptions; Alibaba says when it 'detects a low mood, it responds more slowly and with more empathy.' Price cuts: about 70 percent on text-to-speech, about 85 percent on real-time, up to 95 percent on recognition. Why it matters: voice agents are mostly a cost problem at call-center volume, and this is the price floor moving. API only on Qwen Cloud as far as the report says.
Alphabet's Intrinsic open-sources the robotics stack it uses in real factories, under Apache 2.0
Announced September 22 at ROSCon 2026 in Toronto. Intrinsic Core is a set of ROS-compatible building blocks: a real-time control framework that works across robot brands, pose estimation built on Nvidia's FoundationPose, motion and grasp planning, Gazebo simulation, camera calibration and drivers for supported FANUC and Universal Robots hardware. It runs on local hardware and is on GitHub. Intrinsic also released a reference design for CNC machine tending. Why it matters: the hard, boring parts of a robot cell are now free to use and change, which lowers the cost of trying physical AI in a small shop.

📺 Watch · latest videos

★ Must-WatchI Tested Opus 5.5 vs. GPT-6 Sol on 10 Real Use Cases
Nate Herk, September 23. 'I Tested Opus 5.5 vs. GPT-6 Sol on 10 Real Use Cases.' A head-to-head the morning after both launched, which is exactly the test today's Level Up asks you to run on your own work.
Nate Herk | AI Automation · 70.9K views · 1.4K likes
Introducing Claude Opus 5.5
Claude, September 22. 'Introducing Claude Opus 5.5.' Anthropic's own launch video. Watch it next to the benchmark rows in today's explainer.
Claude · 191.8K views · 3.3K likes
📚 LearnJev CEO: I made ChatGPT, now I'm building what's next
Featured by Latent Space.
AI Engineer
📚 LearnMarc Benioff & Sam Altman | Dreamforce 2026
Featured by Latent Space.
Salesforce
Bring Your MCP Server to Amazon's $190K Hackathon!
Latest from Cole Medin.
Cole Medin · 29 views · 1 likes
Skill issue: stop deploying vision language models, use them with Skills — Merve Noyan, Hugging Face
Merve Noyan wrote a book on vision language models and now wants developers to stop calling them directly. Put one in front of a camera and you will n
AI Engineer · 306 views · 12 likes
GPT 6 + Hyperframe = Crazy combo for expert-level videos
Get free Claude Code AI Coding course: https://clickhubspot.com/9cb13b 🔗 Links - Treg repo (Give agents access to 3000+ API for free): https://github.
AI Jason · 1.1K views · 79 likes
Anthropic went CRAZY (Opus 5.5)
Latest from Matthew Berman.
Matthew Berman · 87.5K views · 1.5K likes
★ Must-WatchMeet Claude Opus 5.5
For most tasks, Claude Opus 5.5 performs at the level of Claude Fable 5.1, our most intelligent model. It writes more clearly and puts the most import
Claude / Anthropic · 6.8K views · 222 likes

🗣 Voices & Blogs

★ Must-ReadAlex Tabarrok: the price of intelligence is falling faster than any technology in history, 725 times cheaper in under 18 months for the same performance
Marginal Revolution, September 23, the morning after the price war. Tabarrok's numbers: an average cost drop of 47 percent per quarter over three years, performance up about 13 times a year, and a 725-fold drop in cost for equal performance going from OpenAI's o3 to GPT-5.6 Luna. His point is that models are getting smarter and cheaper at the same time, which has almost never happened with any other big technology. Why read it: it puts yesterday's launches on a longer line.
📚 LearnJyn, 'Tokens too cheap to meter': once a model costs less than a tool, you start putting models inside tools
Posted September 16, back on the Hacker News front page this morning. Jyn argues the price of machine intelligence is dropping by several orders of magnitude a year, with better GPUs, better models and better inference engines all pushing at once. The example is Jev's pricing: 4.2 cents per million input tokens, output free. The line to keep: 'Once models are cheaper than a tool, it becomes attractive to put models in tools.' Why read it: it is the builder's version of Tabarrok's chart, about what you design differently when a model call costs almost nothing.
📚 LearnDuarte O. Carmo, 'Jev in 25 lines of Python': the new decision-model idea, rebuilt locally on any GGUF model
Posted September 22, 372 points on Hacker News. Jev is TypeSafe AI's model that returns probabilities instead of text. Carmo shows you can get the same shape from any local model by reading its raw scores for a fixed set of answer choices and turning them into probabilities, in about 25 lines. His summary: 'It classifies: it gets a prompt with choices and outputs probabilities.' It runs locally, so the data never leaves your machine. Why read it: it takes the magic out of this week's most hyped release and hands you a working copy.
The MVUEH break: GPT-6 Astra picked, simulated and cracked a 1941 Enigma message that had resisted every attempt since 2005
Cryptocellar, validated September 19, 696 points on Hacker News. Message 172, MVUEH, was a German Army Enigma signal from July 10, 1941. Carter Leffer gave GPT-6 Astra the set of unsolved messages and nothing else. The model chose MVUEH as the most promising, wrote its own Enigma simulator, and broke it using the repeated place name ROSENOW as a crib. The key used an unusual wheel order, 253 instead of 512. The write-up's view: two days of model work that 'would take a human researcher weeks or even months.' Why read it: a clean, checkable example of an agent doing real research end to end, with the human only setting the task.
The software supply chain is the new battlefield. AI just changed the rules.
AI coding tools have seriously accelerated developer speed, but AI has also done the same for attackers — and the The post The software supply chain… · The New Stack
Grok Bot Agents: how to automate your life in 10 Steps (Full-tutorial)
Every AI tool you have used so far waits for you. · Codez (@0xCodez)

🌏 The Wire · Drudge / Breitbart

PENTAGON REVIEW: OVERRELIANCE ON PALANTIR'S MAVEN AI CONTRIBUTED TO THE STRIKE THAT KILLED 123 CHILDREN AT AN IRANIAN SCHOOL
Reported by Bloomberg and Gizmodo, citing an unreleased Pentagon review. On February 28, two Tomahawk missiles hit Shajarah Tayyebeh Elementary School in Minab, Iran. Investigators found staff relied on Palantir's Maven Smart System to catch outdated targeting records, a job it was not built to do. The site was still listed as a Revolutionary Guard facility even though satellite images from 2017 to 2019 showed it had become a school. Civilian harm teams had shrunk by about 90 percent, and no one reviewed the site before launch. Palantir says it 'is not responsible for the underlying data.' Why it matters: this is the clearest case yet of people trusting an AI tool to do a check it was never designed to do. 755 points on Hacker News.
PATCH WORDPRESS NOW: UNAUTHENTICATED FLAW IN 23 VERSIONS CAN LEAD TO REMOTE CODE EXECUTION
CVE-2026-87902, CVSS 9.2, published September 22. Affects WordPress 4.7.0 through 7.1.1, fixed in 7.1.2 and back-ported to every affected branch. An attacker with no login can make the page-template lookup load a PHP file from outside the theme folder. Turning that into code execution needs two things: a theme with a top-level folder whose name starts with 'page-', which includes Twenty Twelve, Twenty Fourteen, Neve, Hestia and Sydney, and a usable PHP file such as pearcmd.php, which is reachable by default in official PHP Docker images and cPanel setups. Why it matters: if you host client sites, this is a patch-today item. Found by Robert Ressl.
TODAY: ALTMAN, AMODEI, BENGIO AND HUGGING FACE'S DELANGUE BRIEF THE UN SECURITY COUNCIL, ITS FIRST MEETING ON AI SAFETY
Wednesday afternoon in New York, chaired by French foreign minister Jean-Noel Barrot. Sam Altman attends in person. Dario Amodei joins remotely. Yoshua Bengio speaks for the UN scientific panel whose brief this week said there is no assurance humans keep control of more capable agents. France's four questions include how AI evaluation and verification can keep up with the technology. No outcome document is expected. Why it matters: it is the first time the Council has put the lab CEOs at the table on safety, one day before the Trump-Xi summit.
META'S NAT FRIEDMAN: MUSE WAS 'HEAVILY INSPIRED AS A PRODUCT BY OPENCLAW'; SAME WORKSPACE FILE NAMES, NEAR-IDENTICAL SOUL.MD
September 22, on X. Friedman, who runs product at Meta Superintelligence Labs, said 'We built Muse from scratch,' but the team 'fell in love with using OpenClaw' in January and set out to build something like it that is safe and can reach billions of people. He called OpenClaw's creator Peter Steinberger 'a genius.' Users had spotted that Muse uses OpenClaw's workspace file names and a SOUL.md personality file with nearly the same text. Why it matters: open source did its job here. The small project set the design and the biggest company in social copied it in the open.
SNORKEL AI TRIPLES TO $3.5 BILLION: REVENUE UP 18 TIMES IN A YEAR SELLING FINISHED TRAINING DATA
September 22. A $350 million Series E led by Insight Partners and S32, at $3.5 billion, up from $1.3 billion 17 months ago. Annualized revenue is $375 million. Snorkel started as data-labeling software and now sells finished datasets and reinforcement learning environments built from synthetic data plus human experts. Why it matters: the labs are paying real money for graded tasks to train agents on, and that is where the next model gains come from.
EMA RAISES $77 MILLION; CEO SAYS CUSTOMERS ARE 'ON THE WAY TO REPLACE LARGE SAAS APPLICATIONS COMPLETELY'
September 23. Series B led by Bengaluru's Creaegis, with Accel, Section 32 and Prosus adding more. Ema runs teams of AI agents across HR, IT and finance work. It reports 50-plus enterprise customers, more than a million users, $150 million-plus in bookings, 180 percent net dollar retention and revenue up 50 times in two years. Customers named include Google, Microsoft, ADP and PwC. Why it matters: agent platforms are now selling against the SaaS seat, which is the line item MSPs manage.
MISTRAL BUYS PARIS AD-TECH STARTUP PIMENTO, ITS THIRD DEAL OF 2026, TO BUILD OUT ITS VIBE ASSISTANT
Confirmed September 23, two weeks after Mistral's 3 billion euro raise. Price reports differ: Mistral says 'several million euros,' L'Informe says about 3.8 million, Sifted says 12.7 million in cash and shares. Pimento makes ad visuals and short animations for more than 180 brands. Its team moves to work on Vibe, the assistant formerly called Le Chat. Earlier 2026 buys were Koyeb for cloud and Emmi AI for physics simulation. Why it matters: Europe's leading lab is buying product teams, not just training models.
ANTHROPIC IS BUILDING A WET LAB WHERE CLAUDE DIRECTS ROBOTS THROUGH DRUG EXPERIMENTS
Reported September 22, citing Reuters. The lab is in the San Francisco area and is led by Eric Kauderer-Abrams, head of Anthropic's life sciences group. Claude runs the experiments through Claude Science and the Model Hardware Standard, with humans still in the loop for safety. The focus is diseases long called undruggable and antibodies that hit more than one target. Anthropic says it will not run its own clinical trials. Why it matters: the model leaves the chat window and starts running physical experiments.
WAYMO WILL PAY YOU A BUS FARE TO RIDE TRANSIT: $2.85 BACK WHEN A WAYMO TRIP AND A TRANSIT TAP LAND WITHIN TWO HOURS
Announced September 22 for the San Francisco Bay Area, employees first and the public in the coming weeks. Link a Visa card, use it on a Waymo ride and on any of 27 Bay Area transit agencies within two hours, and you get $2.85 in Waymo Cash. Waymo is also leasing 40 parking spaces at Caltrain stations to stage cars for riders. It says more than half of its riders in its mature cities also use transit. San Jose Mayor Matt Mahan: it strengthens public transit 'not replace it.' Why it matters: robotaxis pitched as the ride to the train, not the rival to it.

🤖 Trending Models · Hugging Face

prism-ml/Ternary-Bonsai-2-27B-gguf
text-generation · ★1.9K · 2.8M dl
abenzerps/Qwen-Image-2.1-Uncensored-GGUF
text-to-image · ★1.3K · 350.7K dl
deepseek-ai/DeepSeek-V4.1-Flash
image-text-to-text · ★3.6K · 570.9K dl

📈 Markets

NVDA 228.87 ▲0.7%
MSFT 498.00 ▼0.7%
GOOGL 351.16 ▼1.1%
AMZN 254.98 ▼1.3%
META 736.60 ▼0.6%
AMD 615.52 ▲9.9%
AVGO 362.66 ▲1.6%
PLTR 183.09 ▲3.1%
SPCX 151.85 ▼0.6%
TSLA 375.30 ▲3.0%

The BriefAnthropic and OpenAI both shipped new models on Tuesday, about 90 minutes apart, and both cut the price. Claude Opus 5.5 costs $4 per million input tokens and $20 per million output, down from $5 and $25, and Anthropic says it uses fewer tokens per task, so a typical job costs about 40 percent less than on Opus 5. OpenAI's GPT-6 Sol now costs $2 in and $10 out, and GPT-6 Luna costs 10 cents in and 50 cents out. That is half of what the GPT-5.6 versions cost, and OpenAI says these are permanent prices. They are not the same kind of deal. Artificial Analysis scores Opus 5.5 at 58, the highest it has ever recorded, and scores Sol at 48, about where GPT-5.6 Sol already was. So Sol is the same brain for half the price, and Opus 5.5 is a better brain for a little less. One more thing to know if you do security work: Opus 5.5 quietly hands most cybersecurity requests to the older Opus 4.8 unless your organization is approved for more access. If you pay for AI by the token, check which model your tools default to this week. The bill may have just dropped without you doing anything, or it may drop only if you switch.

Level UpTry the new models at medium effort before you try them at max. Effort is the setting that controls how long a model thinks before it answers, and you pay for that thinking as output tokens. Simon Willison ran his usual test on launch day and found Opus 5.5 at max effort thought until it hit its 128,000-token output limit, returned nothing, and still cost $2.56. Anthropic's own default is medium. Pick one real task you run every week, run it once on Opus 5.5 at medium and once on GPT-6 Sol, and write down the cost and whether the answer was right. That one table tells you more than any launch chart. Simon Willison: Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war