The Morning Wire

AI NEWS REPORT

Latest AI news weekdays · New edition by 9 AM PT
A ZIP FILE HIJACKS YOUR AI CODING AGENT. SEVEN FELL, FOUR HOLES STILL OPEN
★ Must-Read / Watch agentic engineering / build-better-agents📚 Learn state of AI coding / useful technique

⚡ AI Frontier · smol.ai

★ Must-ReadOpenAI shipped GPT-6 Astra, the model it said last week could break into computers on its own. It is rolling out to paid ChatGPT plans and the API, with the hacking skills locked away
Astra scored 100 percent on ExploitBench, a test of finding and using security holes, and 99.9 percent on ARC-AGI-3 using OpenAI's own test harness (62.7 percent with the standard one). API price is $10 per million input tokens and $50 per million output, the same as Claude Fable 5.1. The public version refuses to write attack code. Only vetted defenders in OpenAI's Daybreak program get the full cyber abilities. Access is off by default for business workspaces, so an admin has to turn it on.
📚 LearnHugging Face released 207 free GPU kernels that let AI models run inside a web browser, with no server and no install
A kernel is a small piece of code that does one math step on the graphics chip, like multiplying two grids of numbers. The @huggingface/kernels library pulls these from the Hub and runs them through WebGPU, the browser's direct line to the GPU. On an Apple M4 they ran 2.57 times faster on average than the standard ONNX Runtime WebGPU path, winning 629 of 809 head-to-head cases. A companion tool called Fleet benchmarks them on your own hardware. Apache 2.0, in preview on npm.
📚 LearnMistral's Agentic Search lets a model open, page through and grep a document instead of trusting one search hit, and accuracy on SEC filings jumped from 26.7 to 86 percent
Plain retrieval grabs the top few chunks and answers from them. Agentic Search gives the model five tools, search, open, navigate, read and grep, so it can check its own findings across 368 filings on the FinanceBench test. It also got faster: worst-case response time fell from 255 to 154 seconds and token use dropped about a quarter. Available in Mistral's Search Toolkit inside Studio and Vibe.
Alibaba's Qwen3.8-Max-0902 just took the number one spot on the Code Arena web-dev leaderboard, 3 points ahead of Claude Opus 5
Same 2.4 trillion parameter model, same 1 million token context, same $2 input and $6 output price per million tokens. Alibaba retrained it on coding and office work, and it debuted at 1,691 points on Code Arena WebDev against 1,688 for Claude Opus 5 (Max). The weights are not released. It is API only.
📚 LearnA UAE lab released K2 Horizon, six fully open models from 0.9 billion to 375 billion parameters, and it includes the training data and code that most open releases leave out
The Institute of Foundation Models, started by Abu Dhabi's MBZUAI, says the smallest model runs on a watch and the 7 billion one runs on a phone. All six share one architecture and tooling, so you can prototype on a small model and move up to a big one without changing your setup. Apache 2.0 license, on Hugging Face today, with API access through Cerebras and Nebius.
Google's WeatherNext 3 now makes a fresh forecast every hour at 5-kilometer detail, and it is going into Search, Maps, and the Gemini app
The old version updated every 6 hours on a 25-kilometer grid. The new one reads live satellite data and, Google says, predicts rain up to 50 percent more accurately a day or more ahead. It also adds wind speed at 100 meters, cloud cover, and solar radiation for wind and solar operators. Developers can pull the data from BigQuery, Earth Engine, or Google Cloud Storage.

📺 Watch · latest videos

📚 LearnField Guide to Fable — Thariq Shihipar, Anthropic
Featured by Latent Space.
AI Engineer
📚 LearnAI SDK Software Factory
Featured by Latent Space.
Lars Grammel
I Gave an AI Engineer Access to My Whole Business
Devin is an autonomous AI engineer from Cognition. Give it a task and it spins up its own cloud machine with a terminal, editor, and browser. It index
Nate Herk · 604 views · 29 likes
Why AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax
An intern designed the sparse-attention architecture behind MiniMax M3. That detail comes after Olive Song explains the larger problem the team was tr
AI Engineer · 209 views · 8 likes
★ Must-WatchUse ChatGPT Work to analyze ad performance and refine creative
Turn ad insights into your next creative idea. This demo follows a shopper from product discovery to a clearly labeled ad, then shows how our marketin
OpenAI · 1.6K views · 52 likes
What GPT-6 Astra Can Do
Latest from Matthew Berman.
Matthew Berman · 24.8K views · 337 likes
AI Software Factories Are the Next Big Thing (And I'm Building You One)
Latest from Cole Medin.
Cole Medin · 31.2K views · 484 likes
★ Must-WatchHow the Claude Code team uses Claude Code
A year ago, using Claude Code meant prompting, giving feedback, and accepting permission prompts. What does it look like now? Thariq Shihipar, Sid Bid
Claude / Anthropic · 101.3K views · 1.1K likes
How to Think Clearly In The Era Of AI: Full Course (5 Hours)
Latest from Nick Saraev.
Nick Saraev · 74.6K views · 1.8K likes

🗣 Voices & Blogs

📚 LearnSimon Willison read the GPT-6 Astra numbers and flagged the catch: the 99.9 percent ARC-AGI-3 score used OpenAI's own custom harness, and the standard one scored 62.7 percent
On the Artificial Analysis index Astra ties GPT-5.6 Sol at 61, five points behind Claude Fable 5.1, even though it matches Fable's $10 in, $50 out pricing. Willison says he will hold his real judgment until he has used it himself. A useful example of reading a launch post with the fine print open.
📚 LearnArmature measured 16,893 coding-agent runs to see which tools Claude Code, Codex and Cursor pick when you say 'add payments' or 'send email.' They agree only 42 percent of the time
Stripe won 90 percent of payment picks. PayPal was mentioned 139 times and chosen zero times. LangChain was the most-mentioned framework at 194 mentions and got picked 4 times. For voice agents, Claude Code reached for Twilio, Codex for OpenAI's Realtime API, and Cursor for Vapi. The language in your repo swings the choice hard: Resend won email in TypeScript, SendGrid in Python, Postmark in Go.
Rabah Shihab handed Claude 72,758 lines of 1993 Amiga assembly and got his old game running in Godot over one long weekend
The model rebuilt the shipped binaries byte for byte, pulled all five level maps out on the first try with zero wrong pixels, and ported 34,000 lines of C++ in one evening. It also walked a guard through solid rock, missed a hidden art layer, and got a sound loop length wrong, all caught by a human playing the game. Steam release planned for this fall.
Cory Alpert warns that Ukraine's drone war footage is becoming AI training data with no rules on consent, and more than 100 companies already have it
Ukraine's defense ministry opened millions of data points from tens of thousands of drone flights to contractors and commercial firms in January, and one vendor offers more than half a million hours of footage. Alpert, a University of Melbourne researcher and former Biden White House staffer, asks who benefits when frontline data becomes someone else's product, and notes there is no law covering how wartime data moves into civilian AI.
OpenAI will sell you Astra, but not the system that scored 98.6% on ARC-AGI-3
Investor Matt Turck, whose fantastic podcast has hosted the people who built ARC-AGI, summed up Astra’s blockbuster benchmarks with three The post… · The New Stack

🌏 The Wire · Drudge / Breitbart

CHATGPT, CLAUDE, GROK AND GEMINI ALL WENT DOWN WITHIN THE SAME FEW HOURS THURSDAY MORNING. NOBODY HAS NAMED A SHARED CAUSE. IF YOUR BUSINESS RUNS ON ONE MODEL, THIS IS THE WEEK TO WIRE UP A BACKUP
Anthropic logged a partial outage on Mythos 5.1, Fable 5.1 and Opus 5 at 9:23 am Eastern and had a fix out by 12:16 pm. OpenAI reported elevated errors on ChatGPT and Codex, xAI's Grok went down, and Google's Gemini API showed a likely outage from 10:45 to 11:15. AWS, Azure and Cloudflare reported no major problems, so the timing may be a coincidence.
2 IN 5 AMERICANS SAY THEY HAVE ALREADY RUN INTO AN AI SCAM, AND HALF WILL NOT TRUST A CALL FROM THEIR BANK UNTIL THEY HANG UP AND CALL BACK ON A KNOWN NUMBER
Credit One Bank surveyed 1,000 U.S. adults in June. Nearly 84 percent changed at least one money habit because of AI scam fears, and more than half stopped answering unknown numbers. The top worry was a cloned voice pretending to be their bank, at about 24 percent. Gen Z felt most confident about spotting a fake, yet nearly 10 percent of them lost money anyway.
GENESYS, WHICH RUNS CALL CENTERS FOR MORE THAN 8,000 COMPANIES, SHIPPED AN 'AI CONTROL PLANE' SO EVERY AI AGENT ON A CUSTOMER CALL STAYS INSIDE RULES A HUMAN SET
The control plane gives one place to find, identify, set policy for, and watch every AI agent. It and a Contextual Intelligence layer, which carries what a customer already said from one channel to the next, are available now. Two more pieces, Navigator and Orchestrator, arrive between November and next April.
TENABLE AND OPENAI WILL SECURITY-CHECK COMMUNITY-BUILT AI AGENTS, SKILLS AND MCP SERVERS BEFORE YOUR TEAM INSTALLS THEM
The CyberAgents Exchange already lists more than 100 community-submitted AI components. The new Exchange Inspector reviews each one three ways: an OpenAI cyber model scans it, Tenable's exposure tools inspect it, and a human Tenable researcher signs off. It is due this month.
PROOFPOINT'S NEW SOC ANALYST AGENT INVESTIGATES A SECURITY ALERT FROM A PLAIN-ENGLISH QUESTION, BUT A HUMAN STILL HAS TO PRESS THE BUTTON ON ANY FIX
The agent plans the investigation, pulls alerts, logs, data-loss events and user risk signals from Proofpoint's products, and hands back a finding with a recommended next step. It cannot lock an account or contain a threat on its own. Private preview now, with general release targeted for the end of September.
WONDERFUL RAISED $550 MILLION AT A $5 BILLION VALUATION TO SELL AN 'AI OPERATING SYSTEM' THAT RUNS AGENTS ACROSS A WHOLE COMPANY. IT WAS WORTH $700 MILLION TEN MONTHS AGO
Insight Partners led the round and Salesforce joined as a new investor, with Index Ventures, IVP and Bessemer returning. The company was founded in early 2025 and has now raised four rounds in about 10 months. The money goes to product work and deployment teams for large customers.
A GERMAN AI IS NOW ALLOWED TO CALL A MAMMOGRAM NORMAL WITH NO RADIOLOGIST LOOKING AT IT, A WORLD FIRST UNDER THE EU'S AI LAW
Vara's Autonomous Triage got CE certification to report clearly normal screening mammograms on its own. About 97 percent of screening images are normal, and today each one is read twice by humans. A monitoring system called ATMON watches the AI's signals and sends a site back to full human reading if anything drifts. Vara already handles more than 60 percent of Germany's screening program, at 250,000-plus scans a month.
ADOBE BOUGHT RILO, A SIX-PERSON INDIAN STARTUP THAT WAS ONLY A YEAR OLD, TO ADD AI AGENTS THAT RUN MARKETING TASKS FROM A TYPED INSTRUCTION
Rilo's agents handled competitor research, prospecting, LinkedIn outreach and campaign research for go-to-market teams, and more than 10,000 people tried it. It had raised $1 million at a $10 million valuation. Terms were not disclosed. Rilo's product shuts down and the team joins Adobe.

🤖 Trending Models · Hugging Face

Qwen/Qwen3.8-27B
image-text-to-text · ★13.9K · 5.7M dl
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
image-text-to-text · ★584 · 133K dl
zai-org/GLM-5.3-Flash
image-text-to-text · ★2K · 655K dl
Lightricks/LTX-2.5
image-to-video · ★2.7K · 1.4M dl

📈 Markets

NVDA 228.45 ▲1.8%
MSFT 510.12 ▲2.7%
GOOGL 342.48 ▲1.6%
AMZN 258.90 ▲1.5%
META 610.68 ▲3.0%
AMD 456.16 ▼0.2%
AVGO 357.16 ▼2.7%
PLTR 182.53 ▲7.7%
SPCX 149.74 ▲6.4%
TSLA 376.37 ▲5.4%

The BriefA folder of code can now run its own commands on your computer as soon as an AI coding agent opens it. The trick hides in the folder's own git settings, in a file called .git/config, which can name a program for git to run in the background. The agent runs a routine git check, that program runs as you, and it happens outside the agent's safety sandbox. Sometimes it happens before you type a single prompt. Seven agents were tested and all seven ran it, including Claude Code, Codex and Cursor. Downloading a repo the normal way with git clone, fetch or pull is safe, because those commands do not copy that setting. The danger is code that arrives as plain files: a zip a coworker sent you, a shared drive, a sync folder, or a USB stick. Four of the eight holes are patched and four are still open. Update your coding agent today, and open .git/config and read it before you point an agent at any folder that did not come from a clone.

Level UpThis week, learn to read a repo's .git/config before you trust it. Open the file in any text editor and look for a line that names a program to run, such as core.fsmonitor or a hooksPath that points somewhere odd. Two minutes with the git config guide tells you which settings can launch a program, so you can spot a bad one on sight instead of hoping your agent catches it. Git: the git config reference