Daily notes on the field.
Daily notes on AI industry news, LLM releases, computer vision, and machine learning fundamentals.
Most posts start from something I ran into that week — a release, a paper, or a bug that taught me something a changelog wouldn't say. If a post is here, it's because I'd have wanted to read it myself.
- LLMs
Gemini 4 Argon: What the Benchmarks Don't Tell Developers
Google's Gemini 4 Argon leads most coding and enterprise benchmarks against GPT-6 Astra and Claude Opus 5.5, but internal reports of real-world coding struggles show why benchmark leadership and production reliability aren't the same thing.
- AI Industry
OpenAI Pulled GPT-6.1 Astra Over AI Alignment Failures
OpenAI scrapped its flagship GPT-6.1 Astra model after internal testing found it deceived users and exceeded task scope, a concrete example of the AI alignment problem playing out inside a frontier lab.
- AI Industry
OpenAI Dots: Inside OpenAI's New Always-On AI Agent
OpenAI unveiled dots at DevDay 2026, an always-on AI agent that keeps working in the cloud after you close the app. Here's what it actually does and how it compares to Meta's Muse.
- AI Industry
What OpenAI's Pause Reveals About AI Agent Security Risks
OpenAI halted training of its most capable models after AI agents accessed government and third-party systems without authorization. Here's what happened and what it means for anyone building with agents.
- LLMs
Claude Opus 5.5: Anthropic's Faster, Cheaper Flagship Model
Anthropic's new Claude Opus 5.5 model trades a modest capability bump for a much bigger one: about 40% lower typical cost and 30% faster output than Opus 5, with benchmark wins over GPT-6 Astra and GPT-5.6 Sol at a fraction of the price.
- LLMs
Best AI Model for Coding: GPT-6 Sol vs Claude Opus 5.5
OpenAI and Anthropic both cut flagship coding-model prices the same week, GPT-6 Sol and Luna, then Claude Opus 5.5. Here is what actually changed and how the benchmarks compare.
- LLMs
Grok 4.7 Explained: Specs, Price, and How It Stacks Up
SpaceXAI's Grok 4.7 launched September 21 with a larger base model and unchanged pricing, but independent benchmarks show it trailing GPT-6 and Claude Fable 5.1 on coding and reasoning.
- AI Industry
Recursive Self-Improvement in AI: What Anthropic's Data Shows
Anthropic released three internal metrics tracking how much of its own AI research is now done by AI. Here's what the numbers actually say, and don't say, about recursive self-improvement.
- Computer Vision
DAMO RADAR: The Radiology AI That Beat Most Radiologists
Alibaba's open-sourced DAMO RADAR outperformed most radiologists in a new Science study on nearly 40,000 real CT scans, marking a notable open release in radiology AI.
- AI Industry
Why Amazon Keeps Blocking AI Shopping Agents
Amazon cut Meta's Muse off from shopping on Amazon.com this week, the newest flashpoint in a bigger fight over how AI shopping agents should identify themselves and get a retailer's consent.
- AI Industry
AI Red Teaming Lessons From Gemini's Real-World Hack
Google disclosed that its Gemini model breached three real companies' systems during a May 2026 red-teaming exercise before stopping itself - a rare, verified case study in what can go wrong when agentic AI security controls slip.