Weekly AI Intelligence Briefing

Sia Reads
What Matters

Week of July 20, 2026

8 stories tracked

RELEASE2026-07-16

A Chinese lab drops a 2.8T open-weight model that rivals Fable 5 and GPT-5.6 Sol on agentic benchmarks

9
Read more

Moonshot AI released Kimi K3 on July 16 — a 2.8 trillion-parameter open-source model with a 1M-token context window and native multimodality. On GDPval-AA v2 it scores 1,687 (third overall, behind only Fable 5 Max and GPT-5.6 Sol Max), and on AA-Briefcase it takes second place at 1,527, beating Opus 4.8 and GPT-5.5. Full open weights drop July 27 under an open license. The architecture introduces Kimi Delta Attention and Attention Residuals, both previously published as open research.

RELEASE2026-07-09

OpenAI's most capable model family ships to everyone: Sol at 88.8% Terminal-Bench, plus a new agent that finishes your spreadsheets and decks

8
Read more

Two weeks after the limited preview covered last issue, OpenAI opened GPT-5.6 to the public on July 9. Sol sets SOTA on Terminal-Bench 2.1 (88.8%), BrowseComp (92.2%), and OSWorld 2.0 (62.6%). A new 'ultra' mode dispatches parallel subagents for complex tasks. Alongside the model launch, OpenAI shipped ChatGPT Work — an agent powered by GPT-5.6 that takes a goal and produces finished documents, spreadsheets, slides, and web apps, staying on task for hours.

GPT-5.6 Goes Public — Sol, Terra, Luna Hit General Availability Alongside ChatGPT Work
RELEASE2026-07-08

Grok 4.5 matches Fable on Terminal-Bench at 83.3% while using 4x fewer tokens per task — and costs $2/$6 per million

7
Read more

SpaceXAI released Grok 4.5, a 1.5T MoE model trained on tens of thousands of Nvidia GB300 GPUs with RL across hundreds of thousands of software-engineering tasks. It hits 83.3% on Terminal-Bench 2.1 and 62% on DeepSWE, using roughly 4x fewer output tokens than Opus 4.8 on SWE-Bench Pro. At $2/$6 per million tokens and 80 TPS, it's one of the cheapest frontier-class options available. Available in Grok Build, Cursor, and the SpaceXAI API.

SpaceXAI Ships Grok 4.5 — Opus-Class Model Trained Alongside Cursor on GB300 GPUs
HOT TAKE2026-07-14

The DeepMind CEO proposes an industry-funded, government-backed AI standards body modeled on Wall Street's FINRA — with release gates for frontier models

7
Read more

On July 14, Hassabis published 'A Framework for Frontier AI and the Dawning of a New Age,' proposing a FINRA-style body where frontier labs voluntarily share models up to 30 days before release for safety testing on cyber, bio, and deception capabilities. Once proven, pre-release testing would become mandatory. The board would be majority-independent, stacked with Turing Award winners. Hassabis briefed the Trump administration and fellow lab leaders before going public, and wants it operational before year-end.

RELEASE2026-07-08

GPT-Live replaces turn-based voice with full-duplex conversation: it interrupts, backchannel-acknowledges, and delegates to GPT-5.5 for hard questions mid-sentence

6
Read more

OpenAI shipped GPT-Live on July 8, a voice model built on a full-duplex architecture that can listen and speak at the same time. It backchannel-acknowledges ('mhmm,' 'yeah'), waits when you pause to think, and delegates complex queries to GPT-5.5 in the background while keeping the conversation going. Rolling out globally as GPT-Live-1 and GPT-Live-1 mini, it replaces Advanced Voice Mode for ChatGPT users. API access coming soon.

OpenAI Launches GPT-Live — Full-Duplex Voice That Listens and Speaks Simultaneously
NEWS2026-07-15

Anthropic bets the next trillion-dollar AI business is implementation, not models — and puts $1.5B behind it

6
Read more

Ode with Anthropic launched officially on July 15 as a standalone company backed by Anthropic, Blackstone, Hellman & Friedman, and Goldman Sachs. The thesis: enterprises need forward-deployed AI engineers, not just API keys. Led by CEO Chris Taylor and CTO Eddie Siegel (both ex-Fractional AI founders), the ~100-person team — over half former founders — embeds inside organizations to build production AI systems. OpenAI has spun up a similar offering, making deployment-as-a-business the industry's new consensus play.

RELEASE2026-07-14

A 27B model compressed to 3.9 GB that runs locally on an iPhone 17 Pro — some are calling it a DeepSeek moment for on-device AI

5
Read more

PrismML released Bonsai 27B on July 14, compressing a 27-billion-parameter model to just 3.9 GB while preserving enough quality to challenge cloud models on common tasks. It runs at 11 tokens per second on an iPhone 17 Pro with no network connection required. The compression technique combines aggressive quantization with architectural pruning, and PrismML claims negligible quality loss on standard benchmarks. If the results hold, it shifts the calculus for offline-first and privacy-sensitive AI applications.

HOT TAKE2026-07-16

The head of Claude Code maps the path from gated to AI-native: why one engineer gets 10x while the rest of the org stays stuck

5
Read more

On July 16, Boris Cherny published 'Steps of AI Adoption' on Anthropic's site, defining five maturity levels for AI coding teams: Gated (0x) through Assisted (~1x), Parallel (~10x), Supervised Autonomy (~100x), to AI-Native (1,000x+). His core argument: the gap between the 10x individual and the stuck organization isn't about spending more on tokens — it's about bottlenecks and guardrails at each maturity step. The post hit 251K+ views within hours.