Ilya Sutskever · an impressionistic sketch of brain and thought · MIT Technology Review
The first model from SSI, founded by former OpenAI chief scientist Ilya Sutskever, has leaked: a small reasoning engine based on TTT (Test-Time Training), modest in scale but able to "update its own weights while solving problems". It may appear as early as August; Nvidia has invested $5B.
Per the leak, the model learns "how to learn" from specially curated data, then updates parts of its weights in real time during inference. Even at small scale it can trade blows with much larger models. SSI has shipped no model and no product for over two years; this is the first sword it has kept hidden so long.
While everyone else stacks parameters, Ilya is betting that "a model can keep learning after deployment". That may be the first signal of the pre-training era approaching its end.
xAI released Grok Bot (early beta): an AI teammate that "logs into your tools, uses them like you would, and comes back with the work done". A single tweet drew 22.9M views. The new Grok 4.6 model shipped alongside.
Grok 4.6 is a confirmed 1.5T-parameter model scoring 61 on Artificial Analysis' Intelligence Index, tying GPT-5.6 Sol Max; it took 88.4% on Terminal-Bench v2.1, priced at $2/$6 per million tokens. swyx called it "the efficiency king".
When a "tying" model sells at a "discounted" price and ships with a bot that works for you, the next battlefield of AI competition is: who becomes your teammate first.
Grok Bot goes live.
A top lab is in talks to acquire Israeli AI startup Decart AI for about $6B; if closed, it would be the lab's largest acquisition ever. Decart specializes in world models and inference optimization.
As "building it yourself" grows more expensive, top labs have started filling gaps through acquisition. AI competition is moving from "who trains better" to "who buys better".
$6B acquisition of Decart.
A latest Reuters corporate survey finds a "strong majority" of Japanese companies have yet to fully adopt AI, underscoring the social and organizational inertia in technology diffusion.
While America races for models and China races for open source, Japanese companies' AI adoption remains in "wait and see". Not everyone is running; most companies are still standing at the door.
Alibaba released Qwen3.8-Max open weights: 2.4T total / 95B active MoE, priced at $2/$6 per million tokens via API. Community verdict: "one of the largest open-weight releases ever".
Evaluations are strong: Frontend Code Arena #4 (Elo 1668, behind only Claude Opus 5 and Kimi K3), Vision Arena #2, SWE-bench 87.3% ahead of GPT-5.5. vLLM day-0 support plus dedicated 4-bit quantizations for NVIDIA B300 and AMD MI355X; Together and Baseten could run it immediately. One caveat: the initial release is text-only; vision input is not yet unlocked.
When a 2.4T open model scores near closed flagships at a fraction of the price, the "scale gap" between open and closed becomes a "deployment gap".
Qwen3.8-Max weights released.
DeepSeek V4 Pro formally reached general availability: $0.435/M input and $0.87/M output, roughly 57 times cheaper than Fable 5 by Cline's math, with Terminal Bench up 15.8% versus the preview.
The story here is not "strongest" but "cheapest": community evaluations find the capability solid though not comprehensively leading Kimi or Flash. The real signal is price. When inference costs fall to this level, "affording the strongest" becomes "affording the strong-enough".
DeepSeek's next deciding factor may not be scale but RL environments and agent engineering, plus the price that makes everyone look twice.
DeepSeek V4 Pro GA.
Mustafa Suleyman announced MAI-Thinking-1, Microsoft's first reasoning model "built from scratch", now live on Foundry.
The team's posture is interesting: Finbarr Timbers publicly asked developers for feedback on tool use, signaling Microsoft positions it as an "applied reasoning model", not a leaderboard chaser. Now that Microsoft finally has an in-house reasoning model, the OpenAI-Microsoft relationship formally moves from "exclusive supplier" to "frenemies".
Microsoft ships its first in-house reasoning model.
Zed released Delta, a multiplayer coding environment for AI agents, letting multiple agents and humans work in parallel on the same codebase.
When "one person plus one agent" is no longer enough, collaboration environments start being designed for "a swarm of agents plus people". The next battlefield for coding tools is collaboration between agents.
Zed Delta: a multiplayer coding environment for AI agents.
Image: Zed official.
The Claude in Chrome sidebar upgraded into Claude Cowork sessions: you work in the page while it collaborates alongside.
With an AI assistant resident in the browser sidebar, no separate app to open, the barrier to use drops another notch. AI's entrance is turning from "an app" into "part of the browser".
Claude Cowork.
AutoGPT demonstrated how to manage AI-generated pull requests with AGENTS.md and skill gating, making AI pass a "should this change happen" check before touching code.
Once AI files its own PRs, engineering management needs new gates. AGENTS.md has turned from documentation into AI's code of conduct.
AutoGPT with AGENTS.md.
Community discussion surrounds DeepSeek-V4-Pro-0813 stopping early on long-horizon tasks: is it wrapping up early to save tokens, or genuinely done?
When a model learns to "save power", it may also learn to "loaf". Long-horizon agent reliability is becoming a more important metric than "smartness".
DeepSeek-V4-Pro-0813 long-horizon behavior.
Amazon will use Twitch streamers' content to train generative AI by default unless users opt out; Ars Technica reported on the default-on "data collection".
Streams, clips and voice become part of the training set automatically; opting out requires clicking it yourself. When platforms take your content to feed AI by default, the word "default" is the biggest trap of all.
Anthropic signed a 20-year, $9.1B compute agreement with Riot Platforms, locking 191 MW of Texas facility capacity to power Claude.
$6B of acquisitions to fill model gaps on one side, $9.1B of locked compute on the other: competition among top labs has become capital lock-up measured in decades. The AI arms race has shifted from "who is stronger" to "who locks up resources first".
Anthropic locks in 20 years of compute.
"Unless you understand probability, you will not understand why large models occasionally answer wrong."
— Lord Tarassenko · Machine learning researcher · Financial Times
"Wrap up early to save tokens, or genuinely finish the job? Long-horizon agent reliability is becoming a metric that matters more than smartness."
— Community discussion · DeepSeek-V4-Pro long-horizon tasks · SemiAnalysis
Anthropic paid $6B for Decart to fill its world-model gap and locked $9.1B for twenty years of compute.
Grok 4.6 tied GPT-5.6 Sol; long-horizon agents became the new battlefield.
Most Japanese companies have yet to embrace AI; AutoGPT manages AI's code with AGENTS.md.
The leaders grab resources, the edges watch and wait, safety gets weaponized: in AI's 2026, every player is looking for its position.
The Cascade volcanoes, Pacific Northwest, USA — NASA's satellite captured a range of volcanoes shrouded in smoke, a line of snowcapped peaks standing on the horizon.
Cascade Volcanoes · Pacific Northwest · photographed by NASA · 44.0°N, 122.0°W
From Rainier to Shasta, a line of volcanoes stands along the ridge, some snowcapped, some smoking. They live at geology's pace and care nothing for leaderboards. While we argue on screens about whose model is stronger and who is locking up compute, these volcanoes decide on a million-year scale when they will wake. After a day of AI news, Earth still has places that need no GPU cooling. Mountains and seas — stay curious, keep exploring.