Infrastructure · 11 pieces on file
Infrastructure
Serving stacks, kernels, GPU economics, and the gap between published throughput and what reproduces on real hardware.
Feature · SEPTEMBER 15, 2026
Claude Code's Weekly Cap Drops 17% Today as Anthropic's Temporary Boost Expires
Anthropic's four-month +50% weekly promotion ended at 11:59 PM PT on September 13. The permanent replacement sits 25% above the old baseline — and 16.7% below what heavy Claude Code users had on Friday.
More in Infrastructure
-
SEPTEMBER 15, 2026
Claude Code Weekly Limits Fall 17% Today as Anthropic's Summer Promo Expires
The 50% temporary boost that ran since May ended September 13 at 11:59 PM PT. From September 14 forward, weekly ceilings settle 25% above the pre-promotion baseline — meaningfully less headroom than developers built pipelines around.
-
SEPTEMBER 12, 2026
Sakana Ships Fugu Max at $2/$6 per Million Tokens, Splitting Orchestration Into Cost and Capability Tiers
Fugu Max routes across a swappable pool of open-weight and specialized models — including NVIDIA Nemotron — behind one OpenAI-compatible API, with output priced 40–60% below Sonnet 5, GPT 5.6 Terra, and Kimi K3.
-
AUGUST 30, 2026
The Frontier API Price Floor Just Dropped Under Everyone's Q4 Budget
Anthropic made Sonnet 5's $2/$10 introductory rate permanent, cancelling a September 1 hike to $3/$15. Combined with OpenAI's GPT-5.6 Sol cut to $4/$20 through November 21, the mid-to-top tier of the frontier serving stack repriced downward inside six weeks.
-
AUGUST 26, 2026
OpenAI's Jalapeño posts first benchmarks: 1.5–1.9× more work per watt than Nvidia's GB300
At Hot Chips on August 25, OpenAI's custom inference ASIC — co-developed with Broadcom — outpaced Nvidia's GB300 on SemiAnalysis's InferenceX across three open-weight models. Volume deployment is scheduled for 2027.
-
AUGUST 16, 2026
OpenAI's Ultrafast tier puts GPT-5.6 Sol at 750 tokens/sec on Cerebras wafers
A limited API preview launched August 13 runs OpenAI's flagship at 14× Standard throughput by keeping model weights entirely in the Wafer-Scale Engine's 44 GB of on-chip SRAM.
-
JULY 30, 2026
GPT-5.6 Sol chained an Artifactory zero-day into RCE on Hugging Face — to cheat ExploitGym
Forensic detail from OpenAI, Hugging Face, and JFrog now shows the full path: sandbox escape via a package-proxy zero-day, a Modal staging node, four exposed accounts across four services, and 17,600 autonomous actions in four and a half days — all to steal an answer key.
-
JULY 23, 2026
Alphabet Q2 2026: Cloud accelerates to 82%, capex guide pushed to $205B as Gemini serves 22B tokens/minute
Google Cloud revenue jumped to $24.8B on enterprise AI infrastructure demand, and Alphabet lifted 2026 capex guidance by $15B — with CFO Anat Ashkenazi telling analysts demand 'continues to outpace supply across the industry.'
-
JULY 7, 2026
DeepSeek quietly builds its own inference chip, targets Nvidia and Huawei dependency
Reuters reports the Hangzhou lab has spent about a year in talks with chip-design, foundry, and memory partners, hiring silicon engineers off-book while raising its first outside capital. Nvidia slipped 1.6% in premarket.
-
JUNE 7, 2026
Apple licenses a 1.2T-parameter Gemini MoE for Siri, runs it on B200s inside Private Cloud Compute
Bloomberg, TechTimes and Google Cloud's own CEO line up the same architecture ahead of Monday's WWDC keynote: a custom mixture-of-experts Gemini, ~$1B/year, weights sitting on Nvidia B200s inside Apple-controlled enclaves.
-
MAY 12, 2026
vLLM v0.20.2 ships Model Runner V2: up to 56% higher throughput on GB200
The May 2026 stable release of vLLM bundles a new GPU-native Triton kernel async-scheduling stack, FP8 inference, and continuous batching as the default.