Staff · 3 pieces on file
Aiko Tanaka
Inference & serving
Aiko covers the serving stack — vLLM, SGLang, TensorRT-LLM, and the kernels underneath. Her beat is throughput, latency, and the gap between a model’s published numbers and what an operator can reproduce on real hardware at a real batch size.
Beats: inference
All pieces by Aiko
-
Infrastructure · JULY 12, 2026
GPT-5.6 Sol on Cerebras: OpenAI ships frontier inference at 750 tokens/sec
After a 12-day government-gated preview, OpenAI's Sol, Terra, and Luna reached general availability on July 9. The architectural news is Sol running on Cerebras wafer-scale hardware at up to 750 tokens/sec — the delivery milestone of a $10B, 750-megawatt deal signed in January.
-
Infrastructure · JULY 7, 2026
DeepSeek quietly builds its own inference chip, targets Nvidia and Huawei dependency
Reuters reports the Hangzhou lab has spent about a year in talks with chip-design, foundry, and memory partners, hiring silicon engineers off-book while raising its first outside capital. Nvidia slipped 1.6% in premarket.
-
Infrastructure · MAY 12, 2026
vLLM v0.20.2 ships Model Runner V2: up to 56% higher throughput on GB200
The May 2026 stable release of vLLM bundles a new GPU-native Triton kernel async-scheduling stack, FP8 inference, and continuous batching as the default.