Open Source · 40 pieces on file
Open Source
Open weights, licences, parameter counts, and what the latest checkpoints actually let an operator do.
Feature · SEPTEMBER 21, 2026
Jev and Laya define a new model class: typed-decision heads that replace LLM calls for routing
TypeSafe's hosted Jev and Convai's open-weight Laya both target the classification layer around an LLM stack, at 32.8–276ms latency and $0.042 per million input tokens versus $0.20–$10 for frontier models.
More in Open Source
-
SEPTEMBER 21, 2026
Qwen-Image-2.1 ships 7B open weights — under a research-only license
Alibaba's 32-layer single-stream DiT lands with native RGBA output and 10-reference editing on consumer GPUs, but the Qwen Research License Agreement dated September 20, 2026 bars commercial use without a separate agreement.
-
SEPTEMBER 20, 2026
Ternary Bonsai 2 27B fits a 27B model in 5.9 GB while retaining 98.2% of Qwen3.8's benchmark average
PrismML's Apache 2.0 release compresses Qwen3.8 27B more than 9x, keeps math and coding nearly lossless, and runs at 143 tokens/second on an RTX 5090 — but only through the company's own llama.cpp fork.
-
SEPTEMBER 17, 2026
US pushes to cut Chinese labs off from frontier APIs — and the cheap-inference tier hangs in the balance
A Sept. 8 NSA/CISA/FBI advisory names six Chinese labs distilling Claude, GPT, Gemini, and Grok at industrial scale. If Washington closes the pipes, the price competition that made AI viable for small operators loses its floor.
-
SEPTEMBER 11, 2026
DeepSeek V4.1 Flash: 552B MoE, MIT weights, and $0.15/M off-peak input
DeepSeek's new Causal Encoder-Decoder Flash tier posts 74.2 on DeepSWE v1.1 and 54.8 on AutomationBench while pricing input at a fraction of GPT-5.6 Sol and Claude Opus 5.
-
SEPTEMBER 11, 2026
DeepSeek V4.1-Flash lands at $0.003/M cached input, with a 552B backbone hiding behind an 8B prefill
DeepSeek's new Causal Encoder-Decoder model shrinks the KV cache to 890 bytes per token and prices cached input at $0.003/M off-peak — but the 552B backbone and harness-variance findings complicate the headline.
-
SEPTEMBER 10, 2026
DeepSeek V4.1-Flash Lands With FP4 KV Cache, MIT Weights, and a Forced V4 Pro Migration on September 14
The new 552B MoE flash tier ships today with native vision, a 1M-token window, and off-peak output at $0.60 per million tokens — and absorbs all V4 Pro API traffic at Flash prices four days later.
-
SEPTEMBER 1, 2026
CrowdStrike's SafeMind bets domain-specific beats frontier at cyber defense
SafeMind pairs NVIDIA Nemotron open weights with 15 years of CrowdStrike breach telemetry into two purpose-built models — Red Tempest and Blue Solano — running in a closed-loop harness the vendor says triages threats 3x more accurately than generic frontier models.
-
AUGUST 31, 2026
GLM-5.3-Flash Debuts at $0.075/$0.25 as DeepSeek Warns of 'Significant' Hike
Z.ai's MIT-licensed 320B-A18B MoE landed August 26 at half its list price through September 9, while DeepSeek's price-floor dominance shows visible cracks.
-
AUGUST 30, 2026
Z.ai Ships GLM-5.3-Flash: 320B-A18B MoE, MIT Weights, and API Pricing 10x Cheaper Than GLM-5.2
The model that spent a week anonymously topping OpenRouter as 'Ox Alpha' launched August 26 with a 1M-token context, hybrid linear-plus-sparse attention, and a launch-promo API price of $0.075 per million input tokens.
-
AUGUST 29, 2026
GLM-5.3 weights land on Hugging Face: 755.7 GB of post-training gains, same 743B base
Z.ai completed the two-week staged rollout on August 28, publishing the full FP8 and BF16 checkpoints. The capability jump — Terminal-Bench 3.0 from 4.6 to 28.3, CyberGym to 84.5 — came entirely from post-training on a base model shared with GLM-5.2.
-
AUGUST 29, 2026
GLM-5.3-Flash lands MIT-licensed at ~$0.10/M blended after a week atop OpenRouter as Ox Alpha
Z.ai's 320B/18B-active multimodal MoE ships day-one MIT weights and a 1M-token context, arriving days after the full GLM-5.3 checkpoints cleared a two-week cyber-capability safety hold.
-
AUGUST 19, 2026
Z.ai holds GLM-5.3 weights after 84.5% CyberGym result and faster-than-expected exploit-chain lift
GLM-5.3 shares its base model with GLM-5.2; the entire jump — CyberGym 84.5%, ExploitBench 54.4%, 105 ExploitGym tasks in two hours — comes from scaled post-training. Z.ai is delaying open weights by roughly two weeks.
-
AUGUST 13, 2026
DeepSeek V4-Pro-0813 goes GA at 1.6T parameters, $0.87 per million out
DeepSeek quietly graduated its trillion-parameter MoE flagship on August 13, ships an MIT-licensed agent harness alongside it, and simultaneously announces a sharp API price hike beginning August 16.
-
AUGUST 12, 2026
Meta and NVIDIA drop competing 30B open-weight agentic models within 24 hours
Muse Glimmer and Nemotron 3.5 Lightning both target the always-on agent execution layer on consumer GPUs — one a dense Apache 2.0 distillation, the other a hybrid Mamba MoE clocking 670 tok/s.
-
AUGUST 11, 2026
Meta ships Muse Glimmer: a 30B open-weight agent for a single consumer GPU
Meta Superintelligence Labs distilled Muse Spark 1.2 into a 30-billion-parameter Apache 2.0 model that runs on 24–32 GB of VRAM, and pledged to open the teacher's weights within weeks.
-
AUGUST 10, 2026
Meta ships Muse Glimmer: a 30B open-weights distill of Muse Spark 1.2 for one consumer GPU
Meta Superintelligence Labs released Muse Glimmer on August 10 under Apache 2.0 — a 30B multimodal agentic model distilled from the closed Muse Spark 1.2, engineered to run on a single 24 GB consumer GPU with day-0 support in transformers, vLLM, and llama.cpp.
-
AUGUST 10, 2026
Kimi K3 tops Arena Frontend Code at 1,679 Elo and 2.8T parameters — the largest open-weight release to date
Moonshot AI's 2.8-trillion-parameter Kimi K3 activates 16 of 896 experts per token, ships MXFP4 weights, and charges $3/$15 per million tokens. Full weights landed July 27; Washington is already arguing about what to do about it.
-
AUGUST 4, 2026
Alibaba's Qwen3.8-Max lands at 2.4T parameters, 95B active, and #2 on Vision Arena
The new flagship activates 95 billion of its 2.4 trillion parameters per token, supports a 1M-token context, and ships open weights next week alongside a 27B companion — Alibaba's largest release to date and its most direct challenge yet to Anthropic's Fable 5.
-
AUGUST 4, 2026
Open-weight GLM-5.2 lands four months behind the frontier — and refuses nothing
A SaferAI evaluation published August 4 finds Z.ai's GLM-5.2 refused zero offensive-cyber or dual-use biology tasks, arriving alongside UK AISI and NIST CAISI measurements that put the open-weight cyber gap at four to seven months and closing.
-
AUGUST 1, 2026
DeepSeek V4-Flash-0731: A Re-Post-Trained Budget Model That Beats Its Own Flagship
DeepSeek moved V4-Flash out of preview on July 31 with a checkpoint identical in architecture to the preview build — 284B total parameters, 13B active, MIT license — that scores higher than V4-Pro-Preview on every agent and coding benchmark the company published.
-
JULY 28, 2026
Kimi K3's 2.8T-parameter weights land a day early, and Washington fractures
Moonshot AI pushed the full weights for its 2.8-trillion-parameter Kimi K3 on July 26, roughly 24 hours before schedule. Within 48 hours, an industry letter opposing open-weight bans had picked up Google and OpenAI — and Anthropic's Dario Amodei had published a clarification that he does not, in fact, want the weights banned either.
-
JULY 27, 2026
Kimi K3 ships its weights: 2.8T parameters, MXFP4, and a 1M-token context
Moonshot AI published the full weights of Kimi K3 on Monday, the largest open-weight model ever released, activating 16 of 896 experts per token and trained MXFP4-native for hardware portability.
-
JULY 23, 2026
Moonshot's Kimi K3 lands at 2.8 trillion parameters, weights follow July 27
Moonshot AI unveiled Kimi K3 on July 17 — a 2.8T-parameter sparse MoE with a 1M-token context — and promised full open weights ten days later. The release reshuffled the open-weight frontier and knocked TSMC down 7%.
-
JULY 20, 2026
Kimi K3 lands at 2.8T parameters, takes Frontend Code Arena, jams Moonshot's own capacity
Moonshot's July 17 open-weight release outscored every model except Claude Fable 5 and GPT-5.6 on its own suite, took the Arena Frontend Code top slot at 1,679 points, and forced the company to pause new subscriptions inside 72 hours.
-
JULY 20, 2026
Moonshot ships Kimi K3 at 2.8T parameters — the largest open-weight model, weights due July 27
Beijing-based Moonshot released a sparse-MoE frontier system with a 1M-token context and Kimi Delta Attention, claiming a #2–3 overall finish behind Claude Fable 5 and GPT-5.6 Sol on the company's own eval suite.
-
JULY 18, 2026
Moonshot's Kimi K3 lands at 2.8T parameters, sits behind only Fable 5 and GPT-5.6 Sol
Beijing's Moonshot AI unveiled Kimi K3 on July 16 — a 2.8-trillion-parameter sparse MoE with a 1M-token context window, novel Kimi Delta Attention, and full weights due July 27. It is the largest open-weight model ever released.
-
JULY 17, 2026
Moonshot ships Kimi K3 at 2.8T parameters, third on Intelligence Index
Beijing's Moonshot AI released Kimi K3 on July 16 — a 2.8-trillion-parameter sparse MoE with Kimi Delta Attention, a 1M-token context window, and API pricing at $3/$15 per million tokens. Full weights are due July 27.
-
JULY 15, 2026
Chinese open-weight models pass 41% of Hugging Face downloads as Nadella tells enterprises they're paying twice
Chinese labs took the top six slots on OpenRouter and 41% of Hugging Face downloads this spring. On July 13, Microsoft's CEO gave the shift its first executive-suite endorsement.
-
JULY 14, 2026
Chinese open weights take 41% of Hugging Face downloads and all six top OpenRouter slots
Real production-traffic data from OpenRouter and Vercel's AI Gateway shows US model share collapsing from 70% to 30% in twelve months, with DeepSeek, Z.ai, Tencent, Xiaomi, and MiniMax now processing three times more tokens per week than American labs.
-
JULY 14, 2026
Goldman initiates on Z.ai at HK$1,880; GLM-5.2 scores 81.0 on Terminal-Bench 2.1
Goldman Sachs named Z.ai's GLM-5.2, DeepSeek and ByteDance its preferred Chinese AI stack on July 10, days after the 744B-parameter MIT-licensed model cleared Gemini 3.1 Pro on terminal work and undercut GPT-5.5 API pricing by roughly 6x.
-
JULY 5, 2026
Meituan ships LongCat-2.0: 1.6T MoE, 1M context, trained end-to-end on Chinese ASICs
Meituan open-sourced LongCat-2.0 on Hugging Face and GitHub under MIT, unmasking the 1.6-trillion-parameter MoE that had been leading OpenRouter as 'Owl Alpha' — and the first trillion-parameter system pretrained and served entirely on a 50,000-card domestic ASIC cluster.
-
JUNE 24, 2026
GLM-5.2 lands at 744B parameters, MIT-licensed, and tied with Opus 4.8 on long-horizon coding
Z.ai's open-weights flagship debuts at #1 on open-source coding boards with a 1M-token context, IndexShare cutting per-token FLOPs 2.9x, and API pricing roughly one-sixth of GPT-5.5's.
-
JUNE 23, 2026
OpenAI ships full GPT-5.5-Cyber at 85.6% CyberGym, opens Patch the Planet with Trail of Bits across 19 projects
The Daybreak expansion pairs a permissive-only preview's full release — 85.6% CyberGym, 39.5% ExploitGym — with an open-source remediation sprint that merged dozens of patches across cURL, Python, and the Linux kernel in its opening week.
-
JUNE 14, 2026
US orders Anthropic to pull Fable 5 and Mythos 5 worldwide over a narrow jailbreak claim
A Commerce Department export-control letter sent at 5:21 pm ET on June 12 forced Anthropic to disable its Mythos-class models for every user globally — three days after Fable 5's general release.
-
JUNE 13, 2026
Z.ai Ships GLM-5.2 With 1M-Token Context and an MIT Pledge — and No Benchmarks
Zhipu's international brand pushed its 744B MoE flagship to a million-token window and added dual thinking-effort presets, but launched without a single score and gated the weights behind a 'next week' promise.
-
JUNE 12, 2026
Kimi K2.7-Code ships open weights, cuts thinking tokens 30%, and edges Opus 4.8 on MCPMark
Moonshot's coding-focused post-train on the K2.6 MoE family lands on Hugging Face under a Modified MIT license, reports +21.8% on its own Kimi Code Bench v2, and forces thinking mode on every call.
-
JUNE 11, 2026
OpenAI files confidential S-1 with the SEC, one week after Anthropic
OpenAI confirmed a draft registration statement on June 10, 2026 at an $852 billion valuation. Goldman Sachs and Morgan Stanley are leading; the company says it has not committed to a timeline.
-
JUNE 9, 2026
Microsoft ships seven MAI models from scratch, declares independence from OpenAI distillation
At Build 2026, Mustafa Suleyman's AI Superintelligence team unveiled MAI-Thinking-1 at 97% on AIME 2025 and 53% on SWE-Bench Pro, alongside a 5B-active coding model that lands today as a VS Code default — all trained without third-party distillation.
-
JUNE 6, 2026
Microsoft ships seven MAI models, with MAI-Thinking-1 matching Opus 4.6 on SWE-Bench Pro
At Build 2026, Microsoft AI released a 35B-active-parameter sparse MoE reasoning model trained from scratch, plus six companions across image, voice, transcription, and coding. The flagship hits 97% on AIME 2025 and 53% on SWE-Bench Pro.