GLM 5.3 Hosted by Mistral
Chinese and Russian labs shipped capable open weights while DeepSeek signaled even larger training runs.
TL;DR
- Hacker News surfaced a post saying GLM 5.3 is hosted by Mistral.
- r/LocalLLaMA threads pointed to Xiaomi's MiMo-V2.6-Flash-RL on Hugging Face, said DeepSeek is training a 2T model and planning an 8T one, and highlighted Yandex's AliceAI-Foundation-80B as a competitor to Qwen 35B and DeepSeek V4 Flash.
- Two Minute Papers covered DeepSeek's new architecture, and a thread described mini-AGI, a 530M-parameter model trained from scratch on an 8GB-VRAM laptop.
Hacker News surfaced a post saying GLM 5.3 is hosted by Mistral. [1]
Chinese and Russian labs kept shipping open weights: an r/LocalLLaMA thread pointed to Xiaomi's MiMo-V2.6-Flash-RL on Hugging Face; another said DeepSeek is training a 2T model and planning an 8T one; and a third highlighted Yandex's AliceAI-Foundation-80B-A3B-Base as a Russian-developed competitor to Qwen 35B and DeepSeek V4 Flash. [2] [3] [4]
Efficiency was the theme on video too: Two Minute Papers covered DeepSeek's new architecture, and a separate thread described mini-AGI, a 530M-parameter continual-learning model trained from scratch on an 8GB-VRAM laptop from a batch-1 stream. [5] [6]
Why it matters
The open-weight frontier is now being pushed by Chinese and Russian labs shipping capable models that run on consumer hardware, which compresses the gap to closed labs and shifts the moat toward distribution, tooling and trust.
Editor's note
Model capability claims come from community posts and project pages and were not independently benchmarked; the DeepSeek training-size claim is as reported in a community thread.