Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog:
our gpu performance engineering resource list has gained a lot of traffic (now 1.5k stars on github)
so we're excited to share AI Performance Engineering v2.
link in thread 🧵
it now starts with how a single inference request works, then builds through the cuda execution
We just released a massive update on our gpu performance engineering resource list
AI Performance Engineering v2
this is the most comprehensive resource list for learning gpu and ai perf engineering
link in thread 🧵
this version starts with how a single inference request