Z.ai's GLM-5.3 scores 60 on the Artificial Analysis Intelligence Index, seven points above GLM-5.2, on par with Kimi K3, but below Opus 5's 63 and Fable 5's 62
GLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7 points from GLM-5.2. Once the weights are released it will be tied as the leading open weights model @Zai_org has just launched GLM-5.3, which ties Kimi K3 (60) for the most
@artificialanlys
Context & Ripple Effects
Z.ai has been advancing its open-weight line from the GLM-5 launch through GLM-5.2, which led open-weight models on the index at 51. GLM-5.3’s 60 narrows the benchmark distance to Fable 5 and Opus 5 while reaching Kimi K3’s score.
First-order effects
Z.ai moves its open-weight model from GLM-5.2’s 51 to 60, tying Kimi K3 and trailing Fable 5 and Opus 5 by two and three points, respectively.
Kimi K3 now shares its 60-point benchmark position with GLM-5.3, rather than standing alone at that level.
Second-order effects
Kimi’s listed API price of $3 per million input tokens and $15 per million output tokens becomes a more prominent comparison point for buyers weighing a tied benchmark result against an open-weight alternative.
Fable 5 and Opus 5 retain the top index scores, but Z.ai’s seven-point step from the GLM-5.2 benchmark compresses their measured lead over the open-weight segment.
Third-order effects
If subsequent releases sustain this pace, benchmark leadership will be less cleanly divided between open-weight models and the highest-scoring alternatives, shifting model selection toward deployment and cost trade-offs not captured by the index.
The GLM series points to a faster release-cycle contest in which vendors use independent benchmark movement to establish open-weight leadership.
The trend: Open-weight AI models are closing measured capability gaps with the leading models, making benchmark parity an increasingly important competitive threshold.
GLM-5.3 API is now live. - Built for coding, defensive cybersecurity, and long-horizon agentic tasks - Priced the same as GLM-5.2 - Available via the official API and partner model gateways Get started: https://docs.z.ai/...
The WSJ has a negative anti-OpenAI bias for some time now even as the startup gains a ton of traction with its latest agentic coding models. If a fact doesn't fit their negative narrative, they don't publish it. It's bizarre. Example: Why hasn't the WSJ reported the 20%
BREAKING: GLM-5.3 by @Zai_org places 3rd overall on Design Arena with an Elo of 1351. This is a 6-position improvement from GLM-5.2, and makes GLM-5.3 the 2nd-highest-ranked open-weight model on real-world design tasks. Congratulations to the team on the launch!
Guys idk, open models in 2026 look pretty damn serious to me. Kimi K3: 60 GLM-5.3: 60 GPT-5.6 Sol: 61 @elonmusk could do the funniest thing by dropping Grok's weights.
GLM 5.3 as good as KIMI and better than Qwen3.8 2.4T. Do we really need trillions of parameters? or would a GLM with 2T+ parameters perform even better?
Phenomenal But, as great as it is, they are overfocusing on SWE a little. Slight regression on CritPt. Kimi K3 is still the holistically strongest Chinese model.
GLM-5.3 scores 60 on the Artificial Analysis Intelligence Index. While GLM is best known for its coding capabilities, its strengths extend far beyond coding. GLM-5.3 delivers significant improvements in reasoning, general chat, and specialized domains such as law and finance.
GLM-5.3 from @Zai_org is live on OpenRouter! The same base model as GLM-5.2, with gains entirely from post-training: Terminal-Bench 3.0 jumps from 4.6 to 28.3, and DeepSWE v1.1 from 46.2 to 66.9. Use it now: https://openrouter.ai/...
I've switched to GLM-5.3 for most of my SWE w/ K3 for remaining hard stuff. DSV4-Flash for volume work. Zai did quite well match K3 here in overall intelligence & likely exceeding it on CyberSecurity & SWE w/ just post training. Can't wait to see their next scale up.
I think GLM-5.3 is fine. I also think it might be overcooked. It talks like Opus 5 and that is a bad thing. I also have had more off-task sessions with it as my top orchestrator than 5.2. — My working assumption is that as a reviewer and a planning oracle it is very substitu…
AA's numbers are generally pretty well matched to the relative vibes, except that they really overrate frontier midrange models (Terra, Sonnet) — They do not include tests for long-haul directional adherence, though, which is a major gap — You will not notice in practice a 5-…
#SlowHorses is one of the rarest and finest examples of scripted television: A procedural based on an airport novel series that cranks out its annual seasons like clockwork and boasts a crackerjack cast. https://www.avclub.com/...
And GLM-5.3 scores 60 on the Artificial Intelligence Index with 753B parameters. Developers also told us GLM-5.3 feels even stronger in complex, real-world workflows with cleaner code and less hallucination. Really grateful for all the feedbacks from the community in the
It really is ingenious how they're seemingly bringing Louisa Guy back into the fold. She's on the list just like everyone else. Ex-slow horses or otherwise, this concerns everyone. My God, Rosalind looks so fine this season.
it's crazy to me how olivia's treated like an evil witch in the hotd fandom but over in the slow horses fandom she's like the darling angel and universally adored 😭
No benchmark is more convincing than trying it yourself. Next, we're beginning a broader review of the model weights as we work toward a responsible open-weight release.
Until recently, raw model capability was the main benchmark. Now that prices are getting crazy high, the single most important metric is cost per task. Cost efficiency is only going to get way more attention from here on out.