GLM-5.2 is the leading open weights model on Artificial Analysis' Intelligence Index, scoring 51, only behind Fable 5's 60, Opus 4.8's 56, and GPT-5.5's 55
Z ai's GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index scoring 51 and it sits on the Pareto frontier of Intelligence vs Cost per Task
Artificial Analysis
Context & Ripple Effects
Z.ai’s GLM line has progressed from the GLM-5 launch, which the company positioned around reasoning, coding and agentic work, through GLM-5.1 and now GLM-5.2. The latest release adds a 1M context window and an MIT license, while the independent Artificial Analysis result places it at the top of the open-weights cohort.
The benchmark result matters because GLM-5.2 is not only the highest-scoring open-weights model in this index; Artificial Analysis also places it on the intelligence-versus-cost-per-task Pareto frontier. It remains behind the named proprietary leaders, preserving a measurable performance gap while narrowing the open-model comparison.
First-order effects
GLM-5.2 gains independent validation as the leading open-weights model in the Artificial Analysis index, strengthening Z.ai’s performance positioning beyond its own launch claims.
Developers and organizations evaluating deployable open-weight models have a newly benchmarked option that pairs a score of 51 with a favorable intelligence-cost trade-off.
Second-order effects
Other open-weight model providers face a clearer competitive bar on both general intelligence and cost efficiency, rather than on benchmark claims alone.
The result gives buyers a more concrete basis to compare open-weight deployments with higher-scoring proprietary models, especially for coding and longer-horizon workloads highlighted in Z.ai’s release.
Third-order effects
If repeated across independent evaluations, the open-weight market could compete increasingly on the efficiency frontier—usable capability per task cost—rather than treating openness and frontier performance as separate categories.
The remaining gap to Fable 5, Opus 4.8 and GPT-5.5 suggests a bifurcated market may persist: proprietary models at the highest measured capability, with open-weight models becoming more credible for workloads where control, licensing and cost matter.
The trend: Open-weight foundation models are moving from an alternative deployment option toward directly benchmarked competitors on the capability-and-cost frontier.
BREAKING: GLM-5.2 is now 1st on Design Arena. With an Elo of 1360, GLM-5.2 has jumped ahead of the now unavailable Claude Fable 5. And it's open weights. This is an improvement of 4 positions and 27 Elo points to achieve one of the highest Elo scores in our code categories [image…
Just caught up with the recent GLM-5.2 release. The best open-weight model today. Architecture-wise, it's build on the GLM-5 and GLM-5.1 architecture that I covered previously, which means it's reusing the Multi-head Latent Attention (MLA) and DeepSeek Sparse Attention (DSA) [ima…
Wow. @Zai_org GLM 5.2 is a marvel! It is *at least* as good as Opus 4.8 and GPT 5.5. It's super fast, inexpensive, and not too verbose. It responds with nuance and judgement, & handles long context VERY well. I've never experienced an open weights model like this before.
@jietang @teortaxesTex On benchmarks, yes, but as measured by true usefulness even Q1 would be very impressive. Anthropic has rightly focused on maximizing useful intelligence, which does not show up in benchmarks, but definitely shows up in revenue.
A couple of days ago, I claimed Chinese models lag US models by ~7 months. Just spent time with GLM 5.2 (new open model from Zhipu). It has a 1-million token context, excels at long coding projects and agent tasks, and beats GPT-5.5 on several tough benchmarks while matching
For the 100th time: the best Chinese model is open source, and this has been the case for almost every week since R1 1.5 years ago I wonder how doomers who coped that “Alibaba is going closed, the trend is clear” feel now
A standout number in Z ai's GLM-5.2 launch is CritPt, a benchmark of unpublished research-level physics problems where it ties with Claude Opus 4.8 and is well above other open weights models Key takeaways: ➤ @Zai_org 's GLM-5.2 (max reasoning effort) leads open weights by a [ima…
I think GLM 5.2 points to a 7 months gap currently It's around Opus 4.7-4.8 level, all told (modulo vision which in Opus's case is garbage anyway). Mythos reached Preview status (≥ Opus 4.8, functionally) by early Feb 2026. This means full PRC Mythos ("Fable") by Nov-Dec'26.
Exciting news: GLM-5.2 (Max) ranks #2 in Code Arena: Frontend, with +29pt over Claude Opus 4.7 (Thinking) and only behind Fable 5! GLM-5.2 is the best open model vs Kimi-K2.6 and Minimax-M3 by a large margin. - #2 React and #4 HTML sub-leaderboards - Ranks as the top model in [im…
Z ai's GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index scoring 51 and it sits on the Pareto frontier of Intelligence vs Cost per Task @Zai_org's GLM-5.2 is the same size as GLM-5.1 (744B total / 40B active parameters) but scores 11 [ima…
For a long time I've been saying that the gap between open source and closed models is going to widen because of the data gap, hardware gap, and increased restrictions on distillation. I was wrong. https://z.ai/ is on another level. Incredible benchmarks on this
I was wrong I've been saying for months that open source AI models are 6 months behind frontier They caught up. GLM 5.2 is as good as Opus 4.8 This changes everything. If you run GLM 5.2 locally no government can take it away. You become sovereign And even if you run through A…
Open source LLMs are catching up on subsidized, compute hungry flagship models. The gap is closing. From 2-3 years behind to now almost on-par, and soon leading. …