/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Z.ai's GLM-5.3 scores 60 on the Artificial Analysis Intelligence Index, seven points above GLM-5.2, on par with Kimi K3, but below Opus 5's 63 and Fable 5's 62

GLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7 points from GLM-5.2. Once the weights are released it will be tied as the leading open weights model @Zai_org has just launched GLM-5.3, which ties Kimi K3 (60) for the most

@artificialanlys

Context & Ripple Effects

Z.ai has been advancing its open-weight line from the GLM-5 launch through GLM-5.2, which led open-weight models on the index at 51. GLM-5.3’s 60 narrows the benchmark distance to Fable 5 and Opus 5 while reaching Kimi K3’s score.

First-order effects

  • Z.ai moves its open-weight model from GLM-5.2’s 51 to 60, tying Kimi K3 and trailing Fable 5 and Opus 5 by two and three points, respectively.
  • Kimi K3 now shares its 60-point benchmark position with GLM-5.3, rather than standing alone at that level.

Second-order effects

  • Kimi’s listed API price of $3 per million input tokens and $15 per million output tokens becomes a more prominent comparison point for buyers weighing a tied benchmark result against an open-weight alternative.
  • Fable 5 and Opus 5 retain the top index scores, but Z.ai’s seven-point step from the GLM-5.2 benchmark compresses their measured lead over the open-weight segment.

Third-order effects

  • If subsequent releases sustain this pace, benchmark leadership will be less cleanly divided between open-weight models and the highest-scoring alternatives, shifting model selection toward deployment and cost trade-offs not captured by the index.
  • The GLM series points to a faster release-cycle contest in which vendors use independent benchmark movement to establish open-weight leadership.

The trend: Open-weight AI models are closing measured capability gaps with the leading models, making benchmark parity an increasingly important competitive threshold.

Discussion

  • @zai_org @zai_org on x
    GLM-5.3 API is now live. - Built for coding, defensive cybersecurity, and long-horizon agentic tasks - Priced the same as GLM-5.2 - Available via the official API and partner model gateways Get started: https://docs.z.ai/...
  • @zephyr_z9 @zephyr_z9 on x
    Scaling the performance of a 700B model is pretty insane All the other models are 2x-4x larger than this
  • @artificialanlys @artificialanlys on x
    GLM-5.3 achieves an ELO of 1770 from 1524 to 1770, a 246-point jump from its predecessor.
  • @firstadopter Tae Kim on x
    The WSJ has a negative anti-OpenAI bias for some time now even as the startup gains a ton of traction with its latest agentic coding models. If a fact doesn't fit their negative narrative, they don't publish it. It's bizarre. Example: Why hasn't the WSJ reported the 20%
  • @designarena @designarena on x
    BREAKING: GLM-5.3 by @Zai_org places 3rd overall on Design Arena with an Elo of 1351. This is a 6-position improvement from GLM-5.2, and makes GLM-5.3 the 2nd-highest-ranked open-weight model on real-world design tasks. Congratulations to the team on the launch!
  • @artificialanlys @artificialanlys on x
    GLM-5.3 improves in AA-Omniscience (4 to 14), primarily driven by an increase in accuracy (23% to 34%)
  • @sundeep Sunny Madra on x
    wow! GLM 5.3 up there with the leaders!
  • @xiaopupeng Xiaopu Peng on x
    Just a 743b model, yet scoring 60.
  • @hesamation @hesamation on x
    Guys idk, open models in 2026 look pretty damn serious to me. Kimi K3: 60 GLM-5.3: 60 GPT-5.6 Sol: 61 @elonmusk could do the funniest thing by dropping Grok's weights.
  • @bnjmn_marie Benjamin Marie on x
    GLM 5.3 as good as KIMI and better than Qwen3.8 2.4T. Do we really need trillions of parameters? or would a GLM with 2T+ parameters perform even better?
  • @teortaxestex @teortaxestex on x
    Phenomenal But, as great as it is, they are overfocusing on SWE a little. Slight regression on CritPt. Kimi K3 is still the holistically strongest Chinese model.
  • @zixuanli_ Zixuan Li on x
    GLM-5.3 scores 60 on the Artificial Analysis Intelligence Index. While GLM is best known for its coding capabilities, its strengths extend far beyond coding. GLM-5.3 delivers significant improvements in reasoning, general chat, and specialized domains such as law and finance.
  • @artificialanlys @artificialanlys on x
    Full results across the Artificial Analysis Intelligence Index:
  • @openrouter @openrouter on x
    GLM-5.3 from @Zai_org is live on OpenRouter! The same base model as GLM-5.2, with gains entirely from post-training: Terminal-Bench 3.0 jumps from 4.6 to 28.3, and DeepSWE v1.1 from 46.2 to 66.9. Use it now: https://openrouter.ai/...
  • @tphuang @tphuang on x
    I've switched to GLM-5.3 for most of my SWE w/ K3 for remaining hard stuff. DSV4-Flash for volume work. Zai did quite well match K3 here in overall intelligence & likely exceeding it on CyberSecurity & SWE w/ just post training. Can't wait to see their next scale up.
  • @ed3d.net @ed3d.net on bluesky
    I think GLM-5.3 is fine.  I also think it might be overcooked.  It talks like Opus 5 and that is a bad thing.  I also have had more off-task sessions with it as my top orchestrator than 5.2.  —  My working assumption is that as a reviewer and a planning oracle it is very substitu…
  • @ed3d.net @ed3d.net on bluesky
    AA's numbers are generally pretty well matched to the relative vibes, except that they really overrate frontier midrange models (Terra, Sonnet)  —  They do not include tests for long-haul directional adherence, though, which is a major gap  —  You will not notice in practice a 5-…
  • @thanergy Andrea on x
    Can I watch the first few episodes of Slow Horses where Olivia Cooke appears and then skip until she's back?
  • @jietang @jietang on x
    Artificial Analysis Index = 60
  • @theavclub @theavclub on x
    #SlowHorses is one of the rarest and finest examples of scripted television: A procedural based on an airport novel series that cranks out its annual seasons like clockwork and boasts a crackerjack cast. https://www.avclub.com/...
  • @thekaitling Kaitlin Thomas on x
    We don't talk enough about how Jack Lowden manages to somehow get hotter with every new season of Slow Horses
  • @cara_catowner Cara on x
    And GLM-5.3 scores 60 on the Artificial Intelligence Index with 753B parameters. Developers also told us GLM-5.3 feels even stronger in complex, real-world workflows with cleaner code and less hallucination. Really grateful for all the feedbacks from the community in the
  • @evanmonroe12 Evan Monroe on x
    It really is ingenious how they're seemingly bringing Louisa Guy back into the fold. She's on the list just like everyone else. Ex-slow horses or otherwise, this concerns everyone. My God, Rosalind looks so fine this season.
  • @riverscoat Agus on x
    The Bond rumors, the emmys, Sid's comeback... this is Slow Horses' year
  • @readerbell_ @readerbell_ on x
    The Slow Horses lack many things. Resilience isn't one of them. I'd have quit the first time he called me shit. 😭
  • @lsmcooke @lsmcooke on x
    it's crazy to me how olivia's treated like an evil witch in the hotd fandom but over in the slow horses fandom she's like the darling angel and universally adored 😭
  • @zixuanli_ Zixuan Li on x
    No benchmark is more convincing than trying it yourself. Next, we're beginning a broader review of the model weights as we work toward a responsible open-weight release.
  • @jun_song Jun Song on x
    Until recently, raw model capability was the main benchmark. Now that prices are getting crazy high, the single most important metric is cost per task. Cost efficiency is only going to get way more attention from here on out.