/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Ox Alpha, a “stealth model” from an unknown AI lab with a 1M-token multimodal context and capacity for 100T tokens/day, goes viral after launching on OpenRouter

When an unknown AI lab drops an anonymous model - Ox Alpha - for free, while declaring that they have the capacity …

Wccftech Rohail Saleem

Context & Ripple Effects

OpenRouter has quietly become the venue where labs test models anonymously before putting a name on them: in February, sources said Zhipu shipped its unreleased GLM-5 under the alias Pony Alpha, and in March a mystery 1T-parameter model called Hunter Alpha appeared on the marketplace and triggered speculation it was DeepSeek testing its V4. Ox Alpha is the third such drop, but the most aggressive yet — free to use, with a claimed 1M-token multimodal context and enough capacity for 100T tokens per day.

The timing matters because OpenRouter is no longer a side channel: its own analysis of 100T+ tokens showed reasoning models already dominate usage there, and the platform was reportedly running at roughly $140M annualized revenue after nearly tripling month-over-month earlier this year. A viral free model landing on that distribution rail is both a marketing stunt and a stress test of who pays for inference.

First-order effects

  • Developers get unrestricted free access to a frontier-spec multimodal model through OpenRouter, and the unknown lab absorbs whatever the 100T-token/day capacity claim costs — making Ox Alpha an instant benchmark target for anyone comparing context length and price.
  • The anonymity immediately fuels attribution guessing games, repeating the pattern where Zhipu anonymously released GLM-5 as Pony Alpha before its official debut.

Second-order effects

  • Priced frontier models are the direct counterweight: OpenAI charged $150/$600 per million tokens for o1-pro's extra compute, so a free 1M-context rival forces every paid API to justify its premium on reliability and support rather than raw specs.
  • Rival labs now have a proven playbook — seed an anonymous model on OpenRouter, harvest virality and real-world eval data, then reveal — which pressures named-model launches to compete with their own shadow releases.

Third-order effects

  • If stealth drops keep outperforming announced launches for attention, OpenRouter's role shifts from neutral marketplace to the de facto proving ground where model identity is decided by usage telemetry rather than press releases — concentrating gatekeeping power in one intermediary.
  • Sustained free capacity at this scale would push the market toward 'operationally free AI,' where inference is subsidized as customer acquisition and monetization migrates upstream to whoever controls the weights or the routing layer.

The trend: Frontier labs are increasingly launching models anonymously on OpenRouter first, turning the marketplace into the industry's stealth-testing ground and shifting competition toward subsidized free capacity.

Discussion

  • @openrouter @openrouter on x
    🥷 New stealth model: Ox Alpha Ox Alpha is a frontier model built for efficient coding, sustained agentic work, and real-world production use. - 1M token context window - Text, image, and video input Try it now and share feedback to improve the model! https://openrouter.ai/...
  • @opencode @opencode on x
    Ox Alpha (stealth model) is free for the next week - 1M Context - Multi-modal - Zero Data Retention Generous rate limits, near unlimited usage We have capacity for 100T tokens per day, lets see what you can do
  • @patrickc Patrick Collison on x
    $ ori —model stealth/ox-alpha It's very impressive.
  • @haider1 Haider on x
    i'm certain the mysterious model “ox-alpha” is not a GLM release they dropped GLM-5.3, so another big model this soon makes no sense, especially when Zhipu can't reliably serve these models at scale what i don't get is why google suddenly can't stop posting about this thing
  • @andrewcurran_ Andrew Curran on x
    The identity of Ox Alpha remains a mystery. It's a long-horizon multimodal model with a 1M context window. Last night the top candidate was a new GLM release. This morning people seem less sure of anything. I'm making a thread for guesses and test results.
  • @pilvar222 @pilvar222 on x
    OK guys I'm sorry to disappoint but Ox Alpha is overhyped We ran it on our cybersecurity benchmark, it's better than GPT-5.6-Luna, but worse than every other frontier models
  • @jeremiahdillon Jeremiah Dillon on x
    ox-alpha just used the word “dialectical” three times in the last 10 minutes. over all of the tokens I've burned on frontier models I can't recall seeing this word used once. based on this and nothing else 😅 this feels like a Chinese model, probably DeepSeek, Qwen, or Kimi in
  • @moleh1ll Moll on x
    I don't know why so many people decided that Ox Alpha is Gemini. In my language, its discourse-pragmatic coherence is noticeably weaker: the model sometimes mixes up grammatical gender, temporal relations, and the current state of the situation. I don't observe this with Gemini.
  • @kimmonismus @kimmonismus on x
    Quick reminder: zAI founder Jietang wrote back in June that they expect a GLM model at the “Fable” level to arrive before the end of the year. They are very confident, and the leap from GLM-5.2 to 5.3 demonstrated their ability to make significant strides,same base model,
  • @nilaykothari Nilay Kothari on x
    Ox alpha has become talk of the town suddenly. And few Google employees are vague posting since morning which has caused rumours linked with ox alpha. A friend gave it a shot and he called me excitedly to share how awesome Ox alpha is. Hoping for someone to own it soon.
  • @ziwenxu_ Ziwen on x
    Ox Alpha is giving away 1M context, multimodal, 100T tokens/day of capacity. For free. That's millions on millions in compute. Not even DeepSeek does this. The shortlist of who can actually eat that cost: Google. Alibaba. Tencent. Xiaomi. Meta. Nobody's confirmed anything.
  • @selfatonements @selfatonements on x
    A free stealth model showed up on OpenRouter with no company on the box. Ox Alpha. One million tokens. Text, image, video in. Built for coding and long agent runs. Free for about a week. Nobody claimed it. Tokenizer fingerprints keep lining up with GLM 5.3 from
  • @skalskip92 @skalskip92 on x
    Ox Alpha is NOT the new Gemini Gemini models have always been exceptionally good at computer vision Ox Alpha is mediocre at best 1/n aerial and satellite images
  • @tokengremlin @tokengremlin on x
    Isn't it a little humiliating that Google has to use Ox Alpha as marketing just to get people engaged with Gemini? Like, if Ox Alpha isn't Gemini (which we're almost certain it isn't, it looks more like GLM or something along those lines), Google would still be stuck with the
  • @haider1 Haider on x
    more “ox-alpha” vagueposting, now from a DeepMind's research scientist but apparently even at max effort it's still pretty average at coding — realizes what it did wrong, apologizes, explains the mistake, then somehow does the exact same thing again very old gemini vibes
  • @tim_dettmers Tim Dettmers on x
    One dead giveaway is also that Zhipu is one of the only labs that serves models with pretty poor partial prefill tok/s while output tok/s are fast. This model is ~30% faster than GLM 5.3. If the infras is the same, likely ~300-500B or fewer params vs activated params.
  • @swarooprm7 Swaroop Mishra on x
    Anyone tried Ox Alpha?
  • @haider1 Haider on x
    google employees have been vagueposting nonstop could be the delayed gemini 3.5 pro? if it's better and nearly as fast at coding as 3.7 flash, it was probably worth the wait, cuz 3.7 flash is already really good also, ox alpha secretly being gemini would explain quite a bit
  • @be_arsh Arsh on x
    Ox Alpha in Hermes is actually insane, way better than running it in OpenCode. OpenCode kept randomly stopping mid task, but in Hermes it hasn't happened once. It's super independent and takes initiative without you having to ask. That can be hit or miss with agents, but here
  • @tanmaigo Tanmai Gopal on x
    The Ox Alpha hype is insane. 1. Most people think it's GLM fam because the tokenizers footprint. 2. Another set of people thinking it's Gemini because it does video. 3. Bro just entered the chat saying it's Microsoft because of the tokenizer. Who's placing bets?
  • @alexatallah Alex Atallah on x
    Ox Alpha, a new frontier model is live on OpenRouter! This is our first stealth launch in a while, and we expect it to be a big one. Expect SOTA performance and decent throughput. Make sure to share feedback! Use it now: https://openrouter.ai/...
  • @vaibhavsisinty Vaibhav Sisinty on x
    A mystery AI model appeared. Free. 1M context. Multimodal. Outperforming GPT-5.6 Sol and Fable 5 Max on early tests. Nobody knows who built it. 🤯 It is called Ox Alpha. Here is what we know. Tokenizer matches GLM-5.3 exactly. Same video encoder. Same patterns. A developer ran
  • @gaelbreton Gael Breton on x
    I'll throw my hot take in the hat. I think Ox Alpha is Composer 3. Based on GLM (the tokenizer clues), served by xAI (100T tokens/day)
  • @haider1 Haider on x
    woah, apparently this stealth model “Ox-Alpha” just mogged Fable 5 and GPT-5.6 Sol on DeepSwe i haven't tried it on large-scale architecture work yet, but as an orchestrator, executor, and reviewer, it's been more than capable definitely a step above 5.6 luna my guess is MiMo
  • @babayagatwt Baba Yaga on x
    Nobody knows who built Ox Alpha. But this mysterious model has: - 1M context + multimodal capabilities - Zero data retention - Free, nearly unlimited usage for a week - OpenCode claims capacity for 100T tokens/day - Reportedly beats Fable, GPT and Opus on some tests So what
  • @elshayib_ @elshayib_ on x
    How are they offering 100T tokens on Ox Alpha, where the fuck is that compute coming from. The only way for it to make sense is if it's the next Grok model, Elon is the only one who's got that much compute.
  • @vamsibatchuk Vamsi Batchu on x
    Google is back. Trust the process.
  • @leo_linsky Leo Linsky on x
    We don't usually test unlisted endpoints but the hype around Ox Alpha was like the second coming of Fable Himself I ran it through a subset of our coding evals, which gives a strong indication of base fluid intelligence (but not necessarily usability in a harness). If this was a
  • @jrysana John on x
    Lots of talk about “Ox Alpha” on here, so I figure I should contribute some results I found. On a (fairly tough) private benchmark (i.e. with zero chance of contamination in any model), with minimal/low reasoning, it underperforms quite a lot, even with a reasoning advantage:
  • @brandonjcarl Brandon Carl on x
    Whoever is offering Ox Alpha had the ability to offer the entire token capacity of Google immediately. 100 trillion tokens a day.
  • @teortaxestex @teortaxestex on x
    Ox Alpha is frontier. wtf. Seriously wtf. Tencent? Xiaomi? Really? It's more well-done than any other Chinese model. I swear, it's better than Kimi K3. It's good *in ways that Chinese models are typically horrible at*. It's FAST I can half-believe it's a stolen Claude checkpoint.
  • @bindureddy Bindu Reddy on x
    Ox-Alpha Underperforms And Is Worse Than Last Generation Models Given all the hype we decided to evaluate Ox-Alpha and it turned out to be quite bad Its scores alongside Kimi 2.6 which is 2 generations old However, it's pure marketing genius. Everyone is talking about it
  • @heyitsmilac Mila on x
    am I the only who thinks Ox Alpha offering zero data retention and near unlimited usage for free is a classic capability test disguised as generosity. whoever's behind this needs real world stress testing data before a paid launch and giving devs a free week is the fastest way
  • @karanc_12 @karanc_12 on x
    Ox Alpha just hit ~80%+ on a DeepSWE subset. For comparison: • gpt-5.6-sol → 52% • Fable → 65% • Ox Alpha → 80%+ People are starting to get very confused.
  • @kimmonismus @kimmonismus on x
    No way! @synthwavedd says that the secret Model “Ox Alpha” on OpenRouter is the upcoming GLM 5.3 Flash A Flash model (!!) outperforming even the best models like Fable and GPT Sol?! If this turns out to be true, zAI's post training put western models at shame Oh, and looks
  • @mr_salio @mr_salio on x
    🚨 Ox Alpha Benchmarks Are RIDICULOUS This new Stealth Model Actually Beats Fable 5 and GPT-5.6 Sol and It's Free > 80% on DeepSWE > Above the scores of Fable and GPT-5.6 Sol > 1M-token context window > Multimodal capabilities > Zero data retention > Free with near-unlimited
  • @hakmgpt @hakmgpt on x
    so it's confirmed ox alpha = the new Gemini model I'm so freaking happy with google come back , and I was believing they will come back , but not in this way I feel like nerfing all the frontier models is honestly a huge blow no one expected this move , but me
  • @unclecode @unclecode on x
    Everyone guessed who made Ox Alpha today. I fingerprinted it instead. One page, 9 infrastructure probes: tokenizers, error codes, hidden templates. I ran • the mystery model against 12 suspects. One family matches every tokenizer test: GLM. Plumbing does not lie. Tool in reply,
  • @niklasdev Niklas on x
    OX Alpha is hosted in America!! 🇺🇸 I sent 49 identical one-token requests from 7 Cloudflare regions, rotated the order to control for load and matched every response to OpenRouter's per-generation server timing. This isolated the geographic network floor: ATL: 57 ms DFW: 76 ms
  • @timjayas Tim Jayas on x
    Ox Alpha vs Opus 5 i asked both models to build rocket engine using three.js ox-alpha is not even close to opus 5 performance and it took equal or more time for the output prompt used for both: “Create a 3d SpaceX raptor engine using three.js, use white theme background so the
  • @ziwenxu_ Ziwen on x
    Ox Alpha is wired into codex now, and it's free for the whole week!! no api key, no billing. It sits in the picker next to the models we already pay for, and our native models never moved. Vision and a million tokens of context, costing nothing until the window shuts. so we
  • @vidhisharmx Vidhi on x
    wtf: Gemini team started a new war whole team is suddenly hyping Gemini out of nowhere something big is coming i guess, is the pretraining of Gemini 4 finished? i think Gemini team is really confident with their internal results for Gemini 4, it feels right because last
  • @opencode @opencode on x
    Ox Alpha is now available on OpenCode Go too For the next 6 days, usage is near unlimited and completely free It won't count against your Go usage
  • @kimmonismus @kimmonismus on x
    A mysterious new AI model just appeared. Ox Alpha offers a 1M context window, multimodal capabilities, zero data retention, and nearly unlimited usage for an entire week. OpenCode says it has capacity for 100 trillion tokens per day. That's 1.16b tokens per second. Where the
  • @omarsar0 Elvis on x
    Got confirmation on what the Ox Alpha model is. Can't disclose anything yet, but I can confirm it's a spectacular model. All I can say is that we all need to adjust our timelines. Achieving frontier multimodal agentic capabilities is officially the new race.
  • @evanotero Evan Otero on x
    What if the Ox Alpha was the friends we made along the way
  • @altryne Alex Volkov on x
    The mysterious model Ox alpha is now offered for free through: - @OpenRouter - @opencode (with 100T capacity per day) - @NousResearch Portal with... **checks notes*... **rubs eyes and checks notes again** 1 quadrillion tokens... per day. Given that.. all of OpenRouter
  • @scaling01 @scaling01 on x
    Ox Alpha one-shotted the GPU accelerated fluid simulation in a single 1000 line html file it looks absolutely stunning and is 1000x better than what Qwen3.8 27B or Opus 4.5 did when I tested them earlier this week
  • @timkellogg.me Mr. Tim on bluesky
    Ox Alpha, a new stealth model on openrouter.  People are saying it's stronger than Sol & Fable.  No benchmarks yet.  —  Some are hypothesizing it's Chinese.  Everyone seems convinced that it's not OpenAI or Anthropic  —  openrouter.ai/stealth/ox-a...
  • r/Futurology r on reddit
    A mysterious free AI model is impressing developers.  And nobody knows who made it.
  • r/InterstellarKinetics r on reddit
    EXCLUSIVE: A Mysterious AI Model Named “Ox Alpha” Stuns Developers With Free Near-Unlimited Access And 100 Trillion Daily Token Capacity, But Its Creator Still Remains Unknown 🤖
  • @kimmonismus @kimmonismus on x
    Imagine Ox Alpha running on a DGX spark. One can dream, right
  • @teortaxestex @teortaxestex on x
    Ding-ding-ding! Best one so far
  • @lisley494341191 Gaab on x
    @teortaxesTex So it's GLM-Google colab like what mistral did..
  • @teortaxestex @teortaxestex on x
    Few will understand...
  • r/singularity r on reddit
    Ox Alpha can't be the Chinese.