Ox Alpha, a “stealth model” from an unknown AI lab with a 1M-token multimodal context and capacity for 100T tokens/day, goes viral after launching on OpenRouter
When an unknown AI lab drops an anonymous model - Ox Alpha - for free, while declaring that they have the capacity …
Context & Ripple Effects
OpenRouter has quietly become the venue where labs test models anonymously before putting a name on them: in February, sources said Zhipu shipped its unreleased GLM-5 under the alias Pony Alpha, and in March a mystery 1T-parameter model called Hunter Alpha appeared on the marketplace and triggered speculation it was DeepSeek testing its V4. Ox Alpha is the third such drop, but the most aggressive yet — free to use, with a claimed 1M-token multimodal context and enough capacity for 100T tokens per day.
The timing matters because OpenRouter is no longer a side channel: its own analysis of 100T+ tokens showed reasoning models already dominate usage there, and the platform was reportedly running at roughly $140M annualized revenue after nearly tripling month-over-month earlier this year. A viral free model landing on that distribution rail is both a marketing stunt and a stress test of who pays for inference.
First-order effects
- Developers get unrestricted free access to a frontier-spec multimodal model through OpenRouter, and the unknown lab absorbs whatever the 100T-token/day capacity claim costs — making Ox Alpha an instant benchmark target for anyone comparing context length and price.
- The anonymity immediately fuels attribution guessing games, repeating the pattern where Zhipu anonymously released GLM-5 as Pony Alpha before its official debut.
Second-order effects
- Priced frontier models are the direct counterweight: OpenAI charged $150/$600 per million tokens for o1-pro's extra compute, so a free 1M-context rival forces every paid API to justify its premium on reliability and support rather than raw specs.
- Rival labs now have a proven playbook — seed an anonymous model on OpenRouter, harvest virality and real-world eval data, then reveal — which pressures named-model launches to compete with their own shadow releases.
Third-order effects
- If stealth drops keep outperforming announced launches for attention, OpenRouter's role shifts from neutral marketplace to the de facto proving ground where model identity is decided by usage telemetry rather than press releases — concentrating gatekeeping power in one intermediary.
- Sustained free capacity at this scale would push the market toward 'operationally free AI,' where inference is subsidized as customer acquisition and monetization migrates upstream to whoever controls the weights or the routing layer.
The trend: Frontier labs are increasingly launching models anonymously on OpenRouter first, turning the marketplace into the industry's stealth-testing ground and shifting competition toward subsidized free capacity.
Related: Operationally free AI · Frontier capacity allocation · OpenRouter · Ox Alpha · Hunter Alpha mystery model sparks DeepSeek speculation · Zhipu anonymously released GLM-5 as Pony Alpha
Related Coverage
- Ox Alpha — Ox Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. OpenRouter
- A mysterious free AI model is impressing developers. And nobody knows who made it. Business Insider · Lakshmi Varanasi
- Mystery AI Model Ox Alpha Draws Developers With Free Access Bloomberg · Sohee Kim
- An Anonymous AI Called ‘Ox Alpha’ Just Dropped on OpenRouter—And the Internet Thought It Was Gemini NPowerUser · Nayan
- Mystery AI Ox Alpha Surfaces on OpenRouter, Fueling Speculation Over Frontier Chinese Tech Techstrong.ai · Jon Swartz
- Coding Model Ox Alpha Retains Every Prompt: You Cannot Name Company Holding Them Tech Times · Kyle Belmonte
- Ox Alpha Matches Zhipu's GLM Tokenizer in 95 of 95 Tests Implicator.ai · Marcus Schuler
- Who's behind the new ‘stealth model’ Ox Alpha? TechCrunch · Anthony Ha
- OpenRouter gives anonymous Ox Alpha a 1M-token launchpad RuntimeWire
Discussion
-
@openrouter
@openrouter
on x
🥷 New stealth model: Ox Alpha Ox Alpha is a frontier model built for efficient coding, sustained agentic work, and real-world production use. - 1M token context window - Text, image, and video input Try it now and share feedback to improve the model! https://openrouter.ai/...
-
@opencode
@opencode
on x
Ox Alpha (stealth model) is free for the next week - 1M Context - Multi-modal - Zero Data Retention Generous rate limits, near unlimited usage We have capacity for 100T tokens per day, lets see what you can do
-
@patrickc
Patrick Collison
on x
$ ori —model stealth/ox-alpha It's very impressive.
-
@haider1
Haider
on x
i'm certain the mysterious model “ox-alpha” is not a GLM release they dropped GLM-5.3, so another big model this soon makes no sense, especially when Zhipu can't reliably serve these models at scale what i don't get is why google suddenly can't stop posting about this thing
-
@andrewcurran_
Andrew Curran
on x
The identity of Ox Alpha remains a mystery. It's a long-horizon multimodal model with a 1M context window. Last night the top candidate was a new GLM release. This morning people seem less sure of anything. I'm making a thread for guesses and test results.
-
@pilvar222
@pilvar222
on x
OK guys I'm sorry to disappoint but Ox Alpha is overhyped We ran it on our cybersecurity benchmark, it's better than GPT-5.6-Luna, but worse than every other frontier models
-
@jeremiahdillon
Jeremiah Dillon
on x
ox-alpha just used the word “dialectical” three times in the last 10 minutes. over all of the tokens I've burned on frontier models I can't recall seeing this word used once. based on this and nothing else 😅 this feels like a Chinese model, probably DeepSeek, Qwen, or Kimi in
-
@moleh1ll
Moll
on x
I don't know why so many people decided that Ox Alpha is Gemini. In my language, its discourse-pragmatic coherence is noticeably weaker: the model sometimes mixes up grammatical gender, temporal relations, and the current state of the situation. I don't observe this with Gemini.
-
@kimmonismus
@kimmonismus
on x
Quick reminder: zAI founder Jietang wrote back in June that they expect a GLM model at the “Fable” level to arrive before the end of the year. They are very confident, and the leap from GLM-5.2 to 5.3 demonstrated their ability to make significant strides,same base model,
-
@nilaykothari
Nilay Kothari
on x
Ox alpha has become talk of the town suddenly. And few Google employees are vague posting since morning which has caused rumours linked with ox alpha. A friend gave it a shot and he called me excitedly to share how awesome Ox alpha is. Hoping for someone to own it soon.
-
@ziwenxu_
Ziwen
on x
Ox Alpha is giving away 1M context, multimodal, 100T tokens/day of capacity. For free. That's millions on millions in compute. Not even DeepSeek does this. The shortlist of who can actually eat that cost: Google. Alibaba. Tencent. Xiaomi. Meta. Nobody's confirmed anything.
-
@selfatonements
@selfatonements
on x
A free stealth model showed up on OpenRouter with no company on the box. Ox Alpha. One million tokens. Text, image, video in. Built for coding and long agent runs. Free for about a week. Nobody claimed it. Tokenizer fingerprints keep lining up with GLM 5.3 from
-
@skalskip92
@skalskip92
on x
Ox Alpha is NOT the new Gemini Gemini models have always been exceptionally good at computer vision Ox Alpha is mediocre at best 1/n aerial and satellite images
-
@tokengremlin
@tokengremlin
on x
Isn't it a little humiliating that Google has to use Ox Alpha as marketing just to get people engaged with Gemini? Like, if Ox Alpha isn't Gemini (which we're almost certain it isn't, it looks more like GLM or something along those lines), Google would still be stuck with the
-
@haider1
Haider
on x
more “ox-alpha” vagueposting, now from a DeepMind's research scientist but apparently even at max effort it's still pretty average at coding — realizes what it did wrong, apologizes, explains the mistake, then somehow does the exact same thing again very old gemini vibes
-
@tim_dettmers
Tim Dettmers
on x
One dead giveaway is also that Zhipu is one of the only labs that serves models with pretty poor partial prefill tok/s while output tok/s are fast. This model is ~30% faster than GLM 5.3. If the infras is the same, likely ~300-500B or fewer params vs activated params.
-
@swarooprm7
Swaroop Mishra
on x
Anyone tried Ox Alpha?
-
@haider1
Haider
on x
google employees have been vagueposting nonstop could be the delayed gemini 3.5 pro? if it's better and nearly as fast at coding as 3.7 flash, it was probably worth the wait, cuz 3.7 flash is already really good also, ox alpha secretly being gemini would explain quite a bit
-
@be_arsh
Arsh
on x
Ox Alpha in Hermes is actually insane, way better than running it in OpenCode. OpenCode kept randomly stopping mid task, but in Hermes it hasn't happened once. It's super independent and takes initiative without you having to ask. That can be hit or miss with agents, but here
-
@tanmaigo
Tanmai Gopal
on x
The Ox Alpha hype is insane. 1. Most people think it's GLM fam because the tokenizers footprint. 2. Another set of people thinking it's Gemini because it does video. 3. Bro just entered the chat saying it's Microsoft because of the tokenizer. Who's placing bets?
-
@alexatallah
Alex Atallah
on x
Ox Alpha, a new frontier model is live on OpenRouter! This is our first stealth launch in a while, and we expect it to be a big one. Expect SOTA performance and decent throughput. Make sure to share feedback! Use it now: https://openrouter.ai/...
-
@vaibhavsisinty
Vaibhav Sisinty
on x
A mystery AI model appeared. Free. 1M context. Multimodal. Outperforming GPT-5.6 Sol and Fable 5 Max on early tests. Nobody knows who built it. 🤯 It is called Ox Alpha. Here is what we know. Tokenizer matches GLM-5.3 exactly. Same video encoder. Same patterns. A developer ran
-
@gaelbreton
Gael Breton
on x
I'll throw my hot take in the hat. I think Ox Alpha is Composer 3. Based on GLM (the tokenizer clues), served by xAI (100T tokens/day)
-
@haider1
Haider
on x
woah, apparently this stealth model “Ox-Alpha” just mogged Fable 5 and GPT-5.6 Sol on DeepSwe i haven't tried it on large-scale architecture work yet, but as an orchestrator, executor, and reviewer, it's been more than capable definitely a step above 5.6 luna my guess is MiMo
-
@babayagatwt
Baba Yaga
on x
Nobody knows who built Ox Alpha. But this mysterious model has: - 1M context + multimodal capabilities - Zero data retention - Free, nearly unlimited usage for a week - OpenCode claims capacity for 100T tokens/day - Reportedly beats Fable, GPT and Opus on some tests So what
-
@elshayib_
@elshayib_
on x
How are they offering 100T tokens on Ox Alpha, where the fuck is that compute coming from. The only way for it to make sense is if it's the next Grok model, Elon is the only one who's got that much compute.
-
@vamsibatchuk
Vamsi Batchu
on x
Google is back. Trust the process.
-
@leo_linsky
Leo Linsky
on x
We don't usually test unlisted endpoints but the hype around Ox Alpha was like the second coming of Fable Himself I ran it through a subset of our coding evals, which gives a strong indication of base fluid intelligence (but not necessarily usability in a harness). If this was a
-
@jrysana
John
on x
Lots of talk about “Ox Alpha” on here, so I figure I should contribute some results I found. On a (fairly tough) private benchmark (i.e. with zero chance of contamination in any model), with minimal/low reasoning, it underperforms quite a lot, even with a reasoning advantage:
-
@brandonjcarl
Brandon Carl
on x
Whoever is offering Ox Alpha had the ability to offer the entire token capacity of Google immediately. 100 trillion tokens a day.
-
@teortaxestex
@teortaxestex
on x
Ox Alpha is frontier. wtf. Seriously wtf. Tencent? Xiaomi? Really? It's more well-done than any other Chinese model. I swear, it's better than Kimi K3. It's good *in ways that Chinese models are typically horrible at*. It's FAST I can half-believe it's a stolen Claude checkpoint.
-
@bindureddy
Bindu Reddy
on x
Ox-Alpha Underperforms And Is Worse Than Last Generation Models Given all the hype we decided to evaluate Ox-Alpha and it turned out to be quite bad Its scores alongside Kimi 2.6 which is 2 generations old However, it's pure marketing genius. Everyone is talking about it
-
@heyitsmilac
Mila
on x
am I the only who thinks Ox Alpha offering zero data retention and near unlimited usage for free is a classic capability test disguised as generosity. whoever's behind this needs real world stress testing data before a paid launch and giving devs a free week is the fastest way
-
@karanc_12
@karanc_12
on x
Ox Alpha just hit ~80%+ on a DeepSWE subset. For comparison: • gpt-5.6-sol → 52% • Fable → 65% • Ox Alpha → 80%+ People are starting to get very confused.
-
@kimmonismus
@kimmonismus
on x
No way! @synthwavedd says that the secret Model “Ox Alpha” on OpenRouter is the upcoming GLM 5.3 Flash A Flash model (!!) outperforming even the best models like Fable and GPT Sol?! If this turns out to be true, zAI's post training put western models at shame Oh, and looks
-
@mr_salio
@mr_salio
on x
🚨 Ox Alpha Benchmarks Are RIDICULOUS This new Stealth Model Actually Beats Fable 5 and GPT-5.6 Sol and It's Free > 80% on DeepSWE > Above the scores of Fable and GPT-5.6 Sol > 1M-token context window > Multimodal capabilities > Zero data retention > Free with near-unlimited
-
@hakmgpt
@hakmgpt
on x
so it's confirmed ox alpha = the new Gemini model I'm so freaking happy with google come back , and I was believing they will come back , but not in this way I feel like nerfing all the frontier models is honestly a huge blow no one expected this move , but me
-
@unclecode
@unclecode
on x
Everyone guessed who made Ox Alpha today. I fingerprinted it instead. One page, 9 infrastructure probes: tokenizers, error codes, hidden templates. I ran • the mystery model against 12 suspects. One family matches every tokenizer test: GLM. Plumbing does not lie. Tool in reply,
-
@niklasdev
Niklas
on x
OX Alpha is hosted in America!! 🇺🇸 I sent 49 identical one-token requests from 7 Cloudflare regions, rotated the order to control for load and matched every response to OpenRouter's per-generation server timing. This isolated the geographic network floor: ATL: 57 ms DFW: 76 ms
-
@timjayas
Tim Jayas
on x
Ox Alpha vs Opus 5 i asked both models to build rocket engine using three.js ox-alpha is not even close to opus 5 performance and it took equal or more time for the output prompt used for both: “Create a 3d SpaceX raptor engine using three.js, use white theme background so the
-
@ziwenxu_
Ziwen
on x
Ox Alpha is wired into codex now, and it's free for the whole week!! no api key, no billing. It sits in the picker next to the models we already pay for, and our native models never moved. Vision and a million tokens of context, costing nothing until the window shuts. so we
-
@vidhisharmx
Vidhi
on x
wtf: Gemini team started a new war whole team is suddenly hyping Gemini out of nowhere something big is coming i guess, is the pretraining of Gemini 4 finished? i think Gemini team is really confident with their internal results for Gemini 4, it feels right because last
-
@opencode
@opencode
on x
Ox Alpha is now available on OpenCode Go too For the next 6 days, usage is near unlimited and completely free It won't count against your Go usage
-
@kimmonismus
@kimmonismus
on x
A mysterious new AI model just appeared. Ox Alpha offers a 1M context window, multimodal capabilities, zero data retention, and nearly unlimited usage for an entire week. OpenCode says it has capacity for 100 trillion tokens per day. That's 1.16b tokens per second. Where the
-
@omarsar0
Elvis
on x
Got confirmation on what the Ox Alpha model is. Can't disclose anything yet, but I can confirm it's a spectacular model. All I can say is that we all need to adjust our timelines. Achieving frontier multimodal agentic capabilities is officially the new race.
-
@evanotero
Evan Otero
on x
What if the Ox Alpha was the friends we made along the way
-
@altryne
Alex Volkov
on x
The mysterious model Ox alpha is now offered for free through: - @OpenRouter - @opencode (with 100T capacity per day) - @NousResearch Portal with... **checks notes*... **rubs eyes and checks notes again** 1 quadrillion tokens... per day. Given that.. all of OpenRouter
-
@scaling01
@scaling01
on x
Ox Alpha one-shotted the GPU accelerated fluid simulation in a single 1000 line html file it looks absolutely stunning and is 1000x better than what Qwen3.8 27B or Opus 4.5 did when I tested them earlier this week
-
@timkellogg.me
Mr. Tim
on bluesky
Ox Alpha, a new stealth model on openrouter. People are saying it's stronger than Sol & Fable. No benchmarks yet. — Some are hypothesizing it's Chinese. Everyone seems convinced that it's not OpenAI or Anthropic — openrouter.ai/stealth/ox-a...
-
r/Futurology
r
on reddit
A mysterious free AI model is impressing developers. And nobody knows who made it.
-
r/InterstellarKinetics
r
on reddit
EXCLUSIVE: A Mysterious AI Model Named “Ox Alpha” Stuns Developers With Free Near-Unlimited Access And 100 Trillion Daily Token Capacity, But Its Creator Still Remains Unknown 🤖
-
@kimmonismus
@kimmonismus
on x
Imagine Ox Alpha running on a DGX spark. One can dream, right
-
@teortaxestex
@teortaxestex
on x
Ding-ding-ding! Best one so far
-
@lisley494341191
Gaab
on x
@teortaxesTex So it's GLM-Google colab like what mistral did..
-
@teortaxestex
@teortaxestex
on x
Few will understand...
-
r/singularity
r
on reddit
Ox Alpha can't be the Chinese.