Mira Murati's TML launches a research blog called Connectionism, and shares its work on resolving nondeterminism and achieving reproducible results from LLMs
There's been great interest in what Mira Murati's Thinking Machines Lab is building with its $2 billion in seed funding …
Context & Ripple Effects
Thinking Machines Lab was introduced as Mira Murati’s new AI venture before securing a $2B seed round, giving the lab unusual resources to pursue foundational model research. Connectionism is an early public signal of the technical problems it intends to foreground.
The publication arrives ahead of TML’s first product, the Tinker fine-tuning API, linking research on repeatable model behavior to a potential developer-facing need rather than a purely academic exercise.
First-order effects
- TML makes its work on LLM nondeterminism and reproducibility visible to researchers and prospective users, creating a public technical record for how it approaches repeatable outputs.
- Developers evaluating TML’s eventual tools gain a clearer indication that reproducible model behavior is a design priority, not only raw model capability.
Second-order effects
- Competing model labs and fine-tuning platforms face greater pressure to document whether results can be replicated across runs, especially for customers that need dependable evaluation and iteration.
- Public research can turn reproducibility methods into a point of comparison for downstream tooling, including testing, experiment tracking, and managed model deployment.
Third-order effects
- If reproducibility becomes a purchasing criterion, frontier-model competition may shift from benchmark performance alone toward operational consistency and auditable development workflows.
- The pattern supports a more mature AI software stack in which research labs must translate model behavior into controls that developers can validate and repeatedly use.
The trend: Frontier AI labs are increasingly treating reliable, repeatable model behavior as a product and research differentiator alongside scale and capability.
Related: Thinking Machines Lab · Connectionism · Managed Execution Layer · Frontier-lab capital concentration · TML’s $2B seed round · TML launches Tinker
Related Coverage
- Defeating Nondeterminism in LLM Inference Thinking Machines Lab
- Fascinating new work from the start-up of Mila Murati, ex CTO of OpenAI. — The opening sentence sets the stage “Reproducibility is a bedrock of scientific progress. … Giulio Zuanetti
- 💡This AI breakthrough deserves more attention because it solves a critical issue in LLM inference. Thanks to Eduardo Ordax for sharing this here! … Michaël Karpe
- 🚨 Did you know that even with the same model and seed and temperature set to 0, LLMs can give different answers? … Shounak K.
- I finally had a full day blocked off to focus on upleveling and catch up what's new in the space. That gave me time to look at the latest paper … Christina Lin
Discussion
-
@lilianweng
Lilian Weng
on x
Besides the fun fact that Connectionism is connected with the early days of the AI field and highlights similarities between neural networks and human brains, the flagship product of the (first) Thinking Machines is named Connection Machine. — 🧑🎓Enjoy reading and more is coming!
-
@fenbielding
Ben Fielding
on x
strongly agree with the need for determinism in model execution outlined in @thinkymachines first blog post we take this further at @gensynai and build for reproducibility (determinism across devices) more to announce soon but in the meantime, a demo: https://github.com/...
-
@tensorbert
Bert Maher
on x
Read this to the end — the last section is mind-blowing
-
@timlantin
@timlantin
on x
life imitates art (westworld) [image]
-
@yashwanthsai29
Sai Yashwanth
on x
Learnt a lot reading this study. The quality is exceptional! I aspire to contribute at the same level in the coming future.
-
@suhackerr
Suha
on x
awesome research! interesting implications for correctness and assurance here. excited that folks are doing deep dives into this problem and im excited to see what they release next
-
@bilaltwovec
Bilal
on x
I've been using FlexAttention a few months now and I've never felt better. I have more energy. My skin is clearer. Tpot has stopped complaining about my models being quantized 🔥 [image]
-
@dl_insider
Jose Lopez
on x
Thinking Machines has found the reason for the non-deterministic in LLM. Big deal for: * Scientific Reproducibility * High-Stakes & Safety-Critical Fields * Advanced AI Research * Testing and Validation But it has a cost so I expect a toggle : 1. Fast non-deterministic,
-
@an_vo12
An Vo
on x
This blog makes me wonder about the OPPOSITE problem: 👉 Can we make LLMs give uniform random answers when asked (e.g., “randomly pick 0-9")? So far, our work ( https://b-score.github.io/) shown that we can hack it with multi-turn, but I'd love to see this activated in single-turn…
-
@alth0u
@alth0u
on x
the field...worked on this problem for years...and...they just...they just tweeted it out.
-
@okfrankco
Frank
on x
The power of a good question: “Why aren't LLMs deterministic?” In short: computers run floating-point operations (math😝) in parallel, and the order in which each operation comes together can shift. This makes results slightly different each time. This small variance cascades
-
@jxbz
Jeremy Bernstein
on x
Excited about our new research blog!
-
@urmishthakker
Urmish Thakker
on x
Loved this post from @cHHillee ! The posts masks how hard these debugs are. Posts make story telling linear, reality is messy. Lot of wrong paths triaged with no return, sometimes revising old paths. Patiently grinding until you get to the answer. Great read, learnt a lot!
-
@barret_zoph
Barret Zoph
on x
Excited to share our first blog post — one of many to follow!
-
@matthieu_meeus
Matthieu Meeus
on x
I was recently asked during an interview why LLMs were not deterministic even when temperature is 0. This very carefully answers it! My interviewer's answer was different though, that there likely is non determinism in MOE routing for load balancing
-
@davidyin0609
David Yin
on x
Still remember earlier this year, I tried very hard to figure out why vLLM can give very different outputs than huggingface models even with greedy sampling (not sure if it is fixed now). Twisting batch size or number of gpus also makes the output different
-
@ruyimarone
Marc Marone
on x
This post is awesome! Explains why non determinism comes from a lack of batch invariant kernels and why that's hard “Surprisingly, we generate 80 unique completions, with the most common of these occurring 78 times” and then shows a truly deterministic vLLM run
-
@ye_combinator
Zihao Ye
on x
Awesome work from @thinkymachines and @cHHillee! The importance of determinism might be underestimated. Like with LLM-based compression ( https://bellard.org/... - you really need things to work the same way whether you're doing prefill/decode or different batching setups. Here's
-
@sschoenholz
Sam Schoenholz
on x
Looking forward to seeing research / writeups from Thinking Machines get shared with the community. @cHHillee et al.'s determinism work has been fantastic and was fascinating to watch evolve. Our infrastructure keeps getting better!
-
@ehsanshareghi
Ehsan Shareghi
on x
One interesting thing we found was that safety in LLMs is sometimes by pure luck! If the LLM/sampling method choose to generate “I am sorry ...” as initiating tokens, it will reject an unsafe request. Otherwise, it may not. This type of nondeterminism has many implications.
-
@brdkhsrv
Bardia Khosravi
on x
Amazing post on why LLMs are non-deterministic even with a temperature of zero. TLDR; Essentially it boils down to different inference batch sizes, as some kernels are not batch-size invariant. Meaning that their output is influenced by the number of requests🤯
-
@mollehilll
Moll
on x
Thinking Machines published an excellent piece explaining the phenomenon of «nondeterminism in inference». The authors show that the reason isn't «probability magic», but how the server engine itself is built. Requests don't go one by one, they are batched, and the batch size
-
@pmddomingos
Pedro Domingos
on x
Wow, Thinking Machines' research is so varied. It ranges all the way from kernel numerics to prompt engineering. I can't imagine needing anything more to solve AI.
-
@alibaba_qwen
@alibaba_qwen
on x
Awesome work by the @thinkymachines team! Thrilled that Qwen models (Qwen3-235B-A22B & Qwen3-8B & Qwen 2.5-VL) served as a foundation for this experiment. This is exactly why we build—to empower researchers tackling hard problems & unlocking new scientific insights. Can't wait to
-
@shumochu
Shumo Chu
on x
Such a great read! Also the report shows the @thinkymachines folks are not just pure “researchers”, they are also on the ground engineers who deeply understand the system issues of LLMs. We are in the stage where the system side and research side are deeply intertwined.
-
@varshinesri
Varsh Sridharan
on x
One of the best research blogs I've read in the recent times- learned a ton! Such diverse insights ranging all the way from Kernel numerics and reasons for non-determinism! So looking forward to dive into the root causes of our nondeterminism and even solving them!
-
@krinetix1234
Krish Maniar
on x
Very well written blog on how the LLM forward pass is actually deterministic (rarely need atomic_add), but the core source of nondeterminism comes down to variance in batch size Clearly there is still a performance gap (26 vs. 42s on vlllm), but once there's enough eyes on
-
@eltetonoemi
Noémi Éltető
on x
I enjoyed this blog post a lot! I also decided that I will dress as floating-point non-associativity this Halloween and give everyone a good scare. [image]
-
@statusfailed
@statusfailed
on x
super nice to finally see other people give a shit about the “original sin” (nonassociativity). Didn't realise the implications for RL either, this is great stuff.
-
@liyuanlucas
Liyuan Liu
on x
appreciate @thinkymachines taking an open research approach! excited to see the first blog mentioned our work! truly on-policy RL is like RTX3090 for gamers in 2020 - you really want it, but the blockers make your head itch... kernel mismatches, parallelism mismatches, etc. etc.
-
@bidhanxyz
Bidhan
on x
thinky machine has become an intellectual peer to bagel labs by starting a blog blog dot bagel dot com
-
@chhillee
Horace He
on x
Apologies that I haven't written anything since joining Thinking Machines but I hope this blog post on a topic very near and dear to my heart (reproducible floating point numerics in LLM inference) will make up for it!
-
@ashu_1069
Ashutosh Kumar
on x
Learnt a lot while working through this, took a lot of time to really understand the details and not just stay fond of the idea, but get to know it properly. Here's the blog link: https://ashu1069.substack.com/ ... [image]
-
@miramurati
Mira Murati
on x
A big part of our mission at Thinking Machines is to improve people's scientific understanding of AI and work with the broader research community. Introducing Connectionism today to share some of our scientific insights.
-
@abeirami
Ahmad Beirami
on x
This is a great example of what good research looks like. You start with a real problem. You peel it layer by layer to find the root cause. You form a new hypothesis and keep digging. At the end, you have something insightful to share!
-
@manuelfaysse
Manuel Faysse
on x
A big issue we had when serving ColQwen is the non-deterministic output embeddings. More specifically, the embeddings produced for the same images would differ when batch sizes changed at inference, leading to non-zero performance variations. This was surprising to us... I
-
@djpardis
DJ Pardis
on x
I'm excited to read this highly relevant, @thinkymachines post. Also, can we talk about how it's designed and typeset like a Knuth book? [image]
-
@jiayiyuan99
Jiayi Yuan
on x
Thanks for the shout-out to our work—it's great to see more focus on this important problem. Horace's approach is a truly elegant solution. Fantastic work! More reading: https://arxiv.org/... [image]
-
@zhongruiqi
Ruiqi Zhong
on x
I learned so much from this as an ML (not-system-ish) researcher. highly recommend a read!!
-
@thinkymachines
@thinkymachines
on x
Today Thinking Machines Lab is launching our research blog, Connectionism. Our first blog post is “Defeating Nondeterminism in LLM Inference” We believe that science is better when shared. Connectionism will cover topics as varied as our research is: from kernel numerics to [imag…
-
@davidcayjohnston
David Cay Johnston
on bluesky
Artificial Intelligence does weird things, gives delusional responses, because there's a major flaw that makes LLMs unstable, research paper by Mira Murati and her Thinking Machines lab shows. — Who else is asking questions may even affect your results! — thinkingmachines.ai/…
-
@offline.mountainherder.xyz
@offline.mountainherder.xyz
on bluesky
Fascinating article here claiming (and seemingly showing) that non-determinism in LLMs can be attributed to the load conditions of the hardware upon which it runs. — Much like when our brains short circuit due to a bad stomach ache.
-
@phillipcarter.dev
Phillip Carter
on bluesky
Fantastic technical post about experiments in making LLMs deterministic by making kernel operations act invariant of batch sizes: thinkingmachines.ai/blog/defeati... Curious if we'll see something like this roll out more broadly in the future!