/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Mira Murati's TML launches a research blog called Connectionism, and shares its work on resolving nondeterminism and achieving reproducible results from LLMs

There's been great interest in what Mira Murati's Thinking Machines Lab is building with its $2 billion in seed funding …

TechCrunch Maxwell Zeff

Context & Ripple Effects

Thinking Machines Lab was introduced as Mira Murati’s new AI venture before securing a $2B seed round, giving the lab unusual resources to pursue foundational model research. Connectionism is an early public signal of the technical problems it intends to foreground.

The publication arrives ahead of TML’s first product, the Tinker fine-tuning API, linking research on repeatable model behavior to a potential developer-facing need rather than a purely academic exercise.

First-order effects

  • TML makes its work on LLM nondeterminism and reproducibility visible to researchers and prospective users, creating a public technical record for how it approaches repeatable outputs.
  • Developers evaluating TML’s eventual tools gain a clearer indication that reproducible model behavior is a design priority, not only raw model capability.

Second-order effects

  • Competing model labs and fine-tuning platforms face greater pressure to document whether results can be replicated across runs, especially for customers that need dependable evaluation and iteration.
  • Public research can turn reproducibility methods into a point of comparison for downstream tooling, including testing, experiment tracking, and managed model deployment.

Third-order effects

  • If reproducibility becomes a purchasing criterion, frontier-model competition may shift from benchmark performance alone toward operational consistency and auditable development workflows.
  • The pattern supports a more mature AI software stack in which research labs must translate model behavior into controls that developers can validate and repeatedly use.

The trend: Frontier AI labs are increasingly treating reliable, repeatable model behavior as a product and research differentiator alongside scale and capability.

Discussion

  • @lilianweng Lilian Weng on x
    Besides the fun fact that Connectionism is connected with the early days of the AI field and highlights similarities between neural networks and human brains, the flagship product of the (first) Thinking Machines is named Connection Machine. — 🧑‍🎓Enjoy reading and more is coming!
  • @fenbielding Ben Fielding on x
    strongly agree with the need for determinism in model execution outlined in @thinkymachines first blog post we take this further at @gensynai and build for reproducibility (determinism across devices) more to announce soon but in the meantime, a demo: https://github.com/...
  • @tensorbert Bert Maher on x
    Read this to the end — the last section is mind-blowing
  • @timlantin @timlantin on x
    life imitates art (westworld) [image]
  • @yashwanthsai29 Sai Yashwanth on x
    Learnt a lot reading this study. The quality is exceptional! I aspire to contribute at the same level in the coming future.
  • @suhackerr Suha on x
    awesome research! interesting implications for correctness and assurance here. excited that folks are doing deep dives into this problem and im excited to see what they release next
  • @bilaltwovec Bilal on x
    I've been using FlexAttention a few months now and I've never felt better. I have more energy. My skin is clearer. Tpot has stopped complaining about my models being quantized 🔥 [image]
  • @dl_insider Jose Lopez on x
    Thinking Machines has found the reason for the non-deterministic in LLM. Big deal for: * Scientific Reproducibility * High-Stakes & Safety-Critical Fields * Advanced AI Research * Testing and Validation But it has a cost so I expect a toggle : 1. Fast non-deterministic,
  • @an_vo12 An Vo on x
    This blog makes me wonder about the OPPOSITE problem: 👉 Can we make LLMs give uniform random answers when asked (e.g., “randomly pick 0-9")? So far, our work ( https://b-score.github.io/) shown that we can hack it with multi-turn, but I'd love to see this activated in single-turn…
  • @alth0u @alth0u on x
    the field...worked on this problem for years...and...they just...they just tweeted it out.
  • @okfrankco Frank on x
    The power of a good question: “Why aren't LLMs deterministic?” In short: computers run floating-point operations (math😝) in parallel, and the order in which each operation comes together can shift. This makes results slightly different each time. This small variance cascades
  • @jxbz Jeremy Bernstein on x
    Excited about our new research blog!
  • @urmishthakker Urmish Thakker on x
    Loved this post from @cHHillee ! The posts masks how hard these debugs are. Posts make story telling linear, reality is messy. Lot of wrong paths triaged with no return, sometimes revising old paths. Patiently grinding until you get to the answer. Great read, learnt a lot!
  • @barret_zoph Barret Zoph on x
    Excited to share our first blog post — one of many to follow!
  • @matthieu_meeus Matthieu Meeus on x
    I was recently asked during an interview why LLMs were not deterministic even when temperature is 0. This very carefully answers it! My interviewer's answer was different though, that there likely is non determinism in MOE routing for load balancing
  • @davidyin0609 David Yin on x
    Still remember earlier this year, I tried very hard to figure out why vLLM can give very different outputs than huggingface models even with greedy sampling (not sure if it is fixed now). Twisting batch size or number of gpus also makes the output different
  • @ruyimarone Marc Marone on x
    This post is awesome! Explains why non determinism comes from a lack of batch invariant kernels and why that's hard “Surprisingly, we generate 80 unique completions, with the most common of these occurring 78 times” and then shows a truly deterministic vLLM run
  • @ye_combinator Zihao Ye on x
    Awesome work from @thinkymachines and @cHHillee! The importance of determinism might be underestimated. Like with LLM-based compression ( https://bellard.org/... - you really need things to work the same way whether you're doing prefill/decode or different batching setups. Here's
  • @sschoenholz Sam Schoenholz on x
    Looking forward to seeing research / writeups from Thinking Machines get shared with the community. @cHHillee et al.'s determinism work has been fantastic and was fascinating to watch evolve. Our infrastructure keeps getting better!
  • @ehsanshareghi Ehsan Shareghi on x
    One interesting thing we found was that safety in LLMs is sometimes by pure luck! If the LLM/sampling method choose to generate “I am sorry ...” as initiating tokens, it will reject an unsafe request. Otherwise, it may not. This type of nondeterminism has many implications.
  • @brdkhsrv Bardia Khosravi on x
    Amazing post on why LLMs are non-deterministic even with a temperature of zero. TLDR; Essentially it boils down to different inference batch sizes, as some kernels are not batch-size invariant. Meaning that their output is influenced by the number of requests🤯
  • @mollehilll Moll on x
    Thinking Machines published an excellent piece explaining the phenomenon of «nondeterminism in inference». The authors show that the reason isn't «probability magic», but how the server engine itself is built. Requests don't go one by one, they are batched, and the batch size
  • @pmddomingos Pedro Domingos on x
    Wow, Thinking Machines' research is so varied. It ranges all the way from kernel numerics to prompt engineering. I can't imagine needing anything more to solve AI.
  • @alibaba_qwen @alibaba_qwen on x
    Awesome work by the @thinkymachines team! Thrilled that Qwen models (Qwen3-235B-A22B & Qwen3-8B & Qwen 2.5-VL) served as a foundation for this experiment. This is exactly why we build—to empower researchers tackling hard problems & unlocking new scientific insights. Can't wait to
  • @shumochu Shumo Chu on x
    Such a great read! Also the report shows the @thinkymachines folks are not just pure “researchers”, they are also on the ground engineers who deeply understand the system issues of LLMs. We are in the stage where the system side and research side are deeply intertwined.
  • @varshinesri Varsh Sridharan on x
    One of the best research blogs I've read in the recent times- learned a ton! Such diverse insights ranging all the way from Kernel numerics and reasons for non-determinism! So looking forward to dive into the root causes of our nondeterminism and even solving them!
  • @krinetix1234 Krish Maniar on x
    Very well written blog on how the LLM forward pass is actually deterministic (rarely need atomic_add), but the core source of nondeterminism comes down to variance in batch size Clearly there is still a performance gap (26 vs. 42s on vlllm), but once there's enough eyes on
  • @eltetonoemi Noémi Éltető on x
    I enjoyed this blog post a lot! I also decided that I will dress as floating-point non-associativity this Halloween and give everyone a good scare. [image]
  • @statusfailed @statusfailed on x
    super nice to finally see other people give a shit about the “original sin” (nonassociativity). Didn't realise the implications for RL either, this is great stuff.
  • @liyuanlucas Liyuan Liu on x
    appreciate @thinkymachines taking an open research approach! excited to see the first blog mentioned our work! truly on-policy RL is like RTX3090 for gamers in 2020 - you really want it, but the blockers make your head itch... kernel mismatches, parallelism mismatches, etc. etc.
  • @bidhanxyz Bidhan on x
    thinky machine has become an intellectual peer to bagel labs by starting a blog blog dot bagel dot com
  • @chhillee Horace He on x
    Apologies that I haven't written anything since joining Thinking Machines but I hope this blog post on a topic very near and dear to my heart (reproducible floating point numerics in LLM inference) will make up for it!
  • @ashu_1069 Ashutosh Kumar on x
    Learnt a lot while working through this, took a lot of time to really understand the details and not just stay fond of the idea, but get to know it properly. Here's the blog link: https://ashu1069.substack.com/ ... [image]
  • @miramurati Mira Murati on x
    A big part of our mission at Thinking Machines is to improve people's scientific understanding of AI and work with the broader research community. Introducing Connectionism today to share some of our scientific insights.
  • @abeirami Ahmad Beirami on x
    This is a great example of what good research looks like. You start with a real problem. You peel it layer by layer to find the root cause. You form a new hypothesis and keep digging. At the end, you have something insightful to share!
  • @manuelfaysse Manuel Faysse on x
    A big issue we had when serving ColQwen is the non-deterministic output embeddings. More specifically, the embeddings produced for the same images would differ when batch sizes changed at inference, leading to non-zero performance variations. This was surprising to us... I
  • @djpardis DJ Pardis on x
    I'm excited to read this highly relevant, @thinkymachines post. Also, can we talk about how it's designed and typeset like a Knuth book? [image]
  • @jiayiyuan99 Jiayi Yuan on x
    Thanks for the shout-out to our work—it's great to see more focus on this important problem. Horace's approach is a truly elegant solution. Fantastic work! More reading: https://arxiv.org/... [image]
  • @zhongruiqi Ruiqi Zhong on x
    I learned so much from this as an ML (not-system-ish) researcher. highly recommend a read!!
  • @thinkymachines @thinkymachines on x
    Today Thinking Machines Lab is launching our research blog, Connectionism. Our first blog post is “Defeating Nondeterminism in LLM Inference” We believe that science is better when shared. Connectionism will cover topics as varied as our research is: from kernel numerics to [imag…
  • @davidcayjohnston David Cay Johnston on bluesky
    Artificial Intelligence does weird things, gives delusional responses, because there's a major flaw that makes LLMs unstable, research paper by Mira Murati and her Thinking Machines lab shows.  —  Who else is asking questions may even affect your results!  —  thinkingmachines.ai/…
  • @offline.mountainherder.xyz @offline.mountainherder.xyz on bluesky
    Fascinating article here claiming (and seemingly showing) that non-determinism in LLMs can be attributed to the load conditions of the hardware upon which it runs.  —  Much like when our brains short circuit due to a bad stomach ache.
  • @phillipcarter.dev Phillip Carter on bluesky
    Fantastic technical post about experiments in making LLMs deterministic by making kernel operations act invariant of batch sizes: thinkingmachines.ai/blog/defeati...  Curious if we'll see something like this roll out more broadly in the future!