Thinking Machines releases Inkling-Small, an open-weight model with 276B total and 12B active parameters, saying it “achieves comparable performance” to Inkling
Thinking Machines Lab introduced Inkling earlier this month as a broad open-weight MoE with 975B total and 41B active parameters. Inkling-Small is a materially lower-active-parameter companion that the lab says preserves comparable performance, shifting the story from a flagship release to a model-family deployment choice.
The release also extends the company’s effort to pair model weights with Tinker’s generally available fine-tuning API, giving users a route from an adaptable base model to customized workloads.
First-order effects
Developers get an open-weight Inkling option with 276B total parameters and 12B active parameters, rather than having to deploy the earlier 41B-active Inkling model for every workload.
Thinking Machines Lab broadens Inkling from a single flagship into a size-tiered offering; its performance claim will make evaluation on target tasks the immediate decision point for adopters.
Second-order effects
A smaller active model can make deployment economics and capacity planning more central to model selection, even where users value the original model’s broad capabilities.
Other open-weight model providers face added pressure to offer clearer capability-versus-active-parameter trade-offs and practical customization paths, rather than competing only on headline total parameter counts.
Third-order effects
If comparable-performance claims hold across deployments, open-weight model competition may increasingly organize around sparse-model efficiency and task-specific tuning, with total parameter count becoming a weaker standalone signal.
The durable value may shift toward the surrounding tooling, evaluation, and governed deployment layers that help organizations select, adapt, and run portable weights.
The trend: Open-weight AI providers are turning flagship MoE releases into tiered model families, competing on usable inference efficiency and customization as much as raw scale.
Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available. https://thinkingmachines.ai/ ... Fine-tune it on Tinker today, or cha…
Another open-weight release from @thinkymachines 👀 Inkling-Small is here. With native reasoning over audio and images and variable thinking effort, it's a great choice for fine-tuning, with NVIDIA NeMo on NVIDIA DGX Station. NVFP4 checkpoint here: https://huggingface.co/...
Inkling-Small is comparable to Inkling at a quarter the size. Weights are open, fine-tunable on Tinker today. Look forward to seeing what people make with it.
It matches or exceeds Inkling on reasoning and agentic tasks. 31.6% on HLE, ahead of Inkling's 29.7%, and the advantage holds at every thinking budget. On SWEBench-Verified it exceeds 80%.
Inkling-Small began training after its larger counterpart, so it benefits from everything we learned: an improved pre-training data mix, a refined ML recipe, on-policy distillation with Inkling as the teacher, and two further weeks of agentic coding RL.
Efficiency is the point. Across agentic tool use (Terminal-Bench 2.1), reasoning (HLE), and instruction following (IFBench), Inkling-Small delivers more performance per FLOP than Inkling. Variable thinking effort lets you pick your point on the cost/performance curve.
Inkling Small by @thinkymachines is now available on Design Arena! Built as a Mixture-of-Experts model with 276B-parameters (a quarter of the size of Inkling), Inkling Small is a natively multimodal, open-weights model designed for reasoning and agentic tasks. Congratulations to …
Inkling-Small is out. 12B active, 276B total. 🚀 We took what we learned from the larger run to get comparable performance at 1/4 the compute. It has native multimodal reasoning and variable thinking effort. Weights are open. Available on Tinker now. https://thinkingmachines.ai/ .…
We're pushing on efficiency and releasing our next model Inkling-Small, with similar performance to Inkling at just over a quarter of the size. If you have a Tinker account you can test it out today in the new Tinker playground: https://tinker.thinkingmachines.ai/ ...
Whereas I felt like it took a village to release inkling, inkling-small felt much more routine 😆 We just took the pipeline used for Inkling, passed in a smaller model, and voila - new model! Inkling small benefited quite a bit vs Inkling from some minor improvements, but there's …
Thinking Machines released Inkling Small (🦖) + NVFP4 12B active 276B total params, the model performs better than larger Inkling on coding 🤯 > check out our blog covering benchmarks, performance and deployment https://huggingface.co/... > https://huggingface.co/... [image]
our model factory is operational 🙂 excited for everyone to try Inkling-Small. much smaller than Inkling but very comparable performance, and ofc open-weights. [video]