Thinking Machines releases Inkling-Small, an open-weight model with 276B total and 12B active parameters, saying it “achieves comparable performance” to Inkling
Thinking Machines introduced Inkling as a broad open-weight MoE model with 975B total and 41B active parameters; Inkling-Small sharply reduces both figures while retaining the same open-weights distribution model.
Developers can evaluate and deploy an open-weight Inkling variant with 12B active parameters rather than Inkling's 41B active footprint, subject to validating the company's comparable-performance claim for their workloads.
Thinking Machines expands the Inkling family from a single broad model into a size-tiered offering, giving Tinker users another potential base model for fine-tuning.
Second-order effects
Model adopters will have a clearer quality-versus-serving-cost comparison within one model family, increasing pressure on competing open-weight releases to show both total and active parameter requirements.
Fine-tuning and deployment tooling become more valuable differentiators: smaller active models can widen the set of teams able to test customized open weights, while model weights alone are less likely to determine adoption.
Third-order effects
If comparable capability is sustained at lower active parameter counts, open-weight competition may increasingly center on inference efficiency and adaptation ecosystems rather than headline total parameter scale.
The pattern supports a more segmented open-model market, with broad base models and lower-active variants serving different deployment constraints rather than a single model defining the category.
The trend: Open-weight model vendors are turning large MoE research releases into product families differentiated by active inference cost and fine-tuning accessibility.
Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available. https://thinkingmachines.ai/ ... Fine-tune it on Tinker today, or cha…
Whereas I felt like it took a village to release inkling, inkling-small felt much more routine 😆 We just took the pipeline used for Inkling, passed in a smaller model, and voila - new model! Inkling small benefited quite a bit vs Inkling from some minor improvements, but there's …
Inkling-Small began training after its larger counterpart, so it benefits from everything we learned: an improved pre-training data mix, a refined ML recipe, on-policy distillation with Inkling as the teacher, and two further weeks of agentic coding RL.
It matches or exceeds Inkling on reasoning and agentic tasks. 31.6% on HLE, ahead of Inkling's 29.7%, and the advantage holds at every thinking budget. On SWEBench-Verified it exceeds 80%.
Another open-weight release from @thinkymachines 👀 Inkling-Small is here. With native reasoning over audio and images and variable thinking effort, it's a great choice for fine-tuning, with NVIDIA NeMo on NVIDIA DGX Station. NVFP4 checkpoint here: https://huggingface.co/...
Thinking Machines released Inkling Small (🦖) + NVFP4 12B active 276B total params, the model performs better than larger Inkling on coding 🤯 > check out our blog covering benchmarks, performance and deployment https://huggingface.co/... > https://huggingface.co/... [image]
We're pushing on efficiency and releasing our next model Inkling-Small, with similar performance to Inkling at just over a quarter of the size. If you have a Tinker account you can test it out today in the new Tinker playground: https://tinker.thinkingmachines.ai/ ...
Inkling Small by @thinkymachines is now available on Design Arena! Built as a Mixture-of-Experts model with 276B-parameters (a quarter of the size of Inkling), Inkling Small is a natively multimodal, open-weights model designed for reasoning and agentic tasks. Congratulations to …
Efficiency is the point. Across agentic tool use (Terminal-Bench 2.1), reasoning (HLE), and instruction following (IFBench), Inkling-Small delivers more performance per FLOP than Inkling. Variable thinking effort lets you pick your point on the cost/performance curve.
our model factory is operational 🙂 excited for everyone to try Inkling-Small. much smaller than Inkling but very comparable performance, and ofc open-weights. [video]
Inkling-Small is comparable to Inkling at a quarter the size. Weights are open, fine-tunable on Tinker today. Look forward to seeing what people make with it.
Inkling-Small is out. 12B active, 276B total. 🚀 We took what we learned from the larger run to get comparable performance at 1/4 the compute. It has native multimodal reasoning and variable thinking effort. Weights are open. Available on Tinker now. https://thinkingmachines.ai/ .…