/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Thinking Machines releases Inkling-Small, an open-weight model with 276B total and 12B active parameters, saying it “achieves comparable performance” to Inkling

Try on Tinker Model card Hugging Face  —  Today, we are releasing Inkling-Small, an efficient open-weights model …

Thinking Machines Lab

Context & Ripple Effects

Thinking Machines Lab introduced Inkling earlier this month as a broad open-weight MoE with 975B total and 41B active parameters. Inkling-Small is a materially lower-active-parameter companion that the lab says preserves comparable performance, shifting the story from a flagship release to a model-family deployment choice.

The release also extends the company’s effort to pair model weights with Tinker’s generally available fine-tuning API, giving users a route from an adaptable base model to customized workloads.

First-order effects

  • Developers get an open-weight Inkling option with 276B total parameters and 12B active parameters, rather than having to deploy the earlier 41B-active Inkling model for every workload.
  • Thinking Machines Lab broadens Inkling from a single flagship into a size-tiered offering; its performance claim will make evaluation on target tasks the immediate decision point for adopters.

Second-order effects

  • A smaller active model can make deployment economics and capacity planning more central to model selection, even where users value the original model’s broad capabilities.
  • Other open-weight model providers face added pressure to offer clearer capability-versus-active-parameter trade-offs and practical customization paths, rather than competing only on headline total parameter counts.

Third-order effects

  • If comparable-performance claims hold across deployments, open-weight model competition may increasingly organize around sparse-model efficiency and task-specific tuning, with total parameter count becoming a weaker standalone signal.
  • The durable value may shift toward the surrounding tooling, evaluation, and governed deployment layers that help organizations select, adapt, and run portable weights.

The trend: Open-weight AI providers are turning flagship MoE releases into tiered model families, competing on usable inference efficiency and customization as much as raw scale.

Discussion

  • @thinkymachines @thinkymachines on x
    Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available. https://thinkingmachines.ai/ ... Fine-tune it on Tinker today, or cha…
  • @natolambert Nathan Lambert on x
    Let's go, they're cooking. Looks like a great model. Fast follow up on a first model is such a green flag. 🫡
  • @artificialanlys @artificialanlys on x
    Thinking Machines' new Inkling Small scores 40 on the Artificial Analysis Intelligence Index …
  • @nvidiaai @nvidiaai on x
    Another open-weight release from @thinkymachines 👀 Inkling-Small is here. With native reasoning over audio and images and variable thinking effort, it's a great choice for fine-tuning, with NVIDIA NeMo on NVIDIA DGX Station. NVFP4 checkpoint here: https://huggingface.co/...
  • @soumithchintala Soumith Chintala on x
    Inkling-small. 2 weeks after inkling Nearly as good as Inkling but 4x smaller. We're just getting started...🔥
  • @miramurati Mira Murati on x
    Inkling-Small is comparable to Inkling at a quarter the size. Weights are open, fine-tunable on Tinker today. Look forward to seeing what people make with it.
  • @thinkymachines @thinkymachines on x
    It matches or exceeds Inkling on reasoning and agentic tasks. 31.6% on HLE, ahead of Inkling's 29.7%, and the advantage holds at every thinking budget. On SWEBench-Verified it exceeds 80%.
  • @thinkymachines @thinkymachines on x
    Inkling-Small began training after its larger counterpart, so it benefits from everything we learned: an improved pre-training data mix, a refined ML recipe, on-policy distillation with Inkling as the teacher, and two further weeks of agentic coding RL.
  • @thinkymachines @thinkymachines on x
    Efficiency is the point. Across agentic tool use (Terminal-Bench 2.1), reasoning (HLE), and instruction following (IFBench), Inkling-Small delivers more performance per FLOP than Inkling. Variable thinking effort lets you pick your point on the cost/performance curve.
  • r/LocalLLaMA r on reddit
    Inkling-Small by thinkingmachines
  • @lu__jasper Jasper Lu on x
    This model has been fun to train and play play around with.  I've found it to be a nice …
  • @designarena @designarena on x
    Inkling Small by @thinkymachines is now available on Design Arena! Built as a Mixture-of-Experts model with 276B-parameters (a quarter of the size of Inkling), Inkling Small is a natively multimodal, open-weights model designed for reasoning and agentic tasks. Congratulations to …
  • @genai_is_real Chayenne Zhao on x
    a model that matches its 4x larger sibling in capability while fitting full-parameter RL …
  • @teortaxestex @teortaxestex on x
    slightly better than V4-Flash on benchmarks, while smaller good [image]
  • @guangxuan_xiao Guangxuan Xiao on x
    Inkling-Small is out. 12B active, 276B total. 🚀 We took what we learned from the larger run to get comparable performance at 1/4 the compute. It has native multimodal reasoning and variable thinking effort. Weights are open. Available on Tinker now. https://thinkingmachines.ai/ .…
  • @litianleli Tim Li on x
    Two weeks later, as promised, Lil Ink is here!  And it outperforms its bigger sibling …
  • @ziqiao_ma Martin Ziqiao Ma on x
    Thanks @ArtificialAnlys for working together on evaluations! Excited to see Inkling-small on the pareto line🫡 [image]
  • @neal_wu Neal Wu on x
    We're pushing on efficiency and releasing our next model Inkling-Small, with similar performance to Inkling at just over a quarter of the size. If you have a Tinker account you can test it out today in the new Tinker playground: https://tinker.thinkingmachines.ai/ ...
  • @chhillee Horace He on x
    Whereas I felt like it took a village to release inkling, inkling-small felt much more routine 😆 We just took the pipeline used for Inkling, passed in a smaller model, and voila - new model! Inkling small benefited quite a bit vs Inkling from some minor improvements, but there's …
  • @lmsysorg @lmsysorg on x
    Inkling-small is out today!  With SGLang, you can get 648 tok/s decode with DSpark …
  • @mervenoyann Merve on x
    Thinking Machines released Inkling Small (🦖) + NVFP4 12B active 276B total params, the model performs better than larger Inkling on coding 🤯 > check out our blog covering benchmarks, performance and deployment https://huggingface.co/... > https://huggingface.co/... [image]
  • @arena @arena on x
    Inkling-Small (@thinkymachines) debuts at rank ~#88 (1431 pts, AutoEval) in Text Arena and ~#21 among open. …
  • @alhyunsoo Andrew Hyunsoo Lee on x
    our model factory is operational 🙂 excited for everyone to try Inkling-Small. much smaller than Inkling but very comparable performance, and ofc open-weights. [video]