A 975 billion parameter model, most of it dormant at any one time
Inkling is a mixture-of-experts model with 975 billion total parameters and about 41 billion active per task, trained on 45 trillion tokens of text, image, audio and video. It reasons natively across those four input modalities, though its output at release is text only, including code and structured data. Context runs up to one million tokens. [1][2]
Thinking Machines shipped a smaller companion, Inkling Small, at 276 billion total parameters and 12 billion active, trading benchmark performance for lower latency and cost. Full weights for both went up on Hugging Face in standard and NVFP4 formats, alongside API access through Together AI, Fireworks, Modal, Databricks and Baseten, and fine-tuning access through Thinking Machines' own Tinker platform. [1]
The company was explicit that Inkling is not the strongest model available, open or closed. TechCrunch quotes Thinking Machines describing it as tuned for calibrated answers that flag uncertainty rather than guess, with an adjustable thinking effort setting that trades speed against accuracy. [2]
The leading US open-weights model, provisionally
Artificial Analysis scored Inkling at 41 on its Intelligence Index, three points above Nemotron 3 Ultra, the prior leading US open-weights model, and reported it running more token-efficient than that comparison set: an average of 25,000 output tokens per Intelligence Index task versus 43,000, 38,000 and 37,000 for GLM-5.2, Kimi K2.6 and DeepSeek v4 Pro respectively. [3]
On agentic benchmarks, Artificial Analysis reported Inkling ahead of two Chinese open-weights models on GDPval-AA v2, with an Elo of 1238 against Kimi K2.6's 1190 and DeepSeek v4 Flash max's 1189, and narrowly ahead on the tau3-Banking benchmark at 24 percent versus 21 and 23 percent. It scored lower on AA-Omniscience, a knowledge-retention benchmark, than those same international models, while still leading other US open-weights entries there. [3]
Atlas interpretation: The summary above calls this the strongest open-weights model from a US lab, in a category China had been holding, and Artificial Analysis's numbers support that on the aggregate index and on agentic tasks specifically, while showing the lead is not uniform across every benchmark. A single index score compresses a model that trails on raw knowledge but leads on efficiency and multi-step task completion into one number, which is a reason to read the category claim as a snapshot rather than a settled ranking. [3]
The pitch is customization, not the leaderboard
Thinking Machines framed Inkling as a starting point for fine-tuning rather than a finished product, arguing that models organizations can adapt for themselves will outperform the one-size-fits-all models the largest labs sell. TechCrunch reports the company citing Microsoft chief executive Satya Nadella's warning that enterprises on proprietary models effectively pay twice, once in subscription cost and again when their own business knowledge gets embedded in prompts sent to someone else's model, and Hugging Face chief executive Clem Delangue's prediction that frontier models will handle experimentation while production work moves to private or open alternatives. [2]
As evidence, Thinking Machines pointed to a collaboration with the hedge fund Bridgewater Associates, in which researchers fine-tuned an open-source model on Bridgewater's financial expertise and reported 84.7 percent on a financial reasoning test, ahead of the proprietary models compared against, at roughly one fourteenth the cost to run. [2]
Atlas interpretation: Because open weights can be run for free once downloaded, Inkling itself does not meter revenue back to Thinking Machines. TechCrunch reports the company's business model runs through Tinker, its fine-tuning and training platform, and through a cut of the hosting ecosystem built around models like this one, which makes Inkling's release closer to a proof of the pitch than a product with its own price tag. [2]
Post-training borrowed from a rival's open weights, for now
TechCrunch reports Inkling's pre-training was largely self-contained, but its post-training drew in part on other open-weight models, including Moonshot AI's Kimi K2.5, before reinforcement learning took over, and that Thinking Machines has committed to fully self-contained post-training for its next model. Inkling trained on Nvidia GB300 NVL72 systems, delivered under the gigawatt-scaleVera Rubin partnership Nvidia and Thinking Machines announced in March 2026. [2]
Inkling arrived roughly nine months after Thinking Machines' founding, and about two months after the company'sinteraction-focused voice model work in May 2026. TechCrunch notes a reported $50 billion fundraising round was in progress as of the prior November but had stalled by January, and that the company had lost two co-founders to OpenAI in January 2026 while operating with roughly 200 employees. [2]
Sources
- Inkling: Our Open-Weights Model
Thinking Machines Lab · Jul 15, 2026
- Thinking Machines amps up its bet against one-size-fits-all AI with its first open model, Inkling
TechCrunch · Jul 15, 2026
- Thinking Machines has released Inkling, the new leading U.S. open weights model
Artificial Analysis · Jul 15, 2026