Inkling in a Sentence: An AI Built to Be Changed
The Inkling model is a 975-billion-parameter open-weight AI model, built with a mixture-of-experts architecture and a one-million-token context window, that developers can freely download, inspect, and fine-tune for their own applications rather than access only through a closed API. Inkling’s release is less about joining the crowded model leaderboard and more about attacking the lock-in economics of today’s AI giants. Thinking Machines Lab, founded by former leaders from a major AI lab, has made its first in-house model openly available so that researchers, startups, and enterprises can run and modify it on their own infrastructure. That choice is not ideological window dressing; it is a direct challenge to the pay-per-token business model that dominates the industry and a bet that, in the long run, custom AI will beat one-size-fits-all systems.

An Open-Weight Giant That Thinks in Audio, Video, and Text
Technically, Inkling is a high-end experiment in what an open-weight AI model can be. It is a Mixture-of-Experts Transformer with 975 billion total parameters, of which about 41 billion activate on any given prompt, trained from scratch on 45 trillion tokens of text, images, audio, and video. It can natively handle text, images, and audio within a context window of up to one million tokens, and was built in under nine months on Nvidia’s GB300 NVL72 systems. The lab says the model can make sense of audio and video input as well as text, even though its current outputs are text-only. Performance-wise, Inkling scores 41 on the Artificial Analysis Intelligence Index, making it the highest-scoring open-weights model from a U.S. lab while still trailing leading open models from elsewhere. But Thinking Machines is blunt that “Inkling is not the strongest overall model available today.” That honesty signals where they believe the real competitive edge lies.

From API Lock-In to Tinker-Led Customization
Where closed models sell access, Inkling sells autonomy. The full weights are available on Hugging Face so anyone can download, inspect, and modify them for free. Instead of charging per API call like dominant labs, Thinking Machines plans to make money through Tinker, its paid platform for fine-tuning the model to specific business tasks. In other words, the base model is a public good; the business is in guided customization. One analyst put it bluntly: “Thinking Machines is charging for Tinker, the platform that companies will likely want to use to customize Inkling for their specific use cases.” For enterprises, that shift matters. Rather than paying ongoing per-token fees, they can invest in infrastructure they control, fine-tuning Inkling into smaller models optimized for their data and latency needs. That is a direct swipe at the rental economics of traditional API-based AI and a clear signal that customizable AI alternatives are ready to compete on cost, governance, and control, not just raw benchmarks.
Why Customizability Beats Benchmarks for Real Users
Thinking Machines is explicit that it values adaptability over leaderboard glory. Open-source and open-weight models already draw attention because they are cheaper to run than closed models and easier to modify for specific tasks. Inkling doubles down on that dynamic. It offers adjustable “thinking effort” controls so developers can trade speed for accuracy, and it flags its own uncertainty instead of generating confident-sounding guesses. Those features make sense if the model is destined to be embedded inside larger systems and workflows rather than used as a standalone chatbot. A case study with Bridgewater Associates illustrates the strategy: researchers used Tinker to fine-tune an open model with specialized financial data, yielding a lightweight system that scored 84.7% on financial reasoning benchmarks while beating proprietary alternatives at under a tenth of the cost. According to Futurum Group’s Mitch Ashely, this gives enterprises “a credible alternative positioned on customization economics,” shifting spend from per-token API pricing to owned infrastructure.
Inkling’s Edge: Open-Weight Today, AI Substrate Tomorrow
Inkling is not a finished super-assistant; it is a substrate for others to build on. The company openly markets it as a starting point rather than an all-purpose product, and is already previewing Inkling-Small, a compact 276-billion-parameter sibling with 12 billion active parameters that actually outperforms the larger model on the GPQA Diamond benchmark, 88.3% versus 87.2%. Full weights for that variant will be released after testing, reinforcing the open-weight commitment. This launch also fills a glaring gap: Chinese labs have dominated the open-weight ecosystem, and Inkling is the first strong alternative from a Western-founded lab, giving enterprises another path that aligns with their governance needs. More broadly, the release advances Thinking Machines’ stated vision that AI should not be controlled by a small handful of companies but decentralized so organizations can build their own models with their own data. If that vision holds, the future of AI will look less like renting a single central brain and more like running many tailored minds on infrastructure you own.






