Inkling’s Big Bet: Own the Model, Not the API
Inkling is a large open-weights AI model designed so organizations can download its full parameters, fine-tune it on their own data, and control how it reasons across text, images, audio, and video, instead of being locked into a single provider’s general-purpose chatbot. Inkling is a Mixture-of-Experts transformer with 975 billion total parameters, 41 billion active at any time, a 1 million token context window, and pretraining on 45 trillion tokens covering text, images, audio and video. The lab behind it, Thinking Machines, is upfront: this is “not the strongest overall model available today, open or closed,” but it is something you can make your own. That framing is a direct challenge to the dominant model of renting intelligence over an API. Instead of selling peak benchmark scores, Inkling sells control—an alternative to the one-size-fits-all paradigm that now defines most commercial AI.

Tinker, Not Tokens: A Business Model Built on Customization
Thinking Machines’ most disruptive move is not technical but commercial: it gives Inkling’s weights away and charges for customization through its Tinker platform. Once an open-weights AI model is public, no developer is obliged to pay per-token inference fees, so the company’s revenue has to come from fine-tuning services and hosting around them. Rather than pitching Inkling as a polished product, the lab markets it explicitly as “a starting point for fine-tuning through Tinker, its model-customization platform.” That shifts the economic focus from usage to adaptation. Enterprises can train on proprietary codebases, internal documentation and domain-specific data to get higher accuracy on their own tasks without surrendering control of their data to a closed provider. In a market that has moved hard toward cost per finished task instead of sheer peak intelligence, this is a quietly radical bet that tailored models will outcompete general-purpose systems on value.

Multimodal Reasoning: Inkling’s Technical Position Against Closed Giants
Technically, Inkling is designed to sit in the same conversation as models from OpenAI, Anthropic and Google while staying downloadable. It supports multimodal AI reasoning: the model is trained across text, image, audio and video, with multimodality “built in from the start rather than bolted on.” Audio is fed as dMel spectrograms and images as 40×40 pixel patches through a four-layer hMLP, then processed jointly with text tokens in the transformer stack. Despite currently outputting only text, this design gives Inkling credible scores on audio and vision benchmarks, placing it among the strongest open-weights audio models. On coding and agentic workflows, Inkling scores 77.6% on SWE-bench Verified, 63.8% on Terminal Bench 2.1, and 79.8% on IFBench chat, ahead of several named proprietary competitors. The efficient, controllable “thinking effort” setting lets Inkling match stronger closed models on Terminal Bench 2.1 with roughly a third of the tokens, a concrete efficiency win for large-scale deployments.

Distribution and Governance: Hugging Face, Databricks and the Enterprise Angle
Inkling’s accessibility is deliberate. Full weights are available on Hugging Face as both the original checkpoint and an NVFP4 checkpoint tuned for Blackwell hardware, while multiple providers offer API access. A day-zero launch partnership brings the model onto Databricks, making it directly available for enterprise teams who want to apply it to their own data and coding workflows inside governed environments. Databricks emphasizes four benefits: context through fine-tuning on proprietary codebases and documentation, control via its Unity AI Gateway, choice that avoids lock-in to any single model provider, and cost management without per-token pricing. Inkling can already be invoked via REST on Databricks, with SQL querying support coming soon. This is where the open-weights AI model strategy shows its teeth: organizations can combine and switch models, run Inkling at scales that match their workloads, and still retain full ownership of customized weights instead of pushing everything through a closed alternative to ChatGPT.
Mira Murati’s Signal: Credible Opposition to Closed AI Orthodoxy
Inkling is not arriving from an unknown lab. Thinking Machines was founded by former OpenAI CTO Mira Murati, whose decision to build an open-weight model is itself a statement about where she thinks the industry should go. The company trained Inkling in about nine months, far shorter than the multi-year timelines common at major labs, and now employs around 200 people following earlier departures, including co-founders who left for her former employer. The business logic behind giving the weights away is explicitly “not charity” but a belief that organisations should own a model rather than rent one. Inkling-Small, a 276 billion parameter sibling with 12 billion active, is still in testing, with weights due later. Meanwhile, Inkling is already on Tinker with 64K and 256K context options at a limited-time discount and a Playground for quick trials. In a landscape dominated by closed platforms and API-only access, Murati is betting her new lab’s future on customizable AI models as the next competitive frontier.







