Inkling in One Sentence: An Open, Multimodal Bet on Developers
Inkling is a 975-billion-parameter open-weight AI model using a Mixture-of-Experts architecture, trained from scratch on text, images, audio, and video to provide developers and enterprises with a customizable multimodal reasoning and coding system that competes with proprietary models while keeping the weights accessible for fine-tuning and deployment. That starting point matters more than benchmark bragging rights. Thinking Machines Lab, founded by former leaders from a major AI lab, has made its first move in the model race by treating openness as a product feature, not a marketing afterthought. Inkling isn’t marketed as “the best” general-purpose model, and that is the smartest thing about it: the real power here is that teams can shape it into what they need rather than wait for a distant provider to adjust a closed API.

Architecture and Capabilities: A Giant That Thinks in Many Mediums
Inkling is unapologetically big: 975 billion total parameters, with 41 billion active at inference thanks to its Mixture-of-Experts transformer design. That routing-based structure is the only reason such scale is usable in practice, and it signals that open-weight models are no longer small hobby projects but serious competitors to closed systems. The model was pretrained from scratch on 45 trillion tokens spanning text, images, audio, and video, and it natively processes text, image, and audio inputs. In real workflows, that means the same system can handle long-form audio understanding, speech transcription, visual analysis, tool use, and conventional reasoning and coding tasks without bolted-on adapters. Benchmarks back this up: Inkling scores 77.6% on SWE-bench Verified, 97.1% on AIME 2026, 87.2% on GPQA Diamond, and 73.5% on MMMU Pro, making it a credible multimodal model release among open-weight options.
Open Weights, Hugging Face, and Tinker: Customization as the Product
Where closed competitors sell one-size-fits-all APIs, Inkling sells malleability. The complete weights are downloadable as checkpoints suitable for NVIDIA Blackwell systems and are surfaced as a Hugging Face model, letting teams pull them directly into their own infrastructure. Live fine-tuning through the Tinker platform turns that openness into a workflow: developers can adjust Inkling on their own data, with context windows up to one million tokens and exposed controls over “thinking effort” between 0.2 and 0.99 to trade off quality against latency and token usage. That knob is more than a technical curiosity; it is explicit recognition that cost and responsiveness are part of model design. The startup even used Inkling to fine-tune itself, with its chain-of-thought explanations becoming more concise over time while remaining comprehensible and preserving answer quality—a sign that they view the model as a living system, not a static monolith.

Databricks, Agents, and Enterprise Control
Inkling’s day-zero presence on Databricks is not a side note; it is a direct play into the enterprise coding and agent ecosystem. On that platform, Inkling is governed through the Unity AI Gateway, inheriting security controls, permissions, audit logging, and policy enforcement so data stays inside a controlled environment while teams fine-tune on proprietary codebases, internal documentation, and domain-specific data for higher task accuracy. This is the answer to the usual complaint that open models are great until governance and compliance enter the picture. Enterprises can connect Inkling to popular coding agents through the same gateway, or build Inkling-powered agents with Agent Bricks to analyze data, automate complex work, evaluate with custom judges, and deploy at scale. Support for REST APIs is live, with SQL querying promised soon, turning an open-weight AI model into a first-class citizen in enterprise data platforms rather than an experimental side project.

Positioning Against Closed Models: Not the Best, but Maybe the Smartest
Thinking Machines admits Inkling is not the highest-scoring general-purpose model on popular benchmarks, but argues it performs well at many tasks and is capable of advanced reasoning and coding. That candor is refreshing in a field obsessed with leaderboard screenshots. According to the company, “Inkling can match Nemotron 3 Ultra on Terminal Bench 2.1 while using roughly one-third as many generated tokens,” underscoring how efficiency can rival raw score-chasing. In the wider context of ex-open-lab founders building new players, Inkling fits a stated vision that AI should not be controlled by a few companies and should be decentralized so more people can build models with their own data. The model’s strong open-source benchmark results, agent support, and enterprise integrations show a viable alternative to proprietary solutions rather than a knockoff. With Inkling-Small—a 276-billion-parameter Mixture-of-Experts model with 12 billion active parameters—already previewed and expected to bring lower-cost, lower-latency workloads once testing finishes, this looks less like a single launch and more like the start of an open-weight family.






