MilikMilik

Inkling’s Open-Weights Strategy Takes Direct Aim at Frontier AI Pricing

Inkling’s Open-Weights Strategy Takes Direct Aim at Frontier AI Pricing
Interest|High-Quality Software

Inkling’s Bigger Bet: Own the Model, Not the API

Inkling is a 975-billion-parameter open-weights AI model built as a Mixture of Experts transformer with 41 billion active parameters and a one-million-token context window, trained from scratch across text, images, audio, and video to be downloaded, inspected, and customized rather than locked behind an API paywall.

Thinking Machines Lab, founded by former OpenAI CTO Mira Murati after her departure in September 2024 and backed by a multibillion-dollar seed round, has finally shipped its first in-house model: Inkling. On July 15, the lab released the Inkling model as full open weights on Hugging Face, meaning anyone can download, run, and inspect the system without sending traffic through the company’s servers. That move alone would make Inkling notable, but the sharper story is strategic: the company openly admits Inkling is “not the strongest overall model available today, open or closed,” and still chooses to give away the core asset. This is not charity; it is a business model experiment designed to challenge how frontier AI is priced and sold.

Inkling’s Open-Weights Strategy Takes Direct Aim at Frontier AI Pricing

A Frontier-Scale Open Model That Refuses the Leaderboard Game

On paper, the Inkling model release plants Thinking Machines squarely in frontier territory. Inkling is a Mixture of Experts model with 975 billion parameters, of which 41 billion are active per token, backed by a one-million-token context window and 45 trillion tokens of pretraining that span text, images, audio, and video. The full weights are on Hugging Face in both original and NVFP4 formats for efficient deployment on modern accelerators. Unlike many multimodal systems that bolt on separate vision or audio encoders, Inkling was trained from scratch to handle these inputs directly, using spectrograms for audio and patch-based processing for images.

Benchmarks show it as a strong open-weights AI model rather than a clear frontier winner: it posts high scores on math, agentic coding, and multimodal tasks but still sits behind some closed specialists and top-tier proprietary models in certain domains. That honesty matters. By stating at launch that Inkling is not the best model in existence, Thinking Machines sidesteps the usual leaderboard theatrics and reframes the product as a base to be molded. The architecture supports a controllable “thinking effort” dial that trades tokens for reasoning depth, and an upcoming Inkling-Small variant with 276 billion total parameters and 12 billion active is already in preview, with weights due once testing concludes.

Inkling’s Open-Weights Strategy Takes Direct Aim at Frontier AI Pricing

Tinker, Not Tokens: A New Monetization Playbook

The most controversial move is financial, not technical. Inkling’s weights are free to download; once a company has them, there is no obligation to pay Thinking Machines for inference. The business logic is explicit: revenue has to come from Tinker, the fine-tuning platform, and from a slice of the hosting ecosystem forming around it. Inkling is live on Tinker with 64K and 256K context options, offered at a 50% discount for a limited time, plus a playground for teams that want to test chat-style interaction before committing compute.

Instead of selling a single closed model tuned “for everyone,” Thinking Machines sells customization infrastructure. It wants to be both the model maker and the customization layer through Tinker, rather than leaving that work to outside platforms. Deployment partners for fine-tuned checkpoints include Together, Fireworks, Modal, Databricks, and Baseten, with inference support across engines like vLLM, SGLang, llama.cpp, and the standard transformers stack. This means an enterprise can fine-tune Inkling on Tinker, then bring its checkpoint to a preferred infrastructure provider, or even run it entirely in-house. The economic unit is no longer “tokens to the mothership”; it is “runs and fine-tunes that shape a model you own.”

Why Open Weights Matter for Coding, Agents, and Cost

Inkling’s open-weights approach is not a purity play; it is a weapon in the fight over total cost per task. The release lands in a market that has shifted toward cost per finished task rather than chasing the single highest IQ model, with rivals focusing on agent products for enterprise buyers and one major lab reportedly raising at a $965 billion valuation to keep up with demand. Meta’s Llama series, Mistral, DeepSeek, and tools like Ollama have already shown that open weights can pull developers away from closed APIs, especially once they need finely tuned systems for niches like legal review or customer-support analytics.

Thinking Machines is leaning into that trend for coding workflows and AI agents. Inkling posts strong scores on benchmarks like SWE-bench and Terminal Bench, and its agentic coding and general agent performance place it among the more capable open models for building software assistants and tools-driven agents. Deployment-ready checkpoints on Databricks and other partners make it easy to integrate into existing pipelines for orchestration, observability, and compliance. A key proof point comes from work with Bridgewater Associates, where researchers used Tinker to fine-tune an open model on specialized financial data and produced a lightweight system that beat leading proprietary alternatives on financial reasoning benchmarks at under 10% of the cost. For CIOs under budget pressure, that is the real benchmark that matters.

The Future: Open-Source Alternatives to Closed AI, With a Price Twist

Inkling positions itself squarely in the growing field of open-source alternatives ChatGPT users and enterprises are exploring, but with a twist in how money changes hands. It aligns with a world where companies do not want to rent intelligence forever; they want AI assets on their balance sheet, not somebody else’s revenue line. Open weights, multi-modal support for video and audio, and a tunable reasoning dial are all designed to make Inkling the starting point for proprietary internal models rather than the end product.

Thinking Machines frames this release as the first entry in a family of models, with larger and more capable versions expected over time and an Inkling-Small variant already matching or beating the main model on several benchmarks while it completes testing. “Thinking Machines says plainly that Inkling is not the strongest model available, closed or open,” yet it is betting that open-weights plus Tinker’s fine-tuning platform pricing can compete with frontier AI without API fees. If that bet pays off, Inkling will not have to outscore every closed model; it only has to be good enough that the savings from owning and tuning it swamp the convenience of renting a black box.

Milik earns a commission when you shop through our links, at no extra cost to you. This article was generated with AI from published sources and product data.

You May Also Like

Comments
Say something...
No comments yet. Be the first to share your thoughts!