MilikMilik

The 10 Cheapest AI Models Right Now and How Pricing Is Reshaping the Market

The 10 Cheapest AI Models Right Now and How Pricing Is Reshaping the Market
Interest|High-Quality Software

What “cheapest AI models” really means today

The cheapest AI models are language and reasoning systems whose low token prices and efficient architectures let developers run large workloads at minimal cost without losing baseline capability for everyday tasks. Cheap no longer means weak: a wave of open-weight releases, trimmed active parameters, and long-context designs has lowered AI token costs by an order of magnitude while still matching mid‑2020s frontier performance on many workloads. Instead of paying premium rates for email drafting, document summarisation, or lightweight agents, teams can now reach for budget AI models that keep per‑token spend low and throughput high. The result is a market where affordable AI pricing is the main competitive battleground, and where choosing the right pricing tier often matters more than chasing the absolute highest benchmark score for most production applications.

Blended pricing, cache hits, and the real cost of AI tokens

Headline input and output token prices can be misleading, because they ignore how often cached results are reused. Artificial Analysis approaches this by ranking the cheapest AI models using a blended price that assumes a 7:2:1 ratio of cache-hit, input, and output tokens. That ratio reflects common production patterns: a large share of requests repeat context, a smaller share adds fresh prompts, and an even smaller share demands long generations. When you compare AI token costs on this blended basis, you see the true cost of ownership instead of marketing numbers. A model that looks slightly more expensive on raw input price can win once cache hits are counted, especially for document-heavy or repetitive workflows. For buyers, the lesson is clear: evaluate blended rates, not just per‑million input or output prices, when judging budget AI models.

The 10 cheapest AI models and what you get for the price

Pricing data from Artificial Analysis shows the current low-price leaders. DeepSeek V4 Flash tops the chart at a blended USD 0.06 (approx. RM0.28) per million tokens while still offering a 1 million token context window and dual Thinking/Non‑Thinking modes. OpenAI’s GPT‑OSS‑20B sits close behind at USD 0.07 (approx. RM0.32) per million tokens, trading depth for throughput but excelling at high‑volume summarisation and routing. DeepSeek V4 Pro and MiMo‑V2.5‑Pro are both priced at USD 0.18 (approx. RM0.83) per million tokens, pushing near‑frontier reasoning into the budget AI models tier. GPT‑OSS‑120B follows at USD 0.20 (approx. RM0.92) per million tokens, in the range where mid‑tier reasoning and multi‑step agents become viable without premium rates. “The AI pricing war is well and truly over — and developers won,” as Artificial Analysis puts it.

Why cheap models are winning real workloads

OpenRouter’s token consumption leaderboard shows how fast cost‑efficient models are spreading through real applications. DeepSeek V4 Flash leads with 10.9 trillion tokens and nearly 10x month‑on‑month growth, reflecting heavy deployment in production pipelines that prioritise speed and low AI token costs. Tencent’s Hy3 Preview, priced at USD 0.063 (approx. RM0.29) per million input tokens and reaching 67.1% on the BrowseComp benchmark, went from almost zero to near‑parity with Flash in a single month, driven by agentic and long‑context use cases. At the same time, models like Claude Opus 4.7 and Claude Sonnet 4.6 retain strong mid‑tier adoption, suggesting teams route simple workloads to cheaper engines while reserving premium models for complex reasoning. Token consumption patterns make one thing clear: once affordable AI pricing reaches “good enough” quality, volume shifts quickly toward those options.

How to choose the right budget AI model for your use case

Selecting among the cheapest AI models starts with workload profiling. For output‑heavy tasks like document generation, DeepSeek V4 Flash’s USD 0.06 (approx. RM0.28) blended rate can unlock major savings, so long as you can tolerate its tendency to hallucinate when it lacks knowledge. For high‑volume classification or summarisation, GPT‑OSS‑20B’s smaller size and USD 0.07 (approx. RM0.32) rate make it appealing when latency and scale matter more than raw intelligence. If you need strong reasoning without frontier prices, DeepSeek V4 Pro, MiMo‑V2.5‑Pro, or GPT‑OSS‑120B balance cost and depth around the USD 0.18–0.20 (approx. RM0.83–0.92) band. Map each workflow to a cost‑capability tier, account for cache‑hit ratios, and reserve premium closed‑source models for the narrow slice of problems where their extra performance translates into clear business value.

Milik earns a commission when you shop through our links, at no extra cost to you. This article was generated with AI from published sources and product data.

You May Also Like

Comments
Say something...
No comments yet. Be the first to share your thoughts!