DeepSeek's V4 promises substantially lower inference costs — but independent tests report uneven results on creativity, code generation and consistency. The Hangzhou startup has previewed V4 in two flavours: a 1.6 trillion‑parameter Pro and a 284 billion‑parameter Flash, both offered under an open‑source licence and tuned for long‑context and agent‑style tasks.

Release and the numbers

DeepSeek published a V4 preview more than a year after its R1 release. The family comes in two flavours: V4 Pro (public reporting lists it at about 1.6 trillion parameters) and V4 Flash (around 284 billion parameters). DeepSeek maintains an open-source approach — developers can download, run and, in many cases, modify the models locally.

Performance: lower running costs, mixed outputs

Cost is the clearest selling point. Neil Shah, vice-president of research at Counterpoint Research, called the V4 preview "a serious flex" and highlighted lower inference costs than earlier models. Wei Sun, principal AI analyst at Counterpoint Research, said benchmark profiles suggest "excellent agent capability at significantly lower cost."

At the same time, public and independent testing has surfaced limits. Early stress tests reported by third parties flagged inconsistent outputs on tasks requiring creativity, nuanced reasoning or extreme precision. A technical write-up noted V4 lagging some established rivals on code generation and select reasoning benchmarks; the Pro build's large parameter count did not translate into uniform superiority across all task types.

Two-tier strategy: where each model fits

  • V4 Pro (1.6T): pitched for demanding STEM, code generation and workflow automation where extra capacity helps.
  • V4 Flash (284B): designed as a lower-cost alternative for routine workflows, local deployment and edge or prototype use cases.

The split gives enterprises a choice between higher-capacity, lower per-query compute for agent chains and a smaller footprint for cost-sensitive or on-premises deployments. The open-source licence also enables fine-tuning and internal hosting.

Developer and ecosystem implications

  • Open-source distribution can lower cloud API bills by shifting workloads to local or hybrid infrastructure.
  • Reducing vendor lock-in: engineering teams can fork, optimise and host their own stacks.
  • Technical risk remains: reviewers found V4 needs refinement on creative and precise reasoning tasks, which matters where incorrect outputs carry business consequences.

Why this matters

If V4 delivers sustained reductions in inference costs, organisations could shift more machine‑learning workloads from cloud APIs to on‑prem or hybrid deployments. That would force cloud providers and API vendors to rethink pricing and support for open models, and accelerate efforts by firms to control their own model stacks and costs.

Related Articles

Enterprises and investors must weigh the 1.6 trillion‑parameter Pro's lower per‑query costs and the open‑source flexibility of both builds against early tests showing gaps on demanding code and reasoning tasks.

This article was created with AI assistance.