Here are five costly mistakes businesses make when scaling AI, from runaway costs and weak governance to poor accountability plus how leaders can avoid them.
Dynatrace already shipped LLM-as-a-judge evaluation. Its $915 million Arize deal buys something else, a position with AI engineers before an app reaches production.
The pattern is the same everywhere: own data no one else has and sit as close as possible to the point where decisions are made.
Google is rolling out Gemini 3.7 Flash, a new version of its workhorse AI model that puts coding, agentic workflows and knowledge work at the center of the upgrade — while temporarily cutting API prices in half.
The release arrives just three weeks after the release of Gemini 3.6 Flash, an unusually short turnaround that Google attributes to developer feedback and algorithmic improvements.
For enterprise developers, the more consequential story may be the combination of those intelligence gains with lower inference costs: through the end of 2026, Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens.
Starting Jan. 1, 2027, pricing rises to $1.50 per million input tokens and $7.50 per million output tokens. That means the current discount is temporary, but it gives teams deploying high-volume coding and business agents several months to evaluate whether Google’s claimed reductions in retries and manual oversight translate into lower total operating costs.
The launch also underscores Google’s rapid iteration on its Flash line while its next flagship Pro model remains absent. Google did not provide a release date for Gemini 3.5 Pro with Thursday’s announcement, Reuters reported, despite the model having previously been described as undergoing partner testing. Axios similarly noted that 3.7 Flash arrives before the anticipated Pro release.
Google describes Gemini 3.7 Flash as its “most intelligent workhorse model yet for coding and agents.” The company says the model is better at adapting when it encounters roadblocks, clarifying intent when necessary and following instructions with greater fidelity.
Those improvements matter beyond benchmark scores. In an enterprise coding agent, a model that makes fewer unnecessary changes, recovers from errors and executes multi-step plans more reliably can reduce the number of human interventions needed to complete a task. The same principle applies to business agents operating across documents and applications, where an incorrect tool call or poorly interpreted instruction can derail an otherwise useful workflow.
Google says 3.7 Flash “thinks more diligently,” applying more effort to multi-step planning and tool calls. Its stated goal is more disciplined execution with fewer retries and less manual supervision.
That represents an interesting evolution from Gemini 3.6 Flash. Google’s developer documentation described 3.6 as reducing reasoning steps, conversational turns and tool calls compared with earlier models while attempting to limit execution-loop spiraling. With 3.7, the emphasis shifts toward putting sufficient effort into planning while improving the quality of execution — potentially a more useful optimization than simply minimizing the number of steps an agent takes.
Google DeepMind said in a post accompanying the release that 3.7 Flash shows gains in debugging and issue resolution, generates more functional web layouts and applications with fewer prompts, and improves reasoning and accuracy on real-world business workflows.
Google’s benchmarks show a large generational improvement in several software engineering tests.
On FrontierCode 1.1 Main, which measures production code quality, Gemini 3.7 Flash scores 43.6%, up from 34.4% for Gemini 3.6 Flash. That also narrowly exceeds the 42.7% Google reports for Claude Sonnet 5 and 41.3% for GPT-5.6 Terra.
On DeepSWE v1.1, a long-horizon software engineering evaluation, 3.7 Flash reaches 65.3%, compared with 49.0% for its predecessor. GPT-5.6 Terra remains ahead at 69.6% in Google’s table.
Web development shows another notable gain. Gemini 3.7 Flash receives an Elo score of 1588 on Code Arena, versus 1538 for 3.6 Flash, 1541 for Claude Sonnet 5 and 1523 for GPT-5.6 Terra. Google says the new model can produce more functional layouts and feature-complete applications in fewer prompts while more closely following reference screenshots, images and design systems.
The broader benchmark table is more mixed, which is important for enterprises evaluating the model against particular workloads rather than looking for a single “best” model.
Gemini 3.7 Flash scores 85.8% on Terminal-bench 2.1, compared with 87.4% for GPT-5.6 Terra. Terra also leads Google’s comparisons on Terminal-bench 3.0 and OSWorld-2.0. Claude Sonnet 5 leads the Agent’s Last Exam multimodal desktop and operating-system tasks with a 33.3% pass rate, versus 26.3% for Gemini 3.7 Flash.
In other words, Google’s own results do not show 3.7 Flash universally displacing higher-priced competitors. They instead suggest a model that has become substantially more competitive in coding and agent workloads while occupying a lower price tier.
The gains extend beyond software development.
On AutomationBench, which Google describes as measuring enterprise workflow automation, Gemini 3.7 Flash scores 30.4%, up sharply from 17.0% for 3.6 Flash. Google’s table lists Claude Sonnet 5 at 10.7% and GPT-5.6 Terra at 23.6%.
The model also reaches 34.0% on GDP.PDF, an evaluation of complex PDF comprehension, compared with 22.0% for 3.6 Flash, 28.0% for Claude Sonnet 5 and 24.7% for GPT-5.6 Terra.
That combination is relevant for enterprise agents because many practical deployments require more than generating text or code. An agent may need to interpret a long report, identify relevant information, decide which tool to invoke, update another system and produce a document for a human reviewer. Reliability across that chain can matter more than performance on an isolated reasoning benchmark.
Google is putting that thesis into practice with Gemini Spark. Google AI Pro and Ultra subscribers can use 3.7 Flash in Spark, the company’s personal AI agent. Google says the upgrade improves Spark’s knowledge work and tool use across Google Workspace applications, including workflows that consolidate files, draft emails and update status documents.
For enterprises, 3.7 Flash is also available through the Gemini Enterprise Agent Platform and Gemini Enterprise app.
Gemini 3.7 Flash’s introductory pricing is a notable bid to embed the model into enterprise workflows.
Until Dec. 31, developers pay $0.75 per million input tokens and $3.75 per million output tokens. Context caching costs $0.075 per million tokens during the introductory period. Google says standard prices will double on Jan. 1, 2027, to $1.50 for input and $7.50 for output, with context caching rising to $0.15.
For comparison, Gemini 3.6 Flash’s standard API pricing is $1.50 per million input tokens and $7.50 per million output tokens. Google’s benchmark table lists Claude Sonnet 5 at $2 and $10, respectively, while GPT-5.6 Terra is listed at $2 and $12.
The economics become more pronounced for autonomous agents because a single user request can produce a long sequence of model calls, reasoning tokens and tool interactions. A model that costs less per token but requires substantially more retries may not ultimately be cheaper.
Conversely, Google’s combination of lower introductory token pricing and claimed improvements in first-pass accuracy could materially change the cost of running high-volume coding or document-processing agents if those gains carry over to production.
That is the metric enterprise teams will ultimately need to test: not price per million tokens in isolation, but cost per successfully completed task.
Gemini 3.7 Flash arrives amid a broader debate over whether Google is losing ground at the AI frontier. The company has not released Gemini 3.5 Pro, despite saying in May that the flagship model would arrive the following month.
By July, Google said it remained in partner testing and would become broadly available when ready; Thursday’s announcement offered no further timetable. Google’s latest released general-purpose Pro model therefore remains Gemini 3.1 Pro, introduced in February.
Reuters reported in July that Gemini 3.5 Pro missed its original target after falling short of internal goals, particularly in coding, even as Google began training what it calls its most ambitious model yet, Gemini 4.
The delay coincides with a major overhaul of Google’s AI leadership announced last week.
Google DeepMind co-founder and Nobel Prize Winner Demis Hassabis has relinquished day-to-day control of the company’s famed DeepMind AI division to become its chair and, simultaneously, to take on the role of Alphabet’s chief scientist.
Meanwhile, former DeepMind CTO Koray Kavukcuoglu now runs the unit as a senior vice president reporting directly to CEO Sundar Pichai.
Kavukcuoglu controls Gemini model development, frontier research, the Gemini app and developer teams—effectively consolidating the full Gemini chain under a more product-focused operator.
Chief scientist Jeff Dean, Gemini co-lead Oriol Vinyals, Quoc Le and Sanjay Ghemawat left to establish the research startup Discovery Loop.
Those exits followed Gemini co-lead Noam Shazeer’s move to OpenAI and Nobel Prize-winning AlphaFold scientist John Jumper’s departure for Anthropic. Reuters reported that internal disagreements, constrained compute allocation and Google’s bureaucracy contributed to slower releases and weaknesses in coding.
Outside interpretations range from organizational repair to a more fundamental retreat.
SemiAnalysis has argued that Google is increasingly prioritizing the highly profitable business of supplying cloud infrastructure to AI companies—including Gemini competitors—over keeping its own models at the absolute frontier. That analysis also claimed Google had effectively canceled 3.5 Pro, although Google has not confirmed that and continues to describe the model as delayed.
The Verge offered a more measured assessment: the departures and model delays are serious, but Google retains enormous advantages through Search, Workspace, Android, Cloud, custom AI chips and consumer distribution. Google says the Gemini app has surpassed 950 million monthly users, giving it a reach that does not depend entirely on owning the highest-scoring model.
Current benchmarks similarly depict a company behind the overall leaders but still firmly competitive. Artificial Analysis places Claude Opus 5 at 63 on its overall model Intelligence Index, while Google reports a score of 56 for Gemini 3.7 Flash—an improvement from 52 for 3.6 Flash but not a return to the top.
Arena’s early human-preference results are more favorable, provisionally ranking 3.7 Flash ninth overall and eighth for web development.
The resulting picture is not that Google has abandoned advanced AI, but that it has become stronger at rapidly shipping efficient Flash models while struggling to deliver the premium flagship required to reclaim broad leadership. Gemini 4 will now serve as the clearest test of whether the leadership reorganization fixes that execution gap.
Developers can access Gemini 3.7 Flash through the Gemini API in Google AI Studio and Android Studio, as well as Google’s Antigravity environment. Enterprises can deploy it through Gemini Enterprise Agent Platform and Gemini Enterprise, while consumers with Google AI Pro or Ultra subscriptions can access the model through Spark in supported countries.
Google is also shipping updated safeguards covering chemical, biological, radiological and nuclear risks and cyber-offense misuse, according to the company.
The unusually fast jump from Gemini 3.6 Flash to 3.7 Flash points toward a model development cycle in which algorithmic improvements can reach production products without waiting for a new flagship generation. Ars Technica also highlighted the three-week interval between the two releases, while Google says the techniques behind the update will inform future models.
For developers, that faster cadence creates its own operational question. Models can improve quickly, but production teams still have to benchmark new releases against their own repositories, prompts, tool schemas and failure modes before changing a deployment.
Gemini 3.7 Flash gives those teams a particularly strong incentive to run that evaluation. Google’s own numbers show major improvements in production coding, web development, document comprehension and workflow automation without claiming leadership everywhere. At its introductory price, Google is effectively betting that developers will value a model that is competitive enough with more expensive systems while being cheap enough to run repeatedly inside agents.
Whether that advantage survives the return to full pricing in January will depend less on leaderboard positions than on how reliably 3.7 Flash completes real work.
Regardless of which technique(s) leaders pursue, they should do so in a structured, measured way.
DeepSeek is expanding beyond the model layer and deeper into the software developers use to put AI agents to work.
The Chinese AI lab on Thursday launched the official version of DeepSeek-V4-Pro, an updated flagship model focused heavily on agentic workloads, alongside DeepSeek Harness v0.1, a new open-source agent harness that gives developers an alternative to integrated coding-agent environments such as Anthropic’s Claude Code.
Together, the releases amount to a broader developer push from DeepSeek. V4-Pro is now available across DeepSeek’s web interface, mobile app and API, with native support for the OpenAI Responses API and integration with Codex.
DeepSeek Harness, meanwhile, is entering developer preview under the MIT license and the code is available now for download and use on GitHub. It’s built around an unusually modular premise: practically every part of the agent runtime can be swapped out as a plugin.
But developers accessing V4 through DeepSeek’s API will soon pay considerably more for it. DeepSeek is simultaneously abandoning its existing flat API pricing in favor of peak and off-peak rates beginning at 16:00 UTC on Sunday, Aug. 16 (2 am ET).
Even the discounted off-peak cache-miss and output prices will be substantially higher than the prices available today.
The combination is significant because DeepSeek is no longer competing solely over model intelligence and token prices. With Harness, it is moving into the layer that determines how models use tools, manipulate files, maintain sessions and execute long-running agent workflows — territory where Anthropic’s Claude Code and other coding agents have become increasingly important developer products.
DeepSeek describes Harness, or dsh, as an open-source agent harness built on Cordis, a framework designed around composable plugins.
Its guiding principle is simple: “Everything is a plugin.”
That extends to models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration and user interfaces, according to DeepSeek. Rather than making those components fixed pieces of a single coding agent, Harness is designed to let developers mix, replace and extend them.
The project is available under the MIT license and can currently be launched from npm with npx @deepseek-ai/dsh web. DeepSeek also provides instructions for building it directly from source. The repository describes the software explicitly as a developer preview and warns that “THERE WILL BE COMPATIBILITY-BREAKING CHANGES.”
That caveat matters for enterprise developers. Harness is not yet being presented as a stable drop-in production platform. But its architecture points toward a potentially important strategy: DeepSeek can now offer developers not only models but an open framework for assembling the systems that surround them.
That makes Anthropic’s Claude Code and OpenAI’s Codex useful competitive references, although the products should not be treated as functionally identical.
DeepSeek Harness is an open-source, model-agnostic alternative to the agent infrastructure underlying Claude Code and Codex—not yet a full replacement for either product’s broader developer experience.
It can already inspect repositories, edit files, execute shell commands, search files and the web, maintain plans, invoke skills, delegate work to subagents and enforce approval policies. Those are the essential capabilities that make Claude Code and Codex agentic coding tools rather than autocomplete systems.
DeepSeek explicitly describes Standard mode as a full coding agent with file editing, shell access, search, planning, subagents and workflows. Its local web interface lets users select a workspace and approve sensitive operations.
But Claude Code and Codex now extend well beyond that agent loop. Here’s a quick comparison:
|
Dimension |
DeepSeek Harness |
Claude Code |
OpenAI Codex |
|
Read, edit and test a repository |
Yes |
Yes |
Yes |
|
Shell and development tools |
Yes |
Yes |
Yes |
|
Planning and subagents |
Yes |
Yes |
Yes |
|
Permission controls and sandboxing |
Yes, configurable through plugins |
Yes, mature built-in permission and sandbox system |
Yes, granular sandbox and approval controls |
|
Primary interfaces |
Local web UI; headless command; Python SDK |
Terminal, VS Code, JetBrains, desktop, browser, mobile and Slack |
CLI, IDE extension, desktop app, web/cloud and integrations |
|
Hosted background agents |
Not documented as a DeepSeek-managed service |
Yes |
Yes |
|
GitHub-native PR workflow |
Not documented as a finished integration |
GitHub Actions, automatic reviews, issue-to-PR workflows |
Cloud tasks, automatic reviews, PR fixes and GitHub Action |
|
Model choice |
DeepSeek, Anthropic, OpenAI and custom compatible endpoints |
Primarily Claude, including Bedrock, Google Cloud and Microsoft hosting |
Primarily OpenAI models, with configurable providers in the open-source CLI |
|
Extensibility |
Exceptional: virtually every component is replaceable |
Strong: skills, hooks, MCP, plugins and agent teams |
Strong: skills, MCP, custom agents, SDK and app server |
|
Product maturity |
Developer preview; breaking changes expected |
Established commercial product |
Established commercial product plus open-source CLI |
|
License |
MIT |
Commercial product with extensibility interfaces |
Codex CLI is open source; cloud and app services are managed products |
DeepSeek Harness instead emphasizes modularity and replacement: the model itself is another plugin rather than necessarily the center of a vertically integrated stack.
DeepSeek’s repository was already attracting significant developer attention on launch day, showing roughly 27,500 GitHub stars and 2,000 forks as of Aug. 13, although those rapidly changing figures are best viewed as a snapshot rather than an adoption metric.
Harness arrives alongside the general-availability release of DeepSeek-V4-Pro-0813.
DeepSeek originally introduced the V4 family in preview in April. The lineup consists of the 1.6-trillion-parameter V4-Pro, with 49 billion parameters activated per token, and the smaller 284-billion-parameter V4-Flash, with 13 billion activated. Both support context windows of up to one million tokens.
The company’s Aug. 13 release therefore is not the first appearance of V4-Pro. It is the transition from the earlier preview into an updated official version, with DeepSeek emphasizing agent performance.
“The official version of DeepSeek-V4-Pro has been released, featuring significantly enhanced agent capabilities and support for the Responses API and Codex integration,” DeepSeek says on its API website. “It is now fully available across the web, mobile app, and API; we welcome your testing and feedback.”
DeepSeek’s changelog similarly says the general-availability model has “significantly enhanced Agent capabilities,” particularly in production environments. Developers using the API do not have to change model identifiers: deepseek-v4-pro now resolves to the latest V4-Pro version.
The company has also added native OpenAI Responses API support, lowering the amount of integration work required for applications already built around that interface.
DeepSeek says V4-Pro is optimized for OpenAI’s own open source harness, Codex, with one-click setup. Its current API documentation lists Responses API, tool calling, JSON output and an Anthropic-format API among the supported interfaces for both V4-Pro and V4-Flash.
For developers using DeepSeek directly rather than through an API, V4-Pro is now accessible through “Expert Mode” on the company’s app and website.
DeepSeek is also making reasoning effort an explicit control across V4-Pro and V4-Flash.
The V4 model documentation describes three levels: Non-think, designed for fast routine tasks; Think High, intended for more complex problem-solving and planning; and Think Max, which allocates substantially more reasoning to difficult problems.
That distinction can be operationally important for agent systems because maximum reasoning on every step can consume unnecessary time and tokens. A coding agent might use relatively little reasoning to inspect a file or execute a routine tool call, then increase effort when diagnosing a difficult bug or planning a multi-stage code change.
DeepSeek’s latest benchmark table suggests the 0813 model improves substantially on agent-oriented tests, although the figures are company-reported and some results depend on the harness configuration.
DeepSeek reports V4-Pro-0813 scores of 87.9 on Terminal Bench 2.1, 74.1 on Toolathlon-Verified, 71.1 on DSBench-FullStack and 67.2 on DSBench-Hard. It does not lead every comparison in DeepSeek’s own table: Fable 5, for example, scores 77.9 on Toolathlon-Verified and 77.2 on DSBench-FullStack.
There is an especially important qualification buried beneath the benchmark table. For public Code Agent tasks, DeepSeek says V4-Pro-0813 was tested using its upcoming DeepSeek Harness in “minimal mode.”
In other words, some of the agent results arriving alongside Harness are not purely model benchmarks. They measure the model operating inside an agent execution environment — precisely the software layer DeepSeek is now releasing to developers.
The bigger immediate change for teams already running DeepSeek in production may be pricing.
DeepSeek’s current API documentation lists V4-Flash at $0.14 per million cache-miss input tokens and $0.28 per million output tokens, while V4-Pro costs $0.435 for cache-miss input and $0.87 for output. Cache hits are dramatically cheaper at $0.0028 for Flash and $0.003625 for Pro.
Those prices themselves represented a major reduction from V4’s original April launch economics. When V4 arrived in April, V4-Pro was priced at $1.74 per million cache-miss input tokens and $3.48 per million output tokens. By late May, DeepSeek had made a 75% reduction permanent, intensifying its position as an unusually inexpensive option for high-volume agent workloads. Now the pendulum is moving in the other direction.
Beginning Aug. 16 at 16:00 UTC, DeepSeek will charge different rates depending on when API calls occur. Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC (9:00 PM – 12:00 AM ET and 2:00 AM – 6:00 AM ET, respectively) with all other hours classified as off-peak. Off-peak rates are half the corresponding peak prices.
For V4-Flash, off-peak cache-miss input rises from $0.14 to $0.22 per million tokens, while output rises from $0.28 to $0.66. During peak hours those rates reach $0.44 input and $1.32 output.
V4-Pro moves from $0.435 per million cache-miss input tokens and $0.87 output today to $0.66 and $1.98 off-peak, respectively. Peak rates rise to $1.32 input and $3.96 output.
The increases are even more pronounced for cached input. V4-Pro cache hits rise from $0.003625 per million tokens today to $0.022 off-peak and $0.044 at peak. Flash moves from $0.0028 to $0.007 off-peak and $0.014 peak.
|
Model |
Old input (per 1M token) |
Old output (per 1M tok) |
Old total (1M in/1M out) |
|
deepseek-v4-flash |
$0.14 |
$0.28 |
$0.42 |
|
deepseek-v4-pro |
$0.435 |
$0.87 |
$1.305 |
The new prices still position DeepSeek as an affordable alternative via API to Western proprietary labs, but Reuters reported Thursday that, depending on model, token category and time of use, the changes represent increases ranging from 50% to more than 1,100% over existing rates.
|
Model |
Input ($/1M) |
Output ($/1M) |
Total ($/1M) |
Source |
|
Muse Spark 1.2 Contributor |
$0.10 |
$0.20 |
$0.30 |
|
|
MiMo-V2.5 Flash |
$0.10 |
$0.30 |
$0.40 |
|
|
DeepSeek-V4-Flash — off-peak |
$0.22 |
$0.66 |
$0.88 |
|
|
GPT-5.6 Luna |
$0.20 |
$1.20 |
$1.40 |
|
|
MiniMax-M3 |
$0.30 |
$1.20 |
$1.50 |
|
|
LongCat-2.0 — limited-time promo |
$0.30 |
$1.20 |
$1.50 |
|
|
DeepSeek-V4-Flash — peak hours |
$0.44 |
$1.32 |
$1.76 |
|
|
MiMo-V2.5 |
$0.40 |
$2.00 |
$2.40 |
|
|
DeepSeek-V4-Pro — off-peak |
$0.66 |
$1.98 |
$2.64 |
|
|
LongCat-2.0 — standard |
$0.75 |
$2.95 |
$3.70 |
|
|
MiMo-V2.5 Pro (≤256K) |
$1.00 |
$3.00 |
$4.00 |
|
|
DeepSeek-V4-Pro — peak hours |
$1.32 |
$3.96 |
$5.28 |
|
|
Muse Spark 1.1 / 1.2 |
$1.25 |
$4.25 |
$5.50 |
|
|
GLM-5.2 |
$1.40 |
$4.40 |
$5.80 |
|
|
Grok 4.6 — <200K prompt tokens |
$2.00 |
$6.00 |
$8.00 |
|
|
MiMo-V2.5 Pro (>256K) |
$2.00 |
$6.00 |
$8.00 |
|
|
Qwen3.8-Max |
$2.00 |
$6.00 |
$8.00 |
|
|
Gemini 3.6 Flash |
$1.50 |
$7.50 |
$9.00 |
|
|
GPT-5.6 Terra |
$2.00 |
$12.00 |
$14.00 |
|
|
Grok 4.6 — ≥200K prompt tokens |
$4.00 |
$12.00 |
$16.00 |
|
|
GPT-5.4 |
$2.50 |
$15.00 |
$17.50 |
|
|
Kimi K3 |
$3.00 |
$15.00 |
$18.00 |
|
|
Claude Opus 5 |
$5.00 |
$25.00 |
$30.00 |
|
|
Sakana Fugu Ultra (≤272K) |
$5.00 |
$30.00 |
$35.00 |
|
|
GPT-5.6 Sol — Standard mode |
$5.00 |
$30.00 |
$35.00 |
|
|
Claude Fable 5 / Claude Mythos 5 |
$10.00 |
$50.00 |
$60.00 |
|
|
GPT-5.6 Sol — Fast mode |
$10.00 |
$60.00 |
$70.00 |
That makes the “50% lower” off-peak framing potentially misleading without context. Off-peak is 50% cheaper than DeepSeek’s new peak rate; it is not a 50% discount from the API prices developers are paying today.
For a simple workload consisting of one million cache-miss input tokens plus one million output tokens, V4-Pro currently costs $1.305. The same token mix will cost $2.64 off-peak, roughly twice as much, or $5.28 during peak hours, more than four times the current price.
V4-Flash moves from $0.42 under the same simple calculation to $0.88 off-peak and $1.76 peak.
Actual application costs will vary considerably depending on the ratio of cached input, uncached input and generated output, making those combined figures illustrative rather than universal total-cost estimates.
The timing makes the strategic direction difficult to miss.
When DeepSeek released the V4 preview on April 24, the major story was how much frontier-class capability the company could deliver with an unusually efficient architecture.
V4-Pro uses a hybrid attention design combining Compressed Sparse Attention and Heavily Compressed Attention; at a one-million-token context, DeepSeek says it requires only 27% of the single-token inference FLOPs and 10% of the KV cache required by V3.2.
By late May, the discussion had shifted toward what those efficiencies meant economically for high-volume agents, whose repeated context reads can make caching a major component of inference costs. DeepSeek’s steep V4 price cuts amplified that advantage.
The Aug. 13 releases move the competition another layer upward.
DeepSeek now has an updated V4-Pro tuned around agent workloads, standardized interfaces designed to make it easier to connect with existing developer tooling, configurable reasoning effort, and an MIT-licensed harness for controlling the models, tools, sandboxes, filesystems and orchestration surrounding an agent.
At the same time, DeepSeek is demonstrating that developers cannot assume its aggressively low API rates are permanent. For organizations considering the platform, workload scheduling, caching behavior and the option to run open weights on their own infrastructure now become more important parts of the total-cost calculation.
That leaves DeepSeek pursuing two potentially conflicting advantages at once: making its agent stack more accessible and open while making its own hosted API considerably more expensive.
For enterprise developers, Harness may ultimately be the more consequential part of Thursday’s announcement. Models can increasingly be swapped behind standardized interfaces. The harness that controls how an agent reasons, invokes tools, edits software and persists across a workflow can be much harder to replace.
DeepSeek is now competing for that layer, too.
Scoring the control families in your quantum readiness assessment is not a rating of how prepared you feel; it is a measure of where evidence exists and gaps remain.
Modern ringfencing operates across three distinct levels of control, moving from basic visibility to proactive prevention.
We are running sophisticated AI systems that help humans make better decisions about tasks that, in many cases, humans shouldn’t be touching at all.
Age and inefficiency alone do not make a process legacy. A process becomes legacy when its design assumptions no longer reflect what is possible.