# DeepSeek raised prices 1,100% and nobody flinched

URL: https://www.thedeepfeed.ai/posts/2026-08-16-deepseek-raised-prices-and-nobody-flinched/
Category: Business
Published: 2026-08-16
Author: the-deep-feed
Tags: deepseek, pricing, api-economics, open-weights, price-war, ipo, china-ai
Kind: deep

> DeepSeek's peak/off-peak billing takes effect today: V4-Pro output up 4.5x at peak, cache-hit input up as much as 12x. It lands a fortnight after OpenAI cut Luna 80%. The price war was never a race to zero — it is convergence toward the true marginal cost of intelligence, approached from both directions at once.

## TL;DR

- **DeepSeek's new peak/off-peak billing takes effect today**: V4-Pro output goes from a flat $0.87 to **$1.98 off-peak and $3.96 at peak** per million tokens, and cache-hit input rises **as much as 12x** — an increase of roughly 1,100%. Reuters puts V4-Pro at up to **14x the price of V4 Flash**.
- It lands a fortnight after **OpenAI cut Luna 80%** and three days after **Google introduced Gemini 3.7 Flash at half price**. The cheapest player raised while the expensive ones cut: the price war is **convergence**, not a race to zero.
- The contrarian read: Chinese AI was never structurally cheap. The discount was **strategic** — customer acquisition — and DeepSeek just ended it. Bloomberg ties the hike to **IPO preparation**; CIO ties it to **capacity strain**. Both readings agree the subsidy was a choice.
- The same week it raised serving prices, DeepSeek put **V4-Pro-0813's weights on Hugging Face under MIT**. That is not a contradiction — it is the playbook: [give away the weights, charge for the serving](/posts/2026-08-12-the-toll-booth-on-the-open-road/).

Today the meters switch over. DeepSeek's new billing regime, announced on August 13 alongside the general-availability release of V4-Pro and effective this morning, replaces the flat per-token rates that made the company the reference point for cheap intelligence with peak and off-peak tiers. The headline number: V4-Pro output, which cost $0.87 per million tokens under the old card, now costs $1.98 off-peak and $3.96 at peak, per the [official pricing page](https://api-docs.deepseek.com/quick_start/pricing/). Cache-hit input rises by as much as a factor of twelve — an increase of roughly 1,100%. [Reuters](https://www.reuters.com/world/china/deepseek-releases-official-v4-pro-model-it-steps-up-expansion-2026-08-13/) notes that V4-Pro is now priced at up to fourteen times V4 Flash, the company's budget line.

Hold that next to the other side of the board. Seventeen days ago, OpenAI [cut GPT-5.6 Luna's price by 80%](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/) and Terra's by 20%, three weeks after launch. Three days ago, Google shipped Gemini 3.7 Flash with a [50% introductory discount](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/). In the space of one fortnight, the most expensive American labs slashed prices while the cheapest Chinese lab raised them by multiples. If the price war were a race to zero, this sequence would be incoherent.

It is not incoherent. It is convergence. The American cuts and the Chinese hikes are the same event observed from opposite ends: prices moving toward the true marginal cost of serving frontier-adjacent intelligence, a number that was always somewhere between OpenAI's old premium and DeepSeek's old giveaway. We argued in July that [the margin is the message](/posts/2026-07-28-the-margin-is-the-message/) — that pricing moves are the most honest disclosures labs make. DeepSeek just made the most honest disclosure of the year, and the market's reaction, so far, has been close to silence.

![Chart schematic: OpenAI's −80% Luna cut descending and DeepSeek's +1,100% hike ascending toward one band — the true marginal cost.](/post-images/2026-08-16-deepseek-raised-prices-and-nobody-flinched/convergence-band.jpg)

# The rate card that ended an era

The mechanics deserve precision, because "raised prices 1,100%" compresses a more textured change. What DeepSeek actually did, per its [change log](https://api-docs.deepseek.com/updates/) and the breakdowns compiled by [Codersera](https://codersera.com/blog/deepseek-v4-price-change-august-2026/), is three things at once.

First, it introduced time-of-day pricing. Peak-hour rates run roughly double off-peak, which means the effective price of a token now depends on when your agent happens to run. For batch workloads that can shift to off-peak windows, the damage is smaller; for latency-sensitive production traffic, the peak rate is the real rate.

Second, it repriced the cache. Cache-hit input, the mechanism that made long agentic sessions nearly free on DeepSeek because repeated context cost a fraction of a fresh read, rises by up to 12x. This is the stealthy part of the hike and arguably the most consequential: agent workloads are overwhelmingly cache-heavy, so the segment that grew fastest on the old card is the segment that pays most under the new one.

Third, it stratified the lineup. V4-Pro at up to 14x the price of V4 Flash, per Reuters, is not an accident of arithmetic. It is segmentation — a premium tier and a budget tier where a single cheap tier used to be. That is what a business does when it starts caring about revenue per customer rather than customers per dollar.

| DeepSeek V4-Pro | Old (flat) | New off-peak | New peak |
|---|---|---|---|
| Output ($/M tokens) | $0.87 | $1.98 | $3.96 |
| Cache-hit input | baseline | up to 12x | up to 12x |
| vs. V4 Flash | — | up to 14x Flash (Reuters) | up to 14x Flash |

The company that trained the West to assume Chinese inference was permanently, structurally cheap has published a rate card that says otherwise.

# Eighteen months of the meter

The August card is not DeepSeek's first pricing move. It is the fifth act of a documented sequence, and reading the acts in order changes what the finale means. The company has been experimenting with the price of intelligence, in public, since the week it became famous.

Act one: on January 20, 2025, [DeepSeek released R1](https://api-docs.deepseek.com/news/news250120) under MIT, priced at $0.55 per million input tokens and $2.19 output on its own API, roughly a twenty-fifth of what OpenAI charged for o1 at the time. That card triggered the year of assumed-cheap Chinese inference, a $600 billion single-day drawdown in Nvidia's market value, and every procurement decision that followed.

Act two: on February 26, 2025, DeepSeek [introduced off-peak discounts of up to 75%](https://www.reuters.com/technology/chinas-deepseek-cuts-off-peak-pricing-by-up-75-2025-02-26/) for API calls made between 12:30am and 8:30am Beijing time. [SCMP's headline captured the mechanism](https://www.scmp.com/tech/tech-trends/article/3300264/ai-night-chinas-deepseek-offers-peak-75-discount-demand-strains-servers): "AI by night," a discount designed to push traffic into the hours when servers sat idle, because demand was already straining capacity a month after R1 shipped.

Act three: on March 1, 2025, the company published its serving economics, claiming a [theoretical cost-profit ratio of 545% per day](https://www.reuters.com/technology/chinas-deepseek-claims-theoretical-cost-profit-ratio-545-per-day-2025-03-01/) on V3 and R1 inference, with daily GPU costs of about $87,000 against theoretical revenue of $562,000. The asterisk was load-bearing: actual revenue was "significantly lower," the company said, because the web and app products were free, V3 was priced below R1, and the off-peak discounts applied. But the disclosure established a fact nobody has retracted since: at scale, on export-controlled hardware, DeepSeek's serving stack cleared its marginal costs with room to spare.

Act four: on August 21, 2025, the [V3.1 release](https://api-docs.deepseek.com/news/news250821/) quietly ended the off-peak program and unified V3 and R1 into one model with one card. The night discount, born of capacity strain, died in a restructure that raised effective prices for the developers who had scheduled around it. Almost nobody covered it as a price increase. It was one.

Act five is August 16, 2026, and the pattern across the five acts is now legible. Time-of-day pricing is not an innovation of this week's card; it is a tool DeepSeek reached for eighteen months ago, retired, and has now reinstated at double scale, with the discount replaced by a surcharge. Every act responded to the same underlying variable, capacity, and each act moved the price closer to whatever the capacity actually costs. The 2025 acts rationed scarce GPUs with discounts. The 2026 act rations them with margins.

| Date | Move | Direction |
|---|---|---|
| Jan 20, 2025 | R1 launches at $0.55/$2.19 per M | The floor is set |
| Feb 26, 2025 | Off-peak discounts up to 75% (12:30am–8:30am Beijing) | Down, off-peak only |
| Mar 1, 2025 | 545% theoretical margin disclosure ($87K/day GPU cost) | Transparency, not price |
| Aug 21, 2025 | V3.1 unifies V3/R1, ends off-peak discounts | Up, quietly |
| Aug 16, 2026 | Peak/off-peak returns as surcharge; cache-hit up to 12x | Up, loudly |

![Five-act DeepSeek pricing timeline from R1 launch to the August 2026 hike; the complaint line below stays flat near zero](/post-images/2026-08-16-deepseek-raised-prices-and-nobody-flinched/eighteen-months-meter.jpg)

# Two directions, one destination

Now widen the frame to the whole fortnight, because DeepSeek's hike only reads correctly against the moves around it.

| Date | Mover | Move | Direction |
|---|---|---|---|
| Jul 30 | OpenAI GPT-5.6 Luna | $1.00/$6.00 → $0.20/$1.20 per M | −80% |
| Jul 30 | OpenAI GPT-5.6 Terra | across the board | −20% |
| Aug 12 | SpaceXAI Grok 4.6 | headline $2/$6 held; cached input | +67% (quiet) |
| Aug 13 | Google Gemini 3.7 Flash | $0.75/$3.75 intro through Dec 31, then doubles | −50%, temporary |
| Aug 16 | DeepSeek V4-Pro | flat → peak/off-peak; output $0.87 → $1.98/$3.96; cache-hit up to 12x | +127% to +1,100% |

Read the table as one document and the pattern is unmistakable. The cutters are cutting from above, and their cuts come with asterisks: Gemini 3.7 Flash's 50% discount [expires December 31 and then doubles](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/); Grok 4.6 held its headline rate while [quietly raising cached-input prices 67%](https://ccleaks.com/news/grok-4-6-launch-benchmarks-pricing-aug-2026) and confining the $2/$6 rate to contexts under 200K tokens. The riser is rising from below. Everyone is walking toward the same band.

There is a second asymmetry worth naming. OpenAI's cut was funded by engineering: the company says [Sol rewrote its own inference kernels](/posts/2026-07-30-the-model-that-cut-its-own-price/), taking serving costs down 20%, and the Luna cut explicitly passed those gains through. That is a price reduction backed by a cost reduction — sustainable by construction. DeepSeek's old prices, we now know, were not backed by costs at all. They were backed by strategy, and the strategy changed.

# Strategic, not structural

Here is the assumption the new rate card breaks. For eighteen months, Western coverage has treated Chinese AI pricing as a fact of geology: cheaper labor, cheaper power, state subsidy, efficient architectures — pick your explanation, the conclusion was always that the discount was permanent. Procurement teams built multi-model strategies on it. The phrase "cheap Chinese models" stopped being an observation and became a category.

DeepSeek just demonstrated that the discount was a decision, and decisions get reversed. Two readings of why are circulating, and they are worth holding together rather than choosing between.

The first is Bloomberg's: [the hikes are IPO preparation](https://www.bloomberg.com/news/articles/2026-08-13/deepseek-increases-prices-for-ai-services-by-multiple-times). The context makes this hard to dismiss. Anthropic's CFO began [early investor meetings this week](https://www.cnbc.com/2026/08/13/anthropic-cfo-early-ipo-meetings-valuation.html), with investors reportedly expecting an October float at a valuation north of $2 trillion. Moonshot is restructuring for a Hong Kong listing with a pre-IPO round targeting roughly $50 billion. The lab-economics window is open on both sides of the Pacific, and a prospectus wants margins, not market share purchased at a loss. A company priced for growth can subsidize forever; a company priced for a listing has to show what customers actually pay when the subsidy stops. August 16 is that disclosure.

The second is CIO's: [demand is straining capacity](https://www.cio.com/article/4209473/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity-3.html), and the peak/off-peak structure is load management by price signal. This reading has its own supporting evidence — you do not invent time-of-day tiers unless your problem is concentrated at certain times of day, and export controls make adding accelerators harder for DeepSeek than for anyone it competes with.

Notice that the two readings converge on the same conclusion. Whether the motive is a balance sheet being dressed for public markets or a GPU fleet running out of headroom, the old prices were below the level the business can sustain. That is what strategic-not-structural means in practice: the subsidy did its job — it acquired the customers, set the reference price for the industry, and forced the [80% cut out of OpenAI](https://www.scmp.com/tech/tech-trends/article/3362568/openai-blinks-face-chinese-rivals-drops-pricing-some-models-80) that SCMP framed as a blink. Then it ended, on three days' notice.

# Why serving a 1.7-trillion-parameter model has a floor

The word "capacity" is doing a lot of work in the CIO reading, so it is worth opening the machine and looking at where the floor actually comes from. The economics of serving a large mixture-of-experts model are not a black box; DeepSeek itself published them in unusual detail during its open-infra week in early 2025, and independent engineering analyses have filled in the rest.

Start with the structural fact: a sparse MoE model is cheap per token only when it is busy. V4-Pro, like V3 before it, activates a small fraction of its 1.7 trillion parameters per token, but the whole parameter set must be resident in accelerator memory, sharded across dozens of GPUs, with expert layers scattered so that any token's routing can reach any expert. That topology means the minimum viable serving unit is not one GPU but a cluster, and the cluster's cost is fixed whether it processes one request or ten thousand. As the engineer Sean Goedecke laid out in a widely circulated [analysis of DeepSeek's batching economics](https://www.seangoedecke.com/inference-batching-and-deepseek/), models of this shape are efficient at high batch sizes and ruinous at low ones: the same hardware that serves a full batch near a dollar per million tokens serves a nearly empty batch at ten or twenty times that. Throughput and latency trade against each other along a curve, and the operator picks a point on it.

That curve explains every structural feature of the new card. Peak/off-peak pricing exists because the cluster's cost per token depends on utilization, and utilization depends on when customers show up; charging double at peak is a way of paying for the empty batches at 4am with the full ones at 2pm. The cache-hit repricing is the same logic applied to memory instead of compute: every cached context occupies KV-cache space in GPU memory for as long as the session lives, and KV cache is the scarcest resource in long-context serving — a million-token agent session holds memory that could otherwise serve dozens of short requests. Pricing cache hits near zero, as the old card did, gave DeepSeek's heaviest users an incentive to occupy its scarcest resource indefinitely. The 12x correction prices the occupancy.

The [545% theoretical margin disclosure](https://www.reuters.com/technology/chinas-deepseek-claims-theoretical-cost-profit-ratio-545-per-day-2025-03-01/) from March 2025 is the number that ties this together, because it cuts both ways. It proved DeepSeek could serve above marginal cost at the old prices — when the batches were full and the model was V3 at 671 billion parameters. V4-Pro is roughly 2.5 times larger, the context window grew to a million tokens, and the workload mix shifted toward exactly the cache-heavy agent sessions the old card underpriced. Export controls cap the supply response: DeepSeek cannot simply buy more accelerators the way its American competitors can, which is why the February 2025 capacity strain produced a discount schedule instead of a bigger fleet. When your fleet is fixed and your workload triples in memory intensity, the options are queues, degraded service, or prices. DeepSeek chose prices.

# Free weights, metered serving

The move that looks like a contradiction is the tell that it is not one. On August 13, the same day it announced the price hikes, DeepSeek put V4-Pro-0813's weights on Hugging Face under MIT — 1.7 trillion parameters, 1M context, no revenue conditions, no named prohibited uses — alongside DeepSeek Harness v0.1, an MIT-licensed agent framework, per the company's [own change log](https://api-docs.deepseek.com/updates/). The most permissive license in the ecosystem and the steepest price hike of the year, issued by the same lab, in the same announcement cycle.

We mapped the licensing turn four days ago in [the toll booth on the open road](/posts/2026-08-12-the-toll-booth-on-the-open-road/): Kimi K3 carries a 30% revenue share above $20 million in sales, Qwen3.8-Max requires a separate commercial license above $50 million in revenue. Against that backdrop DeepSeek's clean MIT drop looked like doctrinal purity. Read against the rate card, it looks like something more precise: a company that has decided exactly where its business lives. Not in the weights — those are marketing, recruiting, ecosystem capture, and a standing argument that the model layer is commoditized for everyone else too. In the serving. The weights cost DeepSeek nothing to give away that it wasn't already conceding by publishing papers; the serving is where the GPUs, the uptime, the cache infrastructure, and now the margins sit.

This is the maturation of the arc we have tracked since [the giveaway became a doctrine](/posts/2026-07-18-the-giveaway-became-a-doctrine/): first the weights were free and the motive was doctrine, then the licenses grew revenue conditions, and now the purest giveaway in the market coexists with 12x serving hikes from the same vendor. Openness and price increases are not in tension. They are one strategy with two instruments — the free weights hold the floor under every competitor's pricing while the metered API monetizes the customers who were never going to self-host a 1.7-trillion-parameter model anyway. Which, at roughly 1.2 terabytes of VRAM-class hardware even heavily quantized, is nearly all of them.

![Schematic: one machine, two hands — MIT weights given away free on one side, peak/off-peak serving meter charging up to 12x on the other.](/post-images/2026-08-16-deepseek-raised-prices-and-nobody-flinched/two-hands-playbook.jpg)

# Who eats the hike

The increase does not land evenly, and tracing who actually pays is where the MIT release stops being a curiosity and starts being the market's pressure valve.

The direct hit lands on first-party API customers with latency-sensitive, cache-heavy workloads: production agents, coding assistants, retrieval pipelines that hold long contexts across many turns. Run the arithmetic on a mid-sized agent operation to see the scale. A team pushing 2 billion output tokens a month on V4-Pro paid about $1,740 under the flat card. Under the new card, the same traffic costs roughly $3,960 if every token can be shifted off-peak, and $7,920 if none can — a swing from under two thousand dollars to nearly eight. Layer the cache repricing on top: a workload whose input mix is 80% cache hits, which is normal for persistent agents, sees its input bill multiply by up to 12 on precisely the portion that used to round to zero. For the heaviest users, the effective increase is not 127%; it is the full 1,100%, and it compounds monthly.

The indirect hit lands on the aggregators and resellers who built margin structures on DeepSeek's old rates. Any gateway that proxies DeepSeek's first-party API — and the routing layers we mapped when [the agent got a wallet](/posts/2026-07-22-the-agent-has-a-wallet-now/) all list DeepSeek endpoints — faces the same choice a utility reseller faces when wholesale prices spike: absorb the difference and bleed, or pass it through and watch traffic re-route. Gateways exist precisely to make re-routing a config change, which means the pass-through is nearly instant and the demand response shows up in someone else's dashboard within days.

And that is where the escape hatch matters. Because V4-Pro's weights are on Hugging Face under MIT, DeepSeek's first-party API is not the only place to buy V4-Pro tokens. Third-party hosts running the open weights on their own hardware felt no cost change on August 16; their GPUs cost what they cost last week. Their prices become the ceiling on what DeepSeek's own hike can extract: any customer whose workload tolerates a third-party host's latency profile can defect without changing models. DeepSeek knows this, which tells you what the hike is actually pricing. It is not pricing the model — the model is free. It is pricing DeepSeek's own serving quality: the first-party cache infrastructure, the uptime, the proximity to the lab that built the thing. The hike is a bet that for the customers who matter, the serving is worth 4 to 12 times what the subsidy charged for it.

# The case for calling it margin-taking

The strongest argument against this post's framing deserves its own section, because it is genuinely strong. The skeptic's case runs: the capacity story is cover, and this is a monopolist's move dressed in engineering language. DeepSeek proved in [March 2025](https://www.reuters.com/technology/chinas-deepseek-claims-theoretical-cost-profit-ratio-545-per-day-2025-03-01/) that it could clear a 545% theoretical margin at the old prices. If the old prices were already profitable at scale, the new prices are not cost recovery; they are rent extraction from a locked-in customer base, timed for a [prospectus](https://www.bloomberg.com/news/articles/2026-08-13/deepseek-increases-prices-for-ai-services-by-multiple-times), executed on three days' notice precisely because three days is not enough time to migrate a production system.

Take the case seriously and two pieces of it hold. The timing is unambiguously IPO-shaped: no company discovers a capacity crisis the same week its bankers start building a revenue narrative. And the notice period was chosen, not forced — a lab that wanted to cushion its customers could have announced in June for October. The hike is partly margin-taking. It would be naive to write otherwise.

But the core of the skeptic's case fails on the arithmetic it cites. The 545% figure described V3, a 671-billion-parameter model, serving the workload mix of early 2025: short contexts, chat-shaped traffic, batches full of small requests. V4-Pro is 2.5 times the parameters with 8 times the context window, serving agent workloads that pin KV cache for hours. A margin computed on the old machine and the old traffic says nothing about the new machine under the new traffic — and the parts of the new card that would be pure rent if the capacity story were false (the time-of-day structure, the cache repricing) are exactly the parts a rent-seeker would not bother with. A monopolist raises the flat rate. An operator with a utilization problem builds a tariff that moves traffic to 4am. DeepSeek built the tariff. The honest verdict is a ratio, not a binary: some margin, taken because the IPO window rewards it, layered on a real cost floor that the old card genuinely sat below.

# The quiet on the timeline

The social reaction is the strangest part of the story, and it is honest to report it plainly: through today, there is barely any. Our monitoring of the price-hike discourse turned up no post dated on or before August 16 that cleared even modest engagement thresholds — no viral migration thread, no organized developer revolt, no repricing panic. The announcement drew wire coverage on the 13th, Reuters and Bloomberg both moved stories, and then the builder timeline mostly shrugged and went back to benchmark screenshots. Three days between announcement and effect is usually enough for outrage to organize. It did not.

The silence supports a reading the outrage would have contradicted. If DeepSeek's customers believed the old prices were the product, a 12x hike would have produced an exodus loud enough to hear from here. What the quiet suggests instead is that the customers already understood the bargain: the prices were an introductory offer from a lab everyone knew was operating below cost, and even the new rates leave V4 Flash among the cheapest capable models on the board. You flinch at a betrayal. You do not flinch at an expiration date you always suspected was coming. The reaction may yet build — meters that switched on this morning take a billing cycle to hurt — but as of the day the era of assumed-cheap Chinese inference formally ended, the room treated it as a formality.

# The war was always about finding the number

> **The Deep Feed's position:** the price war was never a race to zero, and August 16 is the day that became undeniable. OpenAI cut 80% because self-optimizing inference lowered its real costs; DeepSeek raised up to 12x because subsidy-priced serving was never its real cost. Both moves walk toward the same band — the true marginal cost of intelligence at scale — from opposite directions. The strategic question was never who could charge the least. It was who would still be standing, with margins, when everyone was forced to charge the truth.

For two years the standing Western assumption was that Chinese AI would always be the cheap option, and every pricing strategy in the industry priced against that floor. This morning the floor moved, and it moved up. Meanwhile the weights themselves — the thing the whole price war was supposedly about — are sitting on Hugging Face under MIT, free to anyone with a spare terabyte of accelerator memory.

That is the shape of the market this fortnight revealed: intelligence as an artifact is converging on free, and intelligence as a service is converging on its real cost. DeepSeek was simply the first to publish both prices on the same day.

## Sources

- [DeepSeek API Docs — Models & Pricing (effective Aug 16, 2026)](https://api-docs.deepseek.com/quick_start/pricing/)
- [DeepSeek API Docs — Change Log / Updates (Aug 13, 2026)](https://api-docs.deepseek.com/updates/)
- [Reuters — DeepSeek releases official V4 Pro model as it steps up expansion (Aug 13, 2026)](https://www.reuters.com/world/china/deepseek-releases-official-v4-pro-model-it-steps-up-expansion-2026-08-13/)
- [Bloomberg — DeepSeek Increases Prices for AI Services by Multiple Times (Aug 13, 2026)](https://www.bloomberg.com/news/articles/2026-08-13/deepseek-increases-prices-for-ai-services-by-multiple-times)
- [CIO — DeepSeek raises some V4 prices by more than 10x as AI demand strains capacity (Aug 2026)](https://www.cio.com/article/4209473/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity-3.html)
- [Codersera — DeepSeek V4 price change, August 2026: the exact numbers](https://codersera.com/blog/deepseek-v4-price-change-august-2026/)
- [OpenAI — Advancing the price-performance frontier with GPT-5.6 (Jul 30, 2026)](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/)
- [Google — Introducing Gemini 3.7 Flash (Aug 13, 2026)](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/)
- [SpaceXAI — Grok 4.6 (Aug 12, 2026)](https://x.ai/news/grok-4-6)
- [ccleaks — Grok 4.6 launch: benchmarks and the quiet cached-input price increase (Aug 2026)](https://ccleaks.com/news/grok-4-6-launch-benchmarks-pricing-aug-2026)
- [CNBC — Anthropic CFO holds early IPO meetings as valuation talk builds (Aug 13, 2026)](https://www.cnbc.com/2026/08/13/anthropic-cfo-early-ipo-meetings-valuation.html)
- [SCMP — OpenAI blinks in face of Chinese rivals, drops pricing on some models 80% (Jul 30, 2026)](https://www.scmp.com/tech/tech-trends/article/3362568/openai-blinks-face-chinese-rivals-drops-pricing-some-models-80)
- [DeepSeek — DeepSeek-R1 Release (Jan 20, 2025)](https://api-docs.deepseek.com/news/news250120)
- [Reuters — DeepSeek cuts off-peak pricing for developers by up to 75% (Feb 26, 2025)](https://www.reuters.com/technology/chinas-deepseek-cuts-off-peak-pricing-by-up-75-2025-02-26/)
- [SCMP — AI by night: DeepSeek offers off-peak 75% discount as demand strains servers (Feb 26, 2025)](https://www.scmp.com/tech/tech-trends/article/3300264/ai-night-chinas-deepseek-offers-peak-75-discount-demand-strains-servers)
- [Reuters — China's DeepSeek claims theoretical cost-profit ratio of 545% per day (Mar 1, 2025)](https://www.reuters.com/technology/chinas-deepseek-claims-theoretical-cost-profit-ratio-545-per-day-2025-03-01/)
- [DeepSeek — DeepSeek-V3.1 Release and API pricing adjustment (Aug 21, 2025)](https://api-docs.deepseek.com/news/news250821/)
- [Sean Goedecke — Why DeepSeek is cheap at scale but expensive to run locally (2025)](https://www.seangoedecke.com/inference-batching-and-deepseek/)

---

Canonical: https://www.thedeepfeed.ai/posts/2026-08-16-deepseek-raised-prices-and-nobody-flinched/
Site: https://www.thedeepfeed.ai
Full corpus: https://www.thedeepfeed.ai/llms-full.txt