DeepSeek's Traffic Halved. 2026 Brings a New Architecture. Is mHC a Comeback—or a Last Stand?

DeepSeek's web traffic is not telling a comforting story. It's telling a familiar one: a breakout moment, a peak, and then a sustained decline that looks less like "strategic repositioning" and more like what happens when a consumer entry point loses the distribution war.
From Feb 2025 to Nov 2025, Similarweb shows deepseek.com monthly visits falling from 614.7M to 345.7M—roughly -44% from peak. That kind of drawdown is rarely "fine." It is usually constraining, especially for a consumer-facing product. The site's engagement metrics remain decent, but a shrinking top-of-funnel is still a shrinking top-of-funnel.
Then, at the start of 2026, DeepSeek published a new paper proposing mHC (Manifold-Constrained Hyper-Connections), a scaling architecture aimed at training stability and efficiency.
This piece argues a harsher thesis:DeepSeek likely lost the consumer entry-point battle to ByteDance's Doubao. mHC is what you build when you can't win on distribution and must win on economics.
Executive Summary (for decision-makers)
Traffic halving is not a "planned transition." Similarweb indicates 614.7M monthly visits (Feb 2025) → 345.7M (Nov 2025) for deepseek.com. That is the signature of a consumer funnel that peaked and failed to hold.
Doubao's reported scale matters because it signals habit capture. Multiple outlets reported Doubao crossed 100M DAU in late Dec 2025. The exact number is less important than what it implies: ByteDance can convert AI into a default daily behavior through embedded distribution.
mHC is best read as a defensive weapon for the B-side (developers/enterprise). It is an attempt to preserve the only moat DeepSeek can still plausibly widen: cost curve + iteration speed.
"Efficiency" is not an academic win—it becomes pricing power. If mHC shifts training/inference economics meaningfully, 2026 becomes an API price war year, where DeepSeek can push high-quality inference prices toward competitors' cost lines.
Do not underestimate Doubao's technology. ByteDance is not "just distribution." It has strong applied ML capability and massive infrastructure leverage; reporting on AI-focused capex underscores the risk that ByteDance can brute-force performance gaps when needed.
The real fight is intelligence density vs. capital density.
1. The Data Reality: DeepSeek.com's Monthly Trend (Similarweb)
Here is your monthly series (Similarweb "Monthly Visits"), including the peak-to-decline-to-mini-rebound pattern.
Key inflection points
Peak: 614.7M (Feb 2025)
Trough after peak: 319.3M (Aug 2025), -48.1% vs peak
Partial rebound: 355.3M (Oct 2025), +11.3% vs trough
Latest snapshot: 345.7M (Nov 2025), -43.8% vs peak
Month-by-month (visits in millions, MoM change)
Month | Monthly Visits (M) | MoM |
Nov-24 | 4.5 | — |
Dec-24 | 11.8 | 1.634 |
Jan-25 | 277.9 | 22.561 |
Feb-25 | 614.7 | 1.212 |
Mar-25 | 510.6 | -16.90% |
Apr-25 | 480.1 | -6.00% |
May-25 | 436.2 | -9.10% |
Jun-25 | 385.8 | -11.50% |
Jul-25 | 350.5 | -9.20% |
Aug-25 | 319.3 | -8.90% |
Sep-25 | 333 | 0.043 |
Oct-25 | 355.3 | 0.067 |
Nov-25 | 345.7 | -2.70% |
The cold interpretation
A 44% drawdown from peak is not proof of "collapse," but it is strong evidence of a consumer funnel that failed to convert peak attention into stable habit. If DeepSeek had won the consumer entry-point battle, you would expect either a higher plateau or repeated re-acceleration. Instead, the curve suggests DeepSeek's standalone web entry point became less central as the market consolidated around embedded assistants.This is where Doubao's reported DAU scale becomes relevant: it represents a different product physics—habit via distribution.
2. mHC: Not a "Breakthrough Story," a "Cost Curve Story"
DeepSeek's paper, mHC: Manifold-Constrained Hyper-Connections, proposes a constrained way to increase internal connectivity while preserving training stability. Wenfeng Liang (梁文锋) is listed among the authors; the paper's timeline shows a late-2025 / early-2026 release window.
What mHC is trying to fix (in plain business terms)
Scaling LLMs is increasingly about avoiding failure modes—instability, brittleness, inefficient memory access—when you push architecture beyond conventional residual patterns. mHC constrains the mixing matrices by projecting them into a doubly-stochastic manifold (Birkhoff polytope) via Sinkhorn–Knopp, aiming to keep propagation stable while still increasing representational richness.A decision-relevant detail is that the paper discusses overhead control and is widely summarized as keeping added training time relatively modest in at least one configuration (e.g., an expansion setting with ~6.7% overhead).
The correct reading
mHC is not a consumer growth hack. It is a cost and stability instrument. That means its business value is downstream:
lower cost per capability step,
more iterations per unit compute,
and potentially lower inference cost at a given quality frontier.
And that is exactly the kind of weapon you build when you cannot out-distribute a super-app ecosystem.
3. Translating "Efficiency" into the Only Language Markets Respect: Price War
Most coverage stops at "efficiency improves." That is not the point.If mHC materially shifts DeepSeek's cost curve, it expands DeepSeek's pricing envelope. And when one provider can price below others' effective cost for high-quality inference, two things happen:
API price compression accelerates (even if competitors match prices, margins erode).
A large class of "wrapper-only" products—those without defensible data, workflow lock-in, or proprietary infra—gets squeezed toward irrelevance.
Here is the logic chain that matters in 2026:
mHC helps DeepSeek "force" more capability without proportionally more hardware.
Capability without proportional cost becomes pricing power.
Pricing power becomes the ability to push inference pricing to competitors' cost line (or below, temporarily) to force consolidation.
So the prediction is explicit:2026 is likely to see a renewed API price war, and DeepSeek's best shot at influence is to become the cheapest place to buy high-grade intelligence.
4. The Comparison You Must Not Get Wrong: Doubao Has Tech and Distribution
It is tempting to reduce this to:
DeepSeek = technology
Doubao = distribution
That framing is strategically dangerous.Wired's reporting emphasizes Doubao's rapid rise and surpassing popularity in 2025, with ByteDance integration as a major driver.Reuters reporting on ByteDance's AI infrastructure ambitions (even while ByteDance disputes parts of it) underscores the risk profile: ByteDance can choose a "brute force" strategy when others must be elegant.This is why DeepSeek's moat is temporary by default. DeepSeek's mHC is "leverage." ByteDance's compute and surfaces are "gravity." If ByteDance decides to close the gap, it can.So the real contest is:
Intelligence density (DeepSeek): elegance, systems efficiency, faster iteration per GPU.
Capital density (ByteDance): more compute, more surfaces, more distribution, more feedback loops.
DeepSeek's survival condition is harsh but clear:DeepSeek must compound algorithmic advantages faster than ByteDance can compound infrastructure advantages.
5. So—Can a New Architecture Save DeepSeek?
If "save" means "win the consumer entry point," the answer is: unlikely, because architecture does not buy distribution. Doubao's reported DAU scale signals habit capture; DeepSeek's visit curve signals a consumer funnel that peaked and failed to hold.If "save" means "avoid losing the developer/enterprise market," the answer is: plausible, because cost curves can beat distribution in B2B procurement cycles.If "save" means "reshape the market through pricing," the answer is: it depends on execution, but the direction is coherent. If DeepSeek turns mHC into consistent economics, it can force a pricing regime change that hurts everyone who does not control their stack.
What to Watch in 2026 (Concrete signals)
Q1: Does the DeepSeek cost frontier move?
A1: Look for credible signals of lower inference cost at a given quality level (not just "paper claims").
Q2: Does DeepSeek initiate aggressive pricing?
A2:If not, the "efficiency advantage" may be too small, too uncertain, or reserved for internal use.
Q3: Does web traffic stabilize—without marketing spikes?
A3:Stabilization would suggest a surviving consumer nucleus; continued drift suggests DeepSeek is becoming a "model supplier," not a destination.
Q4: Does ByteDance respond with brute-force releases?
A4:If ByteDance closes gaps quickly, mHC's advantage is likely temporary.


