DeepSeek · AI model

DeepSeek-R1

SupersededReleased: 20 January 2025 (R1-0528: 28 May 2025)Context: 128K tokensPrice: Free download (MIT); retired from DeepSeek's APISize: 671B total / 37B active (MoE)Status: Superseded by V3.1 (Aug 2025) and V4 (2026)

DeepSeek-R1 is the open reasoning model that, in January 2025, showed a Chinese lab could match OpenAI's o1 at a tiny fraction of the cost. It is now superseded by DeepSeek V4, but it remains widely downloaded and historically important. DeepSeek released R1 on 20 January 2025 under the MIT licence: 671 billion parameters, 37 billion active per token, and a 128K-token context. Within a week the DeepSeek app topped the US iPhone App Store and US chip stocks fell sharply.

An update, R1-0528 (28 May 2025), raised its AIME 2025 maths score from 70.0% to 87.5% and GPQA Diamond from 71.5% to 81.0%. In September 2025 the R1 paper was published in Nature, reporting that the reasoning training cost about $294,000 on 512 Nvidia H800 GPUs, on top of roughly $6 million for the V3 base model. DeepSeek folded R1's reasoning into its hybrid V3.1 model in August 2025, and the old deepseek-reasoner API name was retired on 24 July 2026.

6.8/10
Our verdictExpert score 6.8/10

DeepSeek-R1 is a landmark model, but in September 2026 you should only use it for research, teaching or legacy systems.

Why it mattered:

  • It opened up reasoning models. R1 showed its step-by-step thinking and came with a paper explaining how reinforcement learning produced it. Before R1, OpenAI's o1 hid that process.
  • It changed the cost debate. DeepSeek trained it with export-compliant H800 chips and a reported $294,000 reasoning budget. The paper passed peer review at Nature.
  • It is truly open. The MIT licence let thousands of teams build on it and its distilled versions.

Why it has been overtaken:

  • Much weaker than today's models. DeepSeek V4 scores 90.1% on GPQA Diamond against R1-0528's 81.0%, with 1M context instead of 128K.
  • No longer served on DeepSeek's API.
  • Known censorship behaviour on Chinese political topics, trained into the weights.

Use R1 to study how reasoning models work, or to run the small distilled versions on modest hardware. Do not use it for new production work. Pick V4, Qwen3.6 or Kimi K2 instead.

Score breakdown

Reasoning7.0
Coding6.0
Openness9.5
Value8.5
Current relevance4.0

Best for

  • Research into reasoning models
  • Teaching how chain-of-thought works
  • Legacy systems already built on R1
  • Small distilled models for local experiments

What we like

  • MIT licence with full commercial rights
  • Published, peer-reviewed training method
  • Distilled versions from 1.5B to 70B run on ordinary hardware
  • Visible chain of thought, useful for teaching and research

Watch out for

  • Far behind 2026 models on coding and reasoning
  • 128K context and text only
  • Retired from DeepSeek's own API
  • Trained-in censorship on Chinese political topics

Specs at a glance

DeveloperDeepSeek (Hangzhou, China)
Release datesR1 and R1-Zero: 20 January 2025; R1-0528: 28 May 2025
ArchitectureMixture of experts built on DeepSeek-V3
Parameters671B total, 37B active (R1-0528 checkpoint listed as 685B including extra layers)
Context window128K tokens
Input / outputText only
Training methodLarge-scale reinforcement learning for reasoning; R1-Zero used RL with no supervised fine-tuning
LicenceMIT (distilled Llama-based versions also follow Llama licence terms)
Distilled versions1.5B, 7B, 14B, 32B (Qwen2.5 bases) and 8B, 70B (Llama 3 bases); R1-0528-Qwen3-8B
Reported training costAbout $294,000 for reasoning training on 512 H800 GPUs (Nature, 2025), plus about $6M for the base model
API statusdeepseek-reasoner pointed to R1 until V3.1 (August 2025); name retired 24 July 2026

Benchmarks

Benchmarks are standard tests. Vendor-run results are marked as such; independent results are preferred where they exist.

BenchmarkScoreSourceNote
AIME 2025R1: 70.0% / R1-0528: 87.5%DeepSeek (vendor)
AIME 2024R1: 79.8% / R1-0528: 91.4%DeepSeek (vendor)
GPQA DiamondR1: 71.5% / R1-0528: 81.0%DeepSeek (vendor)
LiveCodeBenchR1: 63.5% / R1-0528: 73.3%DeepSeek (vendor)
Codeforces ratingR1: 1530 / R1-0528: 1930DeepSeek (vendor)

Pricing

Plan / tierPriceNotes
Download and self-hostFreeMIT licence; full model needs a multi-GPU server
Distilled modelsFreeRun on a laptop or single GPU
DeepSeek APINot availabledeepseek-reasoner retired 24 July 2026

Why R1 made headlines

R1 arrived on 20 January 2025 with reasoning scores close to OpenAI's o1, open weights and API prices far below OpenAI's. By 27 January the DeepSeek app was the most downloaded free app on the US iPhone App Store, and chip stocks fell sharply on fears that AI would need fewer expensive GPUs.

The Nature paper (September 2025) gave the first peer-reviewed account of a major reasoning model. It reported 80 hours of reinforcement learning on 512 H800 chips at about $294,000, and DeepSeek acknowledged owning A100 chips used in early work. Critics noted this figure excludes the base model and research costs.

R1 vs R1-0528 vs V4

R1 (Jan 2025) R1-0528 (May 2025) V4-Pro (2026)
GPQA Diamond 71.5% 81.0% 90.1%
AIME 2025 70.0% 87.5% n/a
Context 128K 128K 1M
Tool calling No Yes Yes
Licence MIT MIT MIT

R1-0528 thought for longer (about 23,000 tokens per AIME question against 12,000), which drove most of its gains.

Running R1 today

The full model needs a multi-GPU server. Most people who want R1 locally use a distilled version: small Qwen- or Llama-based models trained on R1's answers. DeepSeek-R1-0528-Qwen3-8B scored 86.0% on AIME 2024, DeepSeek says, and runs on a single consumer GPU. The Llama-based distilled models also carry Meta's Llama licence terms.

Privacy, censorship and export controls

When you run R1 yourself, no data goes to DeepSeek. Researchers found that R1 still follows Chinese government positions on topics such as Taiwan and Tiananmen Square even when self-hosted, because the behaviour is trained in. Community fine-tunes attempt to remove this.

R1's H800 training chips were designed for China to comply with US export rules at the time; the US later tightened those rules. None of this restricts you from using the MIT weights.

Expert tips
  1. For local experiments, start with DeepSeek-R1-0528-Qwen3-8B. It gets most of the maths benefit at a size a gaming PC can handle.
  2. If you still call deepseek-reasoner, switch to deepseek-v4-pro or deepseek-flash with thinking enabled. The old name no longer works.
  3. Budget for long outputs. R1-0528 used about 23,000 thinking tokens per hard maths question.
  4. Check the licence of any distilled model: Llama-based versions add Meta's terms on top of MIT.

Jargon explained

Reasoning model
A model that writes out step-by-step thinking before its final answer, which helps on maths, logic and coding.
Reinforcement learning (RL)
Training by trial and reward: the model tries answers and is rewarded when they are correct.
Distillation
Training a small model on a big model's answers so it copies some of its skill.
AIME
A hard US high-school maths competition used as an AI benchmark.
H800
A cut-down Nvidia chip made for China to meet US export rules in 2023.

Alternatives to consider

Frequently asked questions

Is DeepSeek-R1 still available?

The weights remain free on Hugging Face under the MIT licence. DeepSeek's API no longer serves it; the deepseek-reasoner name was retired on 24 July 2026.

Did DeepSeek-R1 really cost $294,000 to train?

That is DeepSeek's figure for the reasoning training step only, published in Nature. It excludes about $6 million for the V3 base model and all earlier research.

Can I run DeepSeek-R1 on my computer?

The full 671B model needs a server. The distilled versions (1.5B to 70B) run on laptops and single GPUs through tools like Ollama and LM Studio.

What replaced DeepSeek-R1?

DeepSeek merged reasoning into V3.1 in August 2025, then V3.2 and V4. V4 is the current family.

Sources

Every fact on this page comes from public information. Vendor figures are labelled as vendor claims.

  1. DeepSeek-R1 model card (Hugging Face / DeepSeek)
  2. DeepSeek-R1-0528 model card (Hugging Face / DeepSeek)
  3. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (arXiv / DeepSeek)
  4. DeepSeek API change log (DeepSeek)
  5. DeepSeek reveals the cost of training the AI model (CNN)
  6. DeepSeek didn't really train its flagship model for $294,000 (The Register)
  7. DeepSeek (chatbot) (Wikipedia)
  8. DeepSeek-V4-Pro model card (Hugging Face / DeepSeek)

More from DeepSeek