AI provider / DeepSeek
DeepSeek
A provider associated with fast-moving open releases, strong reasoning performance and model lines spanning flagship frontier systems, reasoning specialists and earlier coding branches.
DeepSeek is the AI company that proved frontier reasoning does not have to be expensive. Starting from open coding and language models in 2023, it shocked the market with DeepSeek-V3 and the o1-class R1 — models that matched closed flagships at a fraction of the training cost, mostly under permissive licenses.
Its catalog now spans flagship generations from V3 to V4, the reasoning-specialist R1 line and bridge releases like V3.2. A fast, open and reasoning-first cadence has made DeepSeek a defining force in both the open-model ecosystem and the wider race on price-performance.
Model catalog
Every release in this provider’s model line.
Each row is a distinct model release — its date, its type and the traits it is known for. Open a source to dig into any of them.
| Model | Released | Type | Key traits | Source |
|---|---|---|---|---|
| DeepSeek-V4 | Apr 2026 | Flagship |
| DataStudios |
| DeepSeek-V3.2 | Sep 2025 | Reasoning |
| Sebastian Raschka |
| DeepSeek-V3.1 | Aug 2025 | Version update |
| DeepSeek API docs |
| DeepSeek-R1 | Jan 2025 | Reasoning |
| DeepSeek API docs |
| DeepSeek-V3 | Dec 2024 | Flagship |
| DeepSeek API docs |
| DeepSeek-V2.5 | Sep 2024 | Version update |
| DeepSeek API docs |
| DeepSeek-V2 | May 2024 | Open release |
| GitHub |
| DeepSeek LLM 67B | Nov 2023 | Open release |
| GitHub |
| DeepSeek Coder | Nov 2023 | Coding |
| GitHub |
Timeline
A practical timeline of model releases and company milestones.
This page tracks the moments that matter most: major model releases, version jumps and company moves that changed how this provider shows up in the AI market.
DeepSeek-V4 ships as V4-Pro and V4-Flash.
The current flagship generation splits into a deep-reasoning Pro and a fast Flash, both aimed at long-context agent workloads.
DeepSeek-V3.2 introduces sparse attention, later joined by Speciale.
DeepSeek Sparse Attention slashes long-context cost; the harder-thinking V3.2-Speciale variant follows for tougher problems.
DeepSeek-V3.1 folds thinking and non-thinking into one model.
The hybrid design absorbs R1’s reasoning into the main line, trained natively for domestic accelerator hardware.
DeepSeek-R1 releases o1-class reasoning under MIT license.
The launch that shook markets: open chain-of-thought reasoning at a fraction of the price, briefly wiping hundreds of billions off AI stocks.
DeepSeek-V3 matches closed flagships at a $5.6M training cost.
A 671B open MoE trained for a fraction of Western budgets becomes the strongest open-weight model and a geopolitical talking point.
DeepSeek-V2.5 merges the chat and coder branches.
Consolidating two specialist lines into one general model sets up the single-flagship pattern V3 will inherit.
DeepSeek-V2 proves frontier MoE can be radically cheap.
A 236B mixture-of-experts with 21B active parameters triggers the first China LLM price war on API tokens.
DeepSeek LLM 67B takes on Llama 2 with open weights.
The first general-purpose DeepSeek models stake out the strategy the company never abandons: frontier ambition, open distribution.
DeepSeek Coder opens the company’s release history.
A fully open code-model family from 1.3B to 33B appears almost from nowhere and immediately tops open coding benchmarks.