The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
DeepSeek has not disclosed a complete dollar budget for DeepSeek-R1. The often-cited $5.6 million figure is instead DeepSeek’s estimated compute cost for the official training of DeepSeek-V3, a model in R1’s development lineage. It is a useful compute benchmark—not R1’s verified price tag or the total cost of developing and operating DeepSeek’s AI products.
Where the $5.6 million figure comes from
DeepSeek’s V3 technical report reports 2.788 million H800 GPU-hours for the model’s official training process. It estimates the cost by assuming a rental price of $2 per H800 GPU-hour:
2,788,000 GPU-hours × $2 = $5,576,000
That rounds to $5.6 million. The report’s breakdown is:
Free tools Windows power users keep installed
One-click scans. No signup required.
| V3 training stage | H800 GPU-hours | Cost at $2/hour |
|---|---|---|
| Pre-training | 2.664 million | $5.328 million |
| Context-length extension | 119,000 | $238,000 |
| Post-training | 5,000 | $10,000 |
| Total | 2.788 million | $5.576 million |
This is a rental-equivalent estimate, not evidence that DeepSeek received an invoice for exactly that amount. The report states the $2 rate as an assumption. The actual cash cost could differ depending on hardware ownership, reserved capacity, internal accounting, or the price of rented capacity.
#1 Best Overall
Why the number gets attached to R1
V3 and R1 are related, but they are not the same training run. DeepSeek-V3 is a 671-billion-parameter mixture-of-experts model, with about 37 billion parameters activated for each token. DeepSeek-R1 was built from a V3-derived base and added further training stages. DeepSeek’s R1 paper describes the methods, but does not publish a comparable complete dollar budget.
The connection matters: V3 supplied a base model and infrastructure for the R1 pipeline, and later V3 post-training drew on distillation from the R1 series. So V3’s disclosed compute figure is relevant context for the R1 story. It should be described as a related benchmark, not as R1’s final cost.
What the V3 estimate includes—and leaves out
The reported total covers V3’s stated official training process: large-scale pre-training, context-length extension, and post-training. DeepSeek says V3 was trained on 14.8 trillion tokens; its main pre-training phase used 2.664 million H800 GPU-hours.
Recommended Free Tools
Rank #2
- EVOLUTION AMD RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 64GB pool, which is perfect for running LLMs such as Deepseek 32B, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 4% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
| Covered by the reported estimate | Excluded or not established by it |
|---|---|
| GPU-hours for V3’s official training stages | Prior research and architecture, algorithm, and data experiments |
| Pre-training, context extension, and the stated post-training stage | Failed runs and the full history of earlier model generations |
| Compute valued at an assumed $2 per H800 GPU-hour | People, data acquisition and preparation, hardware ownership or depreciation, and datacenter costs |
| Networking, storage, safety and evaluation, product development, serving, support, and corporate overhead |
DeepSeek explicitly says its V3 figure excludes prior research and ablation experiments involving architecture, algorithms, and data. The other items are costs an all-in development or business budget would normally need to account for; the published estimate is not an expense ledger that quantifies them.
That distinction creates three different meanings of “budget”:
- Training-run compute: V3’s $5.576 million estimate at the assumed rate; no equivalent complete figure is public for R1.
- Model R&D: would also account for experimentation, data work, research and engineering staff, infrastructure development, evaluation, and iteration. DeepSeek has not published a complete, auditable total.
- Product and company operations: would add ongoing inference, reliability, moderation, support, legal and compliance work, and other business costs. The V3 compute estimate does not measure this.
What DeepSeek disclosed about R1’s training
R1’s published process is more involved than one training run. It includes R1-Zero, an experimental reinforcement-learning-first system trained without conventional supervised fine-tuning as its preliminary stage, followed by a fuller R1 pipeline using cold-start data, reinforcement learning, rejection sampling, supervised fine-tuning, and further reinforcement-learning stages. DeepSeek also released smaller models distilled from R1.
Reinforcement learning can reduce reliance on large collections of human-labeled reasoning examples, but it does not make development free or establish a low total cost. The stages still require compute, data, engineering, evaluation, and iteration. Nor should the cost of producing distilled variants be silently folded into V3’s final-run estimate.
How DeepSeek made its compute more efficient
DeepSeek’s V3 report describes a collection of architecture and systems choices, rather than one trick that explains the result:
- Mixture of experts (MoE): V3 has 671 billion total parameters, but activates about 37 billion per token. Total parameter count and per-token computation therefore describe different things.
- Multi-head Latent Attention (MLA): reduces key-value-cache memory requirements.
- FP8 mixed-precision training: supports more efficient computation and reduces memory and bandwidth pressure.
- Auxiliary-loss-free load balancing: addresses the distribution of work among experts while limiting the performance trade-off associated with conventional balancing approaches.
- Multi-Token Prediction: adds training signals and can support speculative decoding.
- DualPipe and communication/computation overlap: help reduce distributed-training bottlenecks.
- Hardware/software co-design: the training system was optimized around the available H800 cluster.
The report says main pre-training used 2,048 H800 GPUs and took less than two months, at roughly 3.7 days per trillion tokens on that cluster. These figures describe different things: GPU count is the hardware deployed at a time, GPU-hours are cumulative accelerator use, and the dollar estimate multiplies those hours by an assumed price. None can be substituted for another.
Rank #4
- AMD RYZEN AI MAX+ 395 MINI PC – THE NEXT GENERATION AI WORKSTATION --- GMKtec EVO-X3 introduces the next evolution of desktop AI computing powered by AMD Ryzen AI Max+ 395 processor. Featuring 16 cores and 32 threads, Zen 5 architecture, TSMC 4nm FinFET process, up to 5.1GHz boost frequency, and 64MB L3 cache, EVO-X3 delivers flagship-level performance for AI applications, professional creation, gaming, and demanding multitasking. With up to 126 TOPS AI performance, this compact AI workstation brings powerful local computing to your desktop.
- AMD XDNA 2 NPU – 50 TOPS DEDICATED AI ENGINE FOR LOCAL AI --- Equipped with AMD XDNA 2 architecture NPU delivering up to 50 TOPS AI acceleration, EVO-X3 enables efficient local AI processing for generative AI, AI assistants, image creation, content production, and intelligent workflows. By processing AI tasks directly on-device, it helps reduce cloud dependency, improve response speed, and enhance data privacy. Run advanced AI applications locally with smoother performance and greater control over your data.
- AMD RADEON 8060S GRAPHICS – RDNA 3.5 POWER WITH DESKTOP-CLASS PERFORMANCE --- EVO-X3 features AMD Radeon 8060S Graphics with 40 Compute Units and up to 2900MHz frequency based on advanced RDNA 3.5 architecture. Delivering graphics performance comparable to RTX 4070-class laptop GPUs, it provides smooth 1080P high-quality gaming, accelerated video editing, 3D rendering, and creative workloads. Experience powerful integrated graphics performance without the size and power consumption of a traditional desktop tower.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- 128GB LPDDR5X 8000MT/s MEMORY – MASSIVE BANDWIDTH FOR AI AND CREATIVE WORK --- Equipped with up to 128GB LPDDR5X memory running at 8000MT/s, EVO-X3 provides exceptional bandwidth for large AI models, professional software, content creation, and heavy multitasking. The unified memory architecture allows more flexible resource allocation between CPU and GPU, making it ideal for local AI inference, large model deployment, video production, engineering applications, and advanced creative workflows.
Training is not the same as serving
Training is a large development expense; serving a model to users is an ongoing operating cost. DeepSeek published a separate infrastructure overview for V3 and R1 reporting a measured 24-hour period in February 2025. It estimated combined V3/R1 serving at $87,072 per day, using an assumed $2 per GPU-hour. The calculation used H800s and reported average occupancy of about 226.75 eight-GPU nodes.
That is a historical estimate for combined services, not an R1-only daily bill or a current operating-cost guarantee. The overview also notes that web and app usage was not monetized in the same way as API traffic. Actual serving economics depend on token volumes, cache hits, utilization, batching, peak demand, quantization, and the serving stack. The disclosure is a reminder that a relatively low training-run estimate does not make deployment free.
What the figure does—and does not—show about AI economics
The V3 disclosure is evidence that a frontier-scale model’s official training compute can be reported at a much lower figure than many readers expect. It is not proof that any organization can reproduce R1’s capabilities for $5.6 million. Reproduction would also depend on data, expertise, software, hardware availability, experimentation, evaluation, and—in a distillation approach—suitable teacher models.
Best Value
- 【POWERFUL ESP32‑S3 CONTROLLER】Built‑in Xtensa 32‑bit LX7 dual‑core processor, 512KB SRAM, 8MB PSRAM, 16MB Flash for stable AI voice computing and multitask processing.
- 【Preloaded Dual AI Platforms】Comespre-installed with complete Deepseek and OpenAI voice dialogue projects.Experience intelligent voice interaction instantly. (Note: OpenAI functionality requires your own API key.)
- 【STABLE WIRELESS & CLEAR AUDIO】Integrated 2.4GHz Wi‑Fi + Bluetooth 5 (LE); dedicated audio decoding module for natural, responsive voice interaction.
- 【USER‑FRIENDLY VISUAL & PLUG‑AND‑PLAY】2” TFT‑SPI color screen shows real‑time chat; modular design, no extra wiring, ready to use after setup.
- 【FULL LEARNING SUPPORT】45 programmable GPIOs, rich interfaces, online web tutorials, free technical support for beginners & developers.
Nor is GPU-hours an electricity bill. Converting accelerator time into energy cost would require power draw and utilization for GPUs and host systems, networking and cooling overhead, datacenter efficiency, and local electricity prices. The $5.576 million calculation does not provide those inputs.
Finally, model-development cost, API pricing, and the cost of running an open-weight model are separate questions. API prices tell a user what a provider charges for inference, not what it spent to train a model. Self-hosting avoids per-token API billing but requires suitable hardware, storage, software, monitoring, and engineering. The $5.6 million figure is not a product price or a self-hosting estimate.
The Bottom Line
Verdict: DeepSeek-R1’s complete budget is undisclosed. The $5.576 million figure is DeepSeek’s estimated compute cost for DeepSeek-V3’s official training, calculated at an assumed $2 per H800 GPU-hour. It excludes prior research and ablations and does not establish the all-in cost of R1, DeepSeek’s model development, or ongoing operations.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

