What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

“Coming soon — a fully open reconstruction of DeepSeek-R1” refers to Open R1, Hugging Face’s project to recreate the public, inspectable training pipeline behind DeepSeek’s reasoning model. The project is no longer merely forthcoming: its repository, recipes, datasets, and evaluation tooling are public. But it remains a work in progress. Open R1 has demonstrated parts of the process, particularly with smaller distilled models; it has not produced a verified, byte-for-byte or performance-identical reconstruction of DeepSeek’s original 671-billion-parameter model.

That distinction matters. Running DeepSeek-R1’s released weights, training a smaller model on R1-generated answers, and reproducing DeepSeek’s original training run are three different things.

What is DeepSeek-R1?

DeepSeek-R1 is a reasoning-focused large language model released by DeepSeek in January 2025. Its published training approach combines cold-start supervised data, reasoning-oriented reinforcement learning, rejection sampling, additional supervised fine-tuning, and further reinforcement learning.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A simplified version of the published pipeline looks like this:

DeepSeek-V3-Base
        ↓
cold-start reasoning data and supervised fine-tuning
        ↓
reasoning-focused reinforcement learning
        ↓
rejection sampling and additional supervised data
        ↓
further reinforcement learning
        ↓
DeepSeek-R1

This is a high-level representation, not a complete record of every internal dataset, preprocessing stage, checkpoint, hyperparameter, or infrastructure detail.

R1, R1-Zero, and the distilled models

  • DeepSeek-R1-Zero applied reinforcement learning directly to a base model. It showed significant reasoning gains, but also suffered from problems such as poor readability and language mixing.
  • DeepSeek-R1 added a more carefully staged process intended to produce a more usable model.
  • DeepSeek-R1-Distill models are smaller models fine-tuned using reasoning data generated by R1. DeepSeek lists dense distilled models at 1.5B, 7B, 8B, 14B, 32B, and 70B parameter scales, based on Qwen and Llama model families.

The original method is described in DeepSeek’s technical paper and discussed in a peer-reviewed account published by Nature.

What is Open R1?

Open R1 is Hugging Face’s community reproduction framework and research project. It aims to make the missing parts of the R1-style pipeline available for inspection and reuse, including:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • supervised fine-tuning;
  • GRPO reinforcement learning;
  • synthetic reasoning-data generation;
  • reward and formatting procedures;
  • evaluation recipes; and
  • training and model-sharing workflows built around Hugging Face tools.

It should not be described as a new version of DeepSeek-R1. A more accurate description is an open effort to reproduce and investigate the techniques associated with R1.

Why was a reconstruction needed?

DeepSeek released model weights, code, technical documentation, selected data samples, and distilled models. That is valuable, but it is not the same as releasing a fully rerunnable history of the original training process.

The public materials do not necessarily provide every original training example, preprocessing decision, reward implementation, checkpoint, hidden hyperparameter, or internal systems detail. DeepSeek’s paper also refers to its distributed training framework as an internal HAI-LLM system. Its inference implementation is available through the DeepSeek-V3 repository, but inference code alone does not recreate training.

That gap is the reason Open R1 matters. Researchers can download public weights and run the model, but reproducing the route from a base model to a reasoning model requires public data-generation methods, rollout infrastructure, rewards, training recipes, and evaluation conditions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What has Open R1 actually reproduced?

The strongest public evidence concerns individual components and smaller models rather than the full-scale R1 run. Open R1 documents an SFT recipe using a Qwen base-model family similar to the one used for DeepSeek-R1-Distill-Qwen-7B and the Open R1 Mixture-of-Thoughts dataset.

The repository reports the following results for its OpenR1-Distill-7B model:

Model AIME 2024 MATH-500 GPQA Diamond LiveCodeBench v5
OpenR1-Distill-7B 52.7 89.0 52.8 39.4
DeepSeek-R1-Distill-Qwen-7B 51.3 93.5 52.4 37.4

These are repository-reported results, not an independent benchmark audit. The numbers are meaningful only alongside their evaluation conditions. Open R1 notes that its estimates use different sampling counts for different tests—for example, 64 responses per AIME 2024 question, four for MATH-500, eight for GPQA Diamond, and 16 for LiveCodeBench.

Open R1 also reports reproducing DeepSeek’s published results for several distilled models within roughly one to three standard deviations, depending on the task. That supports partial behavioral reproduction, not proof that the original R1 model or its training history has been recreated.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Four meanings of “reproduction”

When evaluating claims about Open R1, separate these levels:

  1. Component reproduction: implementing stages such as supervised fine-tuning, data generation, or GRPO.
  2. Behavioral reproduction: obtaining similar scores on selected benchmarks.
  3. Model reproduction: training a model with broadly similar architecture and capabilities.
  4. Exact reproduction: recreating DeepSeek’s original weights, data path, and training history.

Open R1’s public evidence supports the first two more clearly than the last two.

How the main Open R1 components work

Supervised fine-tuning

Open R1 provides sft.py recipes using Hugging Face tooling. A configuration controls the base model, dataset, sequence length, learning rate, batch size, precision, and distributed-training settings.

accelerate launch --config_file recipes/accelerate_configs/zero3.yaml 
  src/open_r1/sft.py 
  --config recipes/OpenR1-Distill-7B/sft/config_distill.yaml

The documented example uses a maximum sequence length of 32,768 tokens and BF16 training. Results depend on the repository revision, dataset version, base checkpoint, software environment, and hardware.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

GRPO reinforcement learning

Open R1 uses Group Relative Policy Optimization, or GRPO, through the TRL and vLLM ecosystem. A documented smaller-model example is:

ACCELERATE_LOG_LEVEL=info 
accelerate launch --config_file recipes/accelerate_configs/zero3.yaml 
  src/open_r1/grpo.py 
  --config recipes/DeepSeek-R1-Distill-Qwen-1.5B/grpo/config_demo.yaml 
  --vllm_mode colocate

At a high level, GRPO generates multiple candidate answers for each prompt, scores them with a reward signal, and updates the policy according to the candidates’ relative performance. The approach can use correctness and format rewards without following the same value-model arrangement as conventional PPO implementations.

DeepSeek describes rule-based accuracy and format rewards in its paper. In practice, reward design is crucial: a weak checker can teach a model to exploit formatting or answer-extraction quirks rather than improve its reasoning.

Data generation

Open R1 includes workflows for generating data from smaller distilled models and from DeepSeek-R1. These workflows can produce problems, candidate solutions, verifiable answers, and filtered reasoning traces.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

However, data generated by R1 is teacher-generated synthetic data. It is not evidence that Open R1 has recovered the private data DeepSeek used internally, nor that the smaller model independently discovered the same reasoning behavior.

Evaluation

Open R1 uses LightEval and vLLM-based evaluation workflows. Scores can change substantially with the:

  • number of sampled responses;
  • temperature and top-p;
  • maximum generation length;
  • chat template;
  • answer-extraction rules;
  • benchmark revision; and
  • choice between pass@1 and self-consistency.

A bare benchmark table can therefore give a false impression of equivalence. Comparisons should put the prompts, sampling settings, normalization rules, and software version next to the scores.

Can you run Open R1 yourself?

There are two very different possibilities.

Running an existing model

If the goal is inference, download a released DeepSeek-R1 or distilled checkpoint and use a compatible inference runtime such as vLLM or another supported local tool. This does not train or reconstruct the model; it only runs public weights.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Distilled models are far more practical for local experimentation than the full R1 model. A 7B or 1.5B checkpoint can demonstrate the general behavior on suitable hardware, but it should not be treated as equivalent in capability or training economics to the full model.

Training an Open R1 recipe

Open R1’s documented training examples target a node with eight H100 80GB GPUs. That is a recipe-level reference, not a universal minimum: settings may be adjusted for different hardware and topologies. Even so, long sequences, multi-sample rollouts, distributed training, and evaluation make serious experiments expensive and technically demanding.

The repository also warns that its libraries rely on CUDA 12.4. Before running a command, check the current repository documentation and lockfiles for compatible versions of CUDA, PyTorch, Transformers, TRL, Accelerate, vLLM, and related packages.

A particularly important chat-template issue

Open R1 warns that some distilled DeepSeek chat templates can omit the reasoning content between <think> and </think>, or prefill the assistant response with <think>. That can interfere with format rewards during GRPO. A technically correct command may therefore produce poor results if the tokenizer, EOS token, or chat template does not match the model and configuration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to judge whether a reproduction is credible

Do not rely on one benchmark score. Check whether the project provides:

  1. A documented base model: ideally the same checkpoint family or a clearly justified equivalent.
  2. Open, versioned data: including licenses, preprocessing, and validation splits.
  3. Complete training code: covering preprocessing, SFT, RL, rewards, checkpointing, and evaluation.
  4. Transparent rewards: with correctness and format rules explained.
  5. Compute details: including GPU type and count, sequence length, batch size, duration, and parallelism.
  6. Evaluation parity: with matching prompts, benchmark versions, sampling counts, and answer normalization.
  7. Released checkpoints: so other researchers can inspect and run the result.
  8. Clear licensing: including obligations inherited from the base model.
  9. Independent replication: preferably by researchers unaffiliated with the original project.

What “fully open” does—and does not—mean

“Fully open” describes Open R1’s objective: make the pipeline substantially more inspectable and reproducible. It does not prove that every proprietary detail of DeepSeek’s original run is available.

There is also an important licensing distinction. DeepSeek states that its repository code and model weights are MIT licensed, while distilled models inherit conditions from their Qwen or Llama bases. The MIT license for DeepSeek’s repository does not automatically remove obligations attached to a Llama-derived model. Check the specific checkpoint’s license before redistribution or commercial use.

Open R1 also involves practical trade-offs:

  • Openness versus fidelity: a transparent approximation can be scientifically useful without being the original model.
  • Small-model access versus full-model equivalence: 1.5B–7B experiments demonstrate techniques but cannot establish full R1 capability.
  • Synthetic data versus independence: R1-generated traces enable distillation but create dependence on the teacher.
  • Benchmark similarity versus general capability: mathematics and coding scores do not establish equivalent writing, factuality, multilingual, tool-use, or safety performance.

Common failure modes

  • Out-of-memory errors: long sequences and rollout generation can exhaust GPU memory.
  • CUDA or package incompatibility: mismatched versions can fail before training starts.
  • Incorrect templates or EOS tokens: especially damaging to reasoning-format rewards.
  • Slow rollouts: GRPO can become impractical when generation is slower than policy updates.
  • Reward hacking: models may learn formatting tricks or exploit weak answer checkers.
  • Benchmark leakage: math and coding data may overlap with training or teacher-generated data.
  • Non-comparable evaluations: different decoding and sampling settings can move scores materially.
  • License mismatch: derivative checkpoints may carry base-model terms.
  • Mislabelled distillation: a student trained on R1 outputs is not a reconstruction of R1’s original training process.

How Open R1 compares with nearby options

DeepSeek-R1 itself is the appropriate choice for readers who want the original released weights and official materials.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

DeepSeek-R1-Distill models are better suited to local inference and smaller experiments, but they are derivatives rather than the full R1 model.

Skywork-OR1 is a separate open reasoning-model effort that reports releasing weights, training code, and datasets. It is useful for comparison, but it is not evidence that Open R1 has completed an exact DeepSeek replication.

Why the project matters

The durable importance of Open R1 is not that it has simply “rebuilt DeepSeek.” It turns a celebrated result into a collection of testable engineering questions:

  • How much improvement comes from reinforcement learning?
  • How important is cold-start reasoning data?
  • Can verifiable rewards scale beyond mathematics and code?
  • How much synthetic data is needed?
  • Which parts of the recipe transfer to smaller base models?
  • What compute and rollout infrastructure does each stage require?

Those questions are valuable even when the answer is not an exact clone. Open R1 lowers the barrier to inspecting, modifying, and independently evaluating reasoning-model techniques. For most readers, that is the real story behind the old “coming soon” headline: the pipeline is now public enough to experiment with, but the original DeepSeek-R1 training run has not been completely reconstructed.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.