<p>A GPT-2-small-scale model is a stack of 12 decoder blocks, each with 12 attention heads and 768-dimensional hidden states, sitting on a 50,257-token vocabulary with a 1,024-token context window. Counted with the tied token embedding taken once and bias terms included, that configuration has 124,439,808 trainable parameters. That is the number nanoGPT labels 124M. The original GPT-2 paper lists its smallest model as 117M, and the gap is a counting question rather than a different architecture. This guide walks through what each component does, how to build it in PyTorch, and what training it at full scale would actually demand.</p>
<h2>Why the parameter count is 124M and not 117M</h2>n<p>OpenAI’s 2019 GPT-2 paper lists its smallest model at 117M parameters, with 12 layers and 768 model dimensions. The nanoGPT repository uses the same layer count, head count and width, and labels that configuration 124M. The two numbers describe the same shape of network, so the difference comes from how parameters are tallied. The paper’s table value is not reproduced by the arithmetic below under the configuration as described, and the cited sources do not identify which counting choice yields 117M. When you report a size, state the convention you used.</p>n<p>The table below breaks the nanoGPT-style configuration into its trainable tensors. These figures are derived by arithmetic from the configuration values (n_layer 12, n_head 12, n_embd 768, vocab_size 50257, block_size 1024), not measured from a run.</p>n<table>n<thead><tr><th>Component</th><th>Trainable parameters</th><th>How it is computed</th></tr></thead>n<tbody>n<tr><td>Token embedding (wte)</td><td>38,597,376</td><td>50,257 × 768</td></tr>n<tr><td>Position embedding (wpe)</td><td>786,432</td><td>1,024 × 768</td></tr>n<tr><td>One decoder block</td><td>7,087,872</td><td>Attention projections, MLP, two LayerNorms (detailed below)</td></tr>n<tr><td>12 decoder blocks</td><td>85,054,464</td><td>12 × 7,087,872</td></tr>n<tr><td>Final LayerNorm</td><td>1,536</td><td>2 × 768 (weight and bias)</td></tr>n<tr><td>LM head</td><td>0 additional</td><td>Shares the token embedding matrix (weight tying)</td></tr>n<tr><td><strong>Total, tied head counted once</strong></td><td><strong>124,439,808</strong></td><td>Sum of the rows above</td></tr>n<tr><td>Total, LM head counted as a separate matrix</td><td>163,037,184</td><td>124,439,808 + 38,597,376</td></tr>n</tbody>n</table>n<p>Inside one decoder block, the trainable tensors are:</p>n<table>n<thead><tr><th>Tensor</th><th>Shape</th><th>Parameters including bias</th></tr></thead>n<tbody>n<tr><td>Query, key and value projection (c_attn)</td><td>768 → 2304</td><td>1,771,776</td></tr>n<tr><td>Attention output projection (c_proj)</td><td>768 → 768</td><td>590,592</td></tr>n<tr><td>MLP up-projection (c_fc)</td><td>768 → 3072</td><td>2,362,368</td></tr>n<tr><td>MLP down-projection (c_proj)</td><td>3072 → 768</td><td>2,360,064</td></tr>n<tr><td>Two LayerNorms</td><td>768 each</td><td>3,072</td></tr>n</tbody>n</table>n<p>The common mistake is counting the tied matrix twice. PyTorch’s <code>model.parameters()</code> returns each shared tensor once, so <code>sum(p.numel() for p in model.parameters())</code> on the code below prints the tied total. If you sum parameters from named submodules yourself, you can double count the 38.6M-parameter matrix and land near 163M.</p>n<h2>The data path in one pass</h2>n<p>A decoder-only language model does one thing: given the tokens so far, produce a probability distribution over the next token. Every component exists to make that prediction better. The path through the network is:</p>n<ol>n<li>Text is converted to integer token IDs with the GPT-2 byte-pair encoding.</li>n<li>Each ID is looked up in the token embedding table, and a learned position embedding for its index is added.</li>n<li>The hidden states pass through 12 decoder blocks. Each block lets every position attend to earlier positions and then transforms each position independently with a feed-forward network.</li>n<li>A final LayerNorm normalizes the last hidden states.</li>n<li>The language-model head multiplies each hidden state by the transposed embedding matrix, producing one score (logit) per vocabulary entry at each position.</li>n<li>Training applies cross-entropy between the logits at position t and the actual token at position t+1.</li>n</ol>n<p>The causal mask is what makes this a language model rather than a bidirectional encoder. It restricts what each position can use when forming its prediction. The targets are the same sequence shifted one step left, so the model is never given the answer it is being scored on.</p>n<p>Shapes are written batch-first. These follow from the configuration values above:</p>n<table>n<thead><tr><th>Tensor</th><th>Shape</th><th>Produced by</th></tr>n</thead>n<tbody>n<tr><td>Token IDs</td><td>(B, T), int64</td><td>Tokenizer and batching</td></tr>n<tr><td>Token plus position embeddings</td><td>(B, T, 768)</td><td>wte(idx) + wpe(positions), with positions broadcast across the batch</td></tr>n<tr><td>Queries, keys, values per head</td><td>(B, 12, T, 64)</td><td>Split of 768 channels into 12 heads of 64</td></tr>n<tr><td>Attention output after merging heads</td><td>(B, T, 768)</td><td>Transpose and reshape, then output projection</td></tr>n<tr><td>Block output</td><td>(B, T, 768)</td><td>Residual additions keep the shape</td></tr>n<tr><td>Logits</td><td>(B, T, 50257)</td><td>LM head</td></tr>n<tr><td>Loss</td><td>scalar</td><td>Cross-entropy over flattened positions</td></tr>n</tbody>n</table>n<h2>Build the model in seven steps</h2>n<p>The code below follows the structure of the nanoGPT reference model. It is a teaching sketch: it has not been run as part of this article, and it omits dropout, weight initialization and the fused optimizer options that the full repository uses. It requires PyTorch 2.0 or later, because it calls <code>F.scaled_dot_product_attention</code>.</p>n<h3>Step 1: Fix the configuration</h3>n<pre><code>from dataclasses import dataclassnn@dataclassnclass GPTConfig:n block_size: int = 1024n vocab_size: int = 50257n n_layer: int = 12n n_head: int = 12n n_embd: int = 768n</code></pre>n<p>Keep <code>n_embd</code> divisible by <code>n_head</code>. Here 768 / 12 gives 64 channels per head. For a smaller debugging model, reduce <code>n_layer</code>, <code>n_embd</code> and <code>block_size</code> together and keep the divisibility rule.</p>n<h3>Step 2: Causal self-attention</h3>n<pre><code>import torchnimport torch.nn as nnnimport torch.nn.functional as Fnnclass CausalSelfAttention(nn.Module):n def __init__(self, cfg):n super().__init__()n assert cfg.n_embd % cfg.n_head == 0n self.n_head = cfg.n_headn self.c_attn = nn.Linear(cfg.n_embd, 3 * cfg.n_embd)n self.c_proj = nn.Linear(cfg.n_embd, cfg.n_embd)nn def forward(self, x):n B, T, C = x.shapen q, k, v = self.c_attn(x).split(C, dim=2)n hs = C // self.n_headn q = q.view(B, T, self.n_head, hs).transpose(1, 2)n k = k.view(B, T, self.n_head, hs).transpose(1, 2)n v = v.view(B, T, self.n_head, hs).transpose(1, 2)n y = F.scaled_dot_product_attention(q, k, v, is_causal=True)n y = y.transpose(1, 2).contiguous().view(B, T, C)n return self.c_proj(y)n</code></pre>n<p><code>is_causal=True</code> applies the lower-triangular mask, so position i can attend to positions 0 through i only. Position i’s output is never influenced by a later token in the same sequence.</p>n<h3>Step 3: The feed-forward network and the block</h3>n<pre><code>class MLP(nn.Module):n def __init__(self, cfg):n super().__init__()n self.c_fc = nn.Linear(cfg.n_embd, 4 * cfg.n_embd)n self.gelu = nn.GELU(approximate=’tanh’)n self.c_proj = nn.Linear(4 * cfg.n_embd, cfg.n_embd)nn def forward(self, x):n return self.c_proj(self.gelu(self.c_fc(x)))nnclass Block(nn.Module):n def __init__(self, cfg):n super().__init__()n self.ln_1 = nn.LayerNorm(cfg.n_embd)n self.attn = CausalSelfAttention(cfg)n self.ln_2 = nn.LayerNorm(cfg.n_embd)n self.mlp = MLP(cfg)nn def forward(self, x):n x = x + self.attn(self.ln_1(x))n x = x + self.mlp(self.ln_2(x))n return xn</code></pre>n<p>The block normalizes before each sub-layer and adds the sub-layer’s output back to its input. That is the pre-normalization arrangement the GPT-2 paper describes, with the extra LayerNorm after the final block. The feed-forward inner width is 3,072, four times the model width. Each position passes through this network on its own, so the layer adds per-position computation while attention handles mixing across positions.</p>n<h3>Step 4: Assemble the model and tie the weights</h3>n<pre><code>class GPT(nn.Module):n def __init__(self, cfg):n super().__init__()n self.cfg = cfgn self.wte = nn.Embedding(cfg.vocab_size, cfg.n_embd)n self.wpe = nn.Embedding(cfg.block_size, cfg.n_embd)n self.blocks = nn.ModuleList(Block(cfg) for _ in range(cfg.n_layer))n self.ln_f = nn.LayerNorm(cfg.n_embd)n self.lm_head = nn.Linear(cfg.n_embd, cfg.vocab_size, bias=False)n self.lm_head.weight = self.wte.weight # weight tyingnn def forward(self, idx, targets=None):n B, T = idx.shapen assert T <= self.cfg.block_sizen pos = torch.arange(T, device=idx.device)n x = self.wte(idx) + self.wpe(pos)n for block in self.blocks:n x = block(x)n x = self.ln_f(x)n logits = self.lm_head(x)n loss = Nonen if targets is not None:n loss = F.cross_entropy(logits.view(-1, logits.size(-1)), targets.view(-1))n return logits, lossnnmodel = GPT(GPTConfig())nprint(sum(p.numel() for p in model.parameters())) # 124439808n</code></pre>n<p>Weight tying makes the output projection reuse the embedding matrix. This is why the total is 124,439,808 rather than 163,037,184. Tying is a design choice in this configuration, so a different implementation can report a different total from the same layer sizes.</p>n<h3>Step 5: Build input and target batches</h3>n<pre><code>import numpy as npnndef get_batch(data, batch_size, block_size, device):n # data: 1-D array of token IDs, cast to int64 for nn.Embeddingn ix = torch.randint(len(data) – block_size – 1, (batch_size,))n x = torch.stack([torch.from_numpy(data[i:i+block_size].astype(np.int64)) for i in ix])n y = torch.stack([torch.from_numpy(data[i+1:i+1+block_size].astype(np.int64)) for i in ix])n return x.to(device), y.to(device)n</code></pre>n<p>The targets are the inputs shifted by one token, so for an input window x the target at each position is the next token in the stream. Cast the data to int64 before indexing the embedding table. Token files stored as raw uint16 need that cast; the build-nanoGPT repository documents a workaround that goes through NumPy int32, so check the version you are running rather than assuming one conversion path.</p>n<h3>Step 6: Run a training step</h3>n<pre><code>device = ‘cuda’ if torch.cuda.is_available() else ‘cpu’nmodel = GPT(GPTConfig()).to(device)nopt = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=0.1)nnfor step in range(max_steps):n x, y = get_batch(train_data, batch_size=8, block_size=256, device=device)n logits, loss = model(x, y)n opt.zero_grad(set_to_none=True)n loss.backward()n opt.step()n if step % 100 == 0:n print(step, loss.item())n</code></pre>n<p>The batch size, learning rate and sequence length here are placeholders for a debugging run, not tuned settings. Training at <code>block_size=256</code> works because the model accepts any sequence length up to 1,024. For a real run, add a learning-rate schedule, gradient clipping, a held-out validation split evaluated at intervals, and a checkpoint that stores the model configuration and optimizer state so you can resume.</p>n<h3>Step 7: Sample from the model</h3>n<pre><code>@torch.no_grad()ndef generate(model, idx, max_new_tokens, temperature=1.0):n for _ in range(max_new_tokens):n idx_cond = idx[:, -model.cfg.block_size:]n logits, _ = model(idx_cond)n logits = logits[:, -1, :] / temperaturen probs = F.softmax(logits, dim=-1)n next_id = torch.multinomial(probs, num_samples=1)n idx = torch.cat([idx, next_id], dim=1)n return idxn</code></pre>n<p>Sampling uses only the logits at the final position, appends one token, and repeats until the requested length. Once the context is full, the oldest tokens drop out of the window. The output is a continuation sampled from a next-token distribution, not an answer produced by an instruction-following assistant. Chat fine-tuning is a separate stage, and the build-nanoGPT tutorial states that it does not cover it.</p>n<h2>Preparing data: decisions to make explicitly</h2>n<p>Tokenization and batching are where small builds go wrong quietly. Before training, settle these points and write them down:</p>n<ul>n<li><strong>Tokenizer.</strong> Use the GPT-2 byte-pair encoding so IDs fall within the 50,257-entry vocabulary. Mixing tokenizers leaves the embedding table pointing at the wrong meanings.</li>n<li><strong>Document boundaries.</strong> Decide whether documents are concatenated with an end-of-text token between them or padded. Concatenation is what the nanoGPT OpenWebText pipeline uses, and it avoids wasting positions on padding.</li>n<li><strong>Padding.</strong> If you pad, mask the padded positions out of the loss, otherwise the model learns to predict padding.</li>n<li><strong>Split policy.</strong> Hold out a validation set before any tuning, and split by document rather than by window, so overlapping windows from one document do not appear on both sides.</li>n<li><strong>Context length.</strong> No training window may exceed the configured block size, and the position table has no entries beyond it.</li>n</ul>n<h2>Tutorial run or reproduction: two different projects</h2>n<p>Learning the architecture and reproducing a GPT-2-scale training run are separate goals with very different costs. The nanoGPT README documents one reproduction recipe, and the table keeps its figures attached to that recipe.</p>n<table>n<thead><tr><th>Factor</th><th>Educational build or debug run</th><th>Full reproduction attempt</th></tr></thead>n<tbody>n<tr><td>Goal</td><td>Confirm the forward and backward pass works and that the loss falls on a small corpus</td><td>Approximate the GPT-2-scale OpenWebText training recipe</td></tr>n<tr><td>Compute</td><td>Small batches and short sequences. The sources do not establish a hardware minimum for this.</td><td>Eight A100 40GB GPUs for about four days, per the nanoGPT README</td></tr>n<tr><td>Data</td><td>A tiny corpus you can inspect by hand, with a small validation set</td><td>OpenWebText, which the README describes as a best-effort reproduction of WebText</td></tr>n<tr><td>Appropriate claim</td><td>Implements a GPT-2-style decoder-only Transformer</td><td>Follows the cited nanoGPT reproduction setup. It does not recreate GPT-2 exactly.</td></tr>n</tbody>n</table>n<p>A smaller run can be useful even though it cannot reproduce GPT-2’s behavior. If the loss on a tiny corpus drops steadily and the model memorizes a short passage when overfitted on purpose, the training loop and the masking are probably correct. If the loss does not move, check the target shift and the masking first.</p>n<h2>Reading reported loss numbers</h2>n<p>The nanoGPT README reports a loss of about 2.85 for its OpenWebText reproduction, and places GPT-2’s validation loss on OpenWebText at about 3.11 in the same comparison. Those two figures are not a like-for-like benchmark. The README attributes the gap partly to a domain difference: the original GPT-2 was trained on WebText, and the reproduction uses OpenWebText as a best-effort stand-in. Treat both numbers as descriptions of that repository’s setup, not as guaranteed outcomes or current benchmarks.</p>n<p>The GPT-2 paper’s own result tables are meaningful only alongside their dataset, metric, model size and zero-shot conditions. Do not set them beside evaluations of current models.</p>n<h2>Check repository status before you run anything</h2>n<ul>n<li>The nanoGPT README carries a November 2025 update describing the repository as old and deprecated, and points readers to nanochat. Use it to study the reference architecture and training flow, not as the place for new work.</li>n<li>The minGPT README includes a January 2023 note calling the project semi-archived. Its separation of model, dataset and trainer is still a good model for code organization.</li>n<li>Check the current documentation for the PyTorch version, CUDA build and any data-loading helpers before copying commands. Compatibility notes in older repositories can describe a workaround for a version that has since changed.</li>n</ul>n<h2>What you can honestly claim</h2>n<p>Once the model above trains end to end on your data, you can say you built a GPT-2-style decoder-only Transformer with the 124M convention described here. You can state its parameter count, the tokenizer and context length, and the loss on your validation split. A claim that you reproduced GPT-2 or matched its published results needs the full training recipe, the same data and the same evaluation, and the sources here do not establish that for a small run.</p>
Quick Recap
Rank #4
Rank #3
Rank #2
#1 Best Overall
As an Amazon Associate I earn from qualifying purchases.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →




