Autoregressive language models predict by turning the text so far into scores for possible next tokens. A decoding rule chooses one token, adds it to the context, and the model repeats the process. Training adjusts the model’s parameters to make its predictions fit examples. This describes a common GPT-style approach—not every language model or everything an AI assistant does.
What is a token?
A token is a unit in a model’s vocabulary, not necessarily a complete word. Depending on how text is divided, a token may be a word, part of a word, or a single character. For example, a familiar word might be represented as one token, while an uncommon or longer word could be split into several. That is why “next token” is more precise than “next word.” Google’s Machine Learning Crash Course explains that LLMs predict tokens or sequences of tokens.
How does an autoregressive model make a prediction?
1. It processes the context
The model receives a sequence of tokens and turns them into internal representations. In a transformer, self-attention helps each position incorporate information from other positions in the context. Multiple layers process those representations in succession. Attention is a computational mechanism for relating parts of the input; it is not human attention, and an attention head should not automatically be treated as having one simple, fixed meaning. The AISTATS 2024 paper “Mechanics of Next-Token Prediction with Transformers” describes training transformers to predict the next token from an input sequence.
2. It scores possible next tokens
At the end of the sequence, the model produces a score for each token in its vocabulary. These raw scores are called logits. A probability distribution can be derived from them, but the scores do not identify one uniquely correct continuation: several tokens may be plausible in context. In Hugging Face’s documentation for its OpenAI GPT implementation, the model’s output includes logits, and ordinary text generation uses the logits at the final position.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minute#1 Best Overall
3. A decoding rule chooses a token
The system turns the scores into an output choice. Depending on its decoding method and settings, it might choose a top-scoring token or sample among candidates. Sampling can introduce variation; different decoding settings can also change what is emitted. OpenAI notes that multiple continuations can be plausible and that outputs can vary. Its text-generation guide discusses generating text from model outputs.
4. The process repeats with the expanded context
The chosen token is appended to the sequence. The model then predicts again using the updated context, including the token it just produced. Repeating this cycle generates text one token at a time, though a token may correspond to only part of a written word.
Rank #2
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
How does training teach next-token prediction?
During training, the model sees sequences in which later tokens provide targets for predictions based on earlier context. A loss function measures the difference between its predictions and those targets; an optimization process uses that error to adjust the model’s numerical parameters, often called weights. In the OpenAI GPT implementation documented by Hugging Face, labels are shifted so the model is trained to predict the next token, and the model computes a next-token loss. That is an implementation example, not proof that every system is trained in exactly the same way.
OpenAI describes its models as learning patterns through numerical parameters adjusted during training, then generating from those learned weights. This is not simply a database lookup for a stored next sentence. It also does not establish that memorization can never happen. OpenAI’s account is a description of its own models, not a universal guarantee about all language models. OpenAI’s development explainer discusses training, parameter updates, and variability.
Recommended Free Tools
Rank #3
Is next-token prediction the whole explanation for an AI assistant?
No. Next-token prediction describes a central mechanism for autoregressive generation, but an assistant’s behavior can also reflect post-training and the surrounding product. For example, OpenAI says its GPT-4 base model was trained to predict the next word in a document and then refined with reinforcement learning from human feedback (RLHF) to better follow user intent within guardrails. That is OpenAI’s description of GPT-4; it should not be taken as the training recipe for every provider or model. OpenAI’s GPT-4 research page describes this distinction.
Does every LLM predict the next token?
No. The description above applies to autoregressive models such as GPT-style language models. Other language models can use different training objectives. Google’s Machine Learning Crash Course distinguishes next-token prediction from masked-token prediction, in which a model learns to predict missing tokens within text. “LLM” names a broad family of systems, not a single training method.
Rank #4
Why can the same question get different answers?
A prompt may support multiple plausible continuations, and the decoding policy determines how the model selects among them. When a system samples rather than always choosing the same highest-scoring option, repeated generations can differ. The model’s learned parameters shape its scores; inference-time decoding shapes how those scores become text. These are related but distinct parts of generation.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.




