The Story We All Know
The high-level concept of a decoder-only transformer like GPT is deceptively simple. You feed it a sequence of text, and it generates the next piece of text, one token at a time. A token is a chunk of text, which could be a whole word or just a part of one,
like 'pre' in 'predicting'. This process is called autoregressive generation. The model takes the input, predicts a token, appends that new token to the input sequence, and repeats the process. It’s a sophisticated form of autocomplete, where each new word becomes the context for predicting the one that follows. This is made possible by 'causal self-attention', a mechanism that ensures the model only looks at previous tokens to make its prediction, preventing it from 'cheating' by seeing the future. For most day-to-day purposes, this explanation works just fine.
The Detail Most Engineers Skip
Here’s the crucial detail that gets glossed over: The model doesn't just pick one 'next word'. At its final step, the transformer doesn't output a token at all. Instead, it produces a massive list of raw, unnormalized scores called 'logits'. There is one logit for every single token in the model's entire vocabulary, which can be over 100,000 tokens. These scores represent the model's raw preference for what could come next. A higher score for the token 'mat' after 'The cat sat on the' means the model has a stronger raw preference for 'mat' over, say, 'floor' or 'chair'. The trained network's job ends here. It doesn't make the final choice; it just provides a comprehensive scorecard of every possible option based on the context it was given.
From Scores to Probabilities
These raw logit scores aren't probabilities. They can be positive or negative and don't add up to anything meaningful. To make them useful, a function called 'softmax' is applied. Softmax converts the entire list of logits into a clean probability distribution, where every value is positive and the sum of all probabilities equals 1. Now, instead of just having raw scores, you have a percentage chance for every single token in the vocabulary. The token 'mat' might have a 25% chance, 'floor' a 15% chance, and so on. This step is vital. It transforms the model's raw confidence scores into a quantifiable map of possibilities. The model isn't just saying what comes next; it's revealing how certain it is about every potential path forward.
Why This Unlocks True Control
Understanding that the model outputs a probability distribution—not a single word—is the key to unlocking true control over its behavior. If the model only picked the single most likely word every time (a method called 'greedy decoding'), its output would be deterministic and often repetitive. But because we have a full distribution, we can apply different 'decoding policies'. Techniques like temperature sampling, Top-K, and Top-P (nucleus) sampling all work by manipulating this probability distribution before a final choice is made. Lowering the 'temperature' sharpens the distribution, making the model more confident and less random. Raising it flattens the distribution, encouraging more creative or unexpected outputs. This is why you can get different answers to the same prompt. The underlying model weights are the same, but the sampling strategy applied to its final probability distribution is what introduces variability, creativity, and the nuance that makes these models so powerful.













