A large language model is a statistical system trained to predict what comes next in a sequence of text. Training means adjusting billions of numeric parameters until those predictions become accurate across an enormous range of material.
Why that produces useful behaviour
To predict text well you must implicitly model grammar, factual associations, writing conventions and some reasoning patterns. Those capabilities are side effects of the prediction objective rather than features that were coded directly. This is why abilities appear gradually as models scale, and why nobody can specify in advance exactly what a given model will be able to do.
Where the limits come from
- No persistent memory. Knowledge lives in the parameters; the conversation is only a working context.
- No ground truth. A fluent sentence and a correct sentence look identical to the training objective.
- Training cutoff. Without retrieval, the model cannot know about events after its training data ends.
- Tokenisation artefacts. Counting letters or doing arithmetic in text form is genuinely harder than it appears.
Related terms
Tokens are the sub-word units models read and write. Parameters are the learned weights. Context window is the maximum token budget available in one request. Inference is the act of generating output from a trained model.
Comments (0)
Log in to join the discussion
Log InNo comments yet