How it works
An LLM is a neural network, usually of the transformer type, trained to predict the next token (a chunk of a word) across a huge collection of text. Doing that well at enormous scale teaches it grammar, facts, coding patterns and a working form of reasoning. Chat models are then tuned further on examples and human feedback so that they follow instructions and decline harmful requests.
A model does not remember earlier calls: every request carries the whole conversation, and it only knows its training data (up to a cutoff date) plus whatever is in the prompt. Many models are multimodal, reading images, audio or PDFs as well as text. Providers sell them in tiers, from small, fast, cheap models to large flagship ones that are slower and cost more per token.
Related terms
More in AI and LLMs
Basics