Behind the screen

AI and LLMs

A large language model (LLM) is a program trained on huge amounts of text that can write, summarise, translate, answer questions and use tools when asked in plain language. Most apps reach one through a cloud API and pay per token, a small chunk of text, rather than running it themselves.

The pieces stack up. A provider such as OpenAI, Google or Anthropic, or an open model you run yourself, supplies the model. Embeddings and vector search let it find your own documents, RAG feeds those documents into the prompt, and tool calling, MCP and structured output let it take actions and return data your code can rely on.

22 terms · 4 comparisons · prices checked September 2026

22 terms, click any to open

Basics

Large language model (LLM)

Concept
An AI model trained on vast amounts of text to predict what comes next, which lets it write, summarise, translate, answer questions and hold a conversation.

An LLM is a neural network, usually of the transformer type, trained to predict the next token (a chunk of a word) across a huge collection of text. Doing that well at enormous scale teaches it grammar, facts, coding patterns and a working form of reasoning. Chat models are then tuned further on examples and human feedback so that they follow instructions and decline harmful requests.

A model does not remember earlier calls: every request carries the whole conversation, and it only knows its training data (up to a cutoff date) plus whatever is in the prompt. Many models are multimodal, reading images, audio or PDFs as well as text. Providers sell them in tiers, from small, fast, cheap models to large flagship ones that are slower and cost more per token.

Also called: LLM, language model, foundation model, generative AI, AI model

Open Large language model (LLM) as a page

Tokens and context window

Concept
Tokens are the small chunks of text a model reads and writes, and the context window is how many it can handle in one request. Both decide what a call costs.

Before a model sees your text, a tokeniser splits it into tokens: common words are one token, rarer words are broken into pieces, and spaces and punctuation count too. In English a token averages about four characters, or three quarters of a word; many other languages, and code, need more tokens for the same content.

Providers bill per million tokens, and output tokens usually cost several times more than input. The context window is the total a model can take in one request (instructions, history, documents and the reply), and current models range from tens of thousands to about a million tokens. A full window is slower and dearer, and models can overlook details buried in the middle, so sending only what is relevant still pays off.

Also called: token, context window, context length, tokeniser, max tokens

Open Tokens and context window as a page

Prompts and system prompts

Concept
A prompt is the text you send a model. The system prompt is the standing instruction that sets its role, rules and tone before the conversation starts.

Chat APIs take a list of messages. The system (or developer) message says who the model is and how it must behave, for example 'You are a support assistant for a shoe shop. Answer only from the policy below, in under 100 words.' User messages carry the questions and assistant messages hold earlier replies. Clear instructions, a few worked examples (few-shot prompting) and an explicit output format do more for quality than clever wording.

Treat prompts like code: keep them in version control and test them against known cases (often called evals) before changing them. Long prompts that repeat on every call can be cached by most providers at a large discount. Anything a user types, or any page or document the model reads, may try to override the rules (prompt injection), so the system prompt on its own is never a security boundary.

Also called: prompt engineering, system prompt, system message, few-shot prompting, prompt caching, evals

Open Prompts and system prompts as a page

Temperature

Concept
A setting that controls how adventurous a model's word choices are: low values give focused, repeatable answers and higher values give more varied, creative ones.

At each step a model scores every possible next token. Temperature reshapes those odds before one is picked: near 0 the likeliest token almost always wins, while higher values give less likely tokens a real chance. Top-p (nucleus sampling) is a related setting that limits the choice to the most probable tokens that together reach a given share.

Use a low temperature for extraction, classification, code and anything that must be consistent, and a higher one for brainstorming or creative writing. Even at 0, repeated answers are not guaranteed to match word for word. Ranges differ between providers (some accept 0 to 1, others 0 to 2), and some reasoning models fix the value or ignore it.

Also called: sampling temperature, top_p, top-p, randomness

Open Temperature as a page

Hallucination

Concept
When a model states something false or invented, such as a fake quote, figure or citation, with the same confidence as a true fact.

A model produces text that is statistically plausible, not text that has been checked. When the facts are missing (never in its training data, too recent or too obscure), it can fill the gap with something that sounds right: a product feature that does not exist, a court case that was never heard, a link that leads nowhere.

The risk can be reduced but not removed. Ground answers in your own documents (RAG) and ask for citations, allow the model to say it does not know, validate structured output, let tools do arithmetic and lookups, and keep a person in the loop for anything legal, medical or financial. Treat model output as a draft to check, not as a source.

Also called: confabulation, AI making things up, fabrication

Open Hallucination as a page

Model providers

OpenAI API

ServicePay as you go
OpenAI's paid API for its GPT family of models, the technology behind ChatGPT, used to add text, image, speech and search features to your own app.

You send messages (plus optional images, files or tool definitions) and get a reply, streamed token by token if you like. The line-up runs from small, low-cost models for simple high-volume work to flagship reasoning models for hard problems, alongside models for image generation, speech-to-text, text-to-speech, realtime voice and embeddings. The Responses API is the main interface today; the older Chat Completions format is so widely copied that many other providers accept it too.

There is no free tier: you prepay credit and can set budgets per project. Batch jobs that can wait up to a day cost half as much, and repeated prompt prefixes are billed at a discount when cached. API data is not used to train models by default. The same models are also sold through Microsoft Azure and Amazon Bedrock, for companies that prefer to buy through their existing cloud.

Pros

  • Wide range of models, from cheap and fast to flagship reasoning
  • Mature SDKs, documentation and examples in most languages
  • Its request format is supported by many other providers and tools
  • Built-in tools such as web search, file search and code execution

Cons

  • Pay per token from the first request, with no free tier
  • Models are retired on a schedule, so apps need occasional upgrades
  • Self-serve fine-tuning closed to new customers in 2026

Pick it when

  • You want a well-documented default with broad model choice
  • One vendor should cover text, vision, speech and embeddings
  • Your company already buys through Azure or AWS and wants the same models there

Skip it when

  • Data must never leave your own servers (run an open model instead)
  • You need a free tier for a prototype (Gemini has one)

What it costs · Pay as you go

Per million tokens: from about $0.10 input and $0.50 output for small models to about $10 input and $50 output for the flagship. Batch jobs are half price.

OpenAI API pricing (opens in a new tab)Approximate, checked September 2026.

Also called: OpenAI, GPT API, ChatGPT API, Responses API, Chat Completions

Open OpenAI API as a pageOfficial site (opens in a new tab)

Gemini API

ServiceFree tier
Google's API for its Gemini models, with a free tier for trying ideas and long-context models that read text, images, audio, video and PDFs.

You create a key in Google AI Studio and call models in three broad tiers: Flash-Lite for low-cost, high-volume work, Flash for everyday tasks and Pro for harder reasoning. Many accept inputs of up to about a million tokens and take images, audio, video and PDFs directly. Extras include grounding answers in Google Search, context caching, a Batch API at half price and an OpenAI-compatible endpoint, so existing OpenAI code can often switch with a new base URL and key.

The free tier needs no card but has tight rate limits, covers only some models, and lets Google use your prompts and responses to improve its products. The paid tier raises the limits and does not use your content that way. Larger companies can reach the same models through Google Cloud, which adds enterprise security controls, data residency and reserved capacity.

Pros

  • Free tier with no card needed, handy for prototypes
  • Long context and direct image, audio, video and PDF input
  • Low per-token prices on the Flash and Flash-Lite tiers
  • Grounding with Google Search built in
  • OpenAI-compatible endpoint eases switching

Cons

  • Free-tier prompts may be used to improve Google's products
  • Free-tier rate limits are far too low for production traffic
  • Two routes in (AI Studio and Google Cloud) with different setup and terms

Pick it when

  • You want to prototype an AI feature at no cost
  • Inputs are long documents, videos or recordings
  • High-volume, low-cost jobs such as tagging, extraction or translation

Skip it when

  • Sensitive data would pass through the free tier, where it can be used to improve products
  • Your stack is built around another provider's agent tools and SDKs

What it costs · Free tier

Free tier with rate limits. Paid per million tokens: from about $0.10 input and $0.40 output for small models to about $2 input and $12 output for Pro. Batch is half price.

Gemini API pricing (opens in a new tab)Approximate, checked September 2026.

Also called: Google Gemini API, Google AI Studio, Gemini Developer API, Gemini

Open Gemini API as a pageOfficial site (opens in a new tab)

Claude API

ServicePay as you go
Anthropic's API for its Claude models, widely used for writing, coding assistants and agents that work through long tasks with tools.

Claude models come in tiers: Haiku for fast, low-cost work, Sonnet as the balanced middle, and Opus and larger tiers for demanding reasoning and long-running agent tasks. They read text, images and PDFs, and recent models take up to about a million tokens of context at the standard rate. The Messages API supports tool use, structured outputs, citations, web search, code execution and computer use (operating a browser or desktop on the user's behalf).

There is no free API tier: you buy credit in the Claude Console and pay per token. Prompt caching makes repeated long prompts much cheaper (a cache hit costs a tenth of the normal input price or less), and the Batch API halves prices for work that can wait. Anthropic does not train on commercial API data by default. The models are also sold through Amazon Bedrock, Google Cloud and Microsoft Foundry, and Anthropic created MCP, which its apps and SDKs support closely.

Pros

  • Strong at code, long documents and multi-step agent work
  • Large context windows at the standard per-token rate
  • Prompt caching and batch discounts cut the cost of repeated work
  • Available directly or through AWS, Google Cloud and Microsoft Foundry

Cons

  • No free API tier; every call is paid
  • No embedding model of its own, so vector search needs another provider
  • Tool-heavy agent runs use many tokens and add up quickly

Pick it when

  • Coding assistants, agents and long multi-step tasks
  • Long documents, contracts or codebases in a single prompt
  • You want MCP tool connections to work with little glue code

Skip it when

  • You need a free tier to prototype (Gemini offers one)
  • Simple high-volume tasks where a cheaper small model does the job

What it costs · Pay as you go

Per million tokens: from about $1 input and $5 output (Haiku) to about $10 input and $50 output for the largest models. Batch jobs are half price.

Claude API pricing (opens in a new tab)Approximate, checked September 2026.

Also called: Anthropic API, Claude, Anthropic, Messages API, Claude Console

Open Claude API as a pageOfficial site (opens in a new tab)

OpenRouter

ServicePay as you go
A gateway that gives you one API key and one bill for hundreds of models from many providers, so you can switch or mix models without new integrations.

OpenRouter offers an OpenAI-compatible API in front of models from OpenAI, Anthropic, Google, Meta, Mistral, DeepSeek and many more, so trying a different model means changing one name. For open models served by several hosts it can route by price, speed or uptime and fall back to another host when one fails, and account settings can exclude hosts that log prompts or train on them.

Model prices are passed through at each provider's own rates, and OpenRouter charges a fee when you buy credits. Some models are free with low daily request limits, which suits experiments rather than production, and you can bring your own provider keys. The trade-off is one more company in the path of every request, and hosts of the same open model can differ in speed and quality, for example in how heavily they compress it.

Pros

  • One key, one bill and one API for hundreds of models
  • Compare or switch models by changing a name
  • Automatic fallback when a provider is down
  • Free models for experiments, and privacy filters per account

Cons

  • A fee on every credit purchase, on top of model prices
  • Adds a middleman to your data path and your uptime
  • Quality and speed can vary between hosts of the same open model

Pick it when

  • You want to test many models before settling on one
  • An app needs several providers without separate contracts
  • Fallback across providers matters for uptime

Skip it when

  • You use one provider and want the shortest data path
  • A contract requires a direct agreement with the model vendor

What it costs · Pay as you go

Model prices match the providers' own. Buying credits carries a fee (about 5.5% on the standard plan). Free models allow about 50 requests a day, or 1,000 after $10 of credit.

OpenRouter pricing (opens in a new tab)Approximate, checked September 2026.

Also called: LLM gateway, model router, AI gateway, BYOK

Open OpenRouter as a pageOfficial site (opens in a new tab)

NVIDIA NIM

ServiceFree tier
NVIDIA's packaged AI models: try many open models free through a hosted catalogue, then run the same ones in containers on your own NVIDIA GPUs.

A NIM is a ready-made container that bundles a model with an inference server tuned for NVIDIA GPUs and an OpenAI-compatible API. The catalogue at build.nvidia.com hosts many of them, including open models such as Llama, Mistral, DeepSeek, Qwen and NVIDIA's own Nemotron family, plus speech, vision and embedding models. Members of the free NVIDIA Developer Program get an API key to call them.

The hosted endpoints are meant for prototyping: free, but rate limited and without an uptime promise. For production you download the containers and run them on your own GPUs or in a cloud account; production use generally needs an NVIDIA AI Enterprise licence, though some NIMs are free to deploy. The appeal is control: the same model and API on a workstation, in a data centre or in any major cloud, with data staying where you run it.

Pros

  • Free, rate-limited API access to many models for prototyping
  • OpenAI-compatible API, so existing client code works
  • Containers tuned for NVIDIA GPUs that run wherever those GPUs are
  • Data stays on your own infrastructure when self-hosted

Cons

  • Production licences are priced per GPU and aimed at enterprises
  • Needs NVIDIA GPUs, which are expensive to buy or rent
  • Free hosted endpoints have tight limits and no uptime promise

Pick it when

  • Prototyping with open models without setting up hardware
  • A company must run models on its own GPUs or private cloud

Skip it when

  • A small app that a pay-per-token cloud API would serve for less

What it costs · Free tier

Hosted API free for development, with rate limits. Production generally needs NVIDIA AI Enterprise: about $4,500 per GPU a year, or about $1 per GPU an hour on cloud marketplaces plus instance costs.

NVIDIA NIM pricing (opens in a new tab)Approximate, checked September 2026.

Also called: NIM, NVIDIA Inference Microservices, build.nvidia.com, NVIDIA API catalogue, NVIDIA AI Enterprise

Open NVIDIA NIM as a pageOfficial site (opens in a new tab)

Cloud Translation API

ServiceFree tier
Google Cloud's machine translation API, which turns text and documents from one language into another across more than a hundred languages, priced per character.

You send text with a target language and get the translation back; the source language can be detected at no extra cost. The Basic edition is a simple REST call to Google's neural machine translation model. The Advanced edition adds glossaries that fix how brand names and terms are translated, translation of Word, PowerPoint and PDF files that keeps their layout, batch jobs over files in Cloud Storage, custom models trained on your own examples, and a translation-tuned LLM.

Compared with asking a general chat model to translate, a dedicated API is predictable, keeps document formatting, honours glossaries and is priced simply per character. A chat model adapts tone and context better, at the cost of more variation between runs and per-token billing. DeepL, Microsoft's Azure Translator and Amazon Translate are the main alternatives.

Pros

  • Simple per-character pricing with a free monthly allowance
  • Over a hundred languages, with automatic language detection
  • Glossaries keep product names and terms consistent
  • Translates whole documents while keeping their layout

Cons

  • Literal output can miss tone, idiom and context
  • Needs a Google Cloud project with billing set up
  • Custom models and document translation cost noticeably more

Pick it when

  • Translating user content, listings or support messages at scale
  • Consistent terminology across many languages matters

Skip it when

  • Marketing copy where tone matters more than speed (use a translator, or review LLM output)

What it costs · Free tier

First 500,000 characters a month free, then about $20 per million characters. Documents about $0.08 a page; custom models cost more.

Cloud Translation API pricing (opens in a new tab)Approximate, checked September 2026.

Also called: Google Translate API, Google Cloud Translation, machine translation API, Translation LLM

Open Cloud Translation API as a pageOfficial site (opens in a new tab)

Run it yourself

Ollama

ToolOpen source
A free, open-source app that downloads and runs open-weight language models on your own computer with one command, and serves them through a local API.

After installing Ollama on macOS, Windows or Linux, a command such as 'ollama run' followed by a model name downloads that model from its library and starts a chat. In the background it runs a server on your machine (port 11434) with its own REST API and an OpenAI-compatible one, so apps and editors can point at it instead of a cloud provider. Models are stored in compressed (quantised) form so that smaller ones fit in an ordinary laptop's memory.

Speed and quality depend on the hardware: small models run on a typical laptop, while large ones need a strong GPU or plenty of unified memory. Nothing leaves the machine and there is no per-token bill, which suits private data and offline use. For models too big to run at home, Ollama also sells a cloud service that runs larger open models with the same commands and API.

Pros

  • Free, open source and quick to set up
  • Private: prompts and data never leave your machine
  • Works offline, with no per-token costs
  • OpenAI-compatible local API that many tools already support

Cons

  • Open models that fit on a laptop trail the largest cloud models
  • Needs plenty of memory, and a good GPU for speed
  • You manage updates, model choice and any server hosting yourself

Pick it when

  • Private documents that must not go to a cloud provider
  • Offline tools, local coding helpers and experiments
  • Development and tests without paying for tokens

Skip it when

  • Many users need fast answers from a large model (use a cloud API)
  • Your users' devices are phones or low-powered laptops

What it costs · Open source

Free (MIT) to run locally. Ollama's cloud has a free plan with starter credits, Pro at $20 a month and Max at $100 a month, each with monthly usage credits.

Ollama pricing (opens in a new tab)Approximate, checked September 2026.

Also called: ollama run, local LLM, Ollama Cloud, run models locally

Open Ollama as a pageOfficial site (opens in a new tab)

Open-weight models

Concept
Models whose trained weights are published for anyone to download and run, such as Llama, Gemma, Mistral, Qwen and DeepSeek, instead of being reachable only through one company's API.

The weights are the billions of numbers a model learns in training. When they are published, usually on Hugging Face, you can run the model on your own hardware, host it with any provider, inspect it and fine-tune it. Families from Meta (Llama), Google (Gemma), Mistral, Alibaba (Qwen), DeepSeek, Microsoft (Phi) and OpenAI (gpt-oss) come in sizes from a few billion parameters, which run on a laptop, to hundreds of billions or more, which need a cluster of GPUs.

Open weight is not the same as open source: the training data and code are rarely released, and licences range from permissive (Apache 2.0, MIT) to custom ones with conditions for very large companies or on certain uses. Quantisation stores each weight in 4 or 8 bits so models fit in less memory, at some cost in quality. Tools such as Ollama, llama.cpp and vLLM run them, and hosts such as OpenRouter serve them per token.

Also called: open models, open-source LLMs, local models, Hugging Face models, GGUF, quantisation

Open Open-weight models as a pageOfficial site (opens in a new tab)

ONNX Runtime

RuntimeOpen source
Microsoft's open-source engine for running trained AI models saved in the ONNX format, on servers, on phones, in Node.js and inside a web browser.

ONNX (Open Neural Network Exchange) is an open file format for trained models, so a model built in PyTorch or TensorFlow can be exported once and run elsewhere. ONNX Runtime loads that file and runs it efficiently on whatever hardware is present, through plug-in execution providers for CPUs, NVIDIA GPUs (CUDA, TensorRT), Windows GPUs (DirectML), Apple devices (Core ML) and more.

It suits smaller, focused models rather than big chat models: text embeddings for search, classifiers, image recognition, OCR and speech. onnxruntime-node runs them inside a Node.js server, and onnxruntime-web runs them in the browser with WebAssembly or WebGPU, so data can stay on the user's device. Hugging Face's Transformers.js is built on it and handles downloading and preparing models for you.

What it costs · Open source

Free (MIT).

Approximate, checked September 2026.

Also called: ONNX, onnxruntime, onnxruntime-web, onnxruntime-node, Transformers.js

Open ONNX Runtime as a pageOfficial site (opens in a new tab)

Search and memory

Embeddings

ConceptPay as you go
Lists of numbers that capture the meaning of a piece of text or an image, so a computer can tell that 'cheap flights' and 'low-cost airfare' are about the same thing.

An embedding model turns an input into a vector, a list of a few hundred to a few thousand numbers. Inputs with similar meaning land close together, and closeness is measured with cosine similarity or a dot product. This powers semantic search, 'more like this' recommendations, grouping similar support tickets, spotting duplicates and the retrieval step of RAG.

Hosted models from OpenAI, Google, Cohere and Voyage charge per token, while open models such as BGE, E5 or MiniLM run locally, even in a browser with ONNX Runtime. Documents and queries must be embedded with the same model, and switching models means re-embedding everything. The vectors are stored in a vector index, often inside a database you already run.

What it costs · Pay as you go

Hosted APIs cost roughly $0.02 to $0.20 per million tokens (OpenAI's small model at the low end). Open models run locally for free.

Embeddings pricing (opens in a new tab)Approximate, checked September 2026.

Also called: vector embeddings, text embeddings, embedding model, vectors

Open Embeddings as a page

Vector search

ConceptOpen source
Finding items by meaning rather than exact words, by storing embeddings and returning the ones closest to the embedding of a query.

Each document chunk, product or image is stored with its embedding. At query time the question is embedded too, and the database returns the nearest vectors, usually filtered by metadata such as language, owner or date. Comparing against every vector is exact but gets slow as a collection grows, so most systems build an approximate nearest neighbour (ANN) index such as HNSW, giving up a little accuracy for a large gain in speed.

You rarely need a separate product to start. pgvector adds vector columns and HNSW indexes to PostgreSQL (Supabase and Neon include it), sqlite-vec adds vector search to SQLite, and MongoDB Atlas, Redis and Elasticsearch have it built in. Dedicated services such as Pinecone, Qdrant, Weaviate, Milvus and Cloudflare Vectorize run the index for you and add features for very large collections.

Pros

  • Finds relevant results even when the wording differs
  • Works across languages, and for images as well as text
  • Can live inside a database you already run (pgvector, sqlite-vec)
  • The retrieval backbone of RAG and recommendation features

Cons

  • Misses exact terms such as product codes, names and error numbers
  • Every document and every query needs an embedding call
  • Results are harder to explain than keyword matches
  • Changing the embedding model means re-indexing everything

Pick it when

  • Search over help articles, notes or products where wording varies
  • The retrieval step of a RAG chatbot
  • Related items, similar tickets or duplicate detection

Skip it when

  • Users search for exact codes, names or phrases (keyword search fits)
  • Structured filters such as price, size or date answer the question on their own

What it costs · Open source

pgvector and sqlite-vec are free. Pinecone, a hosted option, has a free Starter plan, a $20 a month Builder plan and a $50 monthly minimum on Standard.

Vector search pricing (opens in a new tab)Approximate, checked September 2026.

Also called: vector database, semantic search, pgvector, sqlite-vec, Pinecone, nearest neighbour search

Open Vector search as a pageOfficial site (opens in a new tab)

Retrieval-augmented generation (RAG)

Pattern
A way to make a language model answer from your own documents: search them for the passages relevant to a question, then hand those passages to the model with the question.

Documents are split into chunks, embedded and indexed ahead of time. When a question arrives, the app retrieves the most relevant chunks (with vector, keyword or hybrid search), places them in the prompt with instructions such as 'answer only from these sources and cite them', and the model writes an answer grounded in that text. The name comes from a 2020 research paper by a team at Facebook AI Research.

RAG keeps answers current without retraining: update a document and the next answer reflects it. It also allows citations and per-user permissions, since you decide what each user's search may return. Most of the quality comes from retrieval rather than the model, so chunking, search method, reranking and good source documents matter more than prompt wording. It reduces hallucination without removing it, so answers still need spot checks and a fallback when nothing relevant is found.

Pros

  • Answers from your own, up-to-date content without training a model
  • Can cite sources, so answers are easier to check
  • Respects permissions when search is filtered per user
  • Works with any model, so you can switch providers later

Cons

  • Only as good as the retrieval; poor chunks give poor answers
  • Extra moving parts: ingestion, embeddings, an index and evaluation
  • Retrieved text adds input tokens to every call
  • Still hallucinates sometimes, especially when sources disagree

Pick it when

  • Support bots, internal knowledge search and document Q&A
  • Content changes often and answers must reflect it
  • Users need to see where an answer came from

Skip it when

  • The model must learn a style or format rather than facts (consider fine-tuning)
  • All the text it needs already fits comfortably in the prompt

Also called: RAG, retrieval augmented generation, chat with your documents, grounding

Open Retrieval-augmented generation (RAG) as a page

Building with models

Fine-tuning

ConceptPay as you go
Training an existing model further on your own examples so it picks up a particular style, format or narrow task, producing a customised version of that model.

You prepare pairs of example inputs and ideal outputs, from dozens to thousands of them, and the provider (or your own GPUs) trains the model on them for a few passes, called epochs. Full fine-tuning updates all the weights; parameter-efficient methods such as LoRA train a small add-on instead, which is far cheaper and is how most open models are tuned. Preference and reinforcement tuning go further by rewarding better answers over worse ones.

Fine-tuning is good at behaviour: a consistent tone, a strict output format, sorting tickets into your categories, or getting a small, cheap model to match a bigger one on one task. It is a poor way to teach facts that change, which is RAG's job. A tuned model is tied to its base model and must be retrained when that model is retired. Availability shifts too: OpenAI closed its self-serve fine-tuning to new customers in 2026, while Google Cloud still offers it and open models can be tuned on your own GPUs.

Pros

  • Locks in a tone, format or task more reliably than prompting
  • Shorter prompts, since the instructions live in the model
  • Lets a small, cheap model handle one narrow job well

Cons

  • Needs a good set of example data, which takes effort to build
  • Training costs money, and tuned models may cost more to run
  • Must be redone when the base model is retired
  • Does not reliably teach new facts

Pick it when

  • A prompt with examples still gives inconsistent format or tone
  • High-volume, narrow tasks where a small tuned model saves money
  • You have hundreds of good examples of the output you want

Skip it when

  • The goal is answering from documents (use RAG)
  • Better prompts or structured output have not been tried yet

What it costs · Pay as you go

Charged per training token (data size times passes) plus use. On Google Cloud about $1.50 to $25 per million training tokens; tuned newer Gemini models cost 1.5 times the base rate to use.

Fine-tuning pricing (opens in a new tab)Approximate, checked September 2026.

Also called: model fine-tuning, LoRA, supervised fine-tuning, SFT, custom models

Open Fine-tuning as a page

AI agents and tool calling

Concept
Tool calling lets a model ask your code to run a function, such as looking up an order or sending an email. An agent is a loop in which the model keeps choosing tools until a task is done.

You describe each tool with a name, a plain-language description and a JSON Schema for its inputs. Instead of answering, the model can reply with a request such as get_order_status with order_id 1042. Your code runs the function, sends the result back and the model continues, possibly calling more tools, until it can give a final answer. The model never runs anything itself: your code decides what actually executes.

An agent wraps this in a loop with a goal, a budget and a stopping rule, and may plan, browse, write code or hand work to other agents. Agents fail in new ways: they can loop, run up token bills, act on a misunderstanding, or follow instructions hidden in a web page or email (prompt injection). Keep tools narrow, require confirmation for anything destructive or costly, log every step and cap the number of turns.

Also called: tool calling, function calling, tool use, agents, agentic AI

Open AI agents and tool calling as a page

Model Context Protocol (MCP)

ProtocolOpen source
An open standard for connecting AI apps to tools and data: build an MCP server once, and any app that speaks MCP (Claude, ChatGPT, VS Code, Cursor and others) can use it.

Anthropic introduced MCP in November 2024 and in December 2025 donated it to the Agentic AI Foundation, a Linux Foundation fund that Anthropic co-founded with OpenAI and Block. An MCP server exposes tools (actions the model can call), resources (data it can read) and prompts (reusable templates). An MCP client inside a host app, such as a chat app, an editor or an agent, discovers them and offers them to the model, exchanging JSON-RPC messages.

Local servers run as a process on your machine and talk over standard input and output; remote servers are reached over HTTP (Streamable HTTP) and usually sign users in with OAuth. Thousands of servers exist, for GitHub, databases, file systems, browsers, Slack and more. A server can do anything its credentials allow, and its tool descriptions go straight into the model's context, so install only servers you trust and give them the narrowest access that works.

Pros

  • Open, vendor-neutral standard backed by the major AI companies
  • Write a tool integration once and use it from many apps
  • Large and growing catalogue of ready-made servers
  • Works locally for private data or remotely as a hosted service

Cons

  • A malicious or careless server can leak data or take harmful actions
  • Many tools in context use up tokens and can confuse the model
  • The specification still changes, and older servers lag behind it

Pick it when

  • You want your product's data or actions usable from AI assistants
  • An internal agent needs the same tools across several apps
  • Connecting a coding assistant to docs, databases or issue trackers

Skip it when

  • One app calls a few tools of its own (plain tool calling is simpler)

What it costs · Open source

Free: the specification and official SDKs are open source.

Approximate, checked September 2026.

Also called: MCP, MCP server, MCP client, MCP tools

Open Model Context Protocol (MCP) as a pageOfficial site (opens in a new tab)

Structured output

Concept
Asking a model to reply in an exact machine-readable shape, usually JSON that matches a schema you define, so your code can use the answer directly.

Rather than parsing free text, you pass a JSON Schema (or a Zod or Pydantic model that the SDK converts) describing the fields you want, such as name, email and a priority from a fixed list. With strict structured outputs, offered by OpenAI, Google and Anthropic, the provider constrains generation so the reply always parses and matches the schema. Older 'JSON mode' only promised valid JSON, not the right fields.

It turns a model into a dependable component: extracting invoice fields, tagging support tickets, filling a form from an email, or returning a plan an agent can carry out. The shape is guaranteed but the content is not, so values still need validation and business rules (an extracted total can be wrong even when it is a valid number). Keep schemas small and descriptive, since field names and descriptions guide the model.

Also called: structured outputs, JSON mode, JSON Schema output, response schema

Open Structured output as a page

Side by side

Differences

How the options in this area compare on the questions that usually decide the choice.

OpenAI vs Gemini vs Claude vs OpenRouter

Open as a page: OpenAI vs Gemini vs Claude vs OpenRouter

Three model makers and one gateway in front of them all. Prices are rough ranges per million input tokens; output tokens cost about four to six times more.

CompareOpenAIGeminiClaudeOpenRouter
What it isOpenAI's own modelsGoogle's own modelsAnthropic's own modelsOne API for hundreds of models
Model tiersSmall, mid-size and flagship reasoningFlash-Lite, Flash and ProHaiku, Sonnet, Opus and aboveWhatever each provider offers
Free tierNoYes, with rate limitsNoFree models with daily limits
Input price per millionAbout $0.10 to $10About $0.10 to $2 (more for long prompts)About $1 to $10Provider price, plus a fee on credits
Long contextYes, priced higher past about 272K tokensUp to about 1M tokensUp to about 1M tokens on recent modelsDepends on the model
Stand-out extrasImage, speech and realtime voice modelsSearch grounding, video and audio inputComputer use, citations, prompt cachingRouting, fallbacks, one bill
Your dataNot used for training by defaultFree-tier prompts may improve Google productsNot used for training by defaultSkips hosts that train on prompts by default
Watch out forModel retirements force upgradesDifferent terms for free and paid usePricier small models than rivalsFee on credits and one more middleman

How to choose

  • Pick OpenAI for a wide range of models and mature tooling across text, images and speech.
  • Pick Gemini to prototype for free, or for long documents, audio and video at low per-token prices.
  • Pick Claude for coding assistants, long documents and agents that lean heavily on tools.
  • Pick OpenRouter to compare models, or to mix providers behind one key and one bill.

Cloud API vs gateway vs running it yourself

Open as a page: Cloud API vs gateway vs running it yourself

Three ways to get a model's answers into an app: sign up with one provider, go through a gateway to many, or run an open model on hardware you control.

CompareA cloud APIA gateway (OpenRouter)Your own machine (Ollama)
SetupAn API key and a few lines of codeOne key for many providersInstall, download a model, run it
CostPay per tokenProvider prices plus a fee on creditsNo token bill; you pay for hardware
Model choiceOne vendor's modelsHundreds, across vendorsOpen-weight models only
Quality ceilingUp to flagship modelsUp to flagship modelsGood, but behind the largest models
PrivacyData goes to the providerData goes to the gateway and a providerData stays on your machine
Works offlineNoNoYes
Serving many usersHandled by the providerHandled by the providersYour job, and it needs GPUs
Watch out forLock-in and price changesAn extra middleman and feesHardware limits and slower answers

How to choose

  • Pick a cloud API when you want capable models with the least setup and one vendor to deal with.
  • Pick a gateway when you want to try or mix many models without separate accounts.
  • Pick your own machine when data must stay private, the app must work offline, or token bills must be zero.

Both make a general model fit your product. RAG hands the model the right information at question time; fine-tuning changes how the model behaves. Many teams start with RAG and fine-tune later, if at all.

CompareRAGFine-tuning
What it changesWhat the model knows for this answerHow the model behaves every time
Good forFacts, documents, policies, cataloguesTone, format, narrow classification tasks
Keeping it currentUpdate a document; the next answer uses itRetrain to change anything
What you needDocuments and a search indexHundreds of good example pairs
CostEmbeddings, storage and extra prompt tokensTraining runs, sometimes higher usage rates
Showing sourcesCan cite the passages it usedCannot point to where an answer came from
Getting startedQuick with managed toolsSlower: building the dataset takes most of the effort
Switching modelsWorks with any modelTied to one base model

How to choose

  • Pick RAG when answers must come from your own content, stay current and show their sources.
  • Pick fine-tuning when prompting cannot get a consistent style, format or narrow skill and you have good examples.
  • Combine them when a tuned model should also answer from fresh documents.

Keyword vs vector vs hybrid search

Open as a page: Keyword vs vector vs hybrid search

Keyword search matches the words people type, vector search matches what they mean, and hybrid search runs both and merges the results.

CompareKeyword searchVector searchHybrid search
Matches onShared words, after stemmingSimilar meaningBoth
Good atNames, codes and exact phrasesParaphrases, synonyms, other languagesMixed, real-world queries
Weak atDifferent wording for the same ideaExact terms and rare wordsLittle, beyond extra complexity
NeedsA full-text indexAn embedding model and a vector indexBoth, plus a merge step
CostCheap, no model callsAn embedding call per document and queryBoth costs combined
Typical toolsSQLite FTS5, PostgreSQL, Meilisearchpgvector, sqlite-vec, PineconeElasticsearch, Weaviate, PostgreSQL with both
Example strength'SKU-4471' finds that exact product'shoes for rainy days' finds waterproof bootsHandles both kinds in one search box

How to choose

  • Pick keyword search for catalogues, codes and names, or as a cheap, predictable baseline.
  • Pick vector search when people describe what they want in their own words.
  • Pick hybrid search for RAG and general site search, where both kinds of query turn up.

Crafted in the dark. Shipped to the world.

Tell us what you are building. You get a private project space with a proposal and a line-by-line quote within a day.