AI and LLMs · Comparison
Cloud API vs gateway vs running it yourself
Three ways to get a model's answers into an app: sign up with one provider, go through a gateway to many, or run an open model on hardware you control.
3 options · 8 questions side by side · updated
| Compare | A cloud API | A gateway (OpenRouter) | Your own machine (Ollama) |
|---|---|---|---|
| Setup | An API key and a few lines of code | One key for many providers | Install, download a model, run it |
| Cost | Pay per token | Provider prices plus a fee on credits | No token bill; you pay for hardware |
| Model choice | One vendor's models | Hundreds, across vendors | Open-weight models only |
| Quality ceiling | Up to flagship models | Up to flagship models | Good, but behind the largest models |
| Privacy | Data goes to the provider | Data goes to the gateway and a provider | Data stays on your machine |
| Works offline | No | No | Yes |
| Serving many users | Handled by the provider | Handled by the providers | Your job, and it needs GPUs |
| Watch out for | Lock-in and price changes | An extra middleman and fees | Hardware limits and slower answers |
How to choose between A cloud API, A gateway (OpenRouter) and Your own machine (Ollama)
- Pick a cloud API when you want capable models with the least setup and one vendor to deal with.
- Pick a gateway when you want to try or mix many models without separate accounts.
- Pick your own machine when data must stay private, the app must work offline, or token bills must be zero.
The options
- A cloud APIOpenAI's paid API for its GPT family of models, the technology behind ChatGPT, used to add text, image, speech and search features to your own app.
- A gateway (OpenRouter)A gateway that gives you one API key and one bill for hundreds of models from many providers, so you can switch or mix models without new integrations.
- Your own machine (Ollama)A free, open-source app that downloads and runs open-weight language models on your own computer with one command, and serves them through a local API.
More comparisons
- OpenAI vs Gemini vs Claude vs OpenRouterThree model makers and one gateway in front of them all. Prices are rough ranges per million input tokens; output tokens cost about four to six times more.
- RAG vs fine-tuningBoth make a general model fit your product. RAG hands the model the right information at question time; fine-tuning changes how the model behaves. Many teams start with RAG and fine-tune later, if at all.
- Keyword vs vector vs hybrid searchKeyword search matches the words people type, vector search matches what they mean, and hybrid search runs both and merges the results.
- Node.js vs Deno vs BunThree runtimes for JavaScript and TypeScript on the server. Much of the same code runs on all three; they differ in built-in tools, security defaults, speed and how long each has been used in production.
- Express vs Fastify vs HonoThree JavaScript web frameworks with a similar feel. Express is the long-standing default, Fastify focuses on throughput and structure, and Hono is built on web standards so it can run almost anywhere.
- FastAPI vs Django vs FlaskThree widely used Python web frameworks. Django includes almost everything, Flask includes almost nothing, and FastAPI focuses on typed, self-documenting APIs.
Crafted in the dark. Shipped to the world.
Tell us what you are building. You get a private project space with a proposal and a line-by-line quote within a day.