How it works
ONNX (Open Neural Network Exchange) is an open file format for trained models, so a model built in PyTorch or TensorFlow can be exported once and run elsewhere. ONNX Runtime loads that file and runs it efficiently on whatever hardware is present, through plug-in execution providers for CPUs, NVIDIA GPUs (CUDA, TensorRT), Windows GPUs (DirectML), Apple devices (Core ML) and more.
It suits smaller, focused models rather than big chat models: text embeddings for search, classifiers, image recognition, OCR and speech. onnxruntime-node runs them inside a Node.js server, and onnxruntime-web runs them in the browser with WebAssembly or WebGPU, so data can stay on the user's device. Hugging Face's Transformers.js is built on it and handles downloading and preparing models for you.
ONNX Runtime pricing
Related terms
More in AI and LLMs
Run it yourself