What Is AI Inference? How Trained Models Actually Run in Production
Every time you ask a chatbot a question, upload a photo for automatic tagging, or get a product recommendation, a trained model is doing one specific job: inference. Training gets most of the headlines, but inference is where AI meets real users, real latency limits, and real bills. If you build with AI, understanding what AI inference is and what drives its speed and cost is one of the most useful skills you can develop.
What Is AI Inference?
AI inference is the process of using a trained model to produce an output from new input. Training teaches the model by adjusting its parameters on large datasets. Inference freezes those parameters and simply applies them: input goes in, a prediction, classification, or generated text comes out.
A simple analogy: training is a chef spending years learning techniques and recipes. Inference is that chef cooking your order tonight. The learning is already done, and now the question is how fast, how consistently, and at what cost the meal reaches your table.
AI Inference vs Training
The two phases have very different profiles:
| Aspect | Training | Inference |
| Goal | Learn parameters from data | Apply learned parameters to new input |
| Frequency | Occasional, large jobs | Continuous, every user request |
| Compute pattern | Forward and backward passes, huge batches | Forward pass only, often small or variable batches |
| Main concern | Total time and cost to train | Latency, throughput, and cost per request |
| Hardware | Large GPU or TPU clusters | Anything from a phone to a GPU server |
For most teams, training happens once or rarely (or not at all, if they use a pretrained model), while inference runs every day. Over a product’s lifetime, inference is often the larger cost.
How LLM Inference Works Step by Step
Large language models generate text one token at a time. A token is a small chunk of text, often a word or part of a word. Here is what happens when you send a prompt:
1. Tokenization. Your text is split into tokens and converted to numbers.
2. Prefill. The model processes your entire prompt in parallel and builds internal state. This phase is compute-heavy, and it determines how long you wait for the first token.
3. Decode. The model generates output tokens one by one. Each new token depends on all the previous ones, so this stage is sequential. It tends to be limited by how quickly the GPU can read model weights from memory, which is why memory bandwidth matters so much.
4. Sampling. At each step the model produces probabilities over possible next tokens, and a sampling strategy (controlled by settings like temperature and top-p) picks one.
5. Detokenization. Tokens are converted back into readable text and streamed to the user.
The KV Cache: Why Memory Matters
During decoding, the model would repeat a lot of work if it recomputed attention over the whole conversation for each new token. To avoid that, inference engines store intermediate results, called the key-value (KV) cache. It makes generation much faster, but it consumes GPU memory that grows with prompt length and the number of concurrent users. This is one reason long context windows and heavy traffic can exhaust memory even when the model itself fits.
The Metrics That Matter
When you evaluate or tune inference, these are the numbers to watch:
- Time to first token (TTFT): how long until the user sees the first piece of output. This drives how responsive a chat app feels.
- Tokens per second: the speed of generation after it starts.
- End-to-end latency: total time from request to complete response.
- Throughput: total tokens or requests handled per second across all users.
- Cost per request or per million tokens: the number your finance team cares about.
- GPU utilization and memory use: how well you are using the hardware you pay for.
Latency and throughput often trade off. Batching many requests together improves throughput and lowers cost per token, but each individual user may wait a little longer.
Real-Time vs Batch Inference
- Real-time (online) inference: responds to each request immediately. Chatbots, search, and fraud detection need this, so latency is critical.
- Batch (offline) inference: processes large volumes of data on a schedule, such as summarizing a million documents overnight. Latency matters little, so you can pack the hardware tightly and save money.
Choosing the right mode for each workload is one of the easiest ways to control cost.
Where Inference Can Run
- On-device or locally: on a laptop, phone, or workstation. Great for privacy and offline use. See our guide to [running LLMs locally].
- In the cloud on your own GPUs: you rent or own servers and deploy an inference engine. See our [step-by-step vLLM deployment tutorial].
- Through managed inference APIs: you send requests to a provider and pay per token, with no infrastructure to manage.
Each option trades control, cost, and effort differently, and many teams mix them.
What Makes Inference Fast or Slow?
Several factors shape performance:
- Model size: more parameters mean more memory and more computation per token.
- Numeric precision: lower-precision formats (like 8-bit or 4-bit) shrink memory and can speed things up. This technique is called quantization.
- Hardware: GPU memory size and bandwidth are often the deciding factors. Our [GPU guide for AI inference] covers this in depth.
- Software stack: engines with optimizations like continuous batching and efficient attention can multiply throughput on the same GPU.
- Prompt and output length: longer prompts and responses use more time and memory.
We break down concrete fixes in [how to reduce LLM inference cost and latency].
Frequently Asked Questions
Is inference cheaper than training? A single inference request is far cheaper than a training run. But because inference happens continuously at scale, its total cost over time can exceed training.
Do I need a GPU for inference? Not always. Small models run acceptably on CPUs and laptops, but GPUs are usually much faster for larger models.
What is the difference between latency and throughput? Latency is how long one request takes. Throughput is how much work the system completes per unit of time. Improving one can sometimes hurt the other.
Key Takeaways
AI inference is the stage where a trained model turns input into output. For language models it involves prefill, decode, and a KV cache, and its performance is governed by model size, precision, hardware, and software. Once you understand these pieces, every optimization and deployment decision becomes easier to reason about.
