How to Deploy an LLM With vLLM: A Step-by-Step Tutorial for Developers
Once a model works on your laptop, the next question is how to serve it to real users. vLLM is one of the most popular open-source inference engines for that job. It is fast, memory efficient, and can expose an API that works with many existing tools. In this tutorial, you will learn how to deploy an LLM with vLLM, from a first test to a production-minded setup.
What Is vLLM and Why Use It?
vLLM is an open-source library for high-throughput LLM inference and serving. Its headline features include:
- PagedAttention: manages the KV cache in memory pages, which reduces waste and lets you serve more concurrent requests on the same GPU.
- Continuous batching: merges incoming requests dynamically to keep the GPU busy.
- OpenAI-compatible server: exposes endpoints similar to popular hosted APIs, so many client libraries work with only a change of base URL.
- Broad model support: works with many open-weight models from Hugging Face.
- Multi-GPU support: can split large models across GPUs with tensor parallelism.
If you want background on why these features matter, read [what AI inference is] and [how to reduce LLM inference cost and latency].
Prerequisites
Before you start, you will need:
- A machine with a supported GPU (a cloud instance is fine). See our [GPU guide for AI inference] for sizing.
- Recent GPU drivers and CUDA support installed (or use a prebuilt container).
- Python (for the pip route) or Docker with the NVIDIA Container Toolkit (for the container route).
- A Hugging Face account and access token if your chosen model is gated.
- A model that fits in your GPU memory. For a first test, choose a small model such as a 1B to 3B instruct model.
Requirements and commands change between versions, so always check the official vLLM documentation for current instructions.
Step 1: Prepare the Environment
Create an isolated Python environment so dependencies do not conflict:
python -m venv vllm-env
source vllm-env/bin/activate
Confirm your GPU is visible:
nvidia-smi
You should see your GPU model and available memory. If not, fix drivers before continuing.
Step 2: Install vLLM
pip install vllm
This downloads vLLM and its dependencies, which can be large. Installing into a clean virtual environment is strongly recommended.
Step 3: Launch an OpenAI-Compatible Server
Start the server with a model name from Hugging Face. This example uses a small instruct model:
vllm serve Qwen/Qwen2.5-1.5B-Instruct
By default, the server listens on port 8000. On first launch, vLLM downloads the model, which takes a few minutes. If the model is gated, log in first or set your token as an environment variable (for example HF_TOKEN).
Step 4: Send Your First Request
In another terminal, test the chat endpoint:
curl http://localhost:8000/v1/chat/completions \
-H “Content-Type: application/json” \
-d ‘{
“model”: “Qwen/Qwen2.5-1.5B-Instruct”,
“messages”: [{“role”: “user”, “content”: “Explain inference in one sentence.”}]
}’
If you get a JSON response containing generated text, your deployment works.
You can also use an OpenAI-style client library by pointing its base URL to http://localhost:8000/v1. Because the API format is compatible, switching between a hosted provider and your own server can require very small code changes.
Step 5: Turn On Streaming
For chat apps, streaming tokens as they are generated makes responses feel much faster. Add “stream”: true to your request body, or enable streaming in your client library. Users see the first words quickly, which improves perceived latency even when total generation time is unchanged.
Step 6: Tune Key Settings
A few parameters have a large impact:
- –max-model-len: limits the maximum context length. Lowering it reduces KV cache memory needs and can prevent out-of-memory errors.
- –gpu-memory-utilization: controls what fraction of GPU memory vLLM may use. Higher values allow a larger KV cache and more concurrency, but leave less margin.
- –dtype: chooses numeric precision. Lower precision saves memory.
- –tensor-parallel-size: splits the model across multiple GPUs when it does not fit on one.
- Quantized models: loading a quantized version of a model can cut memory needs significantly.
Example:
vllm serve Qwen/Qwen2.5-1.5B-Instruct \
–max-model-len 4096 \
–gpu-memory-utilization 0.90
Change one setting at a time and measure the effect on latency, throughput, and stability.
Step 7: Run vLLM With Docker
For reproducible deployments, use the official container image:
docker run –gpus all -p 8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
–env “HF_TOKEN=your_token_here” \
vllm/vllm-openai:latest \
–model Qwen/Qwen2.5-1.5B-Instruct
Mounting the Hugging Face cache directory avoids re-downloading models each time the container restarts. Pin a specific image version in production instead of latest, so updates do not surprise you.
Step 8: Prepare for Production
A working server is not yet a production service. Consider these steps:
- Put it behind a reverse proxy or API gateway such as Nginx or a cloud load balancer, with HTTPS.
- Add authentication. Never expose an unprotected inference endpoint to the public internet. vLLM supports API key options, and gateways can enforce more.
- Monitor it. Track latency, throughput, queue depth, GPU memory, and error rates. vLLM exposes metrics that tools like Prometheus and Grafana can collect.
- Scale horizontally. Run multiple replicas behind a load balancer, and use autoscaling on Kubernetes or your cloud platform to match traffic.
- Set request limits. Cap maximum tokens and concurrent requests to protect stability and cost.
- Test under load. Simulate realistic traffic before launch to find bottlenecks and choose sensible limits.
- Plan for updates. Test new vLLM and model versions in staging before rolling them out.
Troubleshooting Common Issues
Out-of-memory error at startup. Lower –max-model-len, reduce –gpu-memory-utilization slightly, choose a smaller or quantized model, or add GPUs with tensor parallelism.
Model fails to download. Check your Hugging Face token, confirm you accepted the model’s license terms, and verify disk space.
Slow responses. Check that the GPU is actually being used, that prompts are not extremely long, and that concurrency is not overwhelming the GPU. Review the optimizations in [how to reduce LLM inference cost and latency].
Connection refused. Confirm the server is running, the port is correct, and firewall or container port mappings allow access.
vLLM vs Other Inference Options
vLLM is not the only choice. TensorRT-LLM offers deep NVIDIA-specific optimization, Hugging Face Text Generation Inference is another popular server, and tools like Ollama are simpler for single-user local use, as covered in [how to run an LLM locally]. The best option depends on your hardware, scale, and how much tuning effort you want to invest. Benchmark on your own workload before deciding.
Frequently Asked Questions
Is vLLM free? Yes, it is open source. You pay for the hardware or cloud GPUs you run it on.
Can vLLM run on CPU? It is designed primarily for GPUs, with support for some other accelerators. For CPU-only use, tools like llama.cpp are usually a better fit.
Is the vLLM API really compatible with OpenAI clients? It provides OpenAI-compatible endpoints, so many client libraries work by changing the base URL. Not every advanced feature is identical, so test the features you rely on.
How many users can one GPU handle? It depends on the model size, prompt and output lengths, and GPU memory. Load testing is the only reliable way to find out.
Key Takeaways
To deploy an LLM with vLLM, install it, launch vllm serve with your model, test the OpenAI-compatible endpoint, tune context length and memory settings, containerize it, and add authentication, monitoring, and scaling for production. Start small, measure, and improve step by step.
