To run an open-weight model locally, choose an exact checkpoint whose license and capabilities fit the task, estimate memory from total parameters and precision, select a compatible model format and runtime, then benchmark at the target context length. For a simple desktop workflow, Ollama is often convenient; for GGUF and low-level control, llama.cpp is common; for GPU servers with concurrency, vLLM or SGLang may be more appropriate. Local inference keeps the model execution under your control, but connected services and logs still determine the real data path.
1. Choose the exact model, not just the family
Start with the workload. Chat, code, document understanding, long-context retrieval, vision and agentic tool use can favor different architectures. Then choose the exact checkpoint and read its license.
A smaller model that fits entirely in fast local memory can outperform a larger model that spends most of its time moving tensors between RAM and VRAM. Likewise, a strong text model may be the wrong choice if the application requires images or audio.
Use the model directory to shortlist by size, focus and license, then verify the publisher source.
2. Match model size to memory
Estimate weight memory from parameter count and precision, then add KV cache and runtime overhead. For a first local deployment, leaving generous headroom is more useful than forcing the largest possible checkpoint to load.
For example, an 8B model at roughly 4-bit weight precision is often in the range where modern consumer hardware can experiment comfortably. A 32B model at 4-bit may need roughly 16 GB just for raw weight data before overhead. A 70B checkpoint quickly moves into high-memory workstation territory.
On unified-memory systems, total memory is shared. On discrete GPUs, VRAM and system RAM play different performance roles.
3. Pick the model format for the runtime
If the target runtime is llama.cpp or a tool built around the GGUF ecosystem, choose a verified GGUF conversion and quantization. If the target is Transformers, vLLM or SGLang, the publisher’s Safetensors-based repository is often the natural starting point.
Do not download a random conversion only because the filename contains the right model name. Verify source model, quantization, tokenizer, chat template and uploader reputation. For production, publisher-provided or well-documented conversions reduce provenance risk.
4. Select a runtime by operational need
| Need | Good starting point | Why |
|---|---|---|
| Quick desktop/local API | Ollama | Simple install, model workflow and API |
| Portable GGUF inference | llama.cpp | Broad hardware support and low-level control |
| GPU production API | vLLM | High-throughput serving and batching |
| Advanced production serving | SGLang | Prefix caching, distributed serving and high throughput |
5. Treat the local model server as infrastructure
A local model process may expose an HTTP port, load third-party model artifacts and write prompts or traces to logs. Production deployment therefore needs authentication, network boundaries, TLS where appropriate, access controls, secret management and a clear logging policy.
vLLM’s own documentation warns that its API-key mechanism does not protect every server endpoint. The general lesson applies to all model servers: place production inference behind appropriate reverse proxies or service-mesh controls rather than assuming a development default is a security boundary.
If the model handles internal documents, the vector database, embedding service and RAG application need the same review.
6. Benchmark what users actually experience
Measure more than raw tokens per second. Users experience time to first token, output speed, queue delay, context handling and failure rate. Compare short prompts and long prompts. Test the maximum context you genuinely need rather than the maximum advertised context.
For multi-user services, measure concurrent throughput and memory growth. A setup that feels fast for one terminal session can collapse under ten simultaneous long-context requests.
From local experiment to self-hosted service
Promotion checklist
- Pin the exact model and runtime versions.
- Record license and source URL.
- Document quantization and chat template.
- Add health checks, metrics and capacity limits.
- Protect the inference endpoint.
- Define retention for prompts, outputs and logs.
- Test rollback before updating the model.
Once these controls exist, local inference becomes a maintainable service rather than an experiment running on someone’s workstation.
Frequently asked questions
What is the easiest way to run an open-weight model locally?
For many users, Ollama is a convenient starting point because it packages model management and a local API. llama.cpp is useful when you want direct GGUF control.
Do I need a GPU?
No for every model. Smaller and quantized models can run on CPUs, though GPUs or unified accelerators usually improve speed. Model size and performance expectations determine what is practical.
Is local inference automatically private?
No. The inference process may be local, but telemetry, embeddings, RAG services, backups or logs can create external data flows.
How do I know whether a model will fit?
Estimate weight memory from total parameters and precision, then add KV cache and runtime headroom. The exact runtime and context length matter.
Apply this knowledge
Move from the concept to a concrete deployment shortlist.
Primary sources and technical references
OpenWeightModels prefers publisher documentation, standards bodies, official repositories and original research papers. The source material remains authoritative where it changes.