> Bron: https://neuralex.nl/en/blog/hardware-voor-lokale-ai
> RAM, VRAM, quantization and storage determine which LLMs you can run locally. Practical memory rules for 7B to 70B+ models on CPUs, GPUs and Apple Silicon.

[Back to Insights](/en/blog)

On-prem ·19 August 2026 ·6 min read

# What hardware do you need to run a language model locally?

RAM, VRAM, quantization and storage determine which LLMs you can run locally. Practical memory rules for 7B to 70B+ models on CPUs, GPUs and Apple Silicon.

To run a language model locally, the primary hardware requirement is enough memory for the model weights, with additional capacity for the inference runtime, context and KV cache. As a baseline, FP16 weights require roughly **2 GB of memory per billion parameters**. At 8-bit precision that falls to about **1 GB per billion parameters**, while 4-bit quantization needs roughly **0.5 GB per billion parameters**. Those figures cover the weights themselves, not the complete memory footprint.

A 7B model therefore needs roughly 14 GB for FP16 weights, around 7 GB at 8-bit and about 3.5 GB at 4-bit. For a 13B model the corresponding figures are approximately 26, 13 and 6.5 GB; for 70B they are around 140, 70 and 35 GB. Whether that capacity must be available as system RAM, GPU VRAM or unified memory depends on the hardware platform and inference engine — separate from the broader question of [when on-prem AI is worth the investment](/en/blog/on-prem-ai-wanneer-de-moeite).

## Parameter count sets the baseline memory requirement

Parameter count is the most useful starting point when sizing hardware. A model described as 7B contains roughly seven billion parameters; a 70B model has roughly seventy billion. For inference, weight memory can be estimated with a simple rule:

Model size

FP16

8-bit

4-bit

7B

~14 GB

~7 GB

~3.5 GB

13B

~26 GB

~13 GB

~6.5 GB

30B

~60 GB

~30 GB

~15 GB

70B

~140 GB

~70 GB

~35 GB

These numbers are not complete system requirements. Quantized formats may also contain scaling factors, metadata and other auxiliary data, while the inference engine needs memory for buffers and intermediate operations. A machine with exactly the theoretical minimum amount of memory is therefore usually undersized.

## Quantization makes larger models practical

Many model weights are originally distributed or executed in FP16 or BF16, using 16 bits per parameter. At that precision, one billion parameters require roughly 2 GB of storage and memory. With 8-bit quantization, weights are represented using approximately one byte per parameter, cutting weight memory roughly in half. At 4-bit, the theoretical requirement falls to around half a byte per parameter.

Actual quantized files are normally somewhat larger than this theoretical minimum because quantization schemes need additional scaling information and metadata. Formats aimed at CPU and mixed CPU/GPU inference, such as GGUF variants, also differ from GPU-focused quantization methods in their exact overhead. Reducing precision is not purely a memory decision — it can affect model quality and inference performance. Nevertheless, 4-bit quantization is widely used for local inference because it allows models to run on hardware that cannot accommodate the FP16 weights. See also [which open model fits which hardware](/en/blog/lokale-modellen-2026-vergeleken).

## Context length and KV cache also consume memory

Model weights are only part of the total memory requirement. During autoregressive inference, the model stores information related to previously processed tokens in a KV cache. Longer contexts require a larger cache. The exact memory requirement depends on factors including model architecture, number of layers, context length, cache precision and the number of simultaneous requests.

A model that fits comfortably when processing a short prompt can therefore run out of memory when the context window is increased substantially. This becomes even more important on shared inference servers, where multiple requests may occupy memory at the same time. Sizing hardware solely from the size of a quantized model file is therefore insufficient — capacity must remain available for the runtime and the expected context workload.

## CPU inference: more accessible RAM, usually lower speed

A local LLM does not require a GPU. Inference engines optimized for quantized models can load all model weights into normal system RAM and execute them on the CPU. The main advantage is memory capacity: general-purpose systems can often provide substantially more RAM than a single GPU provides VRAM, making it possible to run models that would not fit entirely on one graphics card.

The trade-off is generally inference speed. LLM inference relies heavily on large matrix operations and memory bandwidth, workloads for which modern GPUs are well suited. CPU inference is therefore most useful when low latency is not critical, concurrent usage is limited, or the ability to hold a larger model in memory matters more than maximum token throughput. Memory bandwidth and CPU architecture are important here; core count alone is not a reliable measure of LLM inference performance.

## GPU inference: VRAM is usually the limiting resource

For higher-performance local inference, a GPU is often the most practical platform. Model weights are loaded wholly or partly into VRAM and matrix operations are executed on the GPU. The key specification is therefore not just compute performance but available VRAM.

A consumer GPU with enough VRAM can handle 7B and other relatively compact models well, particularly with 8-bit or 4-bit quantization. As model size increases, VRAM increasingly becomes the hard constraint. If the complete model does not fit in VRAM, some inference engines can keep part of it in system RAM and offload selected layers to the GPU — this extends the range of models that can run, but transfers between CPU memory and GPU memory can reduce performance.

## Apple Silicon and unified memory

Apple Silicon differs from a conventional PC with a discrete GPU because the CPU and GPU share the same unified memory pool rather than using separate system RAM and VRAM. For local LLM inference, this can be useful — a machine with a large unified-memory configuration can make much more memory available to GPU workloads than a discrete consumer GPU with a relatively small VRAM pool.

That entire capacity cannot be allocated to the model, however. macOS, other applications, the inference runtime, the context and the KV cache all use the same memory. Unified-memory systems should therefore be sized for total memory consumption rather than just the model file. The installed memory on Apple Silicon systems also cannot be upgraded later.

## How much storage do you need?

Model files also require local storage. Disk requirements broadly follow weight size: a 70B model in a 4-bit representation occupies tens of gigabytes, whereas FP16 weights for the same parameter count require well over one hundred gigabytes. Real installations often need additional space for alternative quantizations, tokenizer files, inference software and multiple model versions.

An SSD is preferable to a hard disk, especially when loading large models. Once the weights have been loaded into RAM or VRAM, storage performance normally does not determine token generation speed, but faster storage reduces model loading time.

## When is one consumer GPU enough?

A single consumer GPU is generally sufficient when the target model, its quantized representation, runtime overhead and required KV cache all fit comfortably inside the available VRAM. For compact models around 7B to roughly 13B, this is often achievable with an appropriate quantization level. Larger models can also run on a single GPU if enough VRAM is available, but memory requirements should always be checked against the specific model and intended context length.

When a model no longer fits in one VRAM pool, the usual options are to use stronger quantization, offload part of the model to system RAM, or distribute the workload across multiple GPUs.

## When do you need multiple GPUs or a server?

Multiple GPUs become relevant when the weights and runtime memory exceed the capacity of a single GPU, or when a production workload requires high throughput for multiple concurrent users. Inference frameworks can divide model weights across several GPUs, increasing total available VRAM — this also introduces communication between devices, making GPU interconnect bandwidth and software support more significant.

Production systems may additionally require server characteristics that are less important on a desktop: large RAM capacity, sufficient PCIe bandwidth, cooling for sustained workloads, remote management and physical support for several GPUs. A server is therefore not required simply because inference is local — the transition is driven mainly by model size, context length, concurrency and the total amount of memory that must be active at the same time.

## Practical rule of thumb

For inference, start with 2 GB per billion parameters at FP16, 1 GB at 8-bit and 0.5 GB at 4-bit, and treat that only as the memory required for the weights. Leave additional capacity for context, KV cache and runtime overhead. If everything fits comfortably into one GPU's VRAM, a consumer GPU is often sufficient. If the model fits only in system RAM, CPU or partial GPU offloading remains possible. Once the model, context and concurrent requests exceed a single memory pool, larger unified memory, multiple GPUs or a server becomes necessary.

Choosing hardware?

## Not sure which hardware fits your AI project?

Neuralex helps you think through the architecture choice — from consumer GPU to server, on-prem or hybrid, matched to your models and volume.

[View the Lab](/en/lab) [Read our approach](/en/aanpak)

---
Volledige (opgemaakte) versie: https://neuralex.nl/en/blog/hardware-voor-lokale-ai
