Using TraDoc
Hardware estimator
Estimate local-model memory without confusing an estimate with a real measurement.
The Estimator helps compare model size, quantization, context and runtime headroom. Use it when preparing an LM Studio or Ollama configuration.
Main factors
- Parameter count.
- Weight quantization.
- Context size.
- KV cache.
- Concurrency.
- Memory used by the system and inference server.
Interpretation
An estimate that only just fits in VRAM leaves little room for context and spikes. First reduce concurrency or context size, or choose a more compact quantization.
Indicative estimate
Actual usage depends on the engine, model architecture and options. Always confirm it with LM Studio, Ollama or GPU-driver metrics during a test.