kvanta

KVANTA estimates the LLM KV-cache memory used by vLLM

Models 0 selected / 0 total

Add a Hugging Face model to begin comparing total VRAM footprint growth.

Config 32,768 ctx / 1 batch / float16
KV bits
Weights bits
Graph
vLLM & GPUs 92% / No GPUs
vLLM gpu_memory_utilization

Default vLLM is 92%. Higher values such as 95% can work when the GPU is dedicated to one vLLM instance, but leave less room for CUDA graphs, kernels, drivers, and fragmentation.

Add GPUs to compare selected models against vLLM VRAM.

raw VRAM 0.0000 GB / 0.0000 GiB
vLLM % 92%
vLLM VRAM 0.0000 GB / 0.0000 GiB

Select at least one processed model to draw the footprint graph.