Technozor Lab · 02

Planificateur d’infrastructure LLM

Estimez l’empreinte mémoire d’un déploiement LLM avant de le tester sur le matériel cible.

Mémoire de déploiement estimée

12,7 GiB
Poids du modèle
7,5 GiB
Cache KV
2 GiB
Surcharge d’exécution
1,1 GiB
Total avec 20 % de marge
12,7 GiB
Topologie mémoire possible1 × NVIDIA L40S
Méthode et limites

Weights = parameters × precision. KV cache = 2 × layers × context × concurrency × hidden size × grouped-query ratio × precision. Runtime overhead is estimated at 15%, followed by 20% headroom and 10% reserved GPU capacity.

This is a capacity estimate, not a throughput, latency, availability or cost guarantee. Validate kernels, tensor parallelism, framework overhead and production traffic on target hardware.