
Will It Fit? - Opinionated llama.cpp VRAM Estimator
估算llama.cpp模型推理所需的显存大小
hypfer · HN
According to this shitty vibecoded thing "I" built https://hypfer.github.io/will-it-fit-llama-cpp/ (and I guess according to math too), FP16 K/V would give me something like 90k context at the same model quant, which doesn't really fit my usage. But maybe someone else has experience to share there










