Measuring quantized LLMs: what is lost when an open model is compressed
Almost 300 GGUF files from dozens of families, served with llama.cpp and scored on function calling, code and long context, with speed and energy per card. The line that reserves the most …
