So? Doesn’t mean that the moment we are using it, the format isn’t broken. I didn’t say that it wouldn’t be fixed in the future. The reality is, the current 4 bit GGUF are giving us subpar results compared to other quantization method. It’s not a helpful comment telling me that “I don’t understand the basic” rather than telling me the exact flags we should using or it’s being fixed.
Inference with llama.cop is not trivial and I can’t summarise in one post all of them parameters. What I’m saying is that in my opinion. is wrong to assume that changing from one transport to the other is causing degradation.
Llamacpp underwent some major changes last few weeks. And following the commits it took few days to stabilise. Try now , works as bliss. And compared to other inference engines such as tinygrad - is much more versatile in options how to be run.