2.08bpw quant has awesome KLD and PPL
#1
by cpral - opened
Hi!
I cooked up 3.51bpw quant and I'm testing Qwen 3.5 397B exl3 community quants, so yours and those from @NeuroSenko . I noticed that your 2.08bpw quant performs really great considering the size.
table is in the model card of my quant - https://huggingface.co/cpral/Qwen3.5-397B-A17B-exl3
I'm new to making exl3 quants, how did you make this 2.08bpw quant? Did you mix 3bpw and 2bpw with optimize.py or did you do manual tuning?