Does DeepSeek-V4-Flash-0731-MXFP4-MLX work with oMLX 0.5.7

#1
by Polygons - opened

"Recommended: stock oMLX 0.5.4rc2 with mtp_enabled", does the model work with the latest 0.5.7?

@Vontra, by any chance, are you planning on creating a version that can fit on 128GB? Would definitely help the M5 Max crowd. Happy to help of needed.

Also on 0.5.7; I'm getting 25tg on Mac Studio M3U 256GB. I enabled Lightning MTP, Hot Cache with Cold Cache since oMLX makes me set a cold cache limit. What am I doing wrong? Did the model update and somehow I'm using an older one?

Edit: now getting 35-40tg, not sure why it changed now. Are there 2 version of the model by chance or has it always been the initial uploaded model? Are there more recommended settings beyond setting hot cache +presumbly the required cold cache with some sort of limit?

Yes — we've been running this checkpoint under oMLX 0.5.7 on M3 Ultra 512 GB for the past few days. Full context sweep, DSpark characterization, and warm-path delta-preill measurements now published:

Dataset: https://huggingface.co/datasets/guruswami-ai/deepseek-v4-flash-0731-m3-ultra

Key findings:

  • 44.8 tok/s decode with DSpark at moderate context, 27.3 at 262k
  • DSpark multiplier 1.31–1.80×, context-dependent, not a single point
  • Prefix cache worth 27–99× — the harness matters more than the hardware
  • Three silent omlx traps that make the stack look 40× slower than it is (documented in the README)

Recommended config and launchd plist included.

TensorFold org

Also on 0.5.7; I'm getting 25tg on Mac Studio M3U 256GB. I enabled Lightning MTP, Hot Cache with Cold Cache since oMLX makes me set a cold cache limit. What am I doing wrong? Did the model update and somehow I'm using an older one?

Edit: now getting 35-40tg, not sure why it changed now. Are there 2 version of the model by chance or has it always been the initial uploaded model? Are there more recommended settings beyond setting hot cache +presumbly the required cold cache with some sort of limit?

oMLX updated their DFLASH and MTP support.

TensorFold org

@Vontra, by any chance, are you planning on creating a version that can fit on 128GB? Would definitely help the M5 Max crowd. Happy to help of needed.

I can do but I think there are already many 128gb versions available now?

​@Vontra
​I don't really know where people usually leave reviews, but I decided to write it right here
🌟🌟🌟🌟🌟
​I just wanted to say a HUGE 👏 thank you! You fulfilled my wish without even knowing it. When I got my new computer, the very first thing I installed was your version. At first, I thought it might be just a little bit weaker than the cloud, BUT!... It turned out that it even outperforms the cloud in many ways! 🔥 It feels like no knowledge was lost at all, and honestly, it feels as if extra knowledge was added to it so it's even smarter than possible.
​Thank you so much for your hard work. Because of you, we can use such cool versions that don't fall short of cloud versions at all in terms of intelligence and capability.
​Thank you again so much, I wish you all the best! ✨ 🏆🥇💎
Because of you, I am very happy 😀😄🥳

TensorFold org

Honestly, thank you so much for taking the time to write this; it genuinely made my day!

Knowing this was one of the first things you installed on your new machine, and that it’s working so well for you, makes all the effort worthwhile.
I started sharing these conversions so more people could run powerful models locally, so feedback like this means a lot.

I hope you continue enjoying it, and please let me know how you get on with it.

Thanks again for the incredibly kind words! 🙏

Thank you 😀! Since this is my first Mac ever, honestly, finding your version of the model is the best thing that could have happened to me. 😄 I've only been using my new setup for a few days, but I'm absolutely thrilled that everything is running perfectly and showing such amazing results. I'll definitely keep you posted as I explore more! 👏👏👏🏆🌟💎

I'm hoping you build something similar for the DeepSeek-V4-Pro-0813 model next. This Flash version was the best we managed to benchmark and Pro looks to be even more capable model if I can get it working in TP2/TP4 modes.

Sign up or log in to comment