Instructions to use mlx-community/CLM-v0.1-8B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/CLM-v0.1-8B-MLX-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir CLM-v0.1-8B-MLX-4bit mlx-community/CLM-v0.1-8B-MLX-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
CLM-v0.1-8B for MLX (4bit-g32)
An Apple-Silicon (MLX) port of CLM-v0.1-8B, the Contrastive Language Model from Contrastive-LM/CLM: a frozen Qwen3-8B encoder plus two small projection heads that score states against candidate actions. It answers typed questions (yes/no, choice, score) and ranks candidates with a probability distribution, without generating text.
This is an unofficial community port, not released or reviewed by the CLM authors. It was checked
against upstream's reference implementation (the authors' code: vLLM + clm-serve at commit
bb42c6c, which we ran on an RTX 4090) on 778 questions; see Parity below.
The encoder itself is a standard MLX-format Qwen3-8B checkpoint, so its weights are usable with general
MLX tooling such as mlx-workflow and its Mac app
MLXUI for local, on-device inference.
Note that the CLM answer/rank behavior (the projection heads and clm_mlx code in this repo) is specific
to this port and isn't one of MLXUI's built-in model protocols — using it there would need custom
integration on top of the raw encoder.
| Encoder | Qwen3-8B @ b968826d9c46, MLX 4bit-g32, lm_head removed (4.7 GB / 4.4 GiB) |
| Heads | CLM_v0.1-8B projection heads, float32 (distilled for this quantization) |
| Pooling | last token, final hidden state after the last RMSNorm, L2-normalised |
| Max input | 2048 tokens; longer inputs keep the first 2048 (as the upstream server does) |
| Measured on | Apple M3 Pro, 18 GB: 346.5 tokens/s encoder throughput, 5.53 GB peak |
Requirements
An Apple Silicon Mac (MLX), Python 3.10+, about 5 GB of disk and roughly 5.53 GB of free memory while running (a 16 GB Mac is recommended; 8 GB Macs are untested; measured peak above).
Quickstart
pip install mlx mlx-lm transformers huggingface_hub
hf download mlx-community/CLM-v0.1-8B-MLX-4bit --local-dir clm-mlx && cd clm-mlx
Run Python from inside the downloaded folder (the clm_mlx package ships in it):
from clm_mlx.engine import Engine
eng = Engine("encoder", "heads")
out = eng.answer(
"Customer: my invoice was charged twice and nobody answers the phone!",
{
"urgency": {"type": "noul", "instructions": "Is this urgent?"},
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, invoices, refunds", "technical": "Bugs and outages"}},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"]},
},
)
print(out["answers"]["department"]["choice"], out["answers"]["department"]["probabilities"])
print(eng.rank("", ["The Moon's gravitational pull.", "Photosynthesis in plants."], "What causes tides on Earth?"))
Questions and answers use upstream's wire format (noul / choice / score); see the
upstream API reference. Option projections are
cached, so repeated questions over new states only pay for the state.
Accuracy tier: approximate
This is a smaller, approximate build for Macs with less memory. It does not match upstream as closely as the 8-bit build: overall it picks the same top option as upstream 91.4% of the time. The disagreements sit mostly on close calls: when upstream is at least 70% sure of its answer, this build agrees 99.3% of the time. Use it where memory matters more than borderline decisions; use the 8-bit build where parity with upstream matters. The projection heads were re-trained (distilled) for this quantization to recover accuracy.
Parity with the upstream server
Reference: upstream's code, run by us on an RTX 4090 (vLLM 0.30.0, bfloat16), 882 texts (2048-token limit included), 778 typed questions. The yardstick is how much the upstream server disagrees with itself: its live answers compared with answers recomputed from its own vectors for the same inputs.
| this port vs upstream | upstream vs itself | |
|---|---|---|
| top option agrees | 91.4% | 98.6% |
| top option agrees, when upstream is ≥70% sure (577 questions) | 99.3% | — |
| answer probability difference, p95 | 0.328 | 0.060 |
| answer probability difference, median | 0.159 | 0.022 |
| answer probability difference, max | 0.488 | 0.096 |
| encoder embedding cosine, min / mean | 0.8291 / 0.9862 | — |
Verdict: outside upstream noise: this build does not meet the parity bar (top-option agreement within 1 point of upstream's own, and p95 difference at most 1.5× upstream's); see Accuracy tier above.
67 questions changed their top option; on 53 of them upstream's own top choice had under 60% probability (a close call). The full report is in parity.json.
Notes and limitations
- Encoder only:
lm_headis removed, so this checkpoint cannot generate text. - Inputs over 2048 tokens keep the first 2048, matching the upstream server. Upstream
trained the heads on the last tokens of long states;
Engine(..., truncation="tail")does that instead. - This build's quantization adds noticeable noise (see Accuracy tier). Answers upstream is confident about are the reliable ones; treat close calls with care.
Credits and licence
- CLM method, heads and reference implementation: Kwok, Kang, Suresh, Saad-Falcon, Pavone, Ré, Mirhoseini, Contrastive Language Models: A System One Model for Fast and Generalizable Decision-Making (2026), contrastive-lm.notion.site. Heads: Contrastive-LM/CLM-v0.1-8B, Apache-2.0.
- Encoder: Qwen/Qwen3-8B, Apache-2.0.
- This port (MLX conversion,
clm_mlxcode, parity harness): Apache-2.0.
@misc{kwok2026contrastivelanguagemodels,
title={Contrastive Language Models: A System One Model for Fast and Generalizable Decision-Making},
author={Jacky Kwok and Hangoo Kang and Tarun Suresh and Jon Saad-Falcon and Marco Pavone and Christopher Ré and Azalia Mirhoseini},
year={2026}, note={Notion Blog}, url={https://contrastive-lm.notion.site}
}
- Downloads last month
- 19
4-bit