souuzaa/GLiNER2.5-Decide-8bit

MLX-compatible version of fastino/GLiNER2.5-Decide (8-bit), for Macs with Apple Silicon (M1/M2/M3/M4 and later).

This repo runs natively on MLX, Apple's machine-learning framework, on the Mac GPU with unified memory. It needs no PyTorch: install mlx, numpy and tokenizers, and the bundled gliner2_mlx.py loads the weights and exposes the same classify_text / extract_entities API as gliner2. Answers match the original PyTorch model (fp32 parity ~1e-6). On an Apple M4, this runs about 2× faster than PyTorch gliner2 on the same Mac (CPU or MPS) and uses far less memory.

PyTorch gliner2 (MPS) MLX bf16 MLX 4-bit
Median latency (Apple M4, batch 1) 215 ms 105 ms 99 ms
Peak process RAM 4.39 GB 1.23 GB 0.63 GB

Converter, parity check and benchmarks: souuzaa/gliner2-mlx.

GLiNER2.5-Decide is a 340M-parameter schema-driven classifier (DeBERTa-v3-large encoder + GLiNER2 heads): pass any label set at call time and get a decision in a single forward pass, no generated tokens. It is not a language model: mlx_lm / mlx_vlm cannot load it. Use the bundled gliner2_mlx.py, which runs the encoder and the GLiNER2 heads and reproduces the gliner2 pre/post-processing.

Usage

pip install mlx tokenizers huggingface_hub   # no torch needed
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("souuzaa/GLiNER2.5-Decide-8bit")
sys.path.insert(0, path)
import gliner2_mlx

model = gliner2_mlx.load(path)
model.classify_text(
    "Battery dies before lunch, but the keyboard and the screen are the best I have used on a laptop.",
    {
        "sentiment": ["positive", "negative", "mixed", "neutral"],
        "aspects": {
            "labels": ["battery", "keyboard", "screen", "camera", "price", "support"],
            "multi_label": True,
            "cls_threshold": 0.4,
        },
    },
)
# {'sentiment': ..., 'aspects': [...]}

The task syntax is the same as gliner2's classify_text: a list of labels, a {label: description} dict, or a config dict with labels, multi_label, cls_threshold, prompt, examples, class_act. Pass include_confidence=True for scores, or use model.predict(text, tasks) for the full probability distribution of every task. batch_classify_text(texts, tasks, batch_size=8) batches inputs.

Entity extraction (the GLiNER2 span head) is also available:

model.extract_entities("Tim Cook announced the new iPhone in Cupertino.", ["person", "product", "location"])

See the original model card for the full set of examples.

Limitations

  • Not a chat model. Only gliner2_mlx.py can run it.
  • Classification and entities only. gliner2 relation extraction (extract_relations) and structured JSON extraction (extract_json) are not ported; *_long chunking helpers are not ported either.
  • Custom code. Like the clef MLX ports, this repo ships Python (gliner2_mlx.py) that you import and run. Read it before use if that matters in your environment.

Conversion

  • Encoder: Linear layers and word embeddings quantized to 8-bit (affine, group size 64).
  • Heads (classifier, count, span): bfloat16, unquantized.
  • Relative-position embeddings: bfloat16, unquantized.
  • Tokenizer files copied unchanged. Converted with convert.py (mlx 0.32.3); parity and evaluation scripts are in the same repo.

Quality check: fast-decisions (dev split)

Variant Size Accuracy (dev) Same answer as fp32 Mean / max abs Δp Median latency Peak memory
fp32 reference 1.9 GB 63.7% — — 140 ms 3.01 GB
GLiNER2.5-Decide-bf16 973 MB 63.7% 99.5% 0.0025 / 0.038 110 ms 1.67 GB
GLiNER2.5-Decide-8bit 567 MB 63.8% 99.6% 0.0037 / 0.056 118 ms 1.37 GB
GLiNER2.5-Decide-4bit 350 MB 63.7% 97.2% 0.0265 / 0.364 117 ms 1.19 GB

Measured on 1700 rows / 2900 heads of the fastino/fast-decisions dev split (the published benchmark uses a held-out test split, so these accuracies are not comparable to it). Accuracy is head-level exact match averaged over the 17 domains; multi-label heads use threshold 0.5. Latency is per predict() call at batch size 1.

License

Apache-2.0, following fastino/GLiNER2.5-Decide.

Downloads last month
52
Safetensors
Model size
0.5B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for souuzaa/GLiNER2.5-Decide-8bit

Quantized
(12)
this model