Instructions to use souuzaa/GLiNER2.5-Decide-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use souuzaa/GLiNER2.5-Decide-8bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download souuzaa/GLiNER2.5-Decide-8bit --local-dir GLiNER2.5-Decide-8bit
- GLiNER2
How to use souuzaa/GLiNER2.5-Decide-8bit with GLiNER2:
from gliner2 import AutoExtractor extractor = AutoExtractor.from_pretrained("souuzaa/GLiNER2.5-Decide-8bit") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
souuzaa/GLiNER2.5-Decide-8bit
MLX-compatible version of fastino/GLiNER2.5-Decide (8-bit), for Macs with Apple Silicon (M1/M2/M3/M4 and later).
This repo runs natively on MLX, Apple's machine-learning framework, on the Mac GPU
with unified memory. It needs no PyTorch: install mlx, numpy and tokenizers, and the bundled gliner2_mlx.py
loads the weights and exposes the same classify_text / extract_entities API as gliner2. Answers match the
original PyTorch model (fp32 parity ~1e-6). On an Apple M4, this runs about 2× faster than PyTorch gliner2 on the
same Mac (CPU or MPS) and uses far less memory.
PyTorch gliner2 (MPS) |
MLX bf16 | MLX 4-bit | |
|---|---|---|---|
| Median latency (Apple M4, batch 1) | 215 ms | 105 ms | 99 ms |
| Peak process RAM | 4.39 GB | 1.23 GB | 0.63 GB |
Converter, parity check and benchmarks: souuzaa/gliner2-mlx.
GLiNER2.5-Decide is a 340M-parameter schema-driven classifier (DeBERTa-v3-large encoder + GLiNER2 heads):
pass any label set at call time and get a decision in a single forward pass, no generated tokens.
It is not a language model: mlx_lm / mlx_vlm cannot load it. Use the bundled gliner2_mlx.py,
which runs the encoder and the GLiNER2 heads and reproduces the gliner2 pre/post-processing.
Usage
pip install mlx tokenizers huggingface_hub # no torch needed
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("souuzaa/GLiNER2.5-Decide-8bit")
sys.path.insert(0, path)
import gliner2_mlx
model = gliner2_mlx.load(path)
model.classify_text(
"Battery dies before lunch, but the keyboard and the screen are the best I have used on a laptop.",
{
"sentiment": ["positive", "negative", "mixed", "neutral"],
"aspects": {
"labels": ["battery", "keyboard", "screen", "camera", "price", "support"],
"multi_label": True,
"cls_threshold": 0.4,
},
},
)
# {'sentiment': ..., 'aspects': [...]}
The task syntax is the same as gliner2's classify_text: a list of labels, a {label: description} dict,
or a config dict with labels, multi_label, cls_threshold, prompt, examples, class_act.
Pass include_confidence=True for scores, or use model.predict(text, tasks) for the full
probability distribution of every task. batch_classify_text(texts, tasks, batch_size=8) batches inputs.
Entity extraction (the GLiNER2 span head) is also available:
model.extract_entities("Tim Cook announced the new iPhone in Cupertino.", ["person", "product", "location"])
See the original model card for the full set of examples.
Limitations
- Not a chat model. Only
gliner2_mlx.pycan run it. - Classification and entities only.
gliner2relation extraction (extract_relations) and structured JSON extraction (extract_json) are not ported;*_longchunking helpers are not ported either. - Custom code. Like the clef MLX ports, this repo ships Python (
gliner2_mlx.py) that you import and run. Read it before use if that matters in your environment.
Conversion
- Encoder: Linear layers and word embeddings quantized to 8-bit (affine, group size 64).
- Heads (classifier, count, span): bfloat16, unquantized.
- Relative-position embeddings: bfloat16, unquantized.
- Tokenizer files copied unchanged. Converted with
convert.py(mlx 0.32.3); parity and evaluation scripts are in the same repo.
Quality check: fast-decisions (dev split)
| Variant | Size | Accuracy (dev) | Same answer as fp32 | Mean / max abs Δp | Median latency | Peak memory |
|---|---|---|---|---|---|---|
| fp32 reference | 1.9 GB | 63.7% | — | — | 140 ms | 3.01 GB |
| GLiNER2.5-Decide-bf16 | 973 MB | 63.7% | 99.5% | 0.0025 / 0.038 | 110 ms | 1.67 GB |
| GLiNER2.5-Decide-8bit | 567 MB | 63.8% | 99.6% | 0.0037 / 0.056 | 118 ms | 1.37 GB |
| GLiNER2.5-Decide-4bit | 350 MB | 63.7% | 97.2% | 0.0265 / 0.364 | 117 ms | 1.19 GB |
Measured on 1700 rows / 2900 heads of the fastino/fast-decisions dev split (the published benchmark uses a held-out test split, so these accuracies are not comparable to it). Accuracy is head-level exact match averaged over the 17 domains; multi-label heads use threshold 0.5. Latency is per predict() call at batch size 1.
License
Apache-2.0, following fastino/GLiNER2.5-Decide.
- Downloads last month
- 52
Quantized