Instructions to use positron-ai/microsoft_phi-4-ingest-best-gptq with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use positron-ai/microsoft_phi-4-ingest-best-gptq with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="positron-ai/microsoft_phi-4-ingest-best-gptq") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("positron-ai/microsoft_phi-4-ingest-best-gptq") model = AutoModelForCausalLM.from_pretrained("positron-ai/microsoft_phi-4-ingest-best-gptq", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use positron-ai/microsoft_phi-4-ingest-best-gptq with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "positron-ai/microsoft_phi-4-ingest-best-gptq" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "positron-ai/microsoft_phi-4-ingest-best-gptq", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/positron-ai/microsoft_phi-4-ingest-best-gptq
- SGLang
How to use positron-ai/microsoft_phi-4-ingest-best-gptq with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "positron-ai/microsoft_phi-4-ingest-best-gptq" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "positron-ai/microsoft_phi-4-ingest-best-gptq", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "positron-ai/microsoft_phi-4-ingest-best-gptq" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "positron-ai/microsoft_phi-4-ingest-best-gptq", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use positron-ai/microsoft_phi-4-ingest-best-gptq with Docker Model Runner:
docker model run hf.co/positron-ai/microsoft_phi-4-ingest-best-gptq
Positron AI Quantized Build
This repository contains a Positron AI quantized build of microsoft/phi-4 for tron inference on Positron FPGA-serving infrastructure.
Recommended Use
Use this artifact when you need a GPTQ 4-bit build of microsoft/phi-4 optimized for Positron's tron runtime. This build targets Positron's ingest-serving path, where runtime fidelity and FPGA deployability are prioritized over general-purpose GPU portability.
For general-purpose GPU inference, compare against the original model and other quantized formats before deployment.
Artifact Summary
| Field | Value |
|---|---|
| Base model | microsoft/phi-4 |
| Published artifact | microsoft_phi-4-ingest-best-gptq |
| Quantization method | GPTQ |
| Quantization format | gptq |
| Source precision | n/a |
| Target runtime | tron |
| Hardware target | FPGA |
| Release date | 2026-06-30 |
| License | mit |
Quantization Details
| Field | Value |
|---|---|
| Weight precision | 4-bit |
| Activation precision | not quantized |
| Bits | 4 |
| Group size | 64 |
| Symmetric quantization | true |
| Activation ordering / desc_act | false |
| Damp percent | 0.05 |
| Calibration dataset | Universal mixed-domain set |
| Calibration samples | 256 |
| Calibration sequence length | 4096 |
| MoE experts per token | n/a |
| Quantization toolchain | GPTQModel 5.8.0, transformers 4.57.6, torch 2.9.1, CUDA 12.8 |
Validation Results
| Metric | Result | Reference | Notes |
|---|---|---|---|
| Mean KL-divergence | 0.0183 | microsoft/phi-4 | Mean across the prompt suite |
| P95 KL-divergence | 0.0753 | microsoft/phi-4 | Mean of per-prompt 95th-percentile token KL-divergence |
| Top-1 agreement | 0.9523 | microsoft/phi-4 | Greedy top-1 token agreement |
| Perplexity / NLL delta | +1.7% | microsoft/phi-4 | Same prompt suite as KL-divergence |
| MMLU mean | pending | n/a | Evaluation pending |
KL-divergence measures token-distribution drift between this quantized artifact and the BF16 reference model; lower values indicate closer agreement. It was computed on Positron's tron FPGA-serving path, so treat it as a runtime-specific drift measurement rather than a GPU benchmark.
Evaluation Methodology
| Field | Value |
|---|---|
| Evaluation date | 2026-06-30 |
| Evaluation suite | Positron mixed-domain prompt suite |
| Number of prompts | 12 |
| Runtime | tron |
| Device | FPGA |
| Pass criteria | Measurement only (no fixed KL-divergence threshold) |
Known Limitations
- KL-divergence was measured on a 12-prompt Positron validation suite; treat it as a runtime validation signal, not a broad benchmark.
- Results are specific to the Positron tron FPGA-serving path and may differ from GPU-native inference.
- MMLU evaluation is pending; results will be added when available.
Provenance
This artifact was produced by Positron AI from microsoft/phi-4. The original model license and usage restrictions continue to apply.
- Downloads last month
- 19
Model tree for positron-ai/microsoft_phi-4-ingest-best-gptq
Base model
microsoft/phi-4