Instructions to use prithivMLmods/LightOnOCR-3-0.8B-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use prithivMLmods/LightOnOCR-3-0.8B-MLX with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("prithivMLmods/LightOnOCR-3-0.8B-MLX") config = load_config("prithivMLmods/LightOnOCR-3-0.8B-MLX") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use prithivMLmods/LightOnOCR-3-0.8B-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "prithivMLmods/LightOnOCR-3-0.8B-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prithivMLmods/LightOnOCR-3-0.8B-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use prithivMLmods/LightOnOCR-3-0.8B-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "prithivMLmods/LightOnOCR-3-0.8B-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prithivMLmods/LightOnOCR-3-0.8B-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use prithivMLmods/LightOnOCR-3-0.8B-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "prithivMLmods/LightOnOCR-3-0.8B-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prithivMLmods/LightOnOCR-3-0.8B-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
LightOnOCR-3-0.8B-MLX
LightOnOCR-3-0.8B, developed by lightonai, is the lightest and most efficient model in the LightOnOCR-3 family, designed for end-to-end Optical Character Recognition (OCR) and comprehensive document understanding. Built on the Qwen3.5 vision-language architecture, it streamlines complex document pipelines into a single model featuring two primary modes: a default transcription mode that outputs full-page Markdown text, and a "grounding" mode that returns bounding box coordinates and semantic labels for every visual element. Beyond standard text extraction, the model provides advanced visual understanding by automatically describing images and converting charts into structured HTML data tables. Capable of processing diverse formats—including tables, receipts, forms, multi-column layouts, and math notation—LightOnOCR-3-0.8B offers a fast, versatile, and Apache 2.0-licensed solution for both commercial and research document processing tasks.
Directory Structure
prithivMLmods/LightOnOCR-3-0.8B-MLX/
├── 4bit/
├── 8bit/
└── (root: bf16 base model mlx)
Use with mlx
Install the required library:
pip install -U mlx-vlm
BF16 Variant (Base Weights)
The unquantized BF16 mlx weights reside directly in the root repository directory:
CLI (Terminal)
python -m mlx_vlm generate \
--model prithivMLmods/LightOnOCR-3-0.8B-MLX \
--max-tokens 1024 \
--temperature 0.0 \
--prompt "Transcribe the document in full-page Markdown format, extracting all tables and math formulas." \
--image <path_to_image>
Python API
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config
model_path = "prithivMLmods/LightOnOCR-3-0.8B-MLX"
model, processor = load(model_path)
config = load_config(model_path)
image = ["<path_to_image>"]
prompt = "Transcribe the document in full-page Markdown format, extracting all tables and math formulas."
formatted_prompt = apply_chat_template(processor, config, prompt, num_images=len(image))
output = generate(
model,
processor,
formatted_prompt,
image=image,
max_tokens=1024,
temperature=0.0
)
print(output.text)
8-bit Variant
Access the 8-bit quantized files using --subfolder 8bit:
CLI (Terminal)
python -m mlx_vlm generate \
--model prithivMLmods/LightOnOCR-3-0.8B-MLX \
--subfolder 8bit \
--max-tokens 1024 \
--temperature 0.0 \
--prompt "Transcribe the document in full-page Markdown format, extracting all tables and math formulas." \
--image <path_to_image>
Python API
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config
model_path = "prithivMLmods/LightOnOCR-3-0.8B-MLX"
model, processor = load(model_path, subfolder="8bit")
config = load_config(model_path, subfolder="8bit")
image = ["<path_to_image>"]
prompt = "Transcribe the document in full-page Markdown format, extracting all tables and math formulas."
formatted_prompt = apply_chat_template(processor, config, prompt, num_images=len(image))
output = generate(
model,
processor,
formatted_prompt,
image=image,
max_tokens=1024,
temperature=0.0
)
print(output.text)
4-bit Variant
Access the 4-bit quantized files using --subfolder 4bit:
CLI (Terminal)
python -m mlx_vlm generate \
--model prithivMLmods/LightOnOCR-3-0.8B-MLX \
--subfolder 4bit \
--max-tokens 1024 \
--temperature 0.0 \
--prompt "Transcribe the document in full-page Markdown format, extracting all tables and math formulas." \
--image <path_to_image>
Python API
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config
model_path = "prithivMLmods/LightOnOCR-3-0.8B-MLX"
model, processor = load(model_path, subfolder="4bit")
config = load_config(model_path, subfolder="4bit")
image = ["<path_to_image>"]
prompt = "Transcribe the document in full-page Markdown format, extracting all tables and math formulas."
formatted_prompt = apply_chat_template(processor, config, prompt, num_images=len(image))
output = generate(
model,
processor,
formatted_prompt,
image=image,
max_tokens=1024,
temperature=0.0
)
print(output.text)
License and Attribution
This model is based on and/or incorporates the following open-source projects and models:
- LightOnOCR-3-0.8B: https://huggingface.co/lightonai/LightOnOCR-3-0.8B
- mlx-vlm: https://github.com/Blaizzy/mlx-vlm
- MLX: https://github.com/ml-explore/mlx
This model is released under the Apache License 2.0.
- Downloads last month
- 30
4-bit
Model tree for prithivMLmods/LightOnOCR-3-0.8B-MLX
Base model
lightonai/LightOnOCR-3-0.8B