LightOnOCR-3-0.8B-MLX

LightOnOCR-3-0.8B, developed by lightonai, is the lightest and most efficient model in the LightOnOCR-3 family, designed for end-to-end Optical Character Recognition (OCR) and comprehensive document understanding. Built on the Qwen3.5 vision-language architecture, it streamlines complex document pipelines into a single model featuring two primary modes: a default transcription mode that outputs full-page Markdown text, and a "grounding" mode that returns bounding box coordinates and semantic labels for every visual element. Beyond standard text extraction, the model provides advanced visual understanding by automatically describing images and converting charts into structured HTML data tables. Capable of processing diverse formats—including tables, receipts, forms, multi-column layouts, and math notation—LightOnOCR-3-0.8B offers a fast, versatile, and Apache 2.0-licensed solution for both commercial and research document processing tasks.

Directory Structure

prithivMLmods/LightOnOCR-3-0.8B-MLX/
├── 4bit/
├── 8bit/
└── (root: bf16 base model mlx)

Use with mlx

Install the required library:

pip install -U mlx-vlm

BF16 Variant (Base Weights)

The unquantized BF16 mlx weights reside directly in the root repository directory:

CLI (Terminal)

python -m mlx_vlm generate \
  --model prithivMLmods/LightOnOCR-3-0.8B-MLX \
  --max-tokens 1024 \
  --temperature 0.0 \
  --prompt "Transcribe the document in full-page Markdown format, extracting all tables and math formulas." \
  --image <path_to_image>

Python API

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

model_path = "prithivMLmods/LightOnOCR-3-0.8B-MLX"
model, processor = load(model_path)
config = load_config(model_path)

image = ["<path_to_image>"]
prompt = "Transcribe the document in full-page Markdown format, extracting all tables and math formulas."
formatted_prompt = apply_chat_template(processor, config, prompt, num_images=len(image))

output = generate(
    model, 
    processor, 
    formatted_prompt, 
    image=image, 
    max_tokens=1024, 
    temperature=0.0
)
print(output.text)

8-bit Variant

Access the 8-bit quantized files using --subfolder 8bit:

CLI (Terminal)

python -m mlx_vlm generate \
  --model prithivMLmods/LightOnOCR-3-0.8B-MLX \
  --subfolder 8bit \
  --max-tokens 1024 \
  --temperature 0.0 \
  --prompt "Transcribe the document in full-page Markdown format, extracting all tables and math formulas." \
  --image <path_to_image>

Python API

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

model_path = "prithivMLmods/LightOnOCR-3-0.8B-MLX"
model, processor = load(model_path, subfolder="8bit")
config = load_config(model_path, subfolder="8bit")

image = ["<path_to_image>"]
prompt = "Transcribe the document in full-page Markdown format, extracting all tables and math formulas."
formatted_prompt = apply_chat_template(processor, config, prompt, num_images=len(image))

output = generate(
    model, 
    processor, 
    formatted_prompt, 
    image=image, 
    max_tokens=1024, 
    temperature=0.0
)
print(output.text)

4-bit Variant

Access the 4-bit quantized files using --subfolder 4bit:

CLI (Terminal)

python -m mlx_vlm generate \
  --model prithivMLmods/LightOnOCR-3-0.8B-MLX \
  --subfolder 4bit \
  --max-tokens 1024 \
  --temperature 0.0 \
  --prompt "Transcribe the document in full-page Markdown format, extracting all tables and math formulas." \
  --image <path_to_image>

Python API

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

model_path = "prithivMLmods/LightOnOCR-3-0.8B-MLX"
model, processor = load(model_path, subfolder="4bit")
config = load_config(model_path, subfolder="4bit")

image = ["<path_to_image>"]
prompt = "Transcribe the document in full-page Markdown format, extracting all tables and math formulas."
formatted_prompt = apply_chat_template(processor, config, prompt, num_images=len(image))

output = generate(
    model, 
    processor, 
    formatted_prompt, 
    image=image, 
    max_tokens=1024, 
    temperature=0.0
)
print(output.text)

License and Attribution

This model is based on and/or incorporates the following open-source projects and models:

This model is released under the Apache License 2.0.

Downloads last month
30
Safetensors
Model size
0.9B params
Tensor type
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/LightOnOCR-3-0.8B-MLX

Quantized
(9)
this model

Collections including prithivMLmods/LightOnOCR-3-0.8B-MLX