Day-0 Muse Glimmer Support on Intel Platforms with vLLM and Hugging Face

Community Article
Published August 10, 2026

Muse Glimmer is a 30-billion parameter model released by Meta, distilled from Muse Spark, optimized for always-on agent workflows. It exhibits strong performance on key agentic use cases and benchmarks as compared to leading models of this size category. Together, these capabilities make Muse Glimmer a compelling foundation for building responsive, intelligent AI assistants that are fast, private, and efficient.

Intel enables Muse Glimmer across a broad range of hardware, from workstations with Intel® Arc™ Pro discrete GPU configurations to cloud instances with Intel® Xeon® processors, ensuring developers and enterprises can deploy Muse Glimmer wherever their workflows demand.

Thanks to Intel’s “Upstream First” commitment, continuous contributions to the open-source AI community, and close collaboration with Meta, we are excited to share that Muse Glimmer runs seamlessly on Intel GPUs and Xeon CPUs right out of the box (Day 0) via upstream vLLM and Hugging Face Transformers.

Get Started with vLLM on Intel Platforms

Please make sure PR #51655 is merged when you use Muse Glimmer in upstreaming vLLM.

1. Setup Environment

Intel GPU(XPU)

  • Step 1: build vLLM Intel GPU docker images
$ git clone https://github.com/vllm-project/vllm.git
$ cd vllm
$ docker build -f docker/Dockerfile.xpu -t vllm-xpu-env --shm-size=4g .
  • Step 2: launch vLLM Intel GPU container
$ docker run -it --rm --network=host --ipc=host --privileged \
                 --device /dev/dri:/dev/dri \
                 -v /dev/dri/by-path:/dev/dri/by-path \
                 --entrypoint bash \
                 vllm-xpu-env

Intel Xeon CPU

  • Step 1: build vLLM CPU docker images
$ git clone https://github.com/vllm-project/vllm.git
$ cd vllm
$ docker build -f docker/Dockerfile.cpu \
        --tag vllm-cpu-env \
        --target vllm-openai . 
  • Step 2: launch vLLM CPU container
$ docker run -it --rm \
            --privileged \
            --shm-size=4g \
            --entrypoint bash \
            vllm-cpu-env 

2. Launch vLLM server

The following command lines are for demonstration purposes. You can try different model parallelism configurations per your requirements and Xeon CPU or Intel GPU. The commands below were validated on Intel Arc Pro B70, Intel Arc Pro B60 and Intel Xeon 6 CPU.

vLLM provides an HTTP server that implements OpenAI's Completions API, Chat API, and more. It lets users serve models and interact with them using an HTTP client.

Intel GPU(XPU)

$ vllm serve $<MODEL_PATH> \
      --tensor-parallel-size $<TP_SIZE> \
      --reasoning-parser muse_glimmer \
      --enforce-eager

Intel Xeon CPU

$ vllm serve $<MODEL_PATH> \
      --tensor-parallel-size $<TP_SIZE> \
      --reasoning-parser muse_glimmer

To additionally enable Muse Glimmer’s tool calling, append extra options --tool-call-parser muse_glimmer --enable-auto-tool-choice.

3. Play with Muse Glimmer

Text Chat

$ curl -X POST "http://localhost:8000/v1/chat/completions" \ 
     -H "Content-Type: application/json" \ 
     --data '{ 
        "model": "$<MODEL_PATH>", 
        "messages": [ 
          { 
            "role": "user", 
            "content": [ 
              { 
                "type": "text", 
                "text": "How are you?" 
              } 
            ] 
          } 
        ] 
      }'

Image Q&A

$ curl -X POST "http://localhost:8000/v1/chat/completions" \
     -H "Content-Type: application/json" \
     --data '{
        "model": "$<MODEL_PATH>",
        "messages": [
          {
            "role": "user",
            "content": [
              {
                "type": "text",
                "text": "Describe this image in one sentence."
              },
              {
                  "type": "image_url",
                  "image_url": {
                  "url": "$<IMAGE_ADDRESS>"
                }
              }
            ]
          }
        ]
      }'

Get Started with Hugging Face Transformers on Intel Platforms

1. Setup Environment

Intel GPU(XPU)

  • Step 1: update Intel GPU driver to the latest version
$ apt-get update && \
    apt-get install -y software-properties-common && \
    add-apt-repository -y ppa:kobuk-team/intel-graphics && \
    apt-get install -y libze-intel-gpu1 libze1 intel-metrics-discovery intel-opencl-icd clinfo intel-gsc && \
    apt-get install -y intel-media-va-driver-non-free libmfx-gen1 libvpl2 libvpl-tools libva-glx2 va-driver-all vainfo && \
    apt-get install -y libze-dev intel-ocloc && \
    apt-get install -y libze-intel-gpu-raytracing
  • Step 2: Install dependent software and torch packages
$ sudo apt-get update
$ sudo apt-get install -y ffmpeg

$ uv venv .my-env
$ source .my-env/bin/activate
$ uv pip install torch==2.13.0+xpu torchvision==0.28.0+xpu torchaudio==2.11.0+xpu torchao==0.17.0+xpu --index-url https://download.pytorch.org/whl/xpu --no-cache-dir
$ uv pip install torchcodec==0.15.0 --index-url https://download.pytorch.org/whl/cpu

Intel Xeon CPU

  • Install torch packages
$ sudo apt-get update
$ sudo apt-get install -y ffmpeg

$ uv venv .my-env
$ source .my-env/bin/activate
$ uv pip install torch==2.13.0+cpu torchvision==0.28.0+cpu --index-url https://download.pytorch.org/whl/cpu --no-cache-dir
$ uv pip install torchcodec==0.15.0 --index-url https://download.pytorch.org/whl/cpu

2. Install Hugging Face Packages

$ uv pip install transformers accelerate

3. Play with Muse Glimmer

Intel GPU(XPU)

The following command lines are for demonstration purposes. We validated below configurations on Intel Arc® Pro B70 GPUs:

  • text generation on 2 cards
  • image Q&A on 2 cards

You can use muse_glimmer_test.py utility Python script to play with the model.

👇 click to expand muse_glimmer_test.py
from __future__ import annotations

import argparse
import importlib.util
import os
import time

import torch
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor, AutoConfig
from transformers.distributed import DistributedConfig

def rank0_print(*values, **kwargs):
    if int(os.environ.get("RANK", "0")) == 0:
        print(*values, **kwargs)

def build_user_content(
    prompt: str, has_image: bool, has_video: bool
) -> str | list[dict]:
    """Turn a prompt with an <img>/<video> placeholder into chat-template content.

    Text-only -> the raw string. With one media placeholder -> a list of parts
    ({"type": "text"|"image"|"video"}) in document order, so the media lands at the
    placeholder's position; the chat template renders the image/video part as the
    single sentinel that MuseGlimmerProcessor expands.
    """
    if not (has_image or has_video):
        return prompt
    placeholder, media_type = ("<img>", "image") if has_image else ("<video>", "video")
    parts: list[dict] = []
    chunks = prompt.split(placeholder)
    for i, chunk in enumerate(chunks):
        if chunk:
            parts.append({"type": "text", "text": chunk})
        if i < len(chunks) - 1:
            parts.append({"type": media_type})
    return parts

def check_video_requirements():
    if importlib.util.find_spec("torchcodec") is None:
        raise RuntimeError(
            "Video inference requires torchcodec. Install torchcodec in the "
            "active venv."
        )

def move_to_model_device(value, device):
    if isinstance(value, torch.Tensor):
        return value.to(device)
    return value

def extract_assistant_content_ids(generated_ids: torch.Tensor, tokenizer) -> torch.Tensor:
    token_ids = generated_ids.tolist()
    message_id = tokenizer.convert_tokens_to_ids("<|message|>")
    start_id = tokenizer.convert_tokens_to_ids("<|start|>")
    stop_token_ids = {
        tokenizer.convert_tokens_to_ids(token)
        for token in ("<|eom|>", "<|eot|>", "<|end_of_text|>", "<|start|>", "<|message|>")
    }

    for message_pos, token_id in enumerate(token_ids):
        if token_id != message_id:
            continue

        header_start = 0
        for i in range(message_pos - 1, -1, -1):
            if token_ids[i] == start_id:
                header_start = i
                break

        header = tokenizer.decode(token_ids[header_start:message_pos], skip_special_tokens=True)
        if "assistant to=user" not in header:
            continue

        content_start = message_pos + 1
        content_end = len(token_ids)
        for i in range(content_start, len(token_ids)):
            if token_ids[i] in stop_token_ids:
                content_end = i
                break
        return generated_ids[content_start:content_end]

    return generated_ids

def main():
    parser = argparse.ArgumentParser(description="HF inference for MuseGlimmer")
    parser.add_argument("--hf_dir", required=True, help="Path to MuseGlimmer model directory")
    parser.add_argument(
        "--prompt",
        default="The meaning of life is",
        help="Text prompt (use <img> / <video> for vision placeholders)",
    )
    parser.add_argument(
        "--system",
        default="You are a helpful assistant.",
        help="System prompt; pass '' to omit. The SFT model is trained with a "
        "system turn and degenerates on short prompts without one.",
    )
    parser.add_argument("--image", default=None, help="Path to image file")
    parser.add_argument(
        "--video", default=None, help="Path to video file (use <video> placeholder)"
    )
    parser.add_argument("--max_new_tokens", type=int, default=1024)
    parser.add_argument("--temperature", type=float, default=0.0, help="0.0 = greedy")
    args = parser.parse_args()

    tp_plan = {
        "model.language_model.embed_tokens": "embedding_rowwise",
        "model.vision_tower.patch_embedder.patch_embedding": "colwise_gather_output",
        "model.vision_adapter.fc1": "colwise",
        "model.vision_adapter.fc2": "rowwise",
        "model.vision_projection": "colwise_gather_output",
        "model.vision_tower.layers.*.attn.q_proj": "colwise",
        "model.vision_tower.layers.*.attn.k_proj": "colwise",
        "model.vision_tower.layers.*.attn.v_proj": "colwise",
        "model.vision_tower.layers.*.attn.proj": "rowwise",
        "model.vision_tower.layers.*.mlp.fc1": "colwise",
        "model.vision_tower.layers.*.mlp.fc2": "rowwise",
        "model.language_model.layers.*.self_attn.q_proj": "colwise",
        "model.language_model.layers.*.self_attn.k_proj": "colwise",
        "model.language_model.layers.*.self_attn.v_proj": "colwise",
        "model.language_model.layers.*.self_attn.o_proj": "rowwise",
        "model.language_model.layers.*.self_attn.gate_proj": "colwise",
        "model.language_model.layers.*.mlp.gate_proj": "colwise",
        "model.language_model.layers.*.mlp.up_proj": "colwise",
        "model.language_model.layers.*.mlp.down_proj": "rowwise",
        "lm_head": "colwise_gather_output",
    }

    rank0_print(f"Loading model from {args.hf_dir}")
    t0 = time.time()
    model = AutoModelForImageTextToText.from_pretrained(
        args.hf_dir,
        torch_dtype=torch.bfloat16,
        distributed_config=DistributedConfig(tp_size=2, tp_plan=tp_plan),
    )
    model.eval()
    processor = AutoProcessor.from_pretrained(args.hf_dir)
    if hasattr(processor, "video_processor") and not hasattr(
        processor.video_processor, "patch_temporal"
    ):
        processor.video_processor.patch_temporal = processor.video_processor.temporal_patch_size
    rank0_print(f"Model loaded in {time.time() - t0:.1f}s")

    has_image = args.image is not None and "<img>" in args.prompt
    has_video = args.video is not None and "<video>" in args.prompt
    if has_video:
        check_video_requirements()

    # Read media through MuseGlimmerProcessor's shipped sub-processors: images as PIL;
    # video as a path that MuseGlimmerProcessor decodes + groups via MuseGlimmerVideoProcessor
    # (torchcodec, training-faithful) -- the same path validate_nll_hf.py uses.
    images = videos = None
    if has_image:
        images = [Image.open(args.image).convert("RGB")]
        rank0_print(f"Image: {args.image} ({images[0].width}x{images[0].height})")
    elif has_video:
        videos = [args.video]
        rank0_print(f"Video: {args.video}")

    # One user turn: the chat template emits BOS + the trailing assistant
    # generation prompt and renders each media part as a sentinel, which
    # processor(...) expands into the full span and returns pixel_values for.
    messages = []
    if args.system:
        messages.append({"role": "system", "content": args.system})
    messages.append(
        {
            "role": "user",
            "content": build_user_content(args.prompt, has_image, has_video),
        }
    )
    text = processor.apply_chat_template(
        messages, tokenize=False, add_generation_prompt=True
    )
    inputs = processor(
        text=text,
        images=images,
        videos=videos,
        return_tensors="pt",
    )

    input_ids = inputs["input_ids"].to(model.device)
    attention_mask = inputs["attention_mask"].to(model.device)
    # Pass the media tensors and their grids through together; the vision encoder
    # needs *_grid_thw to build patch sequence lengths.
    pixel_values = move_to_model_device(inputs.get("pixel_values"), model.device)
    image_grid_thw = move_to_model_device(inputs.get("image_grid_thw"), model.device)
    pixel_values_videos = move_to_model_device(
        inputs.get("pixel_values_videos"), model.device
    )
    video_grid_thw = move_to_model_device(inputs.get("video_grid_thw"), model.device)

    rank0_print(f"Input: {input_ids.shape[1]} tokens")
    rank0_print("Generating...")
    do_sample = args.temperature > 0
    gen_kwargs = {"max_new_tokens": args.max_new_tokens, "do_sample": do_sample}
    if do_sample:
        gen_kwargs["temperature"] = args.temperature

    t0 = time.time()
    with torch.no_grad():
        output_ids = model.generate(
            input_ids=input_ids,
            attention_mask=attention_mask,
            pixel_values=pixel_values,
            image_grid_thw=image_grid_thw,
            pixel_values_videos=pixel_values_videos,
            video_grid_thw=video_grid_thw,
            **gen_kwargs,
        )
    elapsed = time.time() - t0

    generated_ids = output_ids[0, input_ids.shape[1] :]
    content_ids = extract_assistant_content_ids(generated_ids, processor.tokenizer)
    text_out = processor.tokenizer.decode(content_ids, skip_special_tokens=True)
    n_generated_tokens = len(generated_ids)
    n_content_tokens = len(content_ids)

    rank0_print(f"\n{'=' * 60}")
    rank0_print(f"PROMPT: {args.prompt}")
    rank0_print(f"{'=' * 60}")
    rank0_print(f"OUTPUT: {text_out}")
    rank0_print(f"{'=' * 60}")
    rank0_print(
        f"Generated {n_generated_tokens} raw tokens in {elapsed:.2f}s "
        f"({n_generated_tokens / elapsed:.1f} tok/s); decoded content tokens: {n_content_tokens}"
    )

if __name__ == "__main__":
    main()
  • Text Chat
$ ZE_AFFINITY_MASK=0,1 torchrun --nproc-per-node 2 muse_glimmer_test.py \
                                --hf_dir <MODEL_PATH> \
                                --prompt "The meaning of life is"
  • Image Q&A
$ ZE_AFFINITY_MASK=0,1 torchrun --nproc-per-node 2 muse_glimmer_test.py \
                                --hf_dir <MODEL_PATH> \
                                --image <IMAGE_PATH> \
                                --prompt "Describe this image: <img>"

Intel Xeon CPU

You can use muse_glimmer_test.py utility Python script to play with the model.

👇 click to expand muse_glimmer_test.py
from __future__ import annotations

import argparse
import copy
import importlib.util
import time

import torch
from PIL import Image
from transformers import AutoModel, AutoModelForImageTextToText, AutoProcessor


def build_user_content(
    prompt: str, has_image: bool, has_video: bool
) -> str | list[dict]:
    """Text-only -> the raw prompt. With an <img>/<video> placeholder -> a list of
    chat-template parts in document order, so the media lands at the placeholder.
    """
    if not (has_image or has_video):
        return prompt
    placeholder, media_type = ("<img>", "image") if has_image else ("<video>", "video")
    parts: list[dict] = []
    chunks = prompt.split(placeholder)
    for i, chunk in enumerate(chunks):
        if chunk:
            parts.append({"type": "text", "text": chunk})
        if i < len(chunks) - 1:
            parts.append({"type": media_type})
    return parts


def main():
    parser = argparse.ArgumentParser(
        description="CPU HF inference with optional DFlash speculative decoding"
    )
    parser.add_argument("--hf_dir", required=True, help="Path to HF model directory")
    parser.add_argument(
        "--assistant_dir",
        default=None,
        help=(
            "Path to the DFlash assistant HF directory. Omit to run plain "
            "autoregressive generation."
        ),
    )
    parser.add_argument(
        "--prompt",
        default="The meaning of life is",
        help="Text prompt (use <img> / <video> for vision placeholders)",
    )
    parser.add_argument(
        "--system",
        default="You are a helpful assistant.",
        help="System prompt; pass '' to omit.",
    )
    parser.add_argument("--image", default=None, help="Path to image file")
    parser.add_argument(
        "--video", default=None, help="Path to video file (use <video> placeholder)"
    )

    parser.add_argument("--max_new_tokens", type=int, default=128)
    parser.add_argument("--temperature", type=float, default=0.0, help="0.0 = greedy")
    parser.add_argument(
        "--assistant_confidence_threshold",
        type=float,
        default=0.4,
        help=(
            "DFlash draft-token confidence threshold; use 0.0 to keep full blocks. "
            "Only used with --assistant_dir."
        ),
    )
    parser.add_argument(
        "--assistant_block_size",
        type=int,
        default=None,
        help=(
            "Override the DFlash assistant block size (default: from its config). "
            "Only used with --assistant_dir."
        ),
    )
    args = parser.parse_args()

    print(f"Loading model from {args.hf_dir}")
    t0 = time.time()
    model = AutoModelForImageTextToText.from_pretrained(
        args.hf_dir,
        dtype=torch.bfloat16,
    ).eval()
    print(f"Model loaded in {time.time() - t0:.1f}s")

    generate_kwargs = {}
    if args.assistant_dir:
        print(f"Loading DFlash assistant from {args.assistant_dir}")
        t0 = time.time()
        assistant_model = AutoModel.from_pretrained(
            args.assistant_dir,
            dtype=torch.bfloat16,
        )
        if args.assistant_block_size is not None:
            assistant_model.config.block_size = args.assistant_block_size
        assistant_model.eval()
        print(
            f"Assistant loaded in {time.time() - t0:.1f}s; "
            f"block_size={assistant_model.config.block_size}, "
            f"target_layer_ids={assistant_model.config.target_layer_ids}"
        )
        generate_kwargs["assistant_model"] = assistant_model

    processor = AutoProcessor.from_pretrained(args.hf_dir)
    if hasattr(processor, "video_processor") and not hasattr(
        processor.video_processor, "patch_temporal"
    ):
        processor.video_processor.patch_temporal = processor.video_processor.temporal_patch_size

    has_image = args.image is not None and "<img>" in args.prompt
    has_video = args.video is not None and "<video>" in args.prompt
    if has_video and importlib.util.find_spec("torchcodec") is None:
        raise RuntimeError(
            "Video inference requires torchcodec; install it in the active venv."
        )

    # Images are passed as PIL objects; videos are passed as file paths and
    # decoded by the processor's video processor (requires torchcodec).
    images = videos = None
    if has_image:
        images = [Image.open(args.image).convert("RGB")]
        print(f"Image: {args.image} ({images[0].width}x{images[0].height})")
    elif has_video:
        videos = [args.video]
        print(f"Video: {args.video}")

    # Build a single user turn; the chat template adds BOS and the trailing
    # assistant generation prompt, and the processor expands each media part.
    messages = []
    if args.system:
        messages.append({"role": "system", "content": args.system})

    messages.append(
        {
            "role": "user",
            "content": build_user_content(args.prompt, has_image, has_video),
        }
    )

    text = processor.apply_chat_template(
        messages, tokenize=False, add_generation_prompt=True
    )
    inputs = processor(
        text=text,
        images=images,
        videos=videos,
        return_tensors="pt",
    )

    print(f"Input: {inputs['input_ids'].shape[1]} tokens")

    do_sample = args.temperature > 0
    generation_config = copy.deepcopy(model.generation_config)
    generation_config.max_new_tokens = args.max_new_tokens
    generation_config.do_sample = do_sample
    if do_sample:
        generation_config.temperature = args.temperature
    if args.assistant_dir:
        generation_config.speculation_type = "dflash"
        generation_config.assistant_confidence_threshold = (
            args.assistant_confidence_threshold
        )
        print("Generating with speculation_type='dflash' on CPU...")
    else:
        print("Generating on CPU...")

    generation_started_at = time.time()
    with torch.no_grad():
        output_ids = model.generate(
            **inputs,
            **generate_kwargs,
            generation_config=generation_config,
        )
    elapsed = time.time() - generation_started_at

    generated_ids = output_ids[0, inputs["input_ids"].shape[1] :]
    n_tokens = len(generated_ids)

    # The model may emit a chat channel header before the final message marker.
    # Retain only the tokens that follow the last message marker.
    message_id = processor.tokenizer.convert_tokens_to_ids("<|message|>")
    generated_id_list = generated_ids.tolist()
    if message_id in generated_id_list:
        content_start = len(generated_id_list) - generated_id_list[::-1].index(message_id)
        content_ids = generated_ids[content_start:]
    else:
        content_ids = generated_ids
    text_out = processor.tokenizer.decode(content_ids, skip_special_tokens=True)

    print(f"\n{'=' * 60}")
    print(f"PROMPT: {args.prompt}")
    print(f"{'=' * 60}")
    print(f"OUTPUT: {text_out}")
    print(f"{'=' * 60}")
    print(
        f"Generated {n_tokens} tokens in {elapsed:.2f}s "
        f"({n_tokens / elapsed:.1f} tok/s)"
    )


if __name__ == "__main__":
    main()
  • Text Chat
$ python muse_glimmer_test.py \
    --hf_dir <MODEL_PATH> \
    --prompt "The meaning of life is"
  • Image Q&A
$ python muse_glimmer_test.py \
    --hf_dir <MODEL_PATH> \
    --image <IMAGE_PATH> \
    --prompt "Describe this image: <img>"
  • DFlash Speculative Decoding

You can enable DFlash speculative decoding by adding --assistant_dir <ASSISTANT_PATH>. Here is an example.

$ python test.py \
    --hf_dir <MODEL_PATH> \
    --assistant_dir <ASSISTANT_PATH> \
    --prompt "The meaning of life is"

Happy Hacking!

Community

Sign up or log in to comment