You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

By clicking "Agree" I confirm I have read and agree to the Apache 2․0 license and the Acceptable Use Policy (AUP), available at: https://github.com/swiss-ai/apertus-legal/blob/main/apertus_1.5/USAGE_POLICY.pdf.

By clicking "Agree" you accept to share your contact information (email and username) with the repository authors. The Apertus v1․5 privacy policy relating to the use of your contact information is available at https://github.com/swiss-ai/apertus-legal/blob/main/apertus_1.5/PRIVACY_POLICY.pdf.

Log in or Sign Up to review the conditions and access this model content.

Apertus-v1.5-8B-DPO

Apertus 1.5 is a family of 8B and 70B fully open, multilingual, and multimodal models from the Swiss AI Initiative, trained only on openly available data and released with open weights, open data, and open training documentation. The models cover a large variety of languages, handle contexts up to 262,144 tokens, accept image and audio input, and support a thinking mode for reasoning tasks.

This checkpoint is the preference alignment stage of the 8B post-training pipeline. All nine checkpoints of the pipeline are released.

Stage 8B 70B
Pre-Training Apertus-v1.5-8B-base Apertus-v1.5-70B-base
Supervised finetuning Apertus-v1.5-8B-SFT Apertus-v1.5-70B-SFT
Reinforcement learning Apertus-v1.5-8B-RLVR Apertus-v1.5-70B-RLVR
Preference alignment (DPO) Apertus-v1.5-8B-DPO (this model) Apertus-v1.5-70B
Charter alignment (SDPO) Apertus-v1.5-8B --

Evaluation

The leftmost columns are the Apertus 1.5 8B lineage, each column adding one stage to the one before it. The final row averages all listed benchmarks.

Thinking disabled

Metric SFT +RLVR +DPO +SDPO Qwen3
8B
EuroLLM
9B Instruct
OLMo 3
7B Instruct
Knowledge
MMLU 65.8 66.6 70.7 70.6 78.8 58.0 68.1
MMLU-Pro 40.1 43.0 49.4 49.2 63.4 33.4 53.5
TruthfulQA 48.1 50.8 58.0 56.1 53.1 49.5 58.7
CommonsenseQA 50.8 27.6 75.5 75.3 21.0 70.8 65.4
SQuADv2 10.1 13.5 30.8 32.0 14.3 13.7 16.4
HellaSwag 72.7 71.4 66.9 66.6 58.6 71.6 61.0
Reasoning
BBH 67.9 69.1 71.1 68.2 58.0 52.4 57.4
ACPBench 26.1 31.2 37.9 37.1 73.6 33.7 67.6
DROP 53.2 56.0 55.2 54.1 58.2 44.6 54.6
Math
GSM8K 77.0 79.5 80.7 78.6 89.3 61.6 83.4
MATH 35.2 39.7 44.6 41.4 66.5 19.6 59.7
Minerva MATH 42.9 48.8 54.3 50.8 81.9 21.8 77.5
MathQA 46.1 46.3 46.7 46.1 53.8 35.3 43.2
MATH-500 43.2 47.0 54.2 50.2 79.2 22.0 78.4
Code
MBPP 48.6 51.6 51.2 52.6 67.6 40.8 50.2
Instruction following
IFEval 80.2 85.8 89.6 90.0 86.0 63.0 86.0
Multi-IF 74.0 83.0 84.5 85.0 83.5 60.1 72.0
Chat
AlpacaEval 20.0 20.6 32.7 42.7 44.6 8.4 42.0
Arena-Hard v1 18.9 19.7 33.6 33.8 77.6 7.2 58.1
Arena-Hard v2 2.3 2.1 2.1 2.6 14.4 0.4 7.7
Multilingual
Global-MMLU 55.8 55.4 58.9 59.3 63.9 53.3 45.7
MGSM 77.2 80.4 78.4 75.6 84.4 61.6 81.6
Mult. TruthfulQA 46.7 48.7 52.4 51.7 51.4 46.2 50.6
INCLUDE-44 56.9 56.7 56.6 56.9 60.9 49.9 38.0
INCLUDE-45 40.2 40.9 39.7 39.5 38.8 36.9 31.2
XNLI 45.8 45.3 43.8 43.7 42.3 41.9 39.2
Cultural
BLEnD 65.1 66.5 67.5 67.3 65.9 61.0 59.6
CulturalBench 71.5 72.6 71.6 72.0 75.7 62.1 71.3
SwitzerlandQA 65.7 64.4 64.6 64.6 60.8 60.9 53.4
Safety and alignment
BBQ 62.9 67.4 67.8 67.6 72.9 64.7 77.6
ToxiGen 53.8 73.3 81.9 81.7 84.0 56.2 83.5
WMDP 51.7 50.7 53.4 53.4 49.1 48.5 49.4
Charter Alignment 33.4 34.3 56.8 66.7 53.3 22.2 64.5
Tool use
BFCL v3 32.2 39.8 52.9 48.2 61.3 22.4 52.9
Average 49.5 51.5 56.9 56.8 61.4 42.8 57.6

Thinking enabled

Metric SFT +RLVR +DPO +SDPO Qwen3
8B
OLMo 3
7B Think
Knowledge
MMLU 69.9 70.7 74.2 73.2 81.3 74.2
Reasoning
BBH 67.9 68.1 70.4 71.0 58.1 72.5
ACPBench 40.1 33.8 48.6 49.4 79.7 68.9
DROP 53.3 55.4 55.5 55.4 58.2 53.3
Math
GSM8K 79.8 86.4 84.7 86.7 92.9 89.9
MATH 51.6 55.8 59.3 60.3 74.6 72.0
Minerva MATH 60.9 66.2 73.0 74.3 89.9 86.3
MathQA 45.8 46.0 46.3 46.4 54.1 42.4
MATH-500 61.2 66.2 70.4 69.6 90.4 85.4
AIME 24 3.3 0.0 16.7 23.3 66.7 66.7
AIME 25 16.7 10.0 23.3 20.0 63.3 66.7
Code
MBPP 49.6 51.4 51.6 52.0 66.2 48.8
Multilingual
MGSM 79.2 88.4 89.2 89.6 92.0 90.4
Average 52.3 53.7 58.7 59.3 74.4 70.6

Long context

Length SFT +RLVR +DPO +SDPO Apertus 1.0
8B Instruct
OLMo 3
7B Instruct
Gemma 3
12B
Llama 3.1
8B
Qwen3.5
9B
RULER
8k 86.8 87.1 88.2 87.4 88.3 58.3 92.6 93.8 96.1
16k 82.6 82.8 84.0 85.0 80.1 45.9 87.7 93.4 95.9
32k 76.3 78.6 80.0 80.7 65.3 36.3 80.0 87.3 96.0
64k 65.9 67.1 69.2 68.3 59.6 26.1 70.7 84.6 95.8
128k 59.5 60.1 61.5 61.1 - - 60.9 78.8 94.1
256k 44.9 48.8 50.4 50.5 - - - - 93.6
HELMET
8k 42.9 41.1 48.5 48.3 44.7 39.6 56.2 45.4 59.7
16k 36.7 37.1 46.0 45.9 34.3 35.3 51.8 43.2 59.9
32k 29.2 29.1 36.1 36.3 27.0 32.1 44.4 42.2 59.2
64k 23.8 24.1 28.5 28.7 23.8 26.8 38.2 42.4 56.9
128k 21.1 19.7 23.1 23.1 - - 33.0 33.4 55.4

How to Use

Users of Apertus can find instructions on getting started with desktop software and cloud providers on our website. Please visit the documentation page if you are interested in trying the model.

Deployment of the models is supported in the following open source frameworks: Transformers, vLLM.

vLLM

We are currently working on adding support for our models to upstream vLLM and transformers releases. In the meantime, you can use our modified version of vLLM and transformers to run the models. We have pre-installed the dependencies in a Docker image which is available in the GitHub Container Registry. The source dockerfile used to build the image is available in this repository. You can pull the image with the following command:

  • amd64 architecture:

    docker pull ghcr.io/swiss-ai/vllm_apertus_1.5_release:latest-amd64
    
  • arm64 architecture:

    docker pull ghcr.io/swiss-ai/vllm_apertus_1.5_release:latest-arm64
    

You can use the following command to run the model with vLLM:

vllm serve swiss-ai/Apertus-v1.5-8B-DPO \
  --chat-template-content-format string \
  --gpu-memory-utilization 0.6 \
  --max-model-len 262144 \
  --enable-auto-tool-choice \
  --tool-call-parser apertus

Depending on your hardware, you may need to adjust --tensor-parallel-size, --gpu-memory-utilization, and --max-model-len (e.g. lower --max-model-len if you run out of memory). On some hardware configurations, CUDA Graph capture may fail with --tensor-parallel-size > 1 due to the fused all-reduce RMS optimization. If this occurs, launch vLLM with --compilation-config.pass_config.fuse_allreduce_rms false.

Instructions for Launching Apertus 1.5 with Thinking Mode Enabled

To enable thinking mode, set --reasoning-parser and --default-chat-template-kwargs.enable_thinking as shown below. The tool-call flags are intentionally omitted: tool calling is unsupported in thinking mode, so we don't recommend combining the two.

vllm serve swiss-ai/Apertus-v1.5-8B-DPO \
  --served-model-name swiss-ai/Apertus-v1.5-8B-DPO-thinking \
  --chat-template-content-format string \
  --gpu-memory-utilization 0.6 \
  --max-model-len 262144 \
  --reasoning-parser apertus \
  --default-chat-template-kwargs.enable_thinking true

Transformers

Apertus 1.5 accepts interleaved text, image, and audio inputs and generates text. The model does not generate audio or images.

The integration is not yet part of a released Transformers version (upstreaming is in progress). Until then, install transformers from our branch:

pip install "transformers[torch,vision,audio] @ git+https://github.com/swiss-ai/transformers.git@3797303dda74844e3d1f8977ff5518bb91f818b4"

Load the processor and model once:

import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor

MODEL_ID = "swiss-ai/Apertus-v1.5-8B-DPO"

processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID, dtype="auto", device_map="auto"
).eval()


def generate(messages, max_new_tokens=256, **template_kwargs):
    inputs = processor.apply_chat_template(
        messages,
        add_generation_prompt=True,
        tokenize=True,
        return_dict=True,
        return_tensors="pt",
        **template_kwargs,
    ).to(model.device)
    with torch.inference_mode():
        output_ids = model.generate(**inputs, max_new_tokens=max_new_tokens)
    return processor.decode(
        output_ids[0, inputs["input_ids"].shape[-1]:], skip_special_tokens=True
    )

Text

messages = [
    {"role": "system", "content": "You are a concise and helpful assistant."},
    {"role": "user", "content": "Explain why the sky appears blue in one sentence."},
]

print(generate(messages))

Image

Images can be supplied as URLs, local paths, PIL images, or arrays:

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image",
                "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/coco_sample.png",
            },
            {"type": "text", "text": "Describe this image in detail."},
        ],
    }
]

print(generate(messages))

Audio

Audio files referenced by URL or local path are decoded and resampled to 24 kHz automatically; in-memory waveforms are accepted as mono NumPy arrays already sampled at 24 kHz:

messages = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "Summarize what is said in this audio clip."},
            {
                "type": "audio",
                "url": "https://huggingface.co/datasets/hf-internal-testing/dummy-audio-samples/resolve/main/bcn_weather.mp3",
            },
        ],
    }
]

print(generate(messages))

Batching

A batch may mix prompts with different numbers and kinds of media. The shipped tokenizer defaults to the left padding that batched generation requires:

conversations = [
    [
        {
            "role": "user",
            "content": [
                {
                    "type": "image",
                    "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/bee.jpg",
                },
                {"type": "text", "text": "Describe this image in one sentence."},
            ],
        }
    ],
    [{"role": "user", "content": "Name the four official languages of Switzerland."}],
]

inputs = processor.apply_chat_template(
    conversations,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    processor_kwargs={"padding": True},
).to(model.device)

with torch.inference_mode():
    output_ids = model.generate(**inputs, max_new_tokens=256)

print(processor.batch_decode(output_ids[:, inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

Thinking mode

Thinking is turned off by default. Pass enable_thinking=True to apply_chat_template to activate it; the model then reasons between the <|inner_prefix|> and <|inner_suffix|> tokens before its visible answer.

messages = [
    {"role": "user", "content": "A bat and a ball cost 1.10 CHF together. The bat costs 1 CHF more than the ball. What does the ball cost?"},
]

print(generate(messages, enable_thinking=True, max_new_tokens=2048))

Budget generously for max_new_tokens in thinking mode: the reasoning can be several times longer than the visible answer, and a truncated generation may end before the answer begins. The reasoning markers are deliberately not stripped by skip_special_tokens=True, so the deliberation span can be parsed out of the decoded text.

Input notes

  • Each image or audio content block corresponds to exactly one media item; placeholder/media count mismatches raise an error instead of being silently reassigned.
  • Audio contributes 40 tokens per second.
  • The LM head covers the 131,072 text tokens (output_vocab_size); logits are padded to the full 266,752-token vocabulary with non-selectable scores, so standard generation utilities work unchanged while image and audio tokens are never generated.
  • The vision and audio tokenizers are precision-sensitive and stay in float32 automatically on half-precision loads. Avoid re-casting the loaded model with .half()/.to(dtype).
  • Native video inputs are not supported; applications may extract frames and pass them as images.

Post-Training Pipeline

Each stage starts from the checkpoint produced by the stage before it.

  1. Supervised finetuning. Trains on a mixture of text and multimodal instruction data. Chain-of-thought traces are attached only to the prompts that a non-reasoning model fails under pass@k. A dedicated long-context stage at the 262,144-token window closes the supervised phase.
  2. Reinforcement learning. Optimizes verifiable rewards, computed by checking or executing the model's output: mathematics is verified against ground truth, instruction-following prompts run formal constraint checkers, and code runs against test cases. The model emits final answers through a dedicated answer tool. Training mixes thinking-enabled and thinking-disabled prompts.
  3. Preference alignment. Direct Preference Optimization. Offline DPO trains on MaxMin_Tr_3600-Filtered-Decontaminated; online DPO uses the prompts of that dataset. The 8B model trains with online DPO; the 70B model trains with offline DPO followed by online DPO. Both forms run with thinking disabled.
  4. Charter alignment. A short self-distillation stage (SDPO). The student answers a prompt, a judge writes textual feedback on how well the answer follows the Apertus charter, and the teacher signal is obtained by conditioning the student on that feedback. Applied to the 8B model.

Data mixtures, hyperparameters, and ablations are documented in the technical report.

This checkpoint

Adds online DPO: rollouts are sampled from the policy under training, scored by a judge on helpfulness, and paired by max-min selection.

Limitations

Apertus can produce text on a variety of topics, but the generated content may not always be factually accurate, logically consistent, or free from biases present in the training data. These models should be used as assistive tools rather than definitive sources of information. Users should always verify important information and critically evaluate any generated content.

See the Apertus Charter for the alignment principles that govern the responses of the models.

License

Apache 2.0.

Legal and Compliance

By using the Apertus LLM you agree to indemnify, defend, and hold harmless ETH Zurich and EPFL against any third-party claims arising from your use of Apertus LLM.

EU AI Act documentation (Apertus_1_5_EU_Public_Summary.pdf and Apertus_1_5_EU_Code_of_Practice.pdf) is available at https://github.com/swiss-ai/apertus-legal/, alongside the usage policy and privacy policy.

We strongly advise downloading and updating new model weights from this site every six months.

Data protection contacts: llm-privacy-requests@swiss-ai.org and llm-copyright-requests@swiss-ai.org.

Contact

General inquiries: llm-requests@swiss-ai.org, or https://apertus-ai.org/contact/.

Citation

@techreport{apertus15,
  title  = {Apertus 1.5},
  author = {Project Apertus},
  year   = {2026},
}
Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including swiss-ai/Apertus-v1.5-8B-DPO