leonsarmiento/Ornith-Agents-A1-3.6-35B-A3B-dare_ties-6bit-XL-mlx

This model was converted to MLX format from tepirale/Ornith-Agents-A1-3.6-35B-A3B-dare_ties using BaseQuant_XL 6/8-bit mixed quantization optimized for Apple Silicon. The vision encoder is preserved and quantized at 6-bit, making this a full multimodal model.

BaseQuant_XL keeps the most routing-critical layers in full bf16 precision — the MoE router gate, shared expert gate, shared expert, and lm_head — while applying aggressive quantization to the bulk parameters. This preserves routing accuracy and output quality where it matters most.

About XL Quantization

BaseQuant_XL is a fully data-agnostic, static quantization. No calibration dataset, no sensitivity analysis, no importance matrix. Precision is allocated purely by architectural role — routing-critical layers get higher precision, bulk expert parameters get lower precision. The result is a transparent, faithful capture of the source model.

Data-dependent calibration quantizations (iMatrix, AWQ, GPTQ, oQ, oQ4e, etc.) use a calibration set to guide bit allocation. This can produce a skewed representation of the model: domains well-represented in the calibration data (English, popular topics, public or leaked benchmarks) are preserved better, while underrepresented domains (non-English languages, niche use cases, your own data) are preserved worse. XL avoids this trade-off entirely — it generalizes honestly because it is never fit to any particular data distribution.

Model Description

This is a 50/50 DARE-TIES merge of two complementary Qwen3.5-35B-A3B agentic models:

Source Model Weight Density Focus
InternScience/Agents-A1 0.5 0.6 General agentic abilities: long-horizon search, engineering, scientific research, instruction following, tool-calling
deepreinforce-ai/Ornith-1.0-35B 0.5 0.6 RL-tuned agentic coding (Terminal-Bench 64.2, SWE-bench Verified 75.6)
Qwen/Qwen3.5-35B-A3B base — Base model

The merge combines Ornith's coding and terminal-task strength with Agents-A1's broader tool-use and research capabilities. The architecture is a 35B Mixture-of-Experts model with only ~3B active parameters per token, featuring 256 experts (8 active + 1 shared), hybrid full + linear (Gated DeltaNet) attention, a vision encoder, and an extended 262K context window.

Use with mlx

pip install -U mlx-vlm
python -m mlx_vlm.generate --model leonsarmiento/Ornith-Agents-A1-3.6-35B-A3B-dare_ties-6bit-XL-mlx --max-tokens 256 --temperature 0.85 --top-p 0.95 --top-k 20 --min-p 0.01 --repeat-penalty 1.05 --prompt "Hello"

BaseQuant_XL Quantization Strategy

Bit Depth Layers Rationale
bf16 (unquantized) mlp.gate (router), shared_expert_gate, lm_head, shared_expert Routing decisions and shared computation path — errors here are qualitatively different from precision loss
8-bit embed_tokens, self_attn (full attention), linear_attn (DeltaNet) Every-token layers with moderate sensitivity — 8-bit is near-lossless
6-bit vision_tower, switch_mlp (routed experts) Bulk of parameters, only 8 of 256 experts active per token — natural redundancy tolerates lower precision

Quantization Details

Layer Bits Group Size
mlp.gate (router) bf16 —
shared_expert_gate bf16 —
lm_head bf16 —
shared_expert bf16 —
embed_tokens 8 64
self_attn (full attention) 8 64
linear_attn (DeltaNet) 8 64
vision_tower 6 64
switch_mlp (routed experts) 6 64
Default fallback 8 64
  • Quantization type: BaseQuant_XL mixed (multimodal, vision preserved)
  • Group size: 64
  • Method: Custom quant_predicate via mlx_vlm

Recommended Inference Parameters

Inherited from Agents-A1 (more conservative, broader tool-use optimization):

Parameter Value
temperature 0.85
top_p 0.95
top_k 20
min_p 0.01
repeat_penalty 1.0
presence_penalty 1.1

Note: These use Agents-A1's recommended settings rather than averaging across both parents. Ornith's original values (temp 1.0, top_p 1.0, top_k 40) were optimized for Terminal-Bench specifically — Agents-A1's broader agentic profile is a safer default for general use. Adjust as needed.

Reasoning and Tool-Call Parsing

Parser Value
reasoning_parser qwen3
tool_call_parser qwen3_coder

This is a Qwen3.5-based model — preserve_thinking is not applicable.

Benchmarks (n=30, 5-bit XL MLX)

Benchmark comparison of both merges against parent models and the Qwen3.6-35B-A3B base. Higher bit depth (6-bit) is expected to match or slightly exceed these results:

Benchmark Comparison

Benchmark Agents-A1 (parent) Ornith-1.0 (parent) DARE-TIES merge Task Arithmetic merge Qwen3.5-35B-A3B (base)
MMLU 70.0 63.3 70.0 73.3 56.7
MMLU_PRO 53.3 60.0 56.7 50.0 56.7
HELLASWAG 86.7 83.3 86.7 83.3 83.3
TRUTHFULQA 100.0 100.0 96.7 93.3 100.0
ARC_CHALLENGE 90.0 83.3 90.0 90.0 86.7
WINOGRANDE 76.7 76.7 73.3 76.7 80.0
HUMANEVAL 86.7 66.7 86.7 86.7 60.0
MBPP 80.0 76.7 80.0 76.7 86.7
MATHQA (thinking) 96.7 96.7 93.3 96.7 93.3
LIVECODEBENCH 40.0 43.3 43.3 36.7 40.0

The DARE-TIES merge inherits Agents-A1's strength on HUMANEVAL, MMLU, and ARC_CHALLENGE while partially recovering Ornith's MMLU_PRO edge. Minor regressions on TRUTHFULQA, WINOGRANDE, and MATHQA compared to parents — a typical DARE-TIES trade-off. Both merges clearly dominate the Qwen3.6 base on most knowledge and coding benchmarks.

Downloads last month
111
Safetensors
Model size
35B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for leonsarmiento/Ornith-Agents-A1-3.6-35B-A3B-dare_ties-6bit-XL-mlx

Quantized
(4)
this model

Collection including leonsarmiento/Ornith-Agents-A1-3.6-35B-A3B-dare_ties-6bit-XL-mlx