Umpire 35B-A3B: calibrated decisions and classification on a Mac

Umpire makes quick, calibrated calls. It reads a state and one or more typed questions (choose one option, yes/no, or a score) and returns a probability for every option. It does not generate text to get there. This is the first Umpire release, a fine-tune of Ornith-1.5-35B-A3B, packaged for Splash on Apple silicon and served through Splash's /v1/systemone route.

It was distilled from OpenJev, a 27B decision model. On 10,125 test decisions, frozen before training, it matches OpenJev's accuracy, and it carries Ornith's MoE footprint: about 3B active parameters per token.

This is an independent derivative. Ornith AI, the OpenJev authors and Inco AI have not endorsed it. Licence: CC BY-NC 4.0, non-commercial use only (see Licence).

brew install incoai/tap/splash
splash serve --model ezoushen/umpire-35b-a3b-splash --port 1241
curl -s http://127.0.0.1:1241/v1/systemone -H 'Content-Type: application/json' -d '{
  "model": "ezoushen/umpire-35b-a3b-splash",
  "state": "Customer writes: I was charged twice for my March invoice and want one charge refunded.",
  "questions": {
    "team": {"type": "choice", "instructions": "Which team should own this case?",
             "criteria": {"billing": "charges, refunds, invoices", "legal": "contracts, claims",
                          "engineering": "software, outages"}},
    "escalate": {"type": "noul", "instructions": "Does the case require escalation to a manager?"}
  }
}'
# -> team: billing 0.9990, legal 0.0006, engineering 0.0003; escalate: 0.085

Requirements: Apple M3 or newer, macOS 26.4 or later, and at least 36 GB of unified memory. The package is 20.9 GB (19.5 GiB). Tested on Splash 1.0.2. These are Splash-specific packed binaries, not Transformers, GGUF or MLX.

Results

The test set is 10,125 decisions frozen before training started and opened once, for the final evaluation:

  • 6,400 items from 16 sources that training also drew on, using different items;
  • 3,725 items from 5 sources never used in training.

Every number was served by Splash through /v1/systemone. Rescoring 200 items on a fresh process reproduced the probabilities exactly.

stock Ornith 1.5 OpenJev (Splash repack) Umpire
accuracy, 10,125 test decisions 0.784 0.866 0.865
accuracy change on the 5 held-out sources — — +4.4 pts [3.5, 5.3]
log loss (lower is better) 0.935 0.406 0.398
expected calibration error 0.125 0.018 0.011
JevBench public, 231 decisions 0.835 0.879 0.879

The JevBench row uses our scorer, which breaks exact probability ties with a seeded random pick. JevBench's official harness takes the first tied option and gives stock 0.831 (192), OpenJev 0.883 (204) and Umpire 0.874 (202). Umpire's one difference is an item where two options tied exactly.

  • Umpire closes 99.6% of the accuracy gap between stock Ornith and OpenJev. Its difference from OpenJev is −0.03 points, with a paired bootstrap 95% CI of about [−0.60, +0.52]. That is a match within noise, not a win.

  • Against stock Ornith: 1,183 items only Umpire gets right, 356 only stock gets right (exact McNemar p ≈ 1e-103).

  • Gains on the held-out sources are uneven:

    source gain (pts)
    ARC-Challenge +9.3
    PIQA +6.4
    CLINC +4.2
    SMS Spam +2.0
    SST-2 +0.1
  • Umpire trails OpenJev by 5 points or more on ANLI and MMLU-Pro, and by 3–5 points on CLINC, WinoGrande and TweetEval hate. It leads OpenJev on 11 of 21 sources.

The evaluation followed a plan written before training. One criterion failed:

  • Speed parity was not shown. The plan required tuned latency within 1.10× of stock Ornith on short decisions.
  • The one timing round that passed the validity check came out at 1.17×: 169 ms against 144 ms per decision.
  • Umpire's other rounds ranged from 0.81× to 1.04×, and none of them passed the validity check.
  • Umpire has the same tensor shapes and quantisation as stock Ornith, so no compute difference is expected. The swings look like lane-to-lane noise, but that is not proven.
  • The maintainer accepted the model with this criterion waived. Measure latency on your own machine if it matters to you.

Option order

Splash scores options in the order you send them, with ordinary causal attention. Neither Splash nor this model uses an order-invariant attention mask. To measure how much order matters:

  • 1,000 test choice questions with 3 or more options were each asked again with the options in a different seeded order;
  • the new probabilities were mapped back to the original options and compared with the original-order scores.
stock Ornith 1.5 OpenJev (Splash repack) Umpire
top answer changes when options are reordered 22.3% 8.3% 8.8%
mean total variation distance 0.230 0.084 0.089
accuracy, original → shuffled order 0.757 → 0.747 0.833 → 0.821 0.840 → 0.851
  • Umpire is about as order-stable as its teacher, and 2.5× more stable than stock.
  • Training did not shuffle options, so this stability comes from distilling OpenJev's distributions.
  • OpenJev's own card reports 2.3% under its own prompt layout and readout. Through Splash's /v1/systemone prompt it measures 8.3%, so compare within one serving path.
  • If order matters to you, average the probabilities over a few option orders. That costs one request per order.

How it was made

  • Base: Ornith-1.5-35B-A3B, trained from its MLX 4-bit build.
  • Method: LoRA, rank 16, scale 2.0, on attention, linear attention and the shared-expert projections. The routed experts and router were not trained.
  • Training: learning rate 2e-5, 2 epochs over 6,000 rows, step 1,500 kept.
  • Loss: 0.5 × cross-entropy on the gold label + 0.5 × KL to OpenJev's probability distribution, both at the answer position.
  • Fusion: the adapter was merged at BF16 into the adapted modules only, then requantised to 4-bit. Every other tensor is Ornith's, unchanged.
  • Data: rows from public classification, reasoning and commonsense sets, plus Zefan Cai's CC0 Open-Jev projection. Teacher labels came from OpenJev.
  • Model selection: hyperparameters and the checkpoint were chosen on a separate 2,000-item development set.
    • JevBench public was used in three other ways:
      • Before training, the teacher was chosen partly on its JevBench public score.
      • Each candidate had to pass a JevBench gate: at most 3 net losses against stock.
      • Two candidates were fixed before any test scoring. The second was scored, once, only because the first failed the speed criterion.
    • After the maintainer waived that criterion, the first candidate was chosen as the winner. It led the runner-up on the test set, log loss and JevBench public: 203 against 193 under our scorer, 202 against 193 under the official harness.
    • This final choice broke the plan's own rule that JevBench is never used for a tuned-model decision. It is disclosed here for that reason.
  • Contamination screen: the state text of every training and development item was screened against every test and JevBench public item. The screen used normalised exact match, shared 13-word n-grams (an n-gram found in 5 or more texts counted as template text and was ignored), and token-set Jaccard ≥ 0.8 for states under 13 words. It removed 460 training rows and 111 development rows. Families of related datasets were kept together: no held-out source shares a family with a training source (for example, ANLI, HellaSwag and BoolQ are one family, and none of them is held out).
    • Screening cannot catch paraphrases.
    • A stricter re-screen flagged 77 test items that share template boilerplate with training. Removing them changes nothing that matters: 99.0% of the gap closed.
  • Package integrity: 56 artifacts and 20,949,446,234 bytes, SHA-256 checked against the manifest.

The package differs from ezoushen/ornith-1.5-35b-a3b-splash only in the LoRA-targeted sections: attention, linear attention and the shared expert. The tokenizer, vision tower, embedding, head and draft are byte-identical to it. Everything that card says about the draft and tokenizer applies here too:

  • the DFlash 2 draft was trained on Qwen3.6, not Ornith;
  • the tokenizer omits \p{M} from two pre-tokenizer classes.

Limitations

  • The teacher caps the student. Umpire inherits OpenJev's errors and biases.
  • Benchmark exposure. OpenJev scores 0.94 on ANLI and WinoGrande training rows, which suggests it saw those corpora. That flatters the stock-to-OpenJev gap on those sources. The held-out sources are the check on transfer.
  • Decisions only. Chat still works and was checked for coherence only. Chat quality, safety and refusal behaviour were not evaluated.
  • Public tasks. None of the evaluation uses real routing or production traffic.
  • Speed is unverified against stock; see Results.
  • Later system messages render as system turns since 2026-09-30, the one-line change incoai's Qwen3.6 Splash package makes to the upstream template. Earlier downloads raise on them, and Splash 1.0.2 answers messages could not be rendered; re-download.

Related

Licence

  • The fine-tuned weights: CC BY-NC 4.0, non-commercial use only (LICENSE).
    • Every training target was produced by, or mixed with, OpenJev's probability distribution, and OpenJev is CC BY-NC 4.0, so this release carries the same terms.
    • Whether a model's licence extends to a model distilled from its outputs is legally unsettled. This release takes the cautious reading.
  • Components kept under their own licences (see NOTICE):
    • Ornith's weights, vision tower and tokenizer: MIT, from Ornith AI.
    • The DFlash 2 draft and Splash's packing format: Apache-2.0, from Inco AI (LICENSE-APACHE-2.0).
  • Training sources: listed with their licences in NOTICE. Some are non-commercial (ANLI, Financial PhraseBank), some are share-alike (BoolQ, Financial PhraseBank), and some have no stated licence.

Attribution:

  • OpenJev (openjev/openjev, CC BY-NC 4.0), the teacher;
  • Ornith AI (ornith-ai/Ornith-1.5-35B-A3B, MIT), the base;
  • Inco AI (incoai/Qwen3.6-35B-A3B-Splash, Apache-2.0), the draft and packing format;
  • Zefan Cai (ZefanCai/Open-Jev, CC0), training data.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ezoushen/umpire-35b-a3b-splash

Finetuned
(32)
this model

Datasets used to train ezoushen/umpire-35b-a3b-splash

Collection including ezoushen/umpire-35b-a3b-splash