Torchcast Decision 12B

A text-based decision checkpoint developed by Torchcast AI, based on Gemma-4-12B-it. It supports yes/no probabilities (noul), choice distributions (choice), and ordinal score distributions (score).

Use the supplied serving package for typed decisions. Generic Hub text-generation examples do not reproduce the evaluated interface.

Serving instructions

Model and readout

A LoRA fine-tune of Gemma-4-12B-it, trained on a 50/50 mix of gold labels and a larger instruction model's option distributions (not TypeSafe/Jev outputs), over procedurally generated decisions, the Open-Jev train split and public dataset train splits; sources and their terms are listed in LICENSE. Each decision is one forward pass that reads the option letters at the answer position. Per-type temperatures in readout_config.json were fitted on validation or test splits of public non-JevBench datasets: Choice and Score use the negative-log-likelihood optimum (1.0); the yes/no (noul) temperature (0.2) is the value on a fixed grid (0.2–3.0) with the best chance-corrected fit-set accuracy among those whose fit-set calibration error stays within the unmodified base model's, and is the lowest value on that grid.

Artifact and evaluation access

  • Pinned revision: tag v1.0.0 of this repository; the full commit hash is given in the serving instructions.
  • Weight SHA-256: f8553c9e625fa24853d57e938a1b1475975d9bce4bb79fccff2e8bfaaf6caa4e.
  • Recorded runtime: BF16, vLLM 0.30.0, one NVIDIA H100 80 GB (release check); the same public-item result was recorded on an L40S 48 GB during development.
  • Served context limit: 16,384 tokens including the rendered request.

Follow the linked serving instructions and use the bundled configuration unchanged. The evaluation interface is POST /v1/systemone, with one question named decision and model name torchcast-decision-12b.

Runtime environment

From the serving-repository checkout, create and activate a virtual environment before installing and launching vLLM:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install vllm==0.30.0

Keep this environment activated when running the linked model-launch command. In the second terminal, enter the same checkout and run source .venv/bin/activate before starting the supplied server. Activation puts installed compilation tools such as ninja on PATH; invoking the vLLM binary by its full path alone does not activate the environment.

Public-231 result (JevBench v1.4 public set)

Author-run, not an official JevBench score or rank. On the 231 public items, the checkpoint answered 203/231 correctly (87.88%): easy 48/48, original 70/72, hard 85/111. All answers were valid and no requests failed. On the H100 release check, serial client latency on easy+original items was 30 ms p50 and 43 ms p95, with a mean of 704 input tokens and 1 output token per decision; this is not the official Speed measurement.

The run used JevBench CLI bb05a335bc809e61b20c0f745d25499a82b326fc. Its manifest and per-item results are in the serving repository under runs/h100-release/. The result does not cover the current full suite or sealed set.

Benchmark exposure

Public JevBench results influenced checkpoint and serving-configuration selection. An 8-gram screen of an earlier training set found procedurally generated rows whose template wording overlapped 14 public items. Those rows were removed, and this checkpoint's training data shares no 8-gram with the 231 public items. No item fact patterns or answers were copied, and no sealed or unpublished items were accessed. The public results are development measurements, not an untouched holdout or a zero-contamination claim.

Intended use and limitations

English text-based decision research, including classification, routing and bounded rubric evaluation. The server returns decisions and does not execute actions. Probabilities can be unreliable on new domains. Inputs exceeding the served context limit receive HTTP 422. The public measurements do not establish larger-option or multi-question performance, multilingual capability, image/audio quality or autonomous-agent quality.

Licence and attribution

Model weights are labelled CC BY-NC 4.0, for non-commercial use. This label does not establish clearance of every training-source right. Applicable upstream model conditions and source-specific terms remain relevant; see the supplied licence files. Use is also subject to Google's Gemma Prohibited Use Policy (https://ai.google.dev/gemma/prohibited_use_policy).

Serving code is MIT-licensed. Preserve the supplied copyright, licence and source notices (LICENSE, LICENSE-gemma-Apache-2.0, LICENSE-serving-MIT). Credit Torchcast AI for this checkpoint, Google DeepMind for Gemma, and Cygnet/blockbrain-ai and NInfer for the upstream serving/readout foundation. This project is independent of TypeSafe AI and the JevBench maintainers.

Downloads last month
478
Safetensors
Model size
12B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for torchcast-ai/torchcast-decision-12b

Finetuned
(225)
this model
Quantizations
1 model

Datasets used to train torchcast-ai/torchcast-decision-12b

Space using torchcast-ai/torchcast-decision-12b 1