Hy-MT2-1.8B-StreamRevise-v4 — GGUF

A 4-bit llama.cpp build for live subtitle translation. On every ASR update, the caller passes the current source hypothesis and the last translation of the utterance. The model can preserve and extend a correct prefix or revise it when recognition changes the meaning. Recent bilingual context is also supported.

The model file is 1.07 GB and runs on CPU or GPU. The LoRA adapter, complete prompt renderer, training details, and evaluation methodology are available in Hy-MT2-1.8B-StreamRevise-v4-LoRA. Base model: tencent/Hy-MT2-1.8B.

中文简介:面向实时字幕的 4-bit 本地翻译模型。每次 ASR 更新时,把当前识别文本和上一版译文 一起传入;模型会尽量保留仍然正确的前缀,在新增内容或识别纠正时继续扩展或改写。支持最近几句 双语上下文。模型文件 1.07 GB,可用 llama.cpp 在 CPU 或 GPU 上运行。


Files

file size description
Hy-MT2-1.8B-StreamRevise-v4-Q4_K_M.gguf 1.07 GB Merged Q4_K_M model with Q4_K token embeddings and imatrix calibration
Hy-MT2-1.8B-StreamRevise-v4.imatrix.gguf 2.39 MB Importance matrix for re-quantizing the model at another bit width

The importance matrix was calibrated with held-out prompts drawn from the same streaming-translation input format as the model's intended workload.


Run it

llama-server -m Hy-MT2-1.8B-StreamRevise-v4-Q4_K_M.gguf -c 4096 --port 8080 -ngl 99 --no-webui

Drop -ngl 99 (or set it to 0) on a CPU-only machine. Then POST to /completion:

{
  "prompt": "<|hy_begin▁of▁sentence|><|hy_User|>{PROMPT}<|hy_Assistant|>",
  "n_predict": 256,
  "temperature": 0,
  "cache_prompt": true
}

Three settings matter:

  1. temperature: 0 uses greedy decoding. Sampling creates unrelated changes between nearly identical ASR updates and makes subtitles flicker.
  2. cache_prompt: true lets consecutive updates reuse their shared prompt prefix.
  3. {PROMPT} must use the StreamRevise layout below. A generic translation instruction does not provide the revision state the model needs.

Loader warning

Some llama.cpp builds print these messages:

load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect
load: special_eom_id is not in special_eog_ids - the tokenizer config may be incorrect

This GGUF contains tokenizer.ggml.eom_token_id = 120020, so generation terminates correctly. A successful request normally reports stop_type: eos before reaching n_predict.


Prompt format

Use full English language names. The complete copy-paste renderer is in the LoRA repository.

First chunk of a new utterance:

Translate the following text from {SOURCE_LANGUAGE} into {TARGET_LANGUAGE}. Note that you should only output the translated result without any additional explanation:

{CURRENT_SOURCE}

Subsequent updates:

[Background Information]
Recent utterances and their translations:
{SOURCE_SENTENCE_1} → {TRANSLATION_1}
{SOURCE_SENTENCE_2} → {TRANSLATION_2}
...

Previous version of the current source:
{PREVIOUS_SOURCE}

Previous translation of the current source:
{PREVIOUS_TRANSLATION}

When the source meaning has not changed, preserve the still-correct prefix of the previous translation whenever possible. When content is added or corrected, accuracy and completeness take priority.

Please translate the following text from {SOURCE_LANGUAGE} into {TARGET_LANGUAGE}, taking the provided background information into consideration.

[Source Text]
{CURRENT_SOURCE}

Each background block is optional and blocks are separated by one blank line. Source-only history is also supported by replacing the bilingual history block with:

Recent source utterances:
{SOURCE_SENTENCE_1}
{SOURCE_SENTENCE_2}
...

Use no more than the latest 10 utterances. The caller owns the revision chain and sends it in full on every request; the model itself is stateless.

Very short fragments can intentionally return an empty string when they do not yet contain enough meaning. Keep displaying the last non-empty translation or wait for the next ASR update.


Evaluation

This exact Q4_K_M build was evaluated with greedy decoding. It passed a three-track gate covering general translation quality, real streaming ASR behavior, and targeted translation probes.

FLORES-200 devtest

Six Chinese/English/Japanese directions, 1,012 examples per direction, corpus chrF with word_order=0:

direction chrF
Japanese → Chinese 27.303
Chinese → Japanese 32.953
English → Chinese 38.202
Chinese → English 57.369
Japanese → English 54.671
English → Japanese 40.643
Combined corpus 46.393

FLORES is English-pivoted, so absolute Chinese↔Japanese scores include reference divergence. The table is provided as a reproducible evaluation record, not as a ranking of language-pair difficulty.

Real streaming replay

500 utterance trajectories / 792 ASR states:

metric result
source-language leakage 1.01%
true source-copy rate 0.00%
empty-output rate for source fragments longer than 4 characters 0.35%
punctuation-insensitive prefix retention 0.950
punctuation-insensitive characters erased per transition 0.272
punctuation-insensitive rewrite rate 6.8%
final output/source length ratio, median 0.795

The no-copy, no-added-brackets, and word-sense probes passed 24/24 cases. These streaming metrics are application-oriented heuristics and are not directly comparable with simultaneous-MT paper benchmarks.


Footprint

The model has 32 layers, 4 KV heads, and a head dimension of 128. With an f16 KV cache, cache allocation is about 64 KiB per token:

n_ctx approximate KV cache approximate model + KV weights
1024 64 MiB 1.13 GiB
2048 128 MiB 1.19 GiB
4096 256 MiB 1.31 GiB
8192 512 MiB 1.56 GiB

Runtime process memory will be higher because llama.cpp also allocates compute buffers and, on GPU, a backend context. Reduce -c, -b, or -ub when memory is tight.


Re-quantization note

The published model uses Q4_K_M weights, Q4_K token embeddings, and the included importance matrix. Hy-MT2 uses tied token embeddings and output weights, so embedding precision has a noticeable effect on file size.

When converting Hy-MT2 yourself, ensure the GGUF contains an end-of-message token. Some converter versions do not write it when the source configuration has a single-valued eos_token_id. The required metadata is:

tokenizer.ggml.eom_token_id = 120020

This published file already contains that field.


Limitations

  • Stability is a tendency, not a guarantee. A recognition correction can require rewriting the whole line.
  • Very short fragments can return an empty string. Preserve the last non-empty subtitle until more text arrives.
  • Prompt format matters. Off-format prompts can reduce both translation quality and stability.
  • Core language coverage is Chinese, English, and Japanese. Other directions were not trained or evaluated.
  • Sentence-scoped revision. Finalized earlier subtitles are not revisited.
  • Greedy decoding is assumed.
  • Garbled or highly incomplete ASR text can still produce mistranslations or hallucinations.
  • The model inherits the base model's biases and limitations.

License

Apache 2.0, the same license as tencent/Hy-MT2-1.8B.

Downloads last month
80
GGUF
Model size
2B params
Architecture
hunyuan-dense
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF

Quantized
(1)
this model