Instructions to use febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF:Q4_K_M
Use Docker
docker model run hf.co/febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF with Ollama:
ollama run hf.co/febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF with Docker Model Runner:
docker model run hf.co/febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF:Q4_K_M
- Lemonade
How to use febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Hy-MT2-1.8B-StreamRevise-v4-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Hy-MT2-1.8B-StreamRevise-v4 — GGUF
A 4-bit llama.cpp build for live subtitle translation. On every ASR update, the caller passes the current source hypothesis and the last translation of the utterance. The model can preserve and extend a correct prefix or revise it when recognition changes the meaning. Recent bilingual context is also supported.
The model file is 1.07 GB and runs on CPU or GPU. The LoRA adapter, complete prompt renderer, training
details, and evaluation methodology are available in
Hy-MT2-1.8B-StreamRevise-v4-LoRA.
Base model: tencent/Hy-MT2-1.8B.
中文简介:面向实时字幕的 4-bit 本地翻译模型。每次 ASR 更新时,把当前识别文本和上一版译文 一起传入;模型会尽量保留仍然正确的前缀,在新增内容或识别纠正时继续扩展或改写。支持最近几句 双语上下文。模型文件 1.07 GB,可用 llama.cpp 在 CPU 或 GPU 上运行。
Files
| file | size | description |
|---|---|---|
Hy-MT2-1.8B-StreamRevise-v4-Q4_K_M.gguf |
1.07 GB | Merged Q4_K_M model with Q4_K token embeddings and imatrix calibration |
Hy-MT2-1.8B-StreamRevise-v4.imatrix.gguf |
2.39 MB | Importance matrix for re-quantizing the model at another bit width |
The importance matrix was calibrated with held-out prompts drawn from the same streaming-translation input format as the model's intended workload.
Run it
llama-server -m Hy-MT2-1.8B-StreamRevise-v4-Q4_K_M.gguf -c 4096 --port 8080 -ngl 99 --no-webui
Drop -ngl 99 (or set it to 0) on a CPU-only machine. Then POST to /completion:
{
"prompt": "<|hy_begin▁of▁sentence|><|hy_User|>{PROMPT}<|hy_Assistant|>",
"n_predict": 256,
"temperature": 0,
"cache_prompt": true
}
Three settings matter:
temperature: 0uses greedy decoding. Sampling creates unrelated changes between nearly identical ASR updates and makes subtitles flicker.cache_prompt: truelets consecutive updates reuse their shared prompt prefix.{PROMPT}must use the StreamRevise layout below. A generic translation instruction does not provide the revision state the model needs.
Loader warning
Some llama.cpp builds print these messages:
load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect
load: special_eom_id is not in special_eog_ids - the tokenizer config may be incorrect
This GGUF contains tokenizer.ggml.eom_token_id = 120020, so generation terminates correctly. A successful
request normally reports stop_type: eos before reaching n_predict.
Prompt format
Use full English language names. The complete copy-paste renderer is in the LoRA repository.
First chunk of a new utterance:
Translate the following text from {SOURCE_LANGUAGE} into {TARGET_LANGUAGE}. Note that you should only output the translated result without any additional explanation:
{CURRENT_SOURCE}
Subsequent updates:
[Background Information]
Recent utterances and their translations:
{SOURCE_SENTENCE_1} → {TRANSLATION_1}
{SOURCE_SENTENCE_2} → {TRANSLATION_2}
...
Previous version of the current source:
{PREVIOUS_SOURCE}
Previous translation of the current source:
{PREVIOUS_TRANSLATION}
When the source meaning has not changed, preserve the still-correct prefix of the previous translation whenever possible. When content is added or corrected, accuracy and completeness take priority.
Please translate the following text from {SOURCE_LANGUAGE} into {TARGET_LANGUAGE}, taking the provided background information into consideration.
[Source Text]
{CURRENT_SOURCE}
Each background block is optional and blocks are separated by one blank line. Source-only history is also supported by replacing the bilingual history block with:
Recent source utterances:
{SOURCE_SENTENCE_1}
{SOURCE_SENTENCE_2}
...
Use no more than the latest 10 utterances. The caller owns the revision chain and sends it in full on every request; the model itself is stateless.
Very short fragments can intentionally return an empty string when they do not yet contain enough meaning. Keep displaying the last non-empty translation or wait for the next ASR update.
Evaluation
This exact Q4_K_M build was evaluated with greedy decoding. It passed a three-track gate covering general translation quality, real streaming ASR behavior, and targeted translation probes.
FLORES-200 devtest
Six Chinese/English/Japanese directions, 1,012 examples per direction, corpus chrF with word_order=0:
| direction | chrF |
|---|---|
| Japanese → Chinese | 27.303 |
| Chinese → Japanese | 32.953 |
| English → Chinese | 38.202 |
| Chinese → English | 57.369 |
| Japanese → English | 54.671 |
| English → Japanese | 40.643 |
| Combined corpus | 46.393 |
FLORES is English-pivoted, so absolute Chinese↔Japanese scores include reference divergence. The table is provided as a reproducible evaluation record, not as a ranking of language-pair difficulty.
Real streaming replay
500 utterance trajectories / 792 ASR states:
| metric | result |
|---|---|
| source-language leakage | 1.01% |
| true source-copy rate | 0.00% |
| empty-output rate for source fragments longer than 4 characters | 0.35% |
| punctuation-insensitive prefix retention | 0.950 |
| punctuation-insensitive characters erased per transition | 0.272 |
| punctuation-insensitive rewrite rate | 6.8% |
| final output/source length ratio, median | 0.795 |
The no-copy, no-added-brackets, and word-sense probes passed 24/24 cases. These streaming metrics are application-oriented heuristics and are not directly comparable with simultaneous-MT paper benchmarks.
Footprint
The model has 32 layers, 4 KV heads, and a head dimension of 128. With an f16 KV cache, cache allocation is about 64 KiB per token:
n_ctx |
approximate KV cache | approximate model + KV weights |
|---|---|---|
| 1024 | 64 MiB | 1.13 GiB |
| 2048 | 128 MiB | 1.19 GiB |
| 4096 | 256 MiB | 1.31 GiB |
| 8192 | 512 MiB | 1.56 GiB |
Runtime process memory will be higher because llama.cpp also allocates compute buffers and, on GPU, a
backend context. Reduce -c, -b, or -ub when memory is tight.
Re-quantization note
The published model uses Q4_K_M weights, Q4_K token embeddings, and the included importance matrix. Hy-MT2 uses tied token embeddings and output weights, so embedding precision has a noticeable effect on file size.
When converting Hy-MT2 yourself, ensure the GGUF contains an end-of-message token. Some converter versions
do not write it when the source configuration has a single-valued eos_token_id. The required metadata is:
tokenizer.ggml.eom_token_id = 120020
This published file already contains that field.
Limitations
- Stability is a tendency, not a guarantee. A recognition correction can require rewriting the whole line.
- Very short fragments can return an empty string. Preserve the last non-empty subtitle until more text arrives.
- Prompt format matters. Off-format prompts can reduce both translation quality and stability.
- Core language coverage is Chinese, English, and Japanese. Other directions were not trained or evaluated.
- Sentence-scoped revision. Finalized earlier subtitles are not revisited.
- Greedy decoding is assumed.
- Garbled or highly incomplete ASR text can still produce mistranslations or hallucinations.
- The model inherits the base model's biases and limitations.
License
Apache 2.0, the same license as tencent/Hy-MT2-1.8B.
- Downloads last month
- 80
4-bit
Model tree for febilly/Hy-MT2-1.8B-StreamRevise-v4-GGUF
Base model
tencent/Hy-MT2-1.8B