You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

VoxCPM2 Persian β€” round 3

Continued from markmuller/TTS_POST_trainnig_1_emo / best_checkpoint_of_post_training, 2 epochs (2,917 steps) on 98,061 rows / 270.5 h:

slice rows hours text
gold (TTS_DATA_GOLD) 77,025 141.4 (emotion)[tag] captions, plain text
fidibo50 8,361 49.4 (fidibo) + diacritized text, same-speaker refs
neutral50 2,236 46.6 Gemini long-form (43–150 s), inline per-sentence captions, same-voice refs
nonverbal38k 10,439 33.1 EN + ZH replay with nonverbal tags, zero-shot (CC-BY-NC-4.0 source)

Length-bucketed batches (400 s padded audio per micro-batch, x3 accumulation), LR 1e-5, one B300.

Checkpoints

folder step why
step_0002917 2917 (end of epoch 2) best flow-matching loss on every slice
step_0001000 1000 best stop-head loss β€” try it if endings are cut short or run on

Each folder is a full trainer checkpoint (weights + optimizer + scheduler), so it can be loaded for inference or resumed.

Per-slice validation (full val sets, identical batches and noise per checkpoint)

diffusion loss / stop loss, lower is better:

ckpt gold neutral50 fidibo50 nonverbal38k
init 0.8983 / 0.0137 0.9345 / 0.0023 0.8031 / 0.0038 1.0056 / 0.0255
step 1000 0.8902 / 0.0125 0.9271 / 0.0017 0.8048 / 0.0028 0.9916 / 0.0209
step 2917 0.8837 / 0.0137 0.9235 / 0.0027 0.8035 / 0.0025 0.9901 / 0.0242

run/ holds the generated training config, the training log and the full evaluation JSON.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support