Gooya RizehPizeh v2

Persian single-speaker Piper/VITS2 voice trained on full, unchunked Wikipedia recordings whose audio/transcript agreement passed AvaSanj. It uses a case-sensitive Negara-style phonetic/Finglish frontend and produces 22,050 Hz mono audio.

Current release

This revision replaces the earlier segmented-data v2 model with the user-selected unchunked/AvaSanj-approved 50K run.

  • Full checkpoint: checkpoint/gooya-rizehpizeh-v2-unchunked-avasanj-step50000-val_mel0.4138.ckpt
  • Checkpoint SHA-256: 8b52b3b6060b4698aba0c8bbd2f9b338aa6ba811d28a620e5f492c0fa1ae4d2b
  • Piper/browser model: gooya-rizehpizeh-v2.onnx
  • ONNX SHA-256: 3f85f450fc21415e1eddaa0aa894c667e76d14aa9a342c4e77f1f9009b96f170
  • ONNX config SHA-256: 981770210ac1aea618cb8134597127911051deabf9900824df3388f209a884fa
  • Generator parameters: approximately 23.7M
  • Sample rate: 22,050 Hz
  • Speakers: 1
  • Phoneme type: text
  • Symbol inventory: 256 IDs

Training data

The parent corpus comes from the approved Persian rows in Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731.

  • Train: 1,290 full recordings, 21.1455 hours
  • Parent-disjoint validation: 100 full recordings, 1.6445 hours
  • Untouched sealed-eval agreement slice: 253 recordings, 4.1080 hours
  • Recording duration: 38.68โ€“95.96 seconds; mean 58.94 seconds
  • Phoneme IDs per recording: 1,711โ€“3,423; mean 2,174.65
  • AvaSanj compact CER threshold: <= 0.12
  • Train/validation/sealed-eval parent overlap: zero
  • Audio transform: none
  • Silence trimming: disabled
  • Segmentation/chunking: none

The 253-row sealed-eval agreement slice was not consumed by training or checkpoint selection. The complete selection and hash receipt is in training/unchunked-avasanj-agree-report.json.

Lineage and training

The original v1.5 production ONNX had no matching resumable checkpoint. Its 350 inference-time generator tensors were reconstructed into a compatible trainable Piper checkpoint; training-only branches absent from the ONNX came from the compatible Mana template. Recovery receipts remain under training/.

The promoted run resumed from that recovered v1.5 state and used:

  • BF16 mixed precision
  • Batch size 1
  • Generator learning rate 5e-6
  • Discriminator learning rate 2.5e-6
  • Configured limit: 50,000 optimizer steps
  • Final/best validation mel: 0.4138

The exact continuation script is training/train_unchunked_avasanj_agree_full50k.sh.

Validation

The published ONNX is self-contained and has no external tensor file. Direct ONNX Runtime CPU inference passed on both the canonical and long colloquial probes with finite, positive-length output. The same WAVs used for human listening are included under samples/.

  • Inputs: input, input_lengths, scales
  • Output: output
  • File size: 63,516,051 bytes
  • Canonical sample: 4.7137 seconds
  • Long colloquial sample: 12.6084 seconds

Machine-readable receipts are under browser/ and training/.

Inference

Input must already be case-sensitive Negara-compatible phonetic text:

echo "zurxAneye vahdat yek makAne varzeSi dar Sahre borujerd ast." | \
  piper --model gooya-rizehpizeh-v2.onnx \
        --config gooya-rizehpizeh-v2.onnx.json \
        --output_file output.wav

For long-form narration, split at semantic sentence or clause boundaries into substantial spans and join them with boundary-aware pauses. Tiny fixed chunks tend to sound breathy and robotic.

Limitations

  • G2P/Finglish quality remains the dominant source of pronunciation errors.
  • The approved training labels are flat Negara phoneme sequences; existing internal punctuation was not available in those labels.
  • Very long single-pass phoneme sequences can still collapse or omit content; use substantial semantic spans.
  • Proper nouns may require corrected phonetic input.

License

MIT. See LICENSE.

Downloads last month
43
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Reza2kn/Gooya-RizehPizeh-v2

Quantized
(1)
this model