6 vwa · 19 fraz6 voices · 19 sentencesFatiman

Fatiman

yon vwa ki li kreyòl jan l ekri a, sou pwòp òdinatè w

a voice that reads Haitian Creole the way it is written, on your own computer

Fatiman-TTS-v1 fè tèks kreyòl tounen pawòl. Li mache sou pwosesè yon laptop, san entènèt, epi premye son an soti nan 0.12 segonn, pandan rès fraz la ap fèt.

Fatiman-TTS-v1 turns Kreyòl text into speech. It runs on a laptop's CPU, without the internet, and the first sound comes out in 0.12 seconds while the rest of the sentence is still being made.

0.12s

    KouteListen

    Sis vwa, ven fraz

    Six voices, twenty sentences

    • Chwazi yon vwaPick a voicetwa fanm, twa gasonthree women, three men
    • Klike yon frazClick a sentencedat, pri, lè, tanperati, nondates, prices, times, temperatures, names

    Chak fraz sa yo te li pa Fatiman-TTS-v1, san okenn koreksyon. Se fraz sa yo, ak yon lòt, ki te sèvi pou mezire vwa yo. Chif yo ekri an lèt anvan modèl la li yo ("1804" → "mil ui san kat").

    Every sentence here was read by Fatiman-TTS-v1, with no editing. The voices were scored on these sentences and one more. Numbers are spelled out before the model reads them ("1804" → "mil ui san kat").

      Odyo: MP3 40 kbps, konprese pou paj la. Orijinal la se 24 kHz.Audio: 40 kbps MP3, compressed for this page. The model outputs 24 kHz.

      PoukisaWhy

      Yon vwa ki li kreyòl, pa franse

      A voice that reads Kreyòl, not French

      • ~20%erè lè vwa franse a li kreyòlerror when the French voice reads Kreyòl
      • 2.3%erè Fatiman, pi bon vwa aFatiman's error, best voice
      • 0.12 sanvan premye son anbefore the first sound

      Kreyòl ekri ak menm lèt ak franse, men li pa li menm jan. Yon vwa franse ki li "Mwen pa konnen" ap fè lèt yo pa pwononse, oswa sote lèt ki la. Modèl franse Kyutai a, san fòmasyon, te fè apeprè 20% erè sou karaktè yo. Fatiman desann sa a 2.3%.

      Kreyòl is written with the same letters as French, but it is not read the same way. A French voice reading "Mwen pa konnen" applies French habits: silent final letters, French vowels. Kyutai's French model, untrained, made about 20% character errors. Fatiman brings that down to 2.3%.

      Pou yon asistan vwa, vitès konte tou. Fatiman pale pandan l ap kalkile: premye son an soti apre 0.12 segonn, kidonk asistan an reponn san moun nan pa tann yon silans.

      For a voice assistant, speed matters too. Fatiman speaks while it computes: the first sound comes after 0.12 seconds, so the assistant answers without a silence.

      KonparezonComparison

      Menm 80 fraz, kat vwa

      The same 80 sentences, four voices

      • 80 fraz li pa t janm wè80 sentences it never saw
      • Li awotvwa, epi yon ASR transkriRead aloud, then transcribed by ASRerè sou karaktè, pi ba pi boncharacter error rate, lower is better
      VwaVoiceMounSpeakerWhisperQwen3-ASRPremye sonFirst sound
      Fatiman · man3gasonman2.3%2.0%~0.12 s
      Fatiman · female3fanmwoman3.6%2.1%~0.12 s
      Kokoro–4.4%3.6%apre tout fraz laafter the whole phrase
      Qwen3-TTS 0.6B–9.1%6.5%apre tout fraz laafter the whole phrase

      Vwa yoThe voices

      Moun reyèl, ak dakò yo

      Real people, with their consent

      • Moun yo te peyeSpeakers were paid
      • Yo siyen dakò yoThey signed their consent
      • 10 segonn pou chak vwa10 seconds per voice

      Chak vwa soti nan apeprè 10 segonn anrejistreman yon moun ki te peye epi ki te siyen pou vwa l itilize. Modèl la ka imite yon vwa apati yon ti klip: pa janm sèvi ak vwa yon moun san pèmisyon l, ni pou twonpe moun.

      Each voice comes from about 10 seconds of a recording by a speaker who was paid and signed for their voice to be used. The model can imitate a voice from a short clip: never use someone's voice without their permission, or to deceive.

      VwaVoiceMounSpeakerWhisperQwen3-ASR

      Erè sou 20 fraz egzanp: 19 ki anlè a ak yon lòt. female2: yon sèl fraz ki te koupe twò bonè fè nòt li monte.Error on 20 sample sentences: the 19 above and one more. female2's score comes from one sentence cut short too early.

      Kijan li fètHow it was made

      Yon vwa franse ki aprann kreyòl

      A French voice that learned Kreyòl

      • ~122 hpawòl kreyòlof Kreyòl speech
      • 12,000etap fòmasyontraining steps
      • CC-BY-4.0baze sou Pocket TTS, Kyutaibased on Pocket TTS by Kyutai

      Baz la se Pocket TTS Kyutai, modèl franse 24 kouch la. Paske kreyòl ekri ak lèt franse, nou kenbe tokenizer franse a epi nou anseye l òtograf kreyòl la, ak lekti Bib la an kreyòl ak lòt anrejistreman kreyòl, tout tcheke ak yon ASR.

      The base is Kyutai's Pocket TTS, the French 24-layer model. Because Kreyòl is written with French letters, we kept the French tokenizer and taught it Kreyòl spelling, with Kreyòl Bible readings and other Kreyòl recordings, all checked with an ASR.

      Lekti Bib laBible readings95 h
      Lòt pawòl kreyòl, 24 kHzOther Kreyòl speech, 24 kHz27 h

      Sèvi avè lUse it

      Sou laptop, sou telefòn, san entènèt

      On a laptop, on a phone, offline

      Python · pocket-tts

      import os, tempfile
      import numpy as np, soundfile as sf
      from huggingface_hub import snapshot_download
      from pocket_tts import TTSModel
      
      folder = snapshot_download("jsbeaudry/Fatiman-TTS-v1")
      config = os.path.join(tempfile.gettempdir(), "fatiman.yaml")
      with open(config, "w") as f:
          f.write(open(os.path.join(folder, "config.template.yaml"))
                  .read().replace("@MODEL_DIR@", folder))
      model = TTSModel.load_model(config=config)
      voice = model.get_state_for_audio_prompt(
          os.path.join(folder, "voices", "man3.safetensors"))
      audio = np.concatenate([c.detach().cpu().numpy().reshape(-1)
          for c in model.generate_audio_stream(voice, "Bonjou! Kijan ou ye?")])
      sf.write("bonjou.wav", audio, model.sample_rate)

      Telefòn · ONNXPhones · ONNX

      Yon vèsyon ONNX (int8 ak q4) pou telefòn, menm sis vwa yo. Se li aplikasyon mobil nou an sèvi.An ONNX version (int8 and q4) for phones, with the same six voices. Our mobile app uses it.

      Fatiman-TTS-v1-onnx ↗

      The Trio Serie

      Fatiman se vwa yon asistan vwa kreyòl ki mache san entènèt, ansanm ak Makandal pou reflechi.Fatiman is the voice of a Kreyòl voice assistant that runs offline, with Makandal as its brain.

      Modèl la ↗The model ↗