Fatiman
yon vwa ki li kreyòl jan l ekri a, sou pwòp òdinatè w
a voice that reads Haitian Creole the way it is written, on your own computer
Fatiman-TTS-v1 fè tèks kreyòl tounen pawòl. Li mache sou pwosesè yon laptop, san entènèt, epi premye son an soti nan 0.12 segonn, pandan rès fraz la ap fèt.
Fatiman-TTS-v1 turns Kreyòl text into speech. It runs on a laptop's CPU, without the internet, and the first sound comes out in 0.12 seconds while the rest of the sentence is still being made.
KouteListen
Sis vwa, ven fraz
Six voices, twenty sentences
- Chwazi yon vwaPick a voicetwa fanm, twa gasonthree women, three men
- Klike yon frazClick a sentencedat, pri, lè, tanperati, nondates, prices, times, temperatures, names
Chak fraz sa yo te li pa Fatiman-TTS-v1, san okenn koreksyon. Se fraz sa yo, ak yon lòt, ki te sèvi pou mezire vwa yo. Chif yo ekri an lèt anvan modèl la li yo ("1804" → "mil ui san kat").
Every sentence here was read by Fatiman-TTS-v1, with no editing. The voices were scored on these sentences and one more. Numbers are spelled out before the model reads them ("1804" → "mil ui san kat").
Odyo: MP3 40 kbps, konprese pou paj la. Orijinal la se 24 kHz.Audio: 40 kbps MP3, compressed for this page. The model outputs 24 kHz.
PoukisaWhy
Yon vwa ki li kreyòl, pa franse
A voice that reads Kreyòl, not French
- ~20%erè lè vwa franse a li kreyòlerror when the French voice reads Kreyòl
- 2.3%erè Fatiman, pi bon vwa aFatiman's error, best voice
- 0.12 sanvan premye son anbefore the first sound
Kreyòl ekri ak menm lèt ak franse, men li pa li menm jan. Yon vwa franse ki li "Mwen pa konnen" ap fè lèt yo pa pwononse, oswa sote lèt ki la. Modèl franse Kyutai a, san fòmasyon, te fè apeprè 20% erè sou karaktè yo. Fatiman desann sa a 2.3%.
Kreyòl is written with the same letters as French, but it is not read the same way. A French voice reading "Mwen pa konnen" applies French habits: silent final letters, French vowels. Kyutai's French model, untrained, made about 20% character errors. Fatiman brings that down to 2.3%.
Pou yon asistan vwa, vitès konte tou. Fatiman pale pandan l ap kalkile: premye son an soti apre 0.12 segonn, kidonk asistan an reponn san moun nan pa tann yon silans.
For a voice assistant, speed matters too. Fatiman speaks while it computes: the first sound comes after 0.12 seconds, so the assistant answers without a silence.
KonparezonComparison
Menm 80 fraz, kat vwa
The same 80 sentences, four voices
- 80 fraz li pa t janm wè80 sentences it never saw
- Li awotvwa, epi yon ASR transkriRead aloud, then transcribed by ASRerè sou karaktè, pi ba pi boncharacter error rate, lower is better
| VwaVoice | MounSpeaker | Whisper | Qwen3-ASR | Premye sonFirst sound |
|---|---|---|---|---|
| Fatiman · man3 | gasonman | 2.3% | 2.0% | ~0.12 s |
| Fatiman · female3 | fanmwoman | 3.6% | 2.1% | ~0.12 s |
| Kokoro | – | 4.4% | 3.6% | apre tout fraz laafter the whole phrase |
| Qwen3-TTS 0.6B | – | 9.1% | 6.5% | apre tout fraz laafter the whole phrase |
Vwa yoThe voices
Moun reyèl, ak dakò yo
Real people, with their consent
- Moun yo te peyeSpeakers were paid
- Yo siyen dakò yoThey signed their consent
- 10 segonn pou chak vwa10 seconds per voice
Chak vwa soti nan apeprè 10 segonn anrejistreman yon moun ki te peye epi ki te siyen pou vwa l itilize. Modèl la ka imite yon vwa apati yon ti klip: pa janm sèvi ak vwa yon moun san pèmisyon l, ni pou twonpe moun.
Each voice comes from about 10 seconds of a recording by a speaker who was paid and signed for their voice to be used. The model can imitate a voice from a short clip: never use someone's voice without their permission, or to deceive.
| VwaVoice | MounSpeaker | Whisper | Qwen3-ASR |
|---|
Erè sou 20 fraz egzanp: 19 ki anlè a ak yon lòt. female2: yon sèl fraz ki te koupe twò bonè fè nòt li monte.Error on 20 sample sentences: the 19 above and one more. female2's score comes from one sentence cut short too early.
Kijan li fètHow it was made
Yon vwa franse ki aprann kreyòl
A French voice that learned Kreyòl
- ~122 hpawòl kreyòlof Kreyòl speech
- 12,000etap fòmasyontraining steps
- CC-BY-4.0baze sou Pocket TTS, Kyutaibased on Pocket TTS by Kyutai
Baz la se Pocket TTS Kyutai, modèl franse 24 kouch la. Paske kreyòl ekri ak lèt franse, nou kenbe tokenizer franse a epi nou anseye l òtograf kreyòl la, ak lekti Bib la an kreyòl ak lòt anrejistreman kreyòl, tout tcheke ak yon ASR.
The base is Kyutai's Pocket TTS, the French 24-layer model. Because Kreyòl is written with French letters, we kept the French tokenizer and taught it Kreyòl spelling, with Kreyòl Bible readings and other Kreyòl recordings, all checked with an ASR.
Sèvi avè lUse it
Sou laptop, sou telefòn, san entènèt
On a laptop, on a phone, offline
Python · pocket-tts
import os, tempfile
import numpy as np, soundfile as sf
from huggingface_hub import snapshot_download
from pocket_tts import TTSModel
folder = snapshot_download("jsbeaudry/Fatiman-TTS-v1")
config = os.path.join(tempfile.gettempdir(), "fatiman.yaml")
with open(config, "w") as f:
f.write(open(os.path.join(folder, "config.template.yaml"))
.read().replace("@MODEL_DIR@", folder))
model = TTSModel.load_model(config=config)
voice = model.get_state_for_audio_prompt(
os.path.join(folder, "voices", "man3.safetensors"))
audio = np.concatenate([c.detach().cpu().numpy().reshape(-1)
for c in model.generate_audio_stream(voice, "Bonjou! Kijan ou ye?")])
sf.write("bonjou.wav", audio, model.sample_rate)Telefòn · ONNXPhones · ONNX
Yon vèsyon ONNX (int8 ak q4) pou telefòn, menm sis vwa yo. Se li aplikasyon mobil nou an sèvi.An ONNX version (int8 and q4) for phones, with the same six voices. Our mobile app uses it.
Fatiman-TTS-v1-onnx ↗The Trio Serie
Fatiman se vwa yon asistan vwa kreyòl ki mache san entènèt, ansanm ak Makandal pou reflechi.Fatiman is the voice of a Kreyòl voice assistant that runs offline, with Makandal as its brain.
Modèl la ↗The model ↗