114 klip114 clipsOswald

Oswald

yon zòrèy ki tande kreyòl, epi ki ekri l jan l di a

ears that hear Haitian Creole, and write it the way it is said

Oswald-ASR-0.6-V1 fè pawòl kreyòl tounen tèks. Li mache sou yon laptop, san entènèt, epi li ekri yon fraz nan apeprè 0.2 segonn.

Oswald-ASR-0.6-V1 turns Kreyòl speech into text. It runs on a laptop, without the internet, and writes a sentence down in about 0.2 seconds.

9.8%

    Teste lTest it

    Koute, epi li sa l tande

    Listen, then read what it heard

    • Fraz la, jan l te ekriThe sentence, as writtenan griin grey
    • Sa Oswald ekriWhat Oswald wrotemo ki diferan yo an woujwords that differ in red

    Vwa yo se vwa Fatiman, modèl ki fè tèks tounen pawòl. Oswald tande chak klip sou yon laptop, epi nou kopye sa l ekri a jan l soti. Anpil "erè" se fason pou ekri: kijan pou ki jan, katsan pou kat san.

    The voices are Fatiman's, the text-to-speech model. Oswald heard each clip on a laptop, and its text is copied as it came out. Many "errors" are spelling choices: kijan for ki jan, katsan for kat san.

    Vwa sentetik: Oswald te tande lòt pawòl sentetik pandan fòmasyon an. Nòt ki konte yo se sa ki sou pawòl moun reyèl, pi ba a.Synthetic voices: Oswald heard other synthetic speech in training. The scores that count are the ones on real human speech, below.

      PoukisaWhy

      Modèl la pa t konnen kreyòl

      The base model did not know Kreyòl

      • 30lang Qwen3-ASR sipòte; kreyòl pa ladan yolanguages Qwen3-ASR supports; Kreyòl is not one
      • 90.1% → 9.8%erè sou mo yo, anvan ak apreword error, before and after

      Qwen3-ASR 0.6B konprann 30 lang ak 22 dyalèk chinwa, men pa kreyòl. Lè l tande kreyòl, li ekri l an franse: 90 mo sou 100 soti mal. Sa li konnen byen, se son yo. Kidonk nou kite pati ki tande a jan l te ye, epi nou anseye pati ki ekri a ekri kreyòl.

      Qwen3-ASR 0.6B understands 30 languages and 22 Chinese dialects, but not Kreyòl. Hearing Kreyòl, it writes French: 90 words in 100 come out wrong. What it does know is the sounds. So we left the part that hears untouched, and taught the part that writes to write Kreyòl.

      Rezilta a: 9.8% erè sou mo, sou 875 klip ak tèks moun te ekri, nan Bib, radyo, YouTube, ak anrejistreman sou sante, lalwa, lajan ak istwa.

      The result: 9.8% word error on 875 clips with human-written transcripts, across the Bible, radio, YouTube, and recordings about health, law, money and history.

      KonparezonComparison

      Sou pawòl moun reyèl

      On real human speech

      • 875 klip li pa t janm tande875 clips it never heard
      • Erè sou mo (WER)Word error rate (WER)pi ba pi bonlower is better

      VitèsSpeed

      Yon repons kout, yon tan kout

      A short answer, a short wait

      • 0.21 smedyàn pa klip, sou yon laptopmedian per clip, on a laptop
      • < 0.24 spou 90% klip yofor 90% of the clips
      • ≥ 0.81 sfine-tune Whisper la, pou nenpòt klipthe Whisper fine-tune, for any clip

      Whisper ranpli chak klip jiska 30 segonn, kidonk de mo koute menm tan ak ven mo. Oswald travay sou longè klip la sèlman. Sou yon MacBook M3 Pro, ak fichye Q4_K_M a (660 MB), li ekri 114 fraz yo nan yon medyàn 0.21 segonn chak.

      Whisper pads every clip to 30 seconds, so two words cost what twenty cost. Oswald works on the clip's real length. On a MacBook M3 Pro, with the Q4_K_M file (660 MB), it wrote the 114 sentences at a median 0.21 seconds each.

      Kijan li fètHow it was made

      146 èdtan pawòl kreyòl

      146 hours of Kreyòl speech

      • 65,382klipclips
      • 16sous donedata sources
      • Apache-2.0baze sou Qwen3-ASR 0.6Bbased on Qwen3-ASR 0.6B

      Bib la li awotvwa (limite a 35 èdtan pou yon sèl vwa pa domine), Radyo Ayiti Entè, anrejistreman CMU, videyo YouTube nou verifye, anrejistreman estidyo, ak 18% pawòl sentetik. Yon LoRA sou pati ki ekri a; pati ki tande a pa chanje.

      The Bible read aloud (capped at 35 hours so one voice does not dominate), Radio Haïti-Inter, CMU recordings, checked YouTube videos, studio recordings, and 18% synthetic speech. A LoRA on the part that writes; the part that hears is unchanged.

      Yon leson: pawòl sentetik poukont li fè modèl la vin pi mal (11.4 → 12.6 WER). Melanje ak pawòl moun, li fè chak domèn vin pi bon.

      One lesson: synthetic speech alone made the model worse (11.4 → 12.6 WER). Mixed with human speech, it improved every domain.

      Sèvi avè lUse it

      Sou laptop, sou telefòn, san entènèt

      On a laptop, on a phone, offline

      llama.cpp · GGUF

      Q4_K_M (660 MB ak projektè a) oswa Q8_0 (966 MB).Q4_K_M (660 MB with the projector) or Q8_0 (966 MB).

      llama-mtmd-cli -m qwen3-asr-0.6b-kreyol-v2-Q8_0.gguf \
        --mmproj mmproj-qwen3-asr-0.6b-kreyol-v2-Q8_0.gguf \
        --audio clip.wav -p '' --temp 0
      Oswald-ASR-0.6-V1-GGUF ↗

      The Trio Serie

      Oswald se zòrèy yon asistan vwa kreyòl ki mache san entènèt, ansanm ak Makandal pou reflechi ak Fatiman pou pale. Aplikasyon mobil nou an sèvi ak menm fichye GGUF yo.Oswald is the ears of a Kreyòl voice assistant that runs offline, with Makandal to think and Fatiman to speak. Our mobile app uses the same GGUF files.