Post

Antalia Mini Web: Offline Turkish TTS That Reads Numbers, Dates and Abbreviations Aloud

🇬🇧 Offline Turkish TTS running entirely in the browser, with a smart normalizer that correctly reads numbers, dates and abbreviations.

Antalia Mini Web: Offline Turkish TTS That Reads Numbers, Dates and Abbreviations Aloud

Almost every Turkish text-to-speech tool depends on a paid API or sends your voice to a server. If you want one-click speech synthesis with privacy, offline support, and zero setup, the options are bleak. To close that gap, I ported the Antalia-2 Mini model to run fully in the browser.

Meet Antalia Mini Web: an open-source web app that turns your text into speech on-device with WebGPU/WASM — nothing ever leaves your machine. | 🇹🇷 Türkçe

▶️ Live Demo: fr0stb1rd.github.io/antalia-mini-web | 💻 Repo: github.com/fr0stb1rd/antalia-mini-web

What Is Antalia Mini Web?

Antalia Mini Web is a browser port of cloud0day3/antalia-mini (PatientDesk AI, Antalia-2 Mini): 7.62M parameters, 48 kHz output, one synthetic male voice, Apache-2.0. The PyTorch pipeline is converted to ONNX and executed with onnxruntime-web + WebAudio. No server, no API keys.

The first launch downloads ~50 MB once (3 ONNX models + 4 dictionaries); afterwards it works fully offline via Cache Storage.

✨ Highlights

  • 1:1 Turkish normalizer (normalize.js): a line-by-line port of the original Python normalizer. Reads numbers (1250), decimals (68,5), ordinals (2.), dates (15.10.2026), times (14:30), percentages (%18), money (1.250,75 TL → lira + kuruş), phones, e-mails/URLs, abbreviations (Dr. → doktor, SGK → segeka, THY → teheye), Roman numerals, and English/brand words (GitHub → git hab). Matches Python output on 9,000+ inputs, verified by tests on every CI run.
  • Identical tone & pauses (audio.js): release tone path (+6 dB low-shelf at 210 Hz, −9 dB high-shelf at 4 kHz, soft limiter) and inter-chunk pauses (0.12 s after sentences, 0.06 s otherwise), sample-identical to the original within float32 tolerance.
  • 49 showcase examples in 9 categories (numbers, money, phones, brands, abbreviations…).
  • Streaming synthesis: sentences are produced piece by piece; the first one plays immediately while the rest generate in the background, with live word highlighting.
  • Flow-Matching + Shortcut Self-Consistency: 8-step DiT inference with CFG 2.0, baked into the ONNX model at export. Re-export with --steps 4 or --steps 2 for faster variants.
  • Zero bandwidth on repeat visits: models and dictionaries cached via the Cache Storage API (1-day TTL).
  • Privacy: 100% client-side; text and audio never leave the device.

🏗️ Architecture

The PyTorch pipeline is split into three ONNX stages for browser streaming:

flowchart LR
    A["Raw Turkish Text"] --> B["normalize.js\nNumbers, dates, brands"]
    B --> C["text_stage.onnx\nLetters → Features + Durations"]
    C --> D["Frame Planning (JS)\nDurations → Timeline"]
    D --> E["sound_stage.onnx\nDiT Flow-Matching → Mel"]
    E --> F["decoder.onnx\nConvNeXt Vocoder + iSTFT → 48 kHz"]
    F --> G["audio.js\nTone EQ + Pauses → WebAudio + .wav"]
Stage Input Output
text_stage.onnx ids [B,L], mask [B,L] h [B,L,192], logd [B,L]
Frame planning (JS) durations frame timeline
sound_stage.onnx cond, fmask, noise mel [B,128,T] (5-block DiT, 8 steps, CFG 2.0)
decoder.onnx mel audio [B,S] @48 kHz

Browser dictionaries built at export and served from HuggingFace: foreign_dict.json (80k phonetic respellings), foreign_parts.json (37k CamelCase parts), known_words.json (6k anti-respell shield), english_i.json (14k dotted/dotless-I decisions).

🚀 Usage

  1. Open the live site.
  2. Type text, press Speak. The first sentence plays instantly.
  3. Optionally download the result as .wav.

No installation; any modern browser works (faster with WebGPU, WASM fallback otherwise).

🛠️ Developer Notes: Automated CI/CD

Everything is automated with GitHub Actions: export-onnx.yml produces 3 ONNX models + dictionaries from safetensors, validates with ONNX checker + smoke test + normalizer/audio differential tests, uploads via push_hf.py to HuggingFace Hub (fr0stb1rd/antalia-mini-web-onnx), and pages.yml deploys to GitHub Pages.

1
2
python3 tests/make_fixtures.py   # regenerate oracle fixtures from Python
node --test tests/test_normalize.mjs tests/test_audio.mjs

One-time setup is creating the antalia-mini-web-onnx Hub repo and adding HF_TOKEN; afterwards it is one click on Run workflow.

📄 License and Credits

  • Model weights and original code: Apache-2.0 (cloud0day3/antalia-mini, PatientDesk AI — NOTICE attribution included verbatim).
  • Web app code in this repo (app.js, normalize.js, audio.js, …): Apache-2.0 © 2026 fr0stb1rd.

Inspect the code and contribute: fr0stb1rd/antalia-mini-web — sibling project: EMA Lightning Web

This post is licensed under CC BY 4.0 by the author.