API
ChindaTTS
Thai and English speech, switching language mid-sentence.
- Voices2
- Speaking styles8
- First audio< 1s
- Per GPU6–7×real time
Drop-in compatible with the iApp TTS v3 API. Managed, or on your own servers.
What it does
- Voices and stylesKaitom or Kaimook, and a style per request: neutral, friendly, cheerful, calm, serious, sad, excited or empathetic.
- Thai and English mixedSwitches within a sentence; a sentence wholly in English is read in a native-English mode.
- Numbers read correctlyDates, currency, percentages, phone and ID numbers, abbreviations, emails and URLs, as a person says them.
- Voice cloningFrom about 10 to 20 seconds of consented audio; the clip is used once and not stored.
Hear it
Unedited samples. Every speaking style and live voice cloning can be tried in the API demo; a new account comes with 50 free credits.
Kaitom
Kaimook
Numbers, dates and money
Thai and English together
Measured
Character error rate when the audio is transcribed back; lower is better. Prime is an optional check that re-rolls a garbled take rather than altering the audio.
| Text | ChindaTTS | With Prime |
|---|---|---|
| Everyday | 3.2% | 3.2% |
| Numbers, dates, currency | 1.8% | 1.4% |
| Thai and English mixed | 7.9% | 7.6% |
| Expressive styles | 8.3% | 3.9% |
| Bad-take rate, expressive styles | 15.0% | 6.7% |
Prime takes about 9% longer on easy text and needs 2 GB more GPU memory. One GPU turns 1,000 characters into 90 seconds of audio in 14 seconds.
Limits
- Thai and Latin script only; other scripts are dropped.
- Up to 1,200 characters, about 100 seconds of audio, per request.
- Tone is set by the style, rate by the
speedparameter. - One stream per GPU; concurrency scales by adding GPUs.
On your own servers
One current-generation GPU with 8 GB (10 GB with Prime; 12 to 16 GB recommended), 16 GB RAM and Linux. For a pilot, write to sale@iapp.co.th.