Open weights
OpenThai 2.0
One open model that reads Thai documents and handwriting, knows Thailand and calls tools.
- Parameters27billion
- LicenceApache 2.0
- OpenThaiEval0.842
- Context256Ktokens
Download the weightsCall the hosted API
Released 27 August 2026. The hosted API costs 0.01 / 0.02 IC per 1K input / output tokens.
What it does
- Reads Thai documentsHandwriting at 0.261 character error rate against the base model's 0.649; books and Royal Gazette pages at 0.126.
- Answers about documentsReads a document and answers questions on it in one model, with no separate OCR stage.
- Calls tools0.820 on the Berkeley Function-Calling Leaderboard, ahead of its base model and Typhoon 2.5.
- Decodes fasterA bundled draft head gives 75.2 against 50.1 tokens/s on one H100 stream, with identical output.
Measured
All models were measured by iApp on one suite and one serving setup, each with its own recommended prompts and parameters. Test sets were held out from training by an n-gram leakage check, and the handwriting test shares no text with training. Per-category results are in the Hugging Face repository.
- Thai knowledge
- Document reading
- Tool use

Typhoon 2.5 (30B-A3B), measured the same way, scores 0.742 on OpenThaiEval and 0.749 on IFEval-TH, against 0.842 and 0.795 for OpenThai 2.0.


Run it
- Hosted API
- curl
- Ollama
- vLLM
Call it with an iApp API key: register, then API Keys, Create New API Key. 0.01 / 0.02 IC per 1K input / output tokens, 30 requests a minute per key. Full reference: OpenThai 2.0 API.
import base64
from openai import OpenAI
client = OpenAI(base_url="https://api.iapp.co.th/v3/llm/openthai2p0",
api_key="YOUR_IAPP_API_KEY")
# Text — ask anything in Thai
r = client.chat.completions.create(
model="openthai2.0",
messages=[{"role": "user", "content": "อากรแสตมป์กับภาษีมูลค่าเพิ่มต่างกันอย่างไร"}],
)
print(r.choices[0].message.content)
# Vision — read a Thai document
img = base64.b64encode(open("thai_document.jpg", "rb").read()).decode()
r = client.chat.completions.create(
model="openthai2.0",
messages=[{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{img}"}},
{"type": "text", "text": "อ่านข้อความในเอกสารนี้ทั้งหมด"},
]}],
temperature=0.0,
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
print(r.choices[0].message.content)
curl -s https://api.iapp.co.th/v3/llm/openthai2p0/chat/completions \
-H "Content-Type: application/json" -H "apikey: YOUR_IAPP_API_KEY" \
-d '{"model":"openthai2.0","messages":[{"role":"user","content":"สวัสดี"}]}'
One command on a laptop, from Ollama.
ollama run openthai/openthai2.0-qwen3.8-27b
One 80 GB GPU, about 56 GB in bf16, on vLLM 0.19 or later. --max-num-seqs 128 is required: vLLM's default of 1,024 stops the engine at startup.
vllm serve iapp/openthai2.0-qwen3.8-27b \
--max-model-len 32768 --gpu-memory-utilization 0.85 \
--max-num-seqs 128 --reasoning-parser qwen3 --trust-remote-code
The server is OpenAI-compatible at http://localhost:8000/v1 with the model iapp/openthai2.0-qwen3.8-27b. Add --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":2}' to use the draft head.
Other formats on Hugging Face
| Format | Runs on |
|---|---|
| GGUF Q4_K_M / Q8_0 | llama.cpp on a CPU or consumer GPU, 17 / 29 GB, with the vision mmproj |
| MLX 4-bit | Apple silicon with 24 GB or more of unified memory |
| INT8 W8A8 | vLLM on ~40 GB-class GPUs |
| NVFP4 | vLLM on NVIDIA Blackwell |
Settings that matter
| Setting | Why |
|---|---|
max_tokens unset, or 8,192 or more | The model reasons before it answers; at a 1,024 cap about a third of replies come back empty. |
enable_thinking: false for images only | Keep thinking on for text: forced off, usable text answers fall from 94% to 42%. |
repetition_penalty: 1.05 | For long Thai output. The hosted API applies it already. |
Limits
- Structured extraction is at parity with the base model: 0.568 against 0.570 on judged ThaiOCRBench. It trails specialist models on tables and key fields; test form pipelines on your own documents.
- Scene text in photos is its weakest reading: 0.819 character error rate against the base model's 0.737.
- A transcription-only model reads clean print and isolated handwriting better: Typhoon-OCR 1.5 scores 0.014 and 0.054, but takes only a fixed OCR prompt and answers no questions.
- Answers are explanatory by default and terse on request about 40% of the time. For machine-read output, ask for a format, such as "ตอบเป็น JSON เท่านั้น".
- On hard handwriting about one character in four is still wrong, and illegible input can produce invented text. Not evaluated on Thai dialects or on vertical or rotated text.
For tax, legal or medical answers, pair it with retrieval over authoritative sources and human review. For law, use OpenThai 2.0 Legal.
Licence
Apache 2.0. Fine-tuned from Qwen/Qwen3.8-27B with LoRA (rank 64) on about 143,000 verified rows; the adapter alone (7 GB) is in the repository for serving on top of the base model.
Built by iApp Technology with the Artificial Intelligence Entrepreneur Association of Thailand. Trained and evaluated on 8 NVIDIA H100 GPUs provided by Siam AI Corporation.
Cite
@misc{openthai2_2026,
title = {OpenThai 2.0: An Open Thai Knowledge and Document AI},
author = {iApp Technology and Artificial Intelligence Entrepreneur Association of Thailand},
year = {2026},
url = {https://openthai.aieat.or.th}
}