Service
Continue Pretraining
Your domain knowledge added to a foundation model before it is tuned.
Talk to usSee OpenThai 2.0 Legal
OpenThai 2.0 Legal learned Thai law this way: 360,985 statute drills, about 800M tokens.
What we do
- Knowledge in the weightsThe model learns your field's terms, concepts and style from its raw text.
- Thai and other languagesA model strengthened in Thai, or in any language with enough text.
- Vocabulary expansionOptional: domain terms added to the tokenizer, so they take fewer tokens.
- A base for finetuningA stronger starting point for finetuning on your tasks.
How we work
- Build the corpus. We collect your documents, manuals and papers, then deduplicate and filter them for quality.
- Pretrain. The model trains on your corpus by next-token prediction.
- Check. We test domain knowledge and confirm the model's general abilities are kept.
What you get
- The adapted base model, ready for finetuning.
- Test results for domain knowledge and general ability.
Start
Tell us your field and how much text you have, and we reply with a scope and a quote. Email sale@iapp.co.th or call 02-124-4041.