Skip to main content

Service

Continue Pretraining

Your domain knowledge added to a foundation model before it is tuned.

Talk to usSee OpenThai 2.0 Legal

OpenThai 2.0 Legal learned Thai law this way: 360,985 statute drills, about 800M tokens.

What we do​

  • Knowledge in the weightsThe model learns your field's terms, concepts and style from its raw text.
  • Thai and other languagesA model strengthened in Thai, or in any language with enough text.
  • Vocabulary expansionOptional: domain terms added to the tokenizer, so they take fewer tokens.
  • A base for finetuningA stronger starting point for finetuning on your tasks.

How we work​

  1. Build the corpus. We collect your documents, manuals and papers, then deduplicate and filter them for quality.
  2. Pretrain. The model trains on your corpus by next-token prediction.
  3. Check. We test domain knowledge and confirm the model's general abilities are kept.

What you get​

  • The adapted base model, ready for finetuning.
  • Test results for domain knowledge and general ability.

Start​

Tell us your field and how much text you have, and we reply with a scope and a quote. Email sale@iapp.co.th or call 02-124-4041.