Service
Reinforcement Learning: DPO, GRPO & PPO
A model aligned to the answers your reviewers prefer.
Talk to usSee OpenThai 2.0 Legal
OpenThai 2.0 Legal was aligned with GRPO on 8,568 graded questions, rewarded by citation F1.
What we do
- DPOTrains directly on chosen and rejected answer pairs, with no separate reward model.
- GRPOScores several answers to each prompt against one another; suited to reasoning and maths.
- PPOA reward model learned from preferences, for complex or multi-objective rewards.
How we work
- Collect preferences. Chosen and rejected pairs, or rankings of the model's answers; your data and objectives decide between DPO, GRPO and PPO.
- Align. We train toward those preferences while keeping what the model can already do.
- Evaluate and repeat. Safety and alignment tests after each round, until the target behaviour holds.
What you get
- The aligned model.
- Its safety and alignment test results.
Start
Tell us the behaviour you want and the preference data you have, and we reply with a scope and a quote. Email sale@iapp.co.th or call 02-124-4041.