OpenThai 2.0 Legal
Chat completions on OpenThai 2.0 Legal: answers grounded in the 2,065 Thai laws of its retrieval index, returned with the sections they rest on and checked before they reach you; OpenAI-compatible.
/v3/llm/openthai2p0-legal/chat/completionsTry it- Price0.01 / 0.02 ICper 1k tokens, in / out
- Served42.1Kcalls
- v2.0
- Active
Input
- model
- openthai2.0-legal
- messages
- role: user content: ลักทรัพย์ในเวลากลางค ืน ผิดมาตราใด
- max_tokens
- 1024
OutputPOST /v3/llm/openthai2p0-legal/chat/completions
- id
- chatcmpl-...
- model
- openthai2.0-legal-thaillm-nemotron-3-nano-30b-a3b
- choices
- message: { role, content }
- usage
- prompt_tokens: 2874 completion_tokens: 17
- rag
- true
- retrieved_documents
- 2 × law: ประมวลกฎหมายอาญา section: 335 text: ผู้ใดลักทรัพย์ (๑) ในเวลากลางคืน ... score: 0.9989
- answerable
- true
- answerable_reason
- ok
- content_modified
- —
- citations
- total: 1 grounded: 1 ungrounded: []
Request
JSON, in the OpenAI chat-completions shape; the OpenAI SDK works as-is with the base URL https://api.iapp.co.th/v3/llm/openthai2p0-legal and the key as apikey or Authorization: Bearer. This API's own fields control the retrieval and the checks around the answer.
Headers
apikeystringrequiredBody
modelstringrequiredopenthai2.0-legal.messagesarrayrequiredrole and content. Send the whole history, the assistant turns included; a short follow-up is searched together with the question before it.ragbooleantrue: retrieve the law sections first and ground the answer in them. false asks the bare model, closed-book.rag_top_kintegerrag_inject is system.rag_injectstringuser, the default, puts the sections in the scaffold the model was trained on, best for JSON citation answers; system puts them in the system prompt as advisory reference, best for essays and long analysis.dekabooleanfalse. true also retrieves Supreme Court decisions, 133,000 of them, and adds the closest as analogous rulings; deka_top_k sets how many. For long analysis; leave it off for short citation answers.webstringauto, the default, searches only when the question names a real case or person or reads like news; true always; false never.unanswerable_modestringreplace, the default, returns a standard refusal instead; append keeps the answer and adds a warning; flag leaves it untouched and only sets the verdict fields.guardbooleantrue: check the question before the model runs, for laws, sections or cases that do not exist, fiction, and questions that are not legal. false turns the check off; not recommended.hold_until_verdictbooleanstream, withhold the answer until it has been checked and send it as one chunk, for a UI that cannot swap a replaced answer.chat_template_kwargsobject{"enable_thinking": true} for step-by-step legal reasoning, returned apart from the answer in message.reasoning_content (delta.reasoning_content when streaming).temperaturefloatrepetition_penaltyfloatmax_tokensintegerstreambooleantrue for server-sent events, token by token. The retrieved sections ride on the first chunk, the verdict fields on the last.Price
A call costs 0.01 IC per 1K input tokens and 0.02 IC per 1K output tokens.
Settings for chat
{"temperature": 0.6, "top_p": 0.95, "repetition_penalty": 1.05, "min_p": 0.05, "max_tokens": 2048,
"rag_inject": "system", "deka": true, "deka_top_k": 2, "web": "auto", "unanswerable_mode": "replace"}
Over 134 runs of a question set chosen to provoke loops, these settings produced no repetition loop, with the same grounding: 99% of cited sections were found in the retrieved text. The former default, temperature 0.7 and top_p 0.9 without a penalty, looped to the 4,096-token cap in 1 to 2% of answers. A loop that still forms is cut where it starts, and the response says content_modified: "loop_cut". For citation answers in JSON, use temperature 0.0 with thinking off and rag_inject user; for an essay or analysis, temperature 0.0 with thinking on and rag_inject system.
Response
The OpenAI chat-completion object, as in the panel, plus this API's own fields: what was retrieved, and what the checks decided.
choices[0]choices[0].messagechoices[0].message.contentretrieved_documentsarrayThe sections the answer is grounded in, each with law, section, text and score
retrieved_documents.lawstringretrieved_documents.sectionstringretrieved_documents.textstringretrieved_documents.scorenumberanswerablebooleanfalse when the answer was refused or rewrittenanswerable_reasonstringok, or why not: no_relevant_law_found, ungrounded_citations, model_declined, empty_answer, unknown_law, unknown_section, unknown_case, case_not_found, fictional_premise, not_legal_questioncontent_modifiedstringnull when the answer stands; replaced (a refusal instead of the answer), pruned (ungrounded lines removed), appended (a warning added) or loop_cut. The text before the change is in original_content.replacement_contentstringcontent_modified is replaced or prunedcitationsobjecttotal, grounded, the ungrounded sections, and lookup: sections the model cited from memory and how they scored against the statute textretrieval_confidencenumberguardobjectlabel (LEGAL, CASE, FICTION or OTHER), refused, reason, and glossary, the terms of art it resolvedwebobjectused), and its results, each with title and url; show them as sourcestruncationobjectprompt_tokens_before, prompt_tokens_after, removed_chars, limitragbooleanusageobjectidstringmodelstringchoicesarrayAs in the OpenAI format
choices.messageobjectrole and contentFor a machine-readable answer, use the system prompt the model was trained on, published with its JSON contract on the model page; the model then cites only sections present in the retrieved context.
Streaming
The draft streams token by token. When the check replaces or prunes it, in 1 to 3% of answers, one more content chunk follows with a separator, \n\n---\n, and the replacement; the last chunk, with choices: [], carries usage, the verdict fields and replacement_content. Nothing arrives after data: [DONE]. Show replacement_content in place of the draft when it is set:
- Python
- JavaScript
draft, replacement = "", None
with client.chat.completions.create(model="openthai2.0-legal", messages=msgs, max_tokens=2048, stream=True,
extra_body={"rag_inject": "system", "deka": True, "web": "auto"}) as stream:
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
draft += chunk.choices[0].delta.content
ui.set_answer(draft) # live
extra = chunk.model_extra or {}
if extra.get("replacement_content") is not None:
replacement = extra["replacement_content"] # last chunk: the check replaced or pruned the draft
ui.set_answer(replacement if replacement is not None else draft)
let draft = "", replacement = null;
const res = await fetch(url, { method: "POST", headers, body: JSON.stringify({ ...body, stream: true }) });
const reader = res.body.getReader(), dec = new TextDecoder(); let buf = "";
while (true) {
const { value, done } = await reader.read(); if (done) break;
buf += dec.decode(value, { stream: true });
let i; while ((i = buf.indexOf("\n\n")) >= 0) {
const line = buf.slice(0, i).trim(); buf = buf.slice(i + 2);
if (!line.startsWith("data: ") || line === "data: [DONE]") continue;
const o = JSON.parse(line.slice(6));
for (const c of o.choices || []) if (c.delta?.content) { draft += c.delta.content; ui.setAnswer(draft); }
if (o.replacement_content != null) replacement = o.replacement_content;
}
}
ui.setAnswer(replacement ?? draft);
A UI that cannot swap sends "hold_until_verdict": true; the answer then arrives as one chunk once it has been checked.
Details
Measured
Checks around the answer
- Before the model, the question is checked against the law corpus and the Supreme Court decisions. A law, section or case number that does not exist, a fictional premise, or a question that is not legal gets a standard Thai refusal at once, and no tokens are generated. A greeting gets a short self-introduction, and a greeting followed by a question is answered as a question.
- After the model, every มาตรา in the answer is matched against the retrieved sections; a section cited from memory is looked up and the sentence citing it is checked against the statute text. What cannot be grounded is removed, and an answer left with no grounded citation is replaced by the refusal.
- For a real current case, the API reads the top news reports, takes the charges they name, retrieves the statutes for those charges and answers from them. A named case that no news source reports is refused with
case_not_found. This adds 1 to 3 seconds; searches are cached for an hour. Exam-style hypotheticals are answered from the statutes, not searched.
On a set of 107 questions, covering laws, sections and cases that do not exist, fiction, off-topic questions, typos, current cases and real exam questions, the API behaves correctly on 105, with thinking on or off; before the checks, it did on 75 of 105. Legal terms of art that no statute spells out, such as คดีอุทลุม, ครอบครองปรปักษ์ or ทางจำเป็น, are resolved through a glossary to the sections that define them.
Long prompts
A prompt may run to 262,144 tokens, about 350,000 to 450,000 Thai characters; 131,072 when the backup engine answers. A longer prompt is not rejected: the middle of the longest user message is cut so that the prompt and max_tokens fit, a Thai marker shows where, and the response carries truncation. Accuracy is not even across the window. In a test with real statute text, the model found both planted facts at 73K and 92K tokens, one of two at 115K to 140K, and neither from 165K, so keep text that must be answered precisely under about 90K tokens, 150,000 Thai characters. A 250K-token prompt takes 30 to 35 seconds to read. Request bodies up to 64 MB are accepted; above 500 KB they take a relay path that adds a few hundred milliseconds.
Health
GET https://api.iapp.co.th/v3/llm/openthai2p0-legal/health needs no key and costs nothing: it answers 200 while the API can answer and 503 when it cannot, with status ok, degraded (answering from the backup only) or down. The status page shows 90 days of uptime and recent incidents.
Limits
- 30 requests per minute per IP address.
- Decision support, not legal advice: a qualified professional must check an answer against the current law before anyone relies on it.
- An answer is only as good as the sections retrieved;
retrieved_documentslets you audit every answer.
Data handling
Prompts and completions are processed in memory and not kept after the response. The service is GDPR and PDPA compliant. The weights are on Hugging Face and run on a single 24 GB GPU; a tutorial sets up the same retrieval on your own machines.