Run ChindaLLM 4B with Ollama
By the end, ChindaLLM 4B answers in Thai on your own machine, in the terminal and through a local API.
You need macOS, Linux or Windows, room for a 2.5 GB download, and at least 4 GB of RAM (8 GB recommended).
1. Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
On Windows, use the installer from ollama.com/download; on any system, ollama --version confirms the install.
2. Download the model
ollama pull iapp/chinda-qwen3-4b
The download is about 2.5 GB; when it finishes, ollama list shows iapp/chinda-qwen3-4b:latest.
3. Chat in the terminal
# a chat session: type at the >>> prompt, /bye to leave
ollama run iapp/chinda-qwen3-4b
# one question, then exit
ollama run iapp/chinda-qwen3-4b "ช่วยแก้สมการ 2x + 5 = 15 ให้หน่อย"
4. Call it from your code
The local API listens on http://localhost:11434; if it is not running, start it with ollama serve.
- curl
- Python
curl http://localhost:11434/api/chat -d '{
"model": "iapp/chinda-qwen3-4b",
"messages": [
{"role": "user", "content": "อธิบายเกี่ยวกับปัญญาประดิษฐ์ให้ฟังหน่อย"}
]
}'
import requests
def chat_with_chinda(message):
url = "http://localhost:11434/api/generate"
data = {
"model": "iapp/chinda-qwen3-4b",
"prompt": message,
"stream": False
}
response = requests.post(url, json=data)
return response.json()["response"]
print(chat_with_chinda("สวัสดีครับ"))
Troubleshooting
Error: pull model manifest: file does not exist: the model name is mistyped; it isiapp/chinda-qwen3-4b.Error: could not connect to ollama app, is it running?: start the server withollama serve.- Slow, or out of memory: close other applications; 8 GB of RAM is recommended.
Chinda 4B answers best with the relevant text in the prompt; asked for facts on their own, such as recent news or statistics, it can give wrong answers.
Benchmarks and licence: ChindaLLM 4B · On Ollama: iapp/chinda-qwen3-4b · Weights: Hugging Face