Language model and document search on a Raspberry Pi
A Raspberry Pi answers questions about company documents inside the local network, and no passage reaches anyone who may not read it.
TippelPi is a Raspberry Pi 5 with 16 GB. It answers questions about company documents without a single byte leaving the building. The Qwen3 language model, the embedding and the index all run on the device; a browser on the local network is the interface. The same application runs unchanged on a GPU server and detects by itself which model backend is reachable.
Search combines vector similarity with full-text search and carries access rights down to the individual passage: a subfolder of the document share becomes a group, and anyone outside it never receives its passages as a source. The screenshot shows the interface in the browser with two questions to two invented test documents: one from maintenance to the fault manual of a press, one from a brewery to the brew log of its lager. Each answer gives the value from the document, and a click on the source opens the cited passage. The run dates from 10 October 2026 and used Qwen3-8B on a notebook, not on the Pi; the corpus and the questions are German. The application mounts the document share read-only, and every access lands in the audit log.
- Tests204, all passing
- Answers82 % correct, 90 % fully grounded · 127 questions
- DeviceRaspberry Pi 5 · 16 GB
- ModelQwen3-8B · Ollama or vLLM
- Searchmultilingual-e5-small · ChromaDB + SQLite FTS5
- Test corpus51 public documents · 3,446 passages
- As of10 October 2026
Answer quality was measured with Qwen3-8B on a GPU server, not on the Pi, and 18 % of the answers were not fully correct. Response times on the Pi have not been measured yet. The first embedding pass runs there at about 6.4 passages per second, so 10,000 documents take roughly four hours once.
$ python -m pytest tests/ -q 204 passed, 2 warnings in 14.52s
| Queries | 1,200 — the restricted passage itself as the question |
|---|---|
| Group pairs | 6 — three groups, 200 passages each, asked by the other two |
| Leaked passages | 0 |
| Control without filter | 600 of 600 — the source document is found |
Both search channels are covered, vector and full text. The test covers retrieval, not the model’s answer.
- Interfacebrowser on the local network, streamed answers with sources
- ApplicationFastAPI: sign-in, roles, groups, audit log, backup
- Searchon-device embedding, vector index and full text, fused by rank
- ModelQwen3 via Ollama on the Pi or via vLLM on a GPU
- Sourcesfolder, SMB share or WebDAV, read-only
More projects