
Private RAG on your own machine with turbovec
Search 177K chunks of the Spanish official gazette locally — no cloud, no API keys
- rag
- turbovec
- embeddings
- ollama
- boe
“Chat with your documents” demos usually end with a hosted vector database and an API key. This one doesn’t: the embedding model, the index and the language model all run on your own machine. As a test corpus I used something real and awkward — 2,000 documents from the BOE, Spain’s official gazette, converted to Markdown. The interesting piece is the index, turbovec, which squeezes vectors down to 4 bits per dimension and still searches faster than FAISS.
Why turbovec
- No training step. Most compressed indexes (like FAISS’s) have to learn from your data first. TurboQuant derives its compression analytically — vectors are searchable the moment they’re added.
- Tiny. The whole corpus below — 177K passages — lives in a 94 MB index. The same vectors uncompressed would take 727 MB.
- Scoped search. You can restrict a query to a folder or project at the kernel level, so permissions don’t need a second index.
- Genuinely offline. No service, no telemetry. Paired with a local embedding model, nothing leaves the laptop.
The stack
| Layer | Choice | Why |
|---|---|---|
| Documents | BOE corpus → Markdown | 2,000 laws and regulations; dense legal Spanish is a real stress test |
| Chunks | ~800-token passages | Big enough for context, small enough to stay relevant |
| Vectors | Qwen3-Embedding-0.6B via Ollama | Small, multilingual, runs on CPU |
| Index | turbovec, 4-bit | See above |
| Answers | Gemma 4 12B or Qwen3.8-27B via Ollama | The two models from the previous experiment |
How it’s built
The glue is about a hundred lines of Python and the flow is three steps:
Step 1 — Index once. Chunk every Markdown file into ~800-token passages, embed each one, and add the vectors to a 4-bit index. The sync call makes the write incremental and crash-safe — the 177K-chunk ingest took a few hours and could be resumed after interruptions.
index = turbovec.IdMapIndex(dim=1024, bit_width=4)
index.add_with_ids(vectors, chunk_ids) # vectors from the embedding model
index.sync("boe.tvec") # incremental, crash-safe
Step 2 — Search. Embed the question, get back the six closest passages.
scores, ids = index.search(question_vector, k=6)
Step 3 — Answer with citations. The chat model only sees the retrieved passages and is instructed to cite them. This is the part that keeps it honest.
answer = chat(f"Answer using only these passages, citing [n]:\n{passages}\n\nQ: {question}")
Measured on this corpus
| Value | |
|---|---|
| Corpus | 2,000 BOE documents · 445 MB of Markdown |
| Chunks indexed | 177,580 vectors (1024-d) |
| Index file on disk | 94.2 MB at 4 bits — ~7.7× smaller than the 727 MB float32 equivalent |
| Index load time | 0.08 s |
| Median search, k = 6 | 1.35 ms synthetic queries · ~27 ms end-to-end in the real battery (includes embedding the question) |
| Full answer time | median ~18 s with Gemma 4 12B on the 16 GB Mac |
Then I ran a battery of 11 real legal questions — rent duration, data-protection duties, driving-licence points — through the full pipeline. Every answer came back with citations pointing at actual BOE entries, and retrieval was a rounding error: ~25 ms to find the passages versus seconds for the model to write the answer. My favourite result is the negative test: asked “what does Spanish law say about colonising Mars?”, the system answered plainly that the passages don’t cover it — instead of inventing an article. That refusal is exactly what you want from a retrieval pipeline.
To put the numbers in context: searching half a million words of legal Spanish takes about the same time as rendering a single frame of a video game. Retrieval stopped being the bottleneck long ago — the LLM reading the passages is where the seconds go, which is exactly why pairing this with a fast local model matters.
What I took away
- Legal text is a great RAG benchmark. Dense, formulaic, cross-referenced — if retrieval works here, it works on gentler corpora.
- The hard part is never the index. Chunking well and writing a prompt that forces citations matter far more than which database you pick.
- Citations are the feature. Making the model quote
[n]passages turns hallucinations from invisible to obvious — and lets you click through to the actual law.

