AHG
Abstract cover: documents compressed into a grid of small coloured cells
Local AI Lab

Private RAG on your own machine with turbovec

Search 177K chunks of the Spanish official gazette locally — no cloud, no API keys

  • rag
  • turbovec
  • embeddings
  • ollama
  • boe

Also in: EspañolFrançais

“Chat with your documents” demos usually end with a hosted vector database and an API key. This one doesn’t: the embedding model, the index and the language model all run on your own machine. As a test corpus I used something real and awkward — 2,000 documents from the BOE, Spain’s official gazette, converted to Markdown. The interesting piece is the index, turbovec, which squeezes vectors down to 4 bits per dimension and still searches faster than FAISS.

Why turbovec

  • No training step. Most compressed indexes (like FAISS’s) have to learn from your data first. TurboQuant derives its compression analytically — vectors are searchable the moment they’re added.
  • Tiny. The whole corpus below — 177K passages — lives in a 94 MB index. The same vectors uncompressed would take 727 MB.
  • Scoped search. You can restrict a query to a folder or project at the kernel level, so permissions don’t need a second index.
  • Genuinely offline. No service, no telemetry. Paired with a local embedding model, nothing leaves the laptop.

The stack

Layer Choice Why
Documents BOE corpus → Markdown 2,000 laws and regulations; dense legal Spanish is a real stress test
Chunks ~800-token passages Big enough for context, small enough to stay relevant
Vectors Qwen3-Embedding-0.6B via Ollama Small, multilingual, runs on CPU
Index turbovec, 4-bit See above
Answers Gemma 4 12B or Qwen3.8-27B via Ollama The two models from the previous experiment

How it’s built

The glue is about a hundred lines of Python and the flow is three steps:

Step 1 — Index once. Chunk every Markdown file into ~800-token passages, embed each one, and add the vectors to a 4-bit index. The sync call makes the write incremental and crash-safe — the 177K-chunk ingest took a few hours and could be resumed after interruptions.

index = turbovec.IdMapIndex(dim=1024, bit_width=4)
index.add_with_ids(vectors, chunk_ids)   # vectors from the embedding model
index.sync("boe.tvec")                   # incremental, crash-safe

Step 2 — Search. Embed the question, get back the six closest passages.

scores, ids = index.search(question_vector, k=6)

Step 3 — Answer with citations. The chat model only sees the retrieved passages and is instructed to cite them. This is the part that keeps it honest.

answer = chat(f"Answer using only these passages, citing [n]:\n{passages}\n\nQ: {question}")

Measured on this corpus

Value
Corpus 2,000 BOE documents · 445 MB of Markdown
Chunks indexed 177,580 vectors (1024-d)
Index file on disk 94.2 MB at 4 bits — ~7.7× smaller than the 727 MB float32 equivalent
Index load time 0.08 s
Median search, k = 6 1.35 ms synthetic queries · ~27 ms end-to-end in the real battery (includes embedding the question)
Full answer time median ~18 s with Gemma 4 12B on the 16 GB Mac

Then I ran a battery of 11 real legal questions — rent duration, data-protection duties, driving-licence points — through the full pipeline. Every answer came back with citations pointing at actual BOE entries, and retrieval was a rounding error: ~25 ms to find the passages versus seconds for the model to write the answer. My favourite result is the negative test: asked “what does Spanish law say about colonising Mars?”, the system answered plainly that the passages don’t cover it — instead of inventing an article. That refusal is exactly what you want from a retrieval pipeline.

To put the numbers in context: searching half a million words of legal Spanish takes about the same time as rendering a single frame of a video game. Retrieval stopped being the bottleneck long ago — the LLM reading the passages is where the seconds go, which is exactly why pairing this with a fast local model matters.

What I took away

  • Legal text is a great RAG benchmark. Dense, formulaic, cross-referenced — if retrieval works here, it works on gentler corpora.
  • The hard part is never the index. Chunking well and writing a prompt that forces citations matter far more than which database you pick.
  • Citations are the feature. Making the model quote [n] passages turns hallucinations from invisible to obvious — and lets you click through to the actual law.

Related experiences

← Back to experiences