
Local LLMs: what fits in 16 GB vs 32 GB+
Gemma 4 12B on a laptop against Qwen3.8-27B on an RTX 3090 — measured, not guessed
- gemma-4
- qwen3.8
- lm-studio
- apple-silicon
- rtx-3090
The default answer to “I want to try a document” is to paste it into a cloud chat. My default is different: I run open models on my own machines. Two of them do almost all the work — Gemma 4 12B, which fits in a 16 GB laptop, and Qwen3.8-27B, which needs a 24 GB GPU. Both summarise documents, write in Spanish and English, and code well enough that I stopped reaching for APIs. The question I wanted answered with numbers, not vibes: what does the bigger model actually buy you?
The two models
| Gemma 4 12B | Qwen3.8-27B | |
|---|---|---|
| Made by | Google DeepMind | Alibaba |
| Size | 12B parameters · 7.6 GB download | 27B · 18 GB |
| Understands | Text, images, audio | Text, images, video |
| Memory | 256K tokens of context | 256K, stretchable to 1M |
| License | Gemma Terms of Use | Apache-2.0 (fully open) |
| Runs on | 16 GB laptop | RTX 3090 or 32 GB+ Mac |
| More info | Model card · ollama gemma4 |
Qwen blog · ollama qwen3.8 |
In short: Gemma is the model you carry; Qwen is the model you visit. A 16 GB machine tops out around 12B parameters in good quality — beyond that the weights alone stop fitting. A 24 GB card is where the “serious” tier starts.
The test
Nine tasks, three runs each, same prompts on both sides:
- Read a document — a ~20-page technical paper: write an executive summary, extract structured data as JSON, and answer questions quoting the exact passages they came from.
- Write and translate — a 600-word Spanish post from an English brief, a Slack thread turned into a formal status update, and a translation where ten glossary terms had to be honoured.
- Code — write a Python function plus tests that must actually pass, refactor a TypeScript module without breaking it, and diagnose a stack trace.
Everything is timed (tokens per second, wall clock) and every claim is checked automatically where possible: the JSON must parse, the tests must pass, the glossary terms must appear. The runner is a small Python script that works against Ollama or LM Studio, so the numbers are reproducible on your own hardware.
Results
| Task | Gemma 4 12B (QAT)Mac · Apple Silicon 16 GB | Qwen3.8-27BPC · NVIDIA RTX 3090 24 GB |
|---|---|---|
| Python function + pytest | 8.7 tok/s✓ | 16.9 tok/s✓ |
| TypeScript refactor (behaviour preserved) | 9 tok/s✓ | 14.8 tok/s✓ |
| Explain a stack trace and propose a fix | 9 tok/s | 15.3 tok/s |
| Technical post ES from EN brief | 9.1 tok/s✗ | 18.7 tok/s✓ |
| Informal thread → formal status update | 9.2 tok/s | 14.5 tok/s |
| EN→ES translation with glossary | 8.9 tok/s✓ | 18.3 tok/s✓ |
| Executive summary (200 words) | 10.8 tok/s | 16.9 tok/s |
| Structured extraction to JSON | 8.5 tok/s✓ | 18.3 tok/s✗ |
| Grounded Q&A with citations | 8.9 tok/s | 17.6 tok/s |
What surprised me
- Speed is not close — and not in the direction I expected. The 27B model on the 3090 generates at 15–19 tok/s; the 12B on the laptop does 8.5–11. Twice the size, nearly twice the speed: GPU memory bandwidth beats the size penalty. Wall-clock tells the same story — Qwen rewrote a Slack thread in 30 s where Gemma took 95.
- Code is a tie. Both wrote a correct function plus a test suite that passed on the first try (8 and 7 tests respectively), and both refactored the TypeScript module cleanly enough to satisfy
tsc --strict. The extra 15B parameters bought speed, not correctness. - Reading documents produced the best and the worst moment. Both answered questions with accurate, quoted passages — no hallucinated citations. But Qwen burned ~5,900 tokens on the JSON task and returned empty output on all three runs: its answer went to the hidden reasoning channel instead of the reply. A serving quirk, not a model defect — and a useful lesson if you run thinking models behind OpenAI-compatible endpoints.
- Instructions: the big model is stricter. Both honoured the translation glossary. But on the Spanish post, Qwen followed every formatting constraint while Gemma — clean prose, good content — quietly swapped the requested bullet list for numbered headings. Small miss, but the kind that breaks pipelines.
Verdict. If your machine has 16 GB, Gemma 4 12B is genuinely capable — nothing here was beyond it, and it sips memory. If you can give the task 24 GB, Qwen3.8-27B is the better daily driver: roughly double the throughput and at least as obedient. I keep both.

