AHG
Abstract cover: two chips side by side — one small, one large — representing 16 GB and 32 GB memory tiers
Local AI Lab

Local LLMs: what fits in 16 GB vs 32 GB+

Gemma 4 12B on a laptop against Qwen3.8-27B on an RTX 3090 — measured, not guessed

  • gemma-4
  • qwen3.8
  • lm-studio
  • apple-silicon
  • rtx-3090

Also in: EspañolFrançais

The default answer to “I want to try a document” is to paste it into a cloud chat. My default is different: I run open models on my own machines. Two of them do almost all the work — Gemma 4 12B, which fits in a 16 GB laptop, and Qwen3.8-27B, which needs a 24 GB GPU. Both summarise documents, write in Spanish and English, and code well enough that I stopped reaching for APIs. The question I wanted answered with numbers, not vibes: what does the bigger model actually buy you?

The two models

Gemma 4 12B Qwen3.8-27B
Made by Google DeepMind Alibaba
Size 12B parameters · 7.6 GB download 27B · 18 GB
Understands Text, images, audio Text, images, video
Memory 256K tokens of context 256K, stretchable to 1M
License Gemma Terms of Use Apache-2.0 (fully open)
Runs on 16 GB laptop RTX 3090 or 32 GB+ Mac
More info Model card · ollama gemma4 Qwen blog · ollama qwen3.8

In short: Gemma is the model you carry; Qwen is the model you visit. A 16 GB machine tops out around 12B parameters in good quality — beyond that the weights alone stop fitting. A 24 GB card is where the “serious” tier starts.

The test

Nine tasks, three runs each, same prompts on both sides:

  • Read a document — a ~20-page technical paper: write an executive summary, extract structured data as JSON, and answer questions quoting the exact passages they came from.
  • Write and translate — a 600-word Spanish post from an English brief, a Slack thread turned into a formal status update, and a translation where ten glossary terms had to be honoured.
  • Code — write a Python function plus tests that must actually pass, refactor a TypeScript module without breaking it, and diagnose a stack trace.

Everything is timed (tokens per second, wall clock) and every claim is checked automatically where possible: the JSON must parse, the tests must pass, the glossary terms must appear. The runner is a small Python script that works against Ollama or LM Studio, so the numbers are reproducible on your own hardware.

Results

TaskGemma 4 12B (QAT)Mac · Apple Silicon 16 GBQwen3.8-27BPC · NVIDIA RTX 3090 24 GB
Python function + pytest8.7 tok/s✓16.9 tok/s✓
TypeScript refactor (behaviour preserved)9 tok/s✓14.8 tok/s✓
Explain a stack trace and propose a fix9 tok/s15.3 tok/s
Technical post ES from EN brief9.1 tok/s✗18.7 tok/s✓
Informal thread → formal status update9.2 tok/s14.5 tok/s
EN→ES translation with glossary8.9 tok/s✓18.3 tok/s✓
Executive summary (200 words)10.8 tok/s16.9 tok/s
Structured extraction to JSON8.5 tok/s✓18.3 tok/s✗
Grounded Q&A with citations8.9 tok/s17.6 tok/s
Last updated: 2026-09-30✓ passed the automatic check · ✗ failed it · — no hard check for this task

What surprised me

  • Speed is not close — and not in the direction I expected. The 27B model on the 3090 generates at 15–19 tok/s; the 12B on the laptop does 8.5–11. Twice the size, nearly twice the speed: GPU memory bandwidth beats the size penalty. Wall-clock tells the same story — Qwen rewrote a Slack thread in 30 s where Gemma took 95.
  • Code is a tie. Both wrote a correct function plus a test suite that passed on the first try (8 and 7 tests respectively), and both refactored the TypeScript module cleanly enough to satisfy tsc --strict. The extra 15B parameters bought speed, not correctness.
  • Reading documents produced the best and the worst moment. Both answered questions with accurate, quoted passages — no hallucinated citations. But Qwen burned ~5,900 tokens on the JSON task and returned empty output on all three runs: its answer went to the hidden reasoning channel instead of the reply. A serving quirk, not a model defect — and a useful lesson if you run thinking models behind OpenAI-compatible endpoints.
  • Instructions: the big model is stricter. Both honoured the translation glossary. But on the Spanish post, Qwen followed every formatting constraint while Gemma — clean prose, good content — quietly swapped the requested bullet list for numbered headings. Small miss, but the kind that breaks pipelines.

Verdict. If your machine has 16 GB, Gemma 4 12B is genuinely capable — nothing here was beyond it, and it sips memory. If you can give the task 24 GB, Qwen3.8-27B is the better daily driver: roughly double the throughput and at least as obedient. I keep both.

Related experiences

← Back to experiences