AHG
Local AI Lab

Local image & video generation with FLUX and Wan2.1

From DALL·E prompts in 2023 to a fully offline pipeline on an RTX 3090

  • flux
  • wan2.1
  • comfyui
  • diffusion
  • rtx-3090

Also in: EspañolFrançais

The clip at the top of this page never touched a cloud service. The still was generated with FLUX and then animated with Wan2.1 image-to-video, both running in ComfyUI on a single NVIDIA RTX 3090 (24 GB). Three years ago the same idea meant typing prompts into DALL·E and waiting for a server somewhere else to answer. This page is the story of that shift, and a practical map of what runs on consumer hardware today.

Where it started: DALL·E, 2023

In 2023 I published a small gallery here to introduce text-to-image generation to colleagues. The images came from DALL·E 2 through OpenAI’s web interface: no control over the model, no seeds, no way to fine-tune, and every image left the machine. It was magic, and it was also a black box.

DALL·E 2 generation, 2023
DALL·E 2 · 2023
DALL·E 2 generation, 2023
DALL·E 2 · 2023
DALL·E 2 generation, 2023
DALL·E 2 · 2023
DALL·E 2 generation, 2023
DALL·E 2 · 2023
DALL·E 2 generation, 2023
DALL·E 2 · 2023
DALL·E 2 generation, 2023
DALL·E 2 · 2023

Where it is now: everything local

The pipeline behind the hero clip has three stages, all inside ComfyUI:

  1. Image — FLUX.1 [dev], fp8. 12 billion parameters. In bf16 the transformer alone is ~24 GB, which does not leave room for the T5 text encoder on a 3090, so I run the fp8 checkpoint (~12 GB) with the fp8 T5-XXL encoder. 20–28 steps at 1280×720, Euler sampler, guidance 3.5. The look you see (natural skin, believable neon bloom, coherent Japanese-style signage) is what made FLUX the most used open image model.
  2. Video — Wan2.1 I2V-14B 720p, fp8. The still is fed to the WanImageToVideo node together with a short motion prompt (“slow push-in, rain falling, he turns his head slightly”). The 14B model in fp8 plus the UMT5-XXL encoder in fp8 fit in 24 GB with the text encoder offloaded to CPU. 49 frames at 16 fps → about three seconds of motion, later trimmed to the loop you see.
  3. Post — interpolation and encode. RIFE 2× to 24/30 fps, then H.264 and VP9 exports for the web.

Open-weight image models you can run at home

My working shortlist of open-weight text-to-image models, sortable by any column. Each model name links to its card, where the licence and usage terms live. “VRAM” is a practical figure for a single GPU; with GGUF weights and offload most of them go lower, at the cost of speed.

Stable Diffusion 1.5Runway / Stability AI2022-10≈0.9B4 GBHuge LoRA / ControlNet ecosystem, very fast
SDXL 1.0Stability AI2023-073.5B (base)8 GBMature fine-tunes, good composition
Stable Diffusion 3.5 MediumStability AI2024-102.5B10 GBBalanced quality on mid-range GPUs
Stable Diffusion 3.5 LargeStability AI2024-108B18 GB (bf16) / ~12 GB fp8Prompt adherence, typography
FLUX.1 [schnell]Black Forest Labs2024-0812B24 GB (bf16) / ~12 GB fp81–4 steps, fully open license
FLUX.1 [dev]Black Forest Labs2024-0812B24 GB (bf16) / ~12 GB fp8Photorealism, most popular open model
FLUX.1 Kontext [dev]Black Forest Labs2025-0612B24 GB (bf16) / ~12 GB fp8Instruction-based image editing
FLUX.1 Krea [dev]Black Forest Labs × Krea2025-0712B24 GB (bf16) / ~12 GB fp8Natural look, avoids the "AI sheen"
Qwen-ImageAlibaba Tongyi2025-0820B20 GB fp8 / ~8 GB GGUF Q4Best-in-class text rendering, editing
Z-Image-TurboAlibaba Tongyi2025-116B16 GB8 steps, photoreal, bilingual text
FLUX.2 [dev]Black Forest Labs2025-1132Bfp8 on 24 GB RTX (with offload)Generation + multi-reference editing in one model
FLUX.2 [klein] 4BBlack Forest Labs2026-014B~13 GBSub-second generation, fully open
FLUX.2 [klein] 9BBlack Forest Labs2026-019B~20 GBKlein quality ceiling, editing

Last updated: 2026-09-27. Each model links to its card, where the licence and terms of use are listed.

How I choose. For personal work FLUX.1 [dev] and its Krea variant are still my default for photorealism. Qwen-Image when the picture contains text, Z-Image-Turbo or FLUX.2 [klein] 4B when I need many iterations fast on 16 GB. FLUX.2 [dev] is the current quality ceiling for editing with several reference images, but at 32B it is a stretch on a 3090.

Open-weight video models

Video is where local generation got interesting in 2025. The table lists what I have run or evaluated; the VRAM column is the community figure for 24 GB cards using fp8/GGUF weights and offload, which is often far below the “official” single-GPU requirement.

Wan2.1 T2V-1.3B2025-021.3BText-to-video480p8 GBEntry point; fast on any modern GPU
Wan2.1 I2V-14B (480p / 720p)2025-0214BImage-to-video480p or 720p, 16 fps24 GB (fp8 + offload)The model behind the clip on this page
Wan2.2 TI2V-5B2025-075BText/Image-to-video720p, 24 fps24 GB (with offload)High-compression VAE; fast but softer output
Wan2.2 T2V / I2V-A14B (MoE)2025-0727B total / 14B activeText-to-video, Image-to-video480p or 720p24 GB (fp8/GGUF + offload); official ≥80 GBTwo experts (high-noise / low-noise); best open quality
HunyuanVideo 1.52025-118.3BText-to-video, Image-to-video480p–720p, 5–10 s (+1080p super-res)14 GB (with offload)Lightweight; step-distilled variants available
LTX-22026-0119BText/Image-to-video with synchronized audioMulti-scale pipeline with 2× spatial/temporal upscalersfp8 / nvfp4 checkpoints for consumer GPUsFirst open DiT to generate audio + video jointly

Last updated: 2026-09-27. The VRAM figures are what the community achieves on 24 GB cards with fp8/GGUF and offload in ComfyUI; official requirements are usually far higher.

Why Wan2.1 for the clip? Because in early 2025 it was the first open model with an Apache-2.0 licence, a 14B image-to-video checkpoint and a ComfyUI workflow that actually fit in 24 GB. Wan2.2’s MoE models look better but are slower on a 3090; HunyuanVideo 1.5 is the lightest quality option today; LTX-2 is the one to watch because it generates audio and video together.

What I learned

  • Licences matter more than benchmarks. Some of the best models don’t allow commercial use; each linked model card lists its terms.
  • fp8 changed everything for 24 GB cards. Without it, 12B–14B diffusion transformers would be out of reach on a 3090.
  • Image first, then motion. Getting a strong still from FLUX and animating it with I2V gives far more control than text-to-video, and it is cheaper to iterate.
  • ComfyUI is the operating system of this space. Every model above ships with a ComfyUI workflow within days of release.

If you want the workflow JSON behind the clip on this page, ask me on LinkedIn.

Related experiences

← Back to experiences