So who should jump into running vision generation locally on an RTX first?

People keep asking which RTX card is best for getting into local visual generation. I’ll break it down by VRAM and how much you’re willing to tinker.

If you’ve got an 8G 30-series card, it’s fine for SD1.5 and low-res Flux, but once you stack multiple ControlNet controls plus hi-res upscaling, VRAM starts choking—you’ll have to rely on tiling. With a 4070 or above, 12G is a much nicer starting point; you can run Flux dev and generation speeds are acceptable. If you really want to turn ComfyUI into a production pipeline, the 4090 or 50-series with 24G is the real dividing line—you can load multiple models at once and batch generate without stuttering.

Newbies shouldn’t rush into high-end gear. First, get your current card to run the workflow, figure out which features you actually use, then decide whether it’s worth shelling out for more VRAM. The tool should serve your needs, not the other way around.

Sorting by VRAM is actually useful. 8G really can’t handle ControlNet plus upscaling without blowing up.

“Don’t go straight for the high-end setup as a beginner” — that advice saved me. First, just get the workflow running.

8G is really awkward, I added tiled to barely hold up the upscale, turning on ControlNet still feels shaky