Spent two days getting the GGUF quantized version of Animagine XL 4.0 running on my own machine, here are my thoughts. The bottom line is: your GPU is basically napping most of the time, yet you’re paying monthly to queue up on someone else’s server to generate images—that math just doesn’t add up.
Once quantized to GGUF, the weights can fit into regular RAM. My 16G machine barely handles it—not blazing fast, but usable. On M-series, it runs via Metal with unified memory, so GPU and CPU share the same pool, no constant copying back and forth. On Windows, it’s CUDA, and if you don’t have an Nvidia card, you can fall back to Vulkan. This model uses booru tags, not full sentences—just list the subject, appearance, and quality tags one by one, and throw in a negative prompt to filter out what you don’t want.
The best part is you can tweak any tag you want, lock the seed, change just one word, and the composition barely shifts. You can run hundreds of tests in a single night without feeling guilty. Prompts and images are all local—pull the plug and you can still generate. For me, that’s more important than saving money.