Sharing FollowFox’s training postmortem on Vodka V4 — it’s rare to see someone actually admit their experiments failed, which makes this way more useful than all those success stories.
The problem they were trying to solve: Vodka series’ realistic images always had this “burnt” overexposed look. The diagnosis was pretty clear — this issue didn’t exist in vanilla SD1.5, so it was introduced during training. And it showed up really early, didn’t get worse with more epochs, which means the learning rate was too high and basically cooked the UNET. They even did a cross-validation by swapping the TE and UNET from the new model with the old one’s, confirming it was the UNET’s fault.
The original plan was simple: decouple the learning rates for TE and UNET, keep TE at 1.5e-07, drop UNET to 5e-8. But here’s where it went sideways — they casually upgraded EveryDream2Trainer, bumped PyTorch to 2.0, xformers 0.20, switched the loss calculation from relative to absolute, VRAM usage changed, and the 25-epoch results ended up way more overfitted. The killer part? They didn’t keep the old branch, so they couldn’t even run a controlled comparison.
In the end, they just switched back to the adamw optimizer to get a passable model, but the burnt look was still there. So they made a few blends, mixed V4 with D-adaptation checkpoints, picked the one that looked least terrible, released it as V4, and let users decide for themselves.
The biggest lesson here has nothing to do with AI: when you’re running a series of experiments, freeze your baseline environment first. Don’t casually upgrade your toolchain mid-experiment.