In ComfyUI, you can use LoRA to extract objects from a video, and it runs fine even on 6G VRAM.

Just tested a LoRA called Obscura, based on LTX 2.3 for video-to-video. It’s supposed to remove unwanted stuff from a video using just a single text prompt. I went in expecting to study the workflow, but turns out the barrier to entry is lower than I thought—the author optimized it for low VRAM setups, runs on 6G VRAM + 16G RAM, which is pretty tame for local video processing.

In practice, large objects get cleaned up nicely, but small ones sometimes stick around even cranking the LoRA strength to 2.5. Makes sense from a technical standpoint—smaller targets mean less edge context to reference, so the model can’t just pull a fix out of thin air.

The tutorial covers installation, node connections, prompts, performance tweaks, and before/after comparisons. The workflow itself isn’t complicated. Overall, it’s way less hassle than manually masking frame by frame, but don’t expect a one-click miracle—you’ll still need to touch up small targets by hand.

6G can run this? That’s such a low barrier to entry, mad respect.

Small objects not being removable is a common issue, don’t just crank up the intensity.

Finally, no more random photobombers in the background.

Maxed out the strength at 2.5 and still failed, so it’s not a strength issue.

Video object removal is way more convenient than Photoshopping.