Diffusion Single File
comfyui

Stop wasting resources on models that are not optimized for consumer grid GPUs

#54
by adm1223 - opened

15.7 GB text encoder for a video model that generate 5-20 sec videos + the model is 21 GB !
With the decompression calculated you will need > 50 GB combined VRAM + System RAM .

Despite that the MiniMax models (Video, music or whatever) themselves are not optimized for consumer grid GPUs
No matter how you try to compress or quantize them they will never run efficiently on consumer grid hardware.

Please stop wasting resources on MinMax.

You can use a 4b text encoder with the help of these clip projection matrices: https://huggingface.co/NicoLab28/ClipProj-MiniMax-H3

Sure the model likes a lot of VRAM for higher resolution and longer duration but it's still possible to generate small videos on an old GPU with 4 GB VRAM.

It's not a waste of resources when the model is this good.

Works like a charm on my 16GB card. Couldn't be happier with this solution.

works great on my 2060 6gb vram

Works a treat on my 5090 32gvram. The details is awesome

However, in reality, it can run with just 6GB of VRAM, while 16GB of VRAM is enough to generate high-definition videos. Its multi-reference generation capability is more than 5 times better than LTX, and it even surpasses Grok’s latest video model.

Nah, as if Comfy is just gonna cram the model and TE straight into RAM/VRAM until it's maxed out. In this day and age with free AI available, why not just go ask one? Feed it Comfy's source code and ask it. (And if you're gonna cite Gemini as your source, just get outta here.)

Comfy Org org

This doesn't even need to be discussed, the model is very much consumer GPU grade, and the "wasted" resources is a big reason for that.

Kijai changed discussion status to closed

Sign up or log in to comment