With the native comfy implementation of Hunyuan I have tweaked the workflow to work for 12GB VRAM cards. It looks like you can get at least to 73 frames and probably a bit more. It takes about 8 minutes for a 4070Ti...
Runtime profile
Source description
With the native comfy implementation of Hunyuan I have tweaked the workflow to work for 12GB VRAM cards. It looks like you can get at least to 73 frames and probably a bit more. It takes about 8 minutes for a 4070Ti to run 20 steps.
Do make sure to update your comfy and to get the exact result above the guidance I put up to 10 (I reduced it for the base workflow as it caused some burning for some prompts).
As its not as complete as the wrapper node there is a few less features than that for now. But there is also no crazy special installations that you need to do.
Links to model downloads here:
AI-generated commentary
AI-generated explanation based on source and configuration details. Suggestions are clearly labeled.
This ComfyUI workflow is for generating Hunyuan Video clips from a text prompt: it encodes the text and outputs a combined video, and it is described as tweaked for cards with 12GB of VRAM.
The graph encodes text, creates an empty Hunyuan latent video, samples it, decodes it with tiled VAE processing, and combines the result into a video.
The description says the output can reach at least 73 frames, possibly more.
The listing reports about 8 minutes for a 4070Ti to run 20 steps.
Before using it, update ComfyUI, as the description instructs.
Suggestion · not verified
The native ComfyUI implementation is described as not requiring special installations.
Named model files are clip_l.safetensors, hunyuan_video_t2v_720p_bf16.safetensors, hunyuan_video_vae_bf16.safetensors, and llava_llama3_fp8_scaled.safetensors.
Named nodes include BasicGuider, BasicScheduler, CLIPTextEncode, DualCLIPLoader, EmptyHunyuanLatentVideo, FluxGuidance, KSamplerSelect, and ModelSamplingSD3.
It also names RandomNoise, SamplerCustomAdvanced, UNETLoader, VAEDecodeTiled, VAELoader, and VHS_VideoCombine.
If you are following the described setup, set guidance up to 10; the base workflow uses lower guidance because some prompts caused burning.
Suggestion · not verified
Before use, verify that the listed model files and nodes are available in your ComfyUI setup.
Suggestion · not verified
The description notes that this implementation has fewer features than the wrapper node.
Does this need to be edited?
Sign in to send an edit request.
Sources
1 sourceSource excerpts
1 excerptSource context: 16736 downloads · Type Workflows · Base model Hunyuan Video
Estimated VRAM requirement
Estimate unavailable
235 MB across 1 of 7 model files. Model file total + 25% loading overhead + 2 GB execution buffer, rounded up.
Requirements
18 requirementsText encoder · 235 MB · SAFETENSORS · Unknown
DualCLIPLoader
Not resolvedText encoder · Unknown
hunyuan_video_t2v_720p_bf16.safetensors
File unverifiedCheckpoint · SAFETENSORS · Unknown
hunyuan_video_vae_bf16.safetensors
PossibleVAE · Unknown
llava_llama3_fp8_scaled.safetensors
ConflictText encoder · Unknown
VAEDecodeTiled
Not resolvedVAE · Unknown
VAELoader
Not resolvedVAE · Unknown
BasicGuider
Not resolvedNode pack · Unknown
BasicScheduler
Not resolvedNode pack · Unknown
CLIPTextEncode
Not resolvedNode pack · Unknown
EmptyHunyuanLatentVideo
Node pack · Unknown
FluxGuidance
Not resolvedNode pack · Unknown
KSamplerSelect
Not resolvedNode pack · Unknown
ModelSamplingSD3
Not resolvedNode pack · Unknown
RandomNoise
Not resolvedNode pack · Unknown
SamplerCustomAdvanced
Not resolvedNode pack · Unknown
UNETLoader
Not resolvedNode pack · Unknown
VHS_VideoCombine
Not resolvedNode pack · Unknown