Instructions to use Comfy-Org/MiniMax-H3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusion Single File
How to use Comfy-Org/MiniMax-H3 with Diffusion Single File:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
problem with off-screen voiceover【画外音问题】
After repeated attempts, the model consistently assumes that the voice-over is spoken by the character's mouth and drives the character's lip movements accordingly. Could you please help me check what might be wrong with the prompt? 【反复尝试后,模型始终会认为画外音是角色嘴里说出来的,并驱动角色嘴部说话。请帮我看看,提示词有哪里不对。】
subject_definitions:
<Subject 1> (S1) is the person in <Picture 2> and <Picture 1> - fully preserved for appearance, hairstyle, clothing, and posture.
<Picture 1> is a fully preserved scene reference for the desk area - it provides the desk surface, blue ceramic mug, pen holder, desk lamp, and the map on the desktop, within a windowless, enclosed indoor environment. Behind the character is a sealed wall with a wooden door - no windows are present anywhere in this room.
summary:
[reference generation] A 14-second single shot inside a windowless, enclosed bedroom. <Subject 1> sits at the desk, picks up a pen from the map and marks it, while keeping his mouth completely closed throughout. The camera is positioned on the desk surface, shooting upward at a low angle. Behind the character is a sealed concrete wall with a wooden door - no windows exist in this room.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the character's appearance, clothing, and posture are exactly as in <Picture 2>.
<Picture 1> (desk area): fully_preserved - the desk surface, blue ceramic mug, pen holder, desk lamp, and map on the desktop are fully retained, within a windowless, enclosed indoor setting.
detailed_description:
The target video adopts a photorealistic, minimalist interior style with neutral lighting from overhead fluorescent tubes, set within a windowless, enclosed bedroom environment as defined by <Picture 1>. Behind the character's back is a sealed concrete wall with a wooden door - there are no windows anywhere in the room, and no natural light enters from outside.
[Shot 1] The shot opens with the camera positioned on the desk surface, shooting upward at a low angle toward <Subject 1>. The character dominates the frame, with the blue ceramic mug occupying the lower-right foreground (softly out of focus), its steam visibly rising. On the left side of the frame, the desk lamp is visible, turned on and casting a warm glow. In the background behind the character, a sealed concrete wall and a wooden door are visible - no windows, no exterior light. The camera performs a static shot. The man (S2) says in an off-screen voiceover: [Chinese] 和平铁路的轨道是超级金属做的——应该还可以用。从这里到喜马拉雅山,走和平铁路要六千七百公里,走公路只需要三千八百公里。 while his lips remain completely closed. <Subject 1> reaches his hand toward the map spread on the desk, picks up a pen from its surface, and begins to mark locations on the map while thinking. The pen produces a soft scratching sound against the paper as he draws lines and circles.
overall_soundscape:
Low-level room tone and the continuous hum of the fluorescent lights persist throughout the shot, within the confined, windowless space.
I had no trouble with adding a narrator voice-over and have it not affect lip movements of the subject(s) in the video.
I simply used the following form in my prompt:
Female narrator voice, saying: "Here we see..."
I don't think you should refer to the narrator as 'the man' or it will assume you mean 'the man' seen on screen at the time, and thus animate the lips.
朋友,你解决了吗??我也遇到 ,这个内心独白,时灵时不灵的,一会儿可以 一会儿不可以 ,好烦

