hello. this video is a walkthrough of the workflow I built, MiniMax Director. in this workflow I implemented every documented feature of the MiniMax H3 model.
I planned to knock out a clone of LTX Director in a couple of days, but my perfectionism made me disappear for almost a month. The main idea of the workflow: from a media file you can take whatever is in it. A person, clothing, a place, a style and so on.
And reuse it anywhere across the whole generation. Every such entity has its own number. You paste it into any segment.
That way the same person stays the same person across different segments. One file gives you as many entities as you like. Each has its own card.
You can, for example, move a face from one photograph onto a person from another. the main node to work with. is called MiniMax Director.
the finished prompt. exactly what goes to the model. above it is the report node.
This node only warns, it never stops the run, and a couple of remarks are almost always sitting there. This group switches turbo mode on. With it on, generation is five times faster but the quality is worse.
Use it for debugging. This is how you turn it on and off. Upscale.
Makes the video high resolution. It is one extra step on top of the generation. Turning it on and off.
the first stage comes from the cache. if you generated the video without upscale before that. the settings row.
one for the whole generation. duration is the length of the clip in frames. seconds are next to it for reference.
the main unit is frames. The framerate for this model is always 24 frames per second. Default resize: how large the reference picture goes into the model.
This one, for example. Match. Shrinks the picture to the video size.
Generates faster, because the picture is given to the model compressed. Max. Generates slower, but the model gets the picture at its original quality.
the model generates video with a specific length, worked out by a formula: the first five frames, then blocks of 17 frames. so the length is rounded. the clear button wipes everything except the length and the resolution.
the main tab, timeline. the main track is what is on screen. camera is how it is shot, audio is what you hear.
These buttons describe a segment with a prompt,. and these attach real files. One block on the timeline is called a segment.
One segment, a second segment. this red line. is the frame pointer.
a new segment lands on it. also, if you press the S key. it splits in half.
if a file was attached it stays on the first half. You can move it by clicking on. the timeline, or with the mouse.
Plus and minus are the zoom. It zooms around the red line. Fit brings the whole timeline back into view.
You can also select by holding the mouse down. Sound and video take up their own length. If an empty segment is selected, the file is loaded into it.
A double click edits the prompt. on the segment. The prompt can be edited either by double-clicking the segment and saving with Enter, or here.
In the segment prompt field. Drg the edge to stretch it, drag the middle to move it. A click selects the current segment.
Command-click, or Control-click, adds to the selection. And all the selected segments can be. edited at once, filling in shared settings.
You can copy, paste, and undo your last actions. The Files tab. This is the list of all the clip's files, right under the tracks.
This label. Picture 1. is the name you use to refer to the file in a prompt.
A file that is not added to any segment is shown dashed. A video has two labels on its row: the image and its sound. the add buttons.
put the file on the track. and add it to the Files list. the add file button adds the file to the file list but not to the timeline.
if a file is added to a segment. the segment's prompt is about that file. shot means segment.
You can insert a reference to a file by pressing the matching button. A file nobody refers to. shows up in the report node.
The node with the warnings. Here, for example. while you drag a file.
the track lights up and an outline shows the segment it will become. when you delete a segment the file stays. This file here.
. . here is the segment holding it.
. . I delete it.
the file has a dashed border, which shows it is not used anywhere. To delete a file. you press the cross.
The cross only appears on files that have no segment. There, I deleted the file. Every file also has a Resize option.
It decides what form the file is handed to the model in. Either at its original size, or shrunk to the video size. Default means this setting is used.
This is the default resize setting for. all files. The model has these limits.
At most 9 pictures, 3 videos, 3 sounds, and no more than 12 files together. The Report node will warn you. about the mismatch, and there will be an error message too.
A picture's sides can also be from 256 to 5760 pixels. And one side no more than 2. 5 times longer than the other.
I'll now try to load a file where one side is more than 2. 5 times the other. And we'll see what the error looks like.
Here is the error message. Now let's talk about the Import/Export tab. The tab saves the whole state of your node into one file, which you can load back later.
You can also copy it to the clipboard and paste it. I'll clear the node now and load a file I saved earlier. if you are running ComfyUI on a new machine.
you will need to upload the files, because. it stores only the file names, not the files. You can press this button and pick all the missing files at once.
Or here, loading them one by one. The tab shows which file is in the cache and which ones you need to upload. You can load a different file, not the one that was there before.
Select a segment to see its settings. Segment prompt is what happens in this segment. The buttons under the field insert a reference where the caret is.
The oval is an entity. From the Who & What tab. The rectangle is a file.
Inserted items are highlighted. A segment knows its own file. For another file you can insert a reference.
The global prompt has the same buttons. Enter with -. what the transition into this segment looks like.
The first segment does not have that field. On-screen text is a phrase that should appear in frame. Only Main segments have dialogue.
For an entity to speak, you have to describe its voice. now you can add dialogue. here you choose who exactly is speaking.
how is the way they say it. and in which language. she can whisper, for instance.
An entity can be tied to a file that is not on the timeline and not attached to a segment. Here the woman will be taken. from this picture.
Now you can choose who speaks. A click selects one entity. If you hold Command, or Ctrl on Windows, and click a second one, it is added.
In that case they will speak in unison. If more than one entity is selected. If Off-screen is ticked, that voice will be off camera.
Carries over means the line continues past the end of the segment. The file row appears when the segment has a file on it. Used as is what this file is for.
Reference is a sample. A person, a scene, a style. Storyboard.
Take only the composition from it - where the camera is and where things stand in frame. A storyboard example. .
. This is how you can sketch out where the objects go. and hand that picture over as a storyboard.
Appearance is not part of it, that needs a file of its own. in the Prompt node the first section is Subject Definitions. as you can see.
Picture 1 is the storyboard, a reference for the first segment. First Frame. and Last Frame are ready-made frames for your video.
The picture becomes the first frame or the last one, and the model generates the rest. Keyframe means the frame lands somewhere in the middle. So the image from the picture will be somewhere in the middle of the segment.
Precision is not guaranteed. So the model generates the beginning and the end, and the middle stays. .
. matching the picture. for frames.
there is a fit setting. crop means cut it down, and stretch means stretch or squash it. set width and height.
fits the clip to the picture. here are width and height. I press it and they change.
keep file. how much of the file survives. fully preserved copies it, partially preserved allows the pose and the light to change.
attribute transfer. moves one feature across. weak reference takes only the style.
Sound has values of its own. The output sound is always synthesised, the model only listens to your recording. detach media removes the file, leaving only the prompt.
the list depends on the file. a video's is different. continue from continues the clip.
edit reworks it. describes. shows which entities are taken from this file.
clicking edit takes you to editing, where you can add another entity. here is the description, what it is and who it is, and a voice if you want it to speak. a video's sound does not reach the prompt until you refer to it.
A video has an audio part and a video part. on Main the clip's picture is taken, on Audio only its sound. you can drag a video onto Main.
for the picture. or onto the audio track - then only the sound is taken. the camera has a track.
It has three fields. motion. strength.
and speed. the middle value is the default. for strength that is medium and for speed normal - those are the defaults.
Zoom changes the magnification. Dolly in is the camera moving towards the subject. Or dolly out, away from it.
You can read the description of every value here. There. In this table.
If the motion you want is not in the list, pick the first item. describe it in the prompt. When you select several segments you can edit.
the properties they have in common. You can give them all the same length. Close gaps butts them together.
Selected segments on the Main track can be merged into one. The dialogue and the prompts are merged along with them. A costume, an object, a place, a style - each is a card.
On this tab. You can take several things out of one file. Here is the description of what is taken from the image or from the video.
This is the name you add. to the prompt. From is the file the entity is taken from.
keep it. this is about the extracted entity. motion from - the movement can be taken from a video.
how they sound - here you add a description of the voice, if you want the entities to talk. As you can see, an S1 label has appeared. Speaker number 1.
Because the voice is described,. the card becomes a speaker. If the field is empty, the card is silent.
It cannot be added to a dialogue. Voice from. takes.
the timbre from a recording, not a description. So the voice can either be described in this field, or taken from a media file, either audio or video. I'll show you with a face swap.
I'll load the person to carry the face onto. and the face file. The person and the face are two separate cards.
On the face card set attribute transfer. And choose who to carry it onto. From the list, or you can type it in.
Now this face will be carried. onto this person. Describe the receiver without the face, otherwise the model may draw something of its own.
What comes from a file? A card takes anything from a file - a person, clothing, an object, a place, a style. The question is: is it about one segment, or about the whole clip?
One segment? Onto the track. The whole clip?
Into Files. Weak reference is only a general likeness. The kind of thing, the composition, the mood.
Nothing literal. Now it is enough to state the style once, in the Global Prompt. If the picture is tied to a track, a matching entry appears in the prompt.
And the model will decide the style applies only to the first segment. The Global Prompt applies to the whole video. Global Music is the background music for the whole clip.
Write instruments and tempo, not mood and emotion. The report node. Remarks in the report are not errors, they do not stop the generation.
Read it before you run, it catches what you would not see until the video is finished. You can also copy the contents to the clipboard with this button. an empty audio track has no effect on the result.
an empty audio card will not reach the prompt. the prompt node is the final prompt that goes to the model, with references to the attached files. here's one example.
This person. is also sitting at a desk, but his face is taken from here, from this picture. And he speaks this line.
Right, MiniMax is up. I'm adding our main picture. Stretching it full length.
Adding the recording. I'll leave room at the start and at the end so the video does not cut off abruptly. Full copy means we take this recording whole, exactly as it is.
Now the prompt. for the main picture. You can pause and read it.
Here audio 1 is mentioned, and this file. is highlighted. I'll add the person's face just as a file.
I'll add a card for the person. I'll add a card for the face. Attribute transfer.
Switching to the face. carry the attribute onto the man card. It is shown by this icon, an arrow to main.
Running it. right, let's see what came out. there we go.
so, if you have any questions or requests, get in touch, I'll try to help. If you need more examples, or anything else. That will give me the motivation to keep going.
Goodbye.