The best open-source AI image generator just got better. The Tongi Lab from Alibaba just released the full Z image model. So, in this video, I'm going to go over all the cool things that you can do with it.
Plus, I'm going to compare it with the other leading open-source image models. And of course, I'm going to show you how to download it so you can run it for free and unlimited times offline. Make sure you stick to the end because I'm also going to show you how you can run this with low VRAM.
Plus, I'm going to show you how to do image to image and a lot more. So, let's jump right in. Note that previously we only had a Zimage Turbo model, but yesterday they released the full model.
Here are some example generations. So, you can see it's super realistic and diverse. It can handle a ton of different art styles and photography styles, different emotions and poses.
Here's the really awesome thing about it. Here are some comparisons with Zage Turbo and Zage. So, you can see it's a lot more detailed and versatile in terms of artistic aesthetics.
And here's my favorite thing about this new Zimage full model. So, for Zimage Turbo, you'll often find that if you use the same prompt and even if you change the different seed, you're going to get very similar images. It's really hard to get more variation in your photos using the same prompt.
But with Z image, you can just keep the same prompt and change up the seed and you can get way more variation in your images as you can see here. And if you prompted to do a selfie photo with four girls for Z image turbo, all four of these girls look very similar whereas for Z image base, each of them look slightly different. Here's another really awesome advantage of Zimage base.
So you can also add negative prompts to your generations. For example, without a negative prompt, you would get this. But if you add westerner as the negative prompt, then you would get a more Asian character.
Or here, if you add sad to the negative prompt, then this person would look happier. All right. Now, let's clarify what exactly this new Zimage model is.
I actually prefer to call it Zimage full because it's not really the base model, which is called Zimage Omni base. This is the really raw model before training it further. And this can do both image generation and editing.
So yesterday they only released this Z image full model. Now this is still a full capacity undistilled model according to Tong Lab which means it's really good for fine-tuning and training Lauras much better than the previous Zimage Turbo. However, I wouldn't call it the base model which is actually Zimage Omni base.
Here's another comparison table between this new Zimage full model and Zimage Turbo. So you can see first of all for this new model it does require a lot more steps to generate an image whereas for Z image turbo it's fine-tuned to generate images super quickly. So for this new full model it's going to take longer to generate an image.
However the diversity of its generations are a lot better than Zimage Turbo and it's also way better to fine-tune and train Lauras from. Whereas for Zimage Turbo, it's actually fine-tuned to do really well in like realistic photos and other visual aesthetics. So, the visual quality of Zimage Turbo is actually a bit higher than Zimage.
You can think of Zimage as a more raw and unpolished model, which you can fine-tune further. Now, don't take my word for it. Here are some comparisons between this new Zimage versus Zimage Turbo and the recently released Flux 2 client.
Now, Zimage is known to do really well at recognizing existing people and characters. So, for the first prompt, I put an Hathaway, Jackiechan, and Messi at a nightclub, lowquality amateur photo. And as you can see, both Zimage models were able to get the characters correct.
And then for Flux 2, it cannot do existing people. Now, I tend to find that for the Zimage full, it creates more plasticky looking faces, which makes it less realistic, whereas for Zimage Turbo, it's a lot better. This does look like an amateur lowquality photo.
So that's something to keep in mind. If you're looking for like realistic portraits, then I would actually prefer Z image Turbo. All right, here's another example.
This time I'm testing its ability to recognize existing anime characters. So here we have Miku, Nezuko, Gojo Saturo, and Sasuk taking a selfie. Again, Flux 2 client just sucks at this task, but you can see for Zage and Zimage Turbo, both of them were able to get the characters correct.
Now, there is a slight flaw with Z image full with the fingers over here, but other than that, it does look very nice. And then here's another test prompt. The prompt is super long.
You can read it up here, but this is basically a Vogue travel magazine cover. And here's what we get. You can see for Z image, it's kind of a bit saturated, whereas for Z Image Turbo, again, I like the visual aesthetics a bit more.
Both of them were able to get most of the text correct. In fact, Z image is even able to get the Vogue logo correct. And then for Flux 2 Klein, while it's good at generating realistic photos, unfortunately, the text is all messed up.
All right, here's another test. We have a woman taking a selfie in a cafe, cute expression with first lips, winking, one hand doing a peace sign, irregular angle, amateur, lowquality photo, and all three of these were able to generate this very well. Now, here's a test on generating long text.
The prompt is up here. Feel free to pause the video and read the prompt, but basically none of the models were able to render the text completely correct. I would say it's a fail for all three.
If you want to generate long snippets of text, the best open- source model to use is currently Quen image. All right, here's another tricky prompt which no image generator, including Nano Banana Pro, could get correct. So, the prompt is 11:15 on the clock and a wine glass filled to the top.
You can see none of these models were able to get it correct. Next, I also wanted to test some artistic styles. So, here we have a flat illustration of a deer in a forest, but everything is composed of dots of varying sizes on a white background.
All three of these are not bad, but for Z image, you can see the blades of grass here aren't really composed of dots. Same with Zimage Turbo. So, here the one that really followed my prompt was Flux 2 Klein.
Next, here's a Manet style impressionist painting of a busy train station. For your information, Manet paintings look something like this. It's basically some really rough brush strokes, and you kind of need to step back in order to see the scene.
Like nothing is really defined. And I would say in this example, Zage handles this the best. This looks the most like a Manet impressionist painting.
Whereas for Zagemage Turbo, this comes close, but it's a bit too defined. And then for Flux 2 Klein, this looks even more defined. Next, here's also a test on a minimalist Chinese watercolor painting of a tiger in a forest.
And I would say both Zimage models generated this very well. For Chinese watercolor paintings, you can see there's no like defined outlines. It's quite abstract brush strokes everywhere.
As you can see from the trees that are drawn in these two images, whereas for Flux 2 Klein, the outline is just too defined. This does not look like a minimalist Chinese watercolor painting. So, I would say it's a tie between Z image and Z image Turbo.
But again, notice that for Z image, it tends to add more saturation to the image. All right. Next, here's a test on its UI design.
Again, the prompt is super long. I'm not going to read it out. Basically, all three models got the prompt mostly correct.
However, I did mention that for the third icon, it should be a teal route marker. and the only one that was able to generate a marker icon was Flux 2 Klein. But I would say in terms of the overall design, Flux 2 Klein's generation looked the ugliest.
Let me know in the comments what you think. Here's a test on its prompt understanding. So, we have a tropical beach in Bali at dusk.
We have a woman in a blue and yellow tie-dye sarong and a pink flower crown, which all three models got correct. She's doing yoga, which I would say both Z image models got correct. For Flux 2 Klein, she's just kneeling there.
She's not really doing yoga. And then we have a gray monkey stealing a brown coconut from her green beach bag. Here, I would say Z image full and Flux 2 Klein were able to generate the monkey stealing from the bag.
Whereas for Zagemage Turbo, the monkey isn't really stealing from the bag. And then we have a red surfboard with a white hibiscus leaning against the palm tree. A brown fisherman's boat floats on the blue waves.
A red Bali sunset sign glows at a bar. All three models were able to get these elements, but the only one that was able to get Bali sunset spelled correctly was Zimage full. For Zimage Turbo, you can see a misspelling here.
Same with Flux 2 Klein. So, I would say the winner here goes to Flux image. And then here we have a quick anatomy test.
A woman sitting and showing her palms and soles of feet. You can see all three of them got it correct, but I would say the two Zimage models followed my prompt better. for Flux 2 Klein.
She's not really showing her palms to the camera. Her palms are just facing up. So, those are some of my quick tests.
Let me know in the comments what you think. Next, let's go over how to install this and use it locally on your computer. So, on the official HuggingFace page, they do give you some instructions on how to install this, but this is using raw code, which is not as intuitive.
A better platform to use this is called Comfy UI. This is in fact the most popular platform for running open- source image, video, and audio generators on your computer. The nice thing about Comfy UI is it's very customizable, plus it has auto offloading, so it's especially useful for those of you who don't have high-end hardware.
If you're not familiar with Comfy UI, definitely see this installation tutorial first. Anyways, for this video, I'm going to assume you already have Comfy UI installed. So, the first step is to update Comfy UI.
I'm using the Windows portable version which is the recommended version. And in your folder, you should see this update folder. So, let's double click on this.
And here, you can either click on this one to update it to the latest version, or click on this one, update Comfy Stable, which updates to the latest stable version. For me, I'm going to doubleclick on this one, and then press run. And it should proceed to pull the latest changes and update Comfy UI to the latest version.
All right. Afterwards, let's press any key to continue, which would exit the terminal. And then afterwards, we can proceed to run Comfy UI.
So, I'm going to start up Comfy UI. All right. After you start up Comfy UI, simply click on this templates button on the left sidebar.
And then, if you click on image, you should see Zimage text to image over here. Make sure you're not selecting the Zimage Turbo one, which is a few months old. This is the new one.
And if you don't see this, you can also search Zimage up here. Anyways, let's click on this one. And here's the workflow.
It's as simple as that. And by the way, if you don't see this workflow for some reason, I will also link to this Comfy UI page where you can directly download the workflow by clicking this button here. There are so many AI image and video generators out there.
It can get very overwhelming. Luckily, Higsfield, the sponsor of this video, brings everything together all in one integrated platform. And they just launched a new AI influencer studio, which is super powerful.
This lets you easily design, control, and bring your AI influencers to life. Here's how it works. You can choose from over a 100 configurable parameters, including the character type.
So, let's choose human. And then also the gender, the ethnicity, the skin color, the eye color, age, and more. And it'll proceed to generate your influencer.
If you don't like the look, you can also use Nano Bananets to edit this further. And then afterwards, you can get your influencer to do pretty much anything, like talk about a product or dance. There are a ton of different motions you can choose from in their motion library.
For example, let's get her to dance like this. And then let's press generate. And here's our result.
In addition, Higsfield is the best platform for accessing all the top image and video generators out there. And they have a ton of pre-built templates, including camera presets, VFX, and more. Higsfield is the ultimate platform for creators to easily generate anything they imagine.
Try it for free using the link in the description below. Let's go over this really quickly. It's really straightforward.
So the first step is we need to load the models. Now very conveniently the links to all the models are here already. So first let's download this Z image BF-16.
So let's click on this and this goes in comfy UI in models and then in diffusion models. Note that this BF-16 file is 12 GB in size. So you'll likely need at least 12 GB of VRAM plus offloading to run this.
And then afterwards, we also need to download this text encoder called Quen 34B. So let's click on this. And this goes in Comfy UI in models and then in text encoders.
And this one is 7. 8 GB in size. Note that you should already have this if you downloaded the Z image turbo workflow, so you don't need to download this again.
And then finally, we also need to download the VAE. So, let's click on this. And this goes in Comfy UI in models and then VAE.
Now, this one is 327 megabytes in size. And again, you should already have this if you've downloaded the previous Zimage Turbo workflow. And that's pretty much it.
So, afterwards for step one, we just need to load the models. If you just finished downloading the models, simply press R to refresh your model list. And then for this first dropdown, let's select Z image BF-16.
For the load clip, yes, let's select Quinn 34B. And then for the VAE, yes, let's select AE. safe tensors.
And that's pretty much it. Now for step two, here is where you specify the width and height of your final image. And here for batch size, this is how many images you want to generate at once.
And then over here is where you can enter a positive prompt and a negative prompt. In fact, adding a negative prompt is highly recommended for Z image as I showed you before. If you've been playing around with Z image Turbo, the negative prompt doesn't really work well, which I'll talk about in a second.
But here, a negative prompt actually works very well. So, for example, for our prompt, let's write a woman in the city. And then for the negative prompt, we can add words like blurry, low res, oversaturated, cartoon, etc.
All right. So your prompt will be fed through these two nodes to basically generate the image. Now let's go over these settings really quickly.
The seed is basically the starting configuration of random noise. So basically, if you keep all the settings the same, but you change the seed to a different number, you're going to get a slightly different image. Conversely, if you keep everything the same and you set the same seed, you're going to get the exact same image as before.
And then this setting just basically tells it to change it to a random different seed after generation. or you can also keep the seed fixed for reproducibility. For me, I'm just going to set this to randomize.
And then the step count is how many well steps it takes for the AI to generate the image. Now, down here, it says that for the image full, the recommended number of steps is 30 to 50. So, keep that in mind.
And then for CFG, this is the super important part. This is how literally you want the AI to follow your prompt. In general, a lower CFG would give it more creativity, whereas a higher CFG value would make it follow your prompt more literally.
Here you can see that the recommended CFG settings is 3 to 5. Now, here is why this Z image base model works with a negative prompt, and that's because CFG works well in a range of 3 to 5, whereas for Zimage Turbo, you'll notice that the CFG is often 0 to one, in which case a negative prompt would not work. So, that's the advantage of Zimage.
It gives you more control over the prompt conditioning using a negative prompt. And then anyways, the sampler and theuler is basically the algorithm used to generate the image. You can see there are a ton of different algorithms you can choose from, each with slightly different characteristics.
Feel free to play around with these, but I'm just going to leave it at the default. And that's pretty much it. Let's press run.
And for your information, I'm using an RTX 5000 ADA on my laptop, which has 16 GB of VRAM. All right. And here is what we get.
Notice that this is a save image node. So, it's automatically saved in your Comfy UI's output folder as you can see over here. And if I expand my terminal, that took around 1 minute and 25 seconds.
So, it is like way slower than Zimage Turbo, which only takes like 7 seconds on my GPU. And that's because for the step count of Zimage Turbo, you only need like 7 to nine steps. So, that's the sacrifice you'll need to make if you use Zimage base.
But the advantage of this model is that it has way more variations. So if you run this again, it's going to be a completely different model of a woman in the city. Whereas if you use Zimage Turbo, you often get very similar results even if you change the seed.
So let's run this again and see what we get. All right, so here's my second generation. And you can see this one took a bit faster.
This took like 92 seconds. And if I open up my output folder, you can see that, you know, both images are very different. So that's the advantage of this Z image full model.
You get a lot more variation even if you use the same prompt. All right. Now, this full model is like 12 GB in size.
So next, let's go over how you can use this if you don't have that much VRAM. Fortunately, there are already GGUFS or more compressed versions of Z image which can run on lower VRAM or even like CPU or AMD GPUs. So here are all the versions of different compressions and different sizes.
You can see the smallest one, Zimage Q2K, is only 4 GB in size. You might barely be able to fit this in 4 GB with some offloading, but don't take my word for it. You'll need to try this out.
And then here are some other options like 5 GB, 6 GB, 7 GB. Of course, the larger the model, the better quality it would be. So, you should select the largest model that can fit within your VRAM.
So, for example, if you have 8 GB of VRAM, then I would recommend this one. Anyways, for me, I'm just going to do a quick example with this Q2 one. So, let's click on this to download it.
And this goes in Comfy UI in models and then in unit. Let's click save. All right.
Now, back here, I'm just going to open up a new workflow so we can start from scratch. So, again, you just need to click on templates and then click on Z image. All right.
Afterwards, what we need to do is basically replace this model node with a GGUF node. So let's double click anywhere on this interface and then search for unit and you should see this one unit loader GGUF. So let's click on this and we just basically need to connect this to this model input over here which would automatically disconnect this one.
In fact we can just click on this and press Ctrl +B to disable it. Now, if you don't see this GGUF node, then what you'll need to do is click on manager and then search for GGUF and you'll need to download this one, Comfy UI GGUF by City96. And if you do have it and it still doesn't work, then try updating to the latest version as well.
All right. Afterwards, if you've just downloaded your GGF model, simply press R to refresh your model list. And then down here in the model selector, click on the model that you selected, which in my case is Z image Q2.
And then again for the text encoder, let's select Quen 34B. And then for VAE, let's select AE. Safe tensors.
And that's pretty much it. So it's the same settings as before. Let's press run to generate this with the GGUF.
All right. And here is our result. So that's how you can run Z image full even if you have lower VRAM.
All right. Now, right now the default workflow only accepts text to image, but you can also modify this to do image to image as well. Let's go over how to do this.
So I'm just on the default workflow from scratch. You can also replace this node with a GGF if you're using that. And again to start, I'm just going to select the models that I have.
So let's select this and then select this. Now, basically how this works is right now for text to image, it's using this to create a canvas of random noise, which would then be plugged through this K sampler to gradually reduce some of that noise step by step until it generates your final image. But instead of starting with random noise, what we can do is start with an image.
So, let's double click anywhere here and then search for load image. And let's select this one, load image by Comfy Core. And then here's where you can upload an image.
So let's choose this image. Now, right now we can't just plug this image into the K sampler because this K sampler requires a latent image. So what we need to do is doubleclick anywhere here and then type in VAE and we need to select this one VAE encode.
So this is going to take your VAE which we've loaded over here. So in fact, let's connect this VAE to our encode node and it's going to turn your image. So let's connect the image to this into latent space which can then be plugged into this K sampler.
So next we just need to connect this to the K sampler like this. And that's pretty much it. And it should automatically disconnect this node.
In fact, just to make it more organized, let's select this and press Ctrl +B to disable it. Now let's say we want to convert this image into a realistic image. So for the prompt we can write something like a girl green background leaves and foliage in the foreground realistic photo midshot.
And then for the negative prompt here is everything that we don't want. So we don't want this to be a cartoon or vector style or 2D low res blurry oversaturated. Now there's one more setting that you need to set which is the d noiseise value.
This is basically how much of this input image do we want to change. Right now it's set at 100% which is not good. You're just going to get a completely different image.
So let's first set this to something like 0. 5 and then press okay. So it's going to retain like 50% of the details of this image but it's going to turn it into a realistic photo.
Let's see if this is enough. If it's not, we might need to adjust this value further. So let's press run.
All right. You can see it still keeps too much of the original image. So let's set this to a higher value.
All right. So after increasing this CFG to five and then also increasing the D noiseis to84 here is what we get. So here is the before and here's the after.
Indeed it does turn this into a realistic photo. It's not perfect. This is not an editing model.
This is just regular image to image. Now one more thing I want to show you in addition to image to image is how to do inpainting. So to keep this organized let's just exit this and start from scratch.
All right. So here's the original workflow. Now to do inpainting again we need to replace this latent canvas of noise with an existing image that we want to edit.
So let's double click anywhere here and then click on load image. This time let me upload this image. And again we need to plug this through a VAE encoder first.
So let me double click anywhere here and then search for VAE and then click on VAE encode. And this requires a VAE which we loaded up here. So let's connect it over here and then let's connect the image to over here.
Now on this load image node, you can right click on this and then click on open and mask editor. So this will open up your image over here and here is where you can brush over parts that you want to remove or edit. And here you can adjust various settings like the thickness of the brush, the hardness, opacity, etc.
So let's brush over this notebook like this. It's also important to add a bit more feathering along the edges. All right, let's do something like that.
Again, this is the part that we want to edit. So, let's click save. And then next, what we need to do is doubleclick anywhere here.
And then let's search for set latent. And you should see this option, set latent noise mask. So, let's click on this.
And here is where you connect the mask that we just drew to this node. And then we just need to connect the latent from the VAE encoder over to here. And then finally, this gives us a latent output which we can then connect to this K sampler.
Just like that. And this would automatically disconnect this one which we can just press Ctrl +B to disable. And then for the prompt up here, we basically need to describe what we want to input over here.
So let's say a cat sleeping on the table. And that's pretty much it. Let's press run to generate the image.
One thing to note is that you can also tweak this den noiseise setting here. So right now it's set at 100%. So it's going to completely replace the area that I drew over.
But if you do want to retain some parts of that image, then you can decrease this the noise value to a lower number. All right. And here's what we get.
So let me drag this side by side so you can compare the before and after. So that's how you can inpaint using Z image full. Again, this is just a very basic inpainting workflow.
I don't actually recommend you use this. I would much rather use a real image editor like Flux 2 Klein or Quen imageedit to actually just edit this image using natural language. But if you are interested in doing in painting, here is how you would do so.
All right. Now, final thing I want to mention is Loras. These are basically fine-tuned models of a certain character or art style or position or effect or whatever which you can add to your workflow.
However, note that right now we have a ton of these Zimage Turbo Luras. They don't actually work for Zimage base, at least for this official workflow. So, don't waste your time trying to add a Laura here with the Zimage base model.
It's not going to work well. However, once we do get some Zimagebased models released, here's how you would add a Laura to it. You just basically need to add a Laura right after your diffusion model.
So let's double click anywhere here and then type Laura loader model only. And here is where you can choose the Laura you just downloaded. And then you would basically connect the diffusion model here and then connect this output over to here.
If I look up Z image base Loras. There aren't any available yet. So I can't do any demos on that.
But once they do come out here is how you would link a Laura to your workflow. And then one more additional thing I forgot to mention is the strength here. So this is how much influence you want the lura to have on your final image.
So right now it's 100%. If you want lower influence you can set this to a lower value like8. Now one of the most important strengths of this new Z image full is that it's really good at creating loras and fine-tuning from.
So let's also go over how to create loras from this. The nice thing is one of the main platforms for creating luras is called AI toolkit by Oris and they've already added the ability for you to create Loras from this new Z image full model. I'll link to this GitHub page in the description below which contains all the instructions on how to download and also set up the data set to train Loras from.
It requires a ton of steps so it's beyond the scope of this tutorial but I'll link to this in the description below. Now, just to give you some context, here's how you would regularly train Lauras. You need to gather a ton of photos of a certain person or art style or effect or whatever.
And then you also need to label each photo and then you would plug it through a Laura trainer like AI toolkit which would create a for you. This takes a lot of time and effort and compute. Now, what I think is a even bigger deal is this new model by diff studio called Zimage image to Laura.
Instead of going through the traditional workflow of downloading a ton of images and then manually labeling them and then training it that way here you can just plug in, you know, a handful of images, even just like two or three images of a certain art style or person or whatever you want to transfer and it can create a Laura from this in minutes. So, for example, if these are your input training images, after creating the Laura, if you generate a cat, here's what you get or here's a dog or here's a girl. Now, the art style isn't exactly like the reference images, but it is like a flat vector illustration.
Here are some other examples. If you input these four photos for training your Laura, afterwards, if you generate a cat or a dog or a girl, here is what it would look like. All generations contain the characteristic blue sky and white clouds and flowers.
Or here's another example for your reference. The awesome thing is this is already released. There's no like easy Comfy UI workflow for this yet, so you'll have to work with raw code, but I'll link to this page in the description below, which contains all the instructions on how you can set this up on your computer.
Note that the quality of this method, Z image I toL, is not as good as if you put in the effort to find a lot of images, labeling them, and then plugging it through something like AI toolkit to generate the Laura. So this AI toolkit method produces higher quality Loras, whereas this Z image ITL is just a quick and dirty way for you to generate Loras with just a few images in a few minutes. Anyways, I'll link to both these methods in the description below for your reference.
And that sums up my review and tutorial on Quen image base. Let me know in the comments what you think of this. And if you run into any errors during the installation, welcome to paste the exact error message in the comments below and I'll try to help you troubleshoot as much as possible.
As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week.
I can't possibly cover everything on my YouTube channel. So, to really stay uptodate with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below.
Thanks for watching and I'll see you in the next one.