MiniMax image-01 is the image model behind MiniMax's paid API. In simple words, we send it a prompt, and it sends back a picture. We can also send a photo of a person, and it will try to put that person in a new scene.
In this blog, we will put image-01 next to Qwen-Image 2.1 running at 8-bit on our own RTX 5090. Both got the same 24 prompts at 1024x1024. Eleven used a photo of the author, seven used no reference, and six built a small picture book.
Why Did We Test image-01 and Not MiniMax H3?
We first set out to test MiniMax H3. H3 is MiniMax's new open model. It can make video, and it can also make still images. But MiniMax's API offers H3 only as a video model. It returns clips of 4 to 15 seconds at 768P or 2K. It cannot return a 1024x1024 still image.
The only model on MiniMax's image API is image-01. Running H3 at home was not an option either. It is a 33B model with a 32B text encoder, and that does not fit in 32 GB at 8-bit.
So, this post compares image-01 with Qwen-Image 2.1. It tells us nothing about H3's images.
There was one more reason to run this test. A user on the Qwen-Image 2.1 model page compared Qwen with H3. They said Qwen tends to paste the face from the reference photo, and it struggles with new poses and camera angles. We built three photo prompts and the whole picture book to test that claim.
Our Test Setup
| Item | Setting |
|---|---|
| MiniMax | image-01 on the official API, width and height 1024, prompt_optimizer off |
| MiniMax reference | subject_reference with type: "character" |
| Qwen | Qwen/Qwen-Image-2.1, fp8 weights and fp8 activations (torchao), all on the GPU |
| Qwen sampling | seed 11, 40 steps, 1024x1024 for every image |
| GPU | RTX 5090, 32 GB of VRAM, driver 610.88, Windows 11 |
| Images | 24 prompts per model, one image each, 48 in total |
We turned MiniMax's prompt optimizer off, so both models read the exact same words. MiniMax returned a 1024x1024 JPEG every time. Qwen returned a PNG.
The reference was one portrait of the author in a plaid shirt. It appears in the first column of each photo gallery below. In every gallery, the top row is MiniMax and the bottom row is Qwen.
We checked each image by eye against a short pass list. We wrote that list before we made any image. A pass scores 1, a half-right image scores 0.5, and a miss scores 0. Face match and text were also measured by software, as we will see below.
How Did Each Model Score Overall?

Here, we can see the whole story in one chart. Qwen won with the photo and with no reference. MiniMax won the storybook, and by a clear margin. The next sections show why.
What Happens When We Give Each Model Our Photo?
MiniMax image-01 did well on plain photo edits. The headshot, the kurta, holding a GPU, and the server room all passed.

The trouble started with style. We asked for a Pixar-like 3D character, an anime drawing, and a pencil sketch. MiniMax gave back a photo or a colour painting every time. It kept the person, but it never changed the style.
Qwen did better here. The sketch was a real graphite drawing. Its Pixar and anime tries were weak, though. The Pixar image is a lightly smoothed photo, and the anime image is a flat cartoon.

The YouTube thumbnail shows the text gap. Qwen wrote "LOCAL AI" cleanly and had the person point to the right. MiniMax's text came out as garbled letters, and the person points at the camera.
Can They Change the Pose or the Camera Angle?
This is where we tested the pasting claim. We asked for three new views: a side profile while walking, a high camera looking down at a desk, and a hiking shot with the body turned away.

MiniMax missed the side profile and the high angle. Both came back as close shots facing the camera. Its hiking shot was its best pose, with a real look back over the shoulder.
Qwen got the high-angle desk shot right. But look at its other two images. The body walks sideways with an umbrella, yet the face still looks straight at us. On the mountain, the body faces the camera instead of turning away. Qwen turned the body but kept the face from the photo. That is the pasting problem, and we can see it clearly here.
Does MiniMax Keep Our Face?
To measure the face, we used SFace. SFace is a face-recognition model that scores how alike two faces are. A score of 0.363 or more means "same person". As a guide, our original photo against a cutout of itself scores 0.95.

Both models stayed above the same-person line on all 11 images. But the two bars behave very differently.
MiniMax scored between 0.72 and 0.82 on 10 of 11 images. It never reached 0.9. Its score barely moves between a photo edit and a style prompt, which fits what we saw above: it never changed the style. By eye, its faces look like a close relative of the author, with fuller cheeks. MiniMax seems to draw a new face that looks like the reference.
Qwen scored 0.94 to 0.96 on the four plain edits. It scored 0.91 on the high-angle desk shot. In simple words, Qwen copies the real face almost pixel for pixel. Its score drops only when it truly changes the image: 0.52 for the Pixar try and 0.49 for the small figure on the mountain.
There is one more clue. Qwen kept the author's plaid shirt in 8 of the 11 images. None of our prompts asked for it. That is copying, not redrawing.
Note
SFace rewards staying close to the photo. It cannot tell "same person in a new style" from "same photo, style ignored". So, we always read it together with the pass scores.
Can They Draw a Storybook?
For the storybook, each model first drew a character with no reference: Pip, a small orange fox cub with a green scarf and a yellow backpack. Then it drew five pages, using its own Pip as the reference. Each page asked for a new scene and a line of text in a banner.

Both character sheets were good. The difference shows on the pages.
MiniMax told the story. Pip stretches and yawns in a tree bedroom. Pip walks across a bridge past a frog. Pip stands small and worried in a dark forest. Pip reads a map with an owl by lantern light. Pip hugs the mother fox at the door. The scarf and backpack stayed the same on every page. It missed one detail: page 4 asked for a view from behind, and it drew a side view.
Qwen pasted its character sheet onto every page. Pip stands in the same pose with the same smile five times, and only the background changes. The owl and the map never show up on page 4. The mother fox never shows up on page 5.
Text went the other way. Qwen spelled all five banners exactly. MiniMax spelled two of five correctly. It wrote "wonke" for "woke", "wawed" for "waved", and "show" for "showed".
Which Model Writes Text Better?

Across all 11 lines of text we asked for, Qwen got 11 exactly right. MiniMax got 6.
- The neon sign "MIDNIGHT RAMEN" was right in both.
- On the poster, Qwen matched all four lines. MiniMax matched three. It wrote the subtitle in capitals without the hyphen and added a "WORKSHOP" label nobody asked for.
- On the thumbnail, only Qwen wrote "LOCAL AI" correctly.
How Did They Do With No Reference?

With no reference, Qwen scored 5.5 of 7 and MiniMax scored 4 of 7.
- Both passed the photo portrait, the neon sign, and four graphics cards on a table.
- Both placed the red cube, blue sphere, and cone correctly. MiniMax's cone is grey-olive instead of green.
- We asked for a clock at 3:15. MiniMax drew about 10:10. Qwen pointed both hands at the 3. That looks like 3:15 at a glance, but the hour hand should sit a little past the 3, so we gave it half.
- We asked for four server racks. MiniMax drew five, and Qwen drew six.
Our Qwen-Image 2.1 benchmark saw the 4-bit model draw eight graphics cards instead of four. Here, 8-bit drew exactly four. That is one image each, so it does not prove 8-bit counts better. Counting is simply not reliable yet.
How Fast Is Each One?

For most of the run, both took about 20 seconds per image. Qwen was a little faster and very steady, because it had the GPU to itself. A reference photo adds about 5 seconds to Qwen (16.1 s to 20.9 s), because the photo becomes extra input for the model.
MiniMax's time is the full round trip from our PC. It includes the upload, the wait in MiniMax's queue, and the download. Its last four storybook pages took 86 to 102 seconds each, against about 20 seconds earlier. Those pages ran last, so we cannot tell whether the prompts or a busy server caused it.
On memory, Qwen at 8-bit holds 16.6 GB of weights on the GPU. It peaked at 21.1 GB for a plain image and 23.6 GB with a reference. That fits on a 32 GB card, but only after we closed a small model that was left loaded in Ollama. The GPU used a median of 3.1 Wh per image.
Our earlier benchmark timed a 1024x1024 image at 19.3 s at 8-bit. Here, it took 16.1 s. That earlier test used a long thumbnail prompt, and a longer prompt gives the model more to process. We did not test prompt length on its own, so treat that as the likely reason, not a measured one.
Where Does Each Model Break?
MiniMax image-01:
- It ignores style changes when we give it a photo. Pixar, anime, and a pencil sketch all stayed photos.
- It keeps a front-facing view when we ask for a side profile or a high camera.
- It misspells text. Three of five story banners had errors, and the thumbnail text was garbled.
- Its faces look like the person but do not match closely. SFace never went above 0.82.
Qwen-Image 2.1 at 8-bit:
- It pastes the reference. A character keeps one pose for a whole story, and the photo's shirt shows up where nobody asked for it.
- It cannot add new characters around a pasted one. The owl and the mother fox never appeared.
- It turns the body but not the face for side and turned-away poses.
- It miscounted the server racks and did not move the clock's hour hand past the 3.
Which One Should We Use?
For thumbnails, posters, signs, and any image with exact text, Qwen was the better pick in our test. It wrote 11 of 11 lines correctly. It was also better at keeping a real face while changing the outfit or background.
For a story with a recurring character, MiniMax image-01 was clearly better. Its character moved, acted, and met other characters from page to page. Just check the text, or add it afterwards.
Limitations
- One image per prompt. A second try could flip any single pass or fail, so trust the group totals more than any one image.
- The pass scores are one reviewer's judgment, against a list written before the run.
- We did not test MiniMax H3 in any form.
- We could not measure MiniMax's price per image, because our plan key does not report a cost per call.
- MiniMax's speed depends on the network and server load. It is a snapshot from one afternoon.
- MiniMax's prompt optimizer was off. It might change the style and angle results.
Conclusion
MiniMax image-01 and Qwen-Image 2.1 fail in opposite ways. MiniMax redraws the person and the scene freely. That makes it good at stories and bad at keeping a face or following a style. Qwen copies what it is given. That makes it great at text and faithful photo edits, but it freezes a character in one pose.
This is how the two compare on the same 24 prompts. We started with why H3 was not an option, then gave both models our photo, a set of no-reference prompts, and a five-page picture book. Qwen won on text and faces, and MiniMax won the story.
Note
How we kept the numbers clean. Qwen ran with a guard that stops PyTorch before it spills into system RAM, and no image spilled. Face scores come from SFace and text scores from RapidOCR. OCR misread two story banners, one from each model, so we checked all ten banners by eye. Software: torch 2.11.0+cu128, diffusers 0.41.0.dev0 (git main), transformers 5.17.0 and torchao 0.18.0.