llmman now runs Qwen-Image 2.1

llmman run qwen-image-2.1 turns a prompt into a 1024×1024 picture, straight from the model's Diffusers safetensors, on ggml. Done in collaboration with Unsloth.

llmman run learned to draw with LTX-2.3 earlier this month. As of #547 it also runs Qwen-Image 2.1, straight from the Diffusers safetensors the model is published as. We did this in collaboration with Unsloth.

Quickstart Example

Unsloth’s mascot is a sloth and llmman’s is a manatee, so the test prompt wrote itself.

llmman run qwen-image-2.1 "A manatee in a sunlit lagoon"

The first run pulls ai/qwen-image-2.1 from Docker Hub (33 GB) and saves the result as a PNG next to your shell. With no seed set, run picks a random one each time. The pictures below used fixed seeds, through the images endpoint of llmman serve, so they can be reproduced:

curl -s http://127.0.0.1:17434/v1/images/generations \
  -H 'content-type: application/json' \
  -d '{
    "model": "docker.io/ai/qwen-image-2.1:latest",
    "prompt": "A three-toed sloth resting on a lily pad and a manatee floating next to it in a calm turquoise lagoon, sunlit rainforest in the background, warm golden light, detailed digital illustration. Only animals, no people.",
    "size": "1024x1024", "steps": 40, "seed": 7,
    "response_format": "b64_json"
  }'
A three-toed sloth sitting on a lily pad in a turquoise lagoon, with a manatee floating beside it, generated by Qwen-Image 2.1

Same model, a different prompt and seed 11:

A sloth hanging upside down from a mossy branch over a sunlit lagoon, gazing down at a manatee that has surfaced below with its snout above the water, rainforest, golden hour light, painterly illustration. Only animals, no people.

A sloth hanging upside down from a mossy branch above a lagoon in golden light, looking down at a manatee surfacing below, generated by Qwen-Image 2.1

1024×1024, forty denoising steps, RGBA PNG. Those are the model’s defaults, so run needs no flags for them.

A note on the last sentence of each prompt. The first attempts at both pictures, without it, came with an uninvited man in the water. Adding it was enough for the first picture. It wasn’t for the second, where the sloth was asked to “pat the nose” of the manatee and a man turned up to do it anyway. Rewording the prompt so the sloth gazes down instead, with the same seed, got rid of him. Worth trying if it happens to you.

What happened underneath

ai/qwen-image-2.1 is a Diffusers pack: a root model_index.json next to text_encoder/, transformer/ and vae/. llmman used to hand every such pack to vLLM-Omni. It now reads _class_name from model_index.json, and when that names a pipeline llmman runs itself (QwenImage21Pipeline), the pack resolves as a diffusion model and is served in process, on the same ggml libraries as LTX-2.3. Every other Diffusers pack still goes to vLLM-Omni. The weights are the bf16 safetensors as published; nothing is converted first.

The pipeline is written in Rust on top of ggml graphs:

  • Text encoder. Qwen3-VL’s language model, the same code Cosmos3’s text encoder uses. The vision tower and the LM head are never loaded.
  • Transformer. 32 blocks and no patching, so a 1024×1024 image is 64×64 = 4,096 tokens. Prompt tokens are modulated as if at t = 0 and attend causally, so their keys and values are computed once per prompt and reused by all 40 steps.
  • Sampler. Flow-matching Euler with dynamic exponential shifting. There is no guidance by default; a cfg above 1 adds a negative-prompt pass.
  • VAE. The Wan VAE that Cosmos3 already used, generalized for Qwen-Image’s 2-D decoder, which emits four channels instead of three. That is why the PNG has an alpha channel. It now lives in mediagen/wan_vae.rs, shared by both models, and a convolution too large for its memory budget runs in bands of rows instead.

Checked against diffusers

The PR compares the pieces with the reference diffusers pipeline: the text encoder’s hidden states against fp16 (cosine > 0.9993), one transformer step (cosine 0.99993), the VAE decode (mean difference of 0.04 out of 255) and the sigmas (within 1e-6). The PR ran them on an M4 Max with 36 GB, which is also what made the pictures above.

Memory and speed

The bf16 text encoder is 14.1 GiB, the transformer 13.3 GiB and the VAE decoder 1 GiB. Nothing loads until it is needed, and when the two big ones don’t fit together they take turns. The daemon log for a picture reads: text encoder loaded, prompt encoded in a fraction of a second, transformer loaded, then the steps. On a 36 GB M4 Max that is what happens on every prompt: the text encoder loads again for the next one, and keeping Qwen-Image and LTX-2.3 loaded in one daemon runs out of memory.

On that machine the two pictures took 8 min 49 s and 10 min 35 s from a running daemon, loading the text encoder and the transformer included. The 40 steps averaged 12.6 s and 15.1 s each (from 8 s to 33 s), and the VAE decode took 11 s and 14 s. The Mac was swapping, with the model taking most of its memory, so treat those numbers as the slow end.

The full flag list is in the README. Questions are welcome at github.com/llmmanorg/llmman. Don’t be afraid to give the project a star or open a PR.