Categories: Image

HunyuanImage 3.0 (80B) running natively in ComfyUI on a single 12–24 GB GPU: text-to-image, editing and style transfer, ~30 s per image

I’ve been working on native ComfyUI support for Tencent’s HunyuanImage 3.0, the 80B mixture-of-experts image model (13B active per step). It isn’t a wrapper around Tencent’s pipeline: it uses the normal KSampler, the normal VAE Decode and ComfyUI’s own memory management, which streams the experts from system RAM so the model fits on one consumer GPU.

What’s in the gallery (all Instruct-Distil, 8 steps):

  • 4-bit vs int8, same prompt and seed. Times are the whole generation on an RTX 3090.
  • image editing with the 4-bit weights. The instruction is at the top of each image.
  • style transfer with the 4-bit weights: two input images, the photo and a style reference.

These are picked from a bigger run: 60 prompts × 2 formats, 60 edits and 14 styles, one seed each, no rerolls. Most of the edits worked; a few didn’t (snow that barely shows, a logo it wouldn’t remove, a “make it night” that stayed day). The text-to-image prompts come from popular prompt posts on X.

What you get

  • All three models: Instruct-Distil (8 steps, the one to start with), Instruct (50 steps) and Base
  • Text-to-image, image editing, and multi-image fusion with up to 3 input images (that’s how the style transfer works)
  • Optional prompt rewriting: the model expands your prompt first (slow, about 1 s per token)
  • Optional Spectrum speed-up: about 3.4× faster for the 50-step models
  • Ready-made weights: 4-bit W4A8 (44 GB), int8 (76 GB) and bf16 (150 GB, mostly for comparisons)
  • Example workflows for each model and task
  • Nothing to pip install

Speed (Instruct-Distil, about 1 megapixel):

GPU 4-bit W4A8 int8
RTX 4090 ~22–26 s ~47–49 s
RTX 3090 ~29 s ~54 s

It also runs with only 16 GB or 12 GB of VRAM (~30 s and ~32 s per image on a 4090 limited to that).

What you need

  • An NVIDIA GPU with 12 GB+
  • Lots of system RAM: ComfyUI held about 50 GB with the 4-bit file loaded. This is the real requirement, since the experts live in RAM and stream over PCIe every step.
  • A recent ComfyUI (late September 2026 or newer)

Links

Happy to answer questions. If something breaks, open an issue on GitHub with the traceback.

submitted by /u/LatentSpacer
[link] [comments]

AI Generated Robotic Content

Share
Published by
AI Generated Robotic Content
Tags: ai images

Recent Posts

Synchronous vs. Asynchronous Agent Execution: Architecture Patterns for Production

In this article, you will learn how synchronous and asynchronous execution patterns differ architecturally, and…

35 seconds ago

RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as…

43 seconds ago

Building a context-aware AI assistant on AgentCore and OpenClaw

Off-the-shelf AI assistants answer individual questions well, but they fall short on a different axis:…

1 min ago

I Found the 20 Best Prime Day Tech and Gadget Deals (October 2026)

Never pay full price. Bag yourself some Prime Day tech deals on our favorite WIRED-tested…

1 hour ago

Agentic AI turns simple language into self-guided X-ray scans of microelectronics

Science has increasingly used artificial intelligence (AI) as a kind of microscope—sorting data, analyzing images…

1 hour ago

From 3D layout to compositing: new LTX VFX tools

During VFX Week last week, we released seven open-weight capabilities for LTX-2.5, covering high-res editing,…

1 day ago