Categories: Image

Bytedance release the full safetensor model for UMO – Multi-Identity Consistency for Image Customization . Obligatory beg for a ComfyUI node πŸ™πŸ™

https://huggingface.co/bytedance-research/UMO
https://arxiv.org/pdf/2509.06818

Bytedance have released 3 days ago their image editing/creation model UMO. From their huggingface description:

Recent advancements in image customization exhibit a wide range of application prospects due to stronger customization capabilities. However, since we humans are more sensitive to faces, a significant challenge remains in preserving consistent identity while avoiding identity confusion with multi-reference images, limiting the identity scalability of customization models. To address this, we present UMO, a Unified Multi-identity Optimization framework, designed to maintain high-fidelity identity preservation and alleviate identity confusion with scalability. With β€œmulti-to-multi matching” paradigm, UMO reformulates multi-identity generation as a global assignment optimization problem and unleashes multi-identity consistency for existing image customization methods generally through reinforcement learning on diffusion models. To facilitate the training of UMO, we develop a scalable customization dataset with multi-reference images, consisting of both synthesised and real parts. Additionally, we propose a new metric to measure identity confusion. Extensive experiments demonstrate that UMO not only improves identity consistency significantly, but also reduces identity confusion on several image customization methods, setting a new state-of-the-art among open-source methods along the dimension of identity preserving.

submitted by /u/AgeNo5351
[link] [comments]

AI Generated Robotic Content

Share
Published by
AI Generated Robotic Content
Tags: ai images

Recent Posts

Qwen Image 2.1 is an EDITING Beast

Just a few tests with the new Qwen Image 2.1. Although it is not a…

10 hours ago

RAG vs. Fine-Tuning for Domain Adaptation: When to Use Which

In this article, you will learn the mechanical difference between retrieval-augmented generation and fine-tuning, when…

10 hours ago

How to Guide Your Language Flow

We introduce a new method to guide flow matching models. Our approach, which we call…

10 hours ago

From portal-hopping to instant answers: HEMA’s journey with MCP and Amazon Bedrock

This post is co-written with Mauro Rallo and Patrick van der Plas from HEMA. When…

10 hours ago

Meta Pinky Promises Its Smart Glasses Will Be Private Soon

The company is bringing its Private Processing encryption service to its much-maligned smart glasses.

11 hours ago

When AI disagrees, people change how they perceive the technologyβ€”not their original view

When an AI chatbot agrees with our reasoning in resolving a social dilemma, we may…

11 hours ago