Papers
arxiv:2504.16915

DreamO: A Unified Framework for Image Customization

Published on Apr 23, 2025
· Submitted by
Yanze Wu
on Apr 24, 2025
Authors:
,
,
,
,
,

Abstract

DreamO, an image customization framework using diffusion transformers, supports various tasks and integrates multiple conditions through a unified approach.

Recently, extensive research on image customization (e.g., identity, subject, style, background, etc.) demonstrates strong customization capabilities in large-scale generative models. However, most approaches are designed for specific tasks, restricting their generalizability to combine different types of condition. Developing a unified framework for image customization remains an open challenge. In this paper, we present DreamO, an image customization framework designed to support a wide range of tasks while facilitating seamless integration of multiple conditions. Specifically, DreamO utilizes a diffusion transformer (DiT) framework to uniformly process input of different types. During training, we construct a large-scale training dataset that includes various customization tasks, and we introduce a feature routing constraint to facilitate the precise querying of relevant information from reference images. Additionally, we design a placeholder strategy that associates specific placeholders with conditions at particular positions, enabling control over the placement of conditions in the generated results. Moreover, we employ a progressive training strategy consisting of three stages: an initial stage focused on simple tasks with limited data to establish baseline consistency, a full-scale training stage to comprehensively enhance the customization capabilities, and a final quality alignment stage to correct quality biases introduced by low-quality data. Extensive experiments demonstrate that the proposed DreamO can effectively perform various image customization tasks with high quality and flexibly integrate different types of control conditions.

Community

Paper author Paper submitter

We propose DreamO, a unified image customization framework, which covers ID, IP, Tryon, and style tasks. DreamO performs well in character fidelity and multi-subject confusion. The model will be open sourced at https://github.com/bytedance/DreamO (within 1 week), please stay tuned.

·

My name is Victor, and I have been studying DreamO closely, particularly its virtual try-on capabilities. Thank you for making the project and research available—it has been very useful for my work.

I am currently working on extending the Try-On capability of DreamO, with a focus on improving garment fidelity, supporting higher-resolution garment references, and eventually exploring more complex garment combinations.

I would be extremely grateful if you could share any additional details about the training recipe used for DreamO’s Try-On task, beyond what is currently available in the paper/repository.

In particular, I am trying to understand:

The approximate number and composition of Try-On training samples
Garment/person image preprocessing and training resolutions
Whether garment references were always compressed/resized to 512 or trained at multiple resolutions
Number of training steps and effective batch size
Optimizer, learning rate, and LR schedule
Which parts of the model were trained/frozen
The sampling ratio between Try-On and DreamO’s other tasks
Loss functions and their relative weights for Try-On
Whether a Try-On-specific checkpoint, LoRA, training config, or training script exists
I completely understand if the datasets or full training code cannot be released. Even an approximate recipe, configuration file, or guidance on reproducing the Try-On training would be incredibly helpful.

My intention is to build upon the research rather than simply reproduce it, and I would of course properly acknowledge and cite DreamO in any resulting research or technical work.

Thank you very much for your time and for the work your team has put into DreamO. I would greatly appreciate any guidance you are able to provide.

Best regards,
Victor

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

My name is Victor, and I have been studying DreamO closely, particularly its virtual try-on capabilities. Thank you for making the project and research available—it has been very useful for my work.

I am currently working on extending the Try-On capability of DreamO, with a focus on improving garment fidelity, supporting higher-resolution garment references, and eventually exploring more complex garment combinations.

I would be extremely grateful if you could share any additional details about the training recipe used for DreamO’s Try-On task, beyond what is currently available in the paper/repository.

In particular, I am trying to understand:

The approximate number and composition of Try-On training samples
Garment/person image preprocessing and training resolutions
Whether garment references were always compressed/resized to 512 or trained at multiple resolutions
Number of training steps and effective batch size
Optimizer, learning rate, and LR schedule
Which parts of the model were trained/frozen
The sampling ratio between Try-On and DreamO’s other tasks
Loss functions and their relative weights for Try-On
Whether a Try-On-specific checkpoint, LoRA, training config, or training script exists
I completely understand if the datasets or full training code cannot be released. Even an approximate recipe, configuration file, or guidance on reproducing the Try-On training would be incredibly helpful.

My intention is to build upon the research rather than simply reproduce it, and I would of course properly acknowledge and cite DreamO in any resulting research or technical work.

Thank you very much for your time and for the work your team has put into DreamO. I would greatly appreciate any guidance you are able to provide.

Best regards,
Victor

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2504.16915
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 2

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2504.16915 in a dataset README.md to link it from this page.

Spaces citing this paper 18

Browse 18 spaces citing this paper

Collections including this paper 9