Please use this identifier to cite or link to this item:
http://thuvienso.dut.udn.vn/handle/DUT/25177| Title: | Personalized text-to-image generation | Authors: | Phạm, Thị Thùy Dương | Advisor: | TS. Nguyễn, Quang Như Quỳnh GS. Wen, Nung Lie |
Issue Date: | 2025 | Publisher: | Trường Đại học Bách Khoa, Đại học Đà Nẵng | Abstract: | To address this, prior methods such as Textual Inversion [4] rely on random initialization, resulting in weak identity retention. Cross Initialization [5] improves this by using semantic information from the subject’s name, but remains limited to known identities within the training data. To overcome these limitations, I propose an improved token initialization strategy that integrates linguistic information with visual features extracted from a reference image. Specifically, I introduce fusion_cross_init-a novel method that combines both textual and visual information at the token initialization stage. This method incorporates a Fusion Transformer module to enhance the alignment between the visual features of the reference image and the semantic information related to the subject. Furthermore, it leverages CLIP’s [6] multi-modal encoders to synchronize semantic and visual signals from the outset, thereby significantly improving identity consistency in personalized text-to-image generation. The method is implemented on top of the pre-trained Stable Diffusion [3] model, ensuring high image quality while optimizing computational efficiency. |
Description: | DA.FA.25.177 93 Tr. |
URI: | http://thuvienso.dut.udn.vn/handle/DUT/25177 |
| Appears in Collections: | Khoa Khoa học Công nghệ tiên tiến - Tin học Công nghiệp |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| 2.DA.FA.177.Pham Thi Thuy Duong.pdf | Thuyết minh | 7.17 MB | Adobe PDF | ![]() View/Open |
CORE Recommender
Google ScholarTM
Check
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.
