Please use this identifier to cite or link to this item: http://thuvienso.dut.udn.vn/handle/DUT/25177
Title: Personalized text-to-image generation
Authors: Phạm, Thị Thùy Dương
Advisor: TS. Nguyễn, Quang Như Quỳnh
GS. Wen, Nung Lie
Issue Date: 2025
Publisher: Trường Đại học Bách Khoa, Đại học Đà Nẵng
Abstract: 
To address this, prior methods such as Textual Inversion [4] rely on random initialization, resulting in weak identity retention. Cross Initialization [5] improves this by using semantic information from the subject’s name, but remains limited to known identities within the training data. To overcome these limitations, I propose an improved token initialization strategy that integrates linguistic information with visual features extracted from a reference image. Specifically, I introduce fusion_cross_init-a novel method that combines both textual and visual information at the token initialization stage. This method incorporates a Fusion Transformer module to enhance the alignment between the visual features of the reference image and the semantic information related to the subject. Furthermore, it leverages CLIP’s [6] multi-modal encoders to synchronize semantic and visual signals from the outset, thereby significantly improving identity consistency in personalized text-to-image generation. The method is implemented on top of the pre-trained Stable Diffusion [3] model, ensuring high image quality while optimizing computational efficiency.
Description: 
DA.FA.25.177
93 Tr.
URI: http://thuvienso.dut.udn.vn/handle/DUT/25177
Appears in Collections:Khoa Khoa học Công nghệ tiên tiến - Tin học Công nghiệp

Files in This Item:
File Description SizeFormat
2.DA.FA.177.Pham Thi Thuy Duong.pdfThuyết minh7.17 MBAdobe PDFThumbnail
View/Open
Show full item record

CORE Recommender

Google ScholarTM

Check


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.