Vui lòng dùng định danh này để trích dẫn hoặc liên kết đến tài liệu này: http://thuvienso.dut.udn.vn/handle/DUT/25177
Nhan đề: Personalized text-to-image generation
Tác giả: Phạm, Thị Thùy Dương
Người hướng dẫn: TS. Nguyễn, Quang Như Quỳnh
GS. Wen, Nung Lie
Năm xuất bản: 2025
Nhà xuất bản: Trường Đại học Bách Khoa, Đại học Đà Nẵng
Tóm tắt: 
To address this, prior methods such as Textual Inversion [4] rely on random initialization, resulting in weak identity retention. Cross Initialization [5] improves this by using semantic information from the subject’s name, but remains limited to known identities within the training data. To overcome these limitations, I propose an improved token initialization strategy that integrates linguistic information with visual features extracted from a reference image. Specifically, I introduce fusion_cross_init-a novel method that combines both textual and visual information at the token initialization stage. This method incorporates a Fusion Transformer module to enhance the alignment between the visual features of the reference image and the semantic information related to the subject. Furthermore, it leverages CLIP’s [6] multi-modal encoders to synchronize semantic and visual signals from the outset, thereby significantly improving identity consistency in personalized text-to-image generation. The method is implemented on top of the pre-trained Stable Diffusion [3] model, ensuring high image quality while optimizing computational efficiency.
Mô tả: 
DA.FA.25.177
93 Tr.
Định danh: http://thuvienso.dut.udn.vn/handle/DUT/25177
Bộ sưu tập: Khoa Khoa học Công nghệ tiên tiến - Tin học Công nghiệp

Các tập tin trong tài liệu này:
Tập tin Mô tả Kích thước Định dạng
2.DA.FA.177.Pham Thi Thuy Duong.pdfThuyết minh7.17 MBAdobe PDFHình minh họa
Xem/Tải về
Hiển thị đầy đủ biểu ghi tài liệu

Các đề xuất từ CORE

Google Scholar TM

Kiểm tra...


Khi sử dụng các tài liệu trong Hệ thống quản lý thông tin nghiên cứu phải tuân thủ Luật bản quyền.