Vui lòng dùng định danh này để trích dẫn hoặc liên kết đến tài liệu này:
http://thuvienso.dut.udn.vn/handle/DUT/25177| Nhan đề: | Personalized text-to-image generation | Tác giả: | Phạm, Thị Thùy Dương | Người hướng dẫn: | TS. Nguyễn, Quang Như Quỳnh GS. Wen, Nung Lie |
Năm xuất bản: | 2025 | Nhà xuất bản: | Trường Đại học Bách Khoa, Đại học Đà Nẵng | Tóm tắt: | To address this, prior methods such as Textual Inversion [4] rely on random initialization, resulting in weak identity retention. Cross Initialization [5] improves this by using semantic information from the subject’s name, but remains limited to known identities within the training data. To overcome these limitations, I propose an improved token initialization strategy that integrates linguistic information with visual features extracted from a reference image. Specifically, I introduce fusion_cross_init-a novel method that combines both textual and visual information at the token initialization stage. This method incorporates a Fusion Transformer module to enhance the alignment between the visual features of the reference image and the semantic information related to the subject. Furthermore, it leverages CLIP’s [6] multi-modal encoders to synchronize semantic and visual signals from the outset, thereby significantly improving identity consistency in personalized text-to-image generation. The method is implemented on top of the pre-trained Stable Diffusion [3] model, ensuring high image quality while optimizing computational efficiency. |
Mô tả: | DA.FA.25.177 93 Tr. |
Định danh: | http://thuvienso.dut.udn.vn/handle/DUT/25177 |
| Bộ sưu tập: | Khoa Khoa học Công nghệ tiên tiến - Tin học Công nghiệp |
Các tập tin trong tài liệu này:
| Tập tin | Mô tả | Kích thước | Định dạng | |
|---|---|---|---|---|
| 2.DA.FA.177.Pham Thi Thuy Duong.pdf | Thuyết minh | 7.17 MB | Adobe PDF | ![]() Xem/Tải về |
Các đề xuất từ CORE
Google Scholar TM
Kiểm tra...
Khi sử dụng các tài liệu trong Hệ thống quản lý thông tin nghiên cứu phải tuân thủ Luật bản quyền.
