RUS  ENG
Full version
JOURNALS // Doklady Rossijskoj Akademii Nauk. Mathematika, Informatika, Processy Upravlenia // Archive

Dokl. RAN. Math. Inf. Proc. Upr., 2024 Volume 520, Number 2, Pages 154–168 (Mi danma597)

SPECIAL ISSUE: ARTIFICIAL INTELLIGENCE AND MACHINE LEARNING TECHNOLOGIES

CRAFT: Cultural Russian-oriented dataset adaptation for focused text-to-image generation

V. A. Vasilevab, V. S. Arkhipkina, J. D. Agafonovaac, T. V. Nikulinaa, È. O. Mironovaa, A. A. Shichaninaa, N. A. Gerasimenkoa, M. Shoytova, D. V. Dimitrovad

a Sber AI, Moscow, Russia
b Moscow Institute of Physics and Technology (National Research University), Dolgoprudny, Moscow Region
c St. Petersburg National Research University of Information Technologies, Mechanics and Optics, Saint-Petersburg, Russia
d Artificial Intelligence Research Institute, Moscow, Russia

Abstract: Despite the fact that popular text-to-image generation models cope well with international and general cultural queries, they have a significant knowledge gap regarding individual cultures. This is due to the content of existing large training datasets collected on the Internet, which are predominantly based on Western European or American popular culture. Meanwhile, the lack of cultural adaptation of the model can lead to incorrect results, a decrease in the generation quality, and the spread of stereotypes and offensive content. In an effort to address this issue, we examine the concept of cultural code and recognize the critical importance of its understanding by modern image generation models, an issue that has not been sufficiently addressed in the research community to date. We propose the methodology for collecting and processing the data necessary to form a dataset based on the cultural code, in particular the Russian one. We explore how the collected data affects the quality of generations in the national domain and analyze the effectiveness of our approach using the Kandinsky 3.1 text-to-image model. Human evaluation results demonstrate an increase in the level of awareness of Russian culture in the model.

Keywords: cultural code, Russian culture, dataset adaptation, text-to-image generation, fine-tuning, captioning, diffusion models, data engineering.

UDC: 004.89

Received: 20.09.2024
Accepted: 02.10.2024

DOI: 10.31857/S2686954324700474


 English version:
Doklady Mathematics, 2024, 110:suppl. 1, S137–S150

Bibliographic databases:


© Steklov Math. Inst. of RAS, 2025