Abstract: Against the backdrop of cultural digitalization, aiming to address the issues of insufficient interaction in the digital communication of intangible cultural heritage (ICH) operas, inadequate fusion of multimodal information, and low level of in-depth cognition of Huangmei Opera among young groups, this study takes the living digital inheritance of Huangmei Opera as the objective, constructs a framework of Huangmei Opera intelligent digital human system based on multimodal interaction, and conducts technical verification with Feng Suzhen, a classic character from Huangmei Opera, as a case study. The system integrates three core modalities, namely speech, vision, and motion: in the speech dimension, it achieves dialect recognition and aria analysis; in the vision dimension, it integrates a high-resolution network (HRNet) to ensure the accuracy of face and facial expression recognition; in the motion dimension, it relies on optical-inertial capture and dynamic time warping (DTW) algorithm to match stylized movements such as "orchid fingers". Furthermore, it realizes multimodal information fusion based on the attention mechanism, supporting voice, gesture, and immersive interaction, and forms a technical path characterized by technology empowerment, multimodal fusion, and scenario-driven application. This study provides technical reference for the digitalization of cultural heritage and holds theoretical and application value for advancing the research on multimodal human-computer interaction.
Keywords: intelligent digital human; multimodal interaction; Huangmei Opera; speech recognition; face detection; motion capture