From Action Units to Emotions: A Two-Stage Fine-Tuning Framework with Decision-Level Fusion for Facial Expression Recognition

This article has 0 evaluations Published on
Read the full article Related papers
This article on Sciety

Abstract

Facial expressions serve as a primary non-verbal channel through which humans convey emo- tional states and behavioral intentions, holding significant value across diverse application scenarios including digital mental health, human-computer interaction, educational engage- ment analysis, and consumer behavior understanding. However, facial expression recognition (FER) technology remains constrained by practical challenges such as illumination variations, pose differences, and scarcity of annotated data, with existing models generally suffering from poor cross-domain adaptability, overfitting, and insufficient interpretability. To address these issues, this paper proposes a multi-modal large-model enhancement strategy based on two-stage fine-tuning. First, a question-answer paired facial action unit (AU) fine-tuning dataset is constructed from the CAS(ME)³ dataset, and the Qwen3-VL-32B- Instruct multi-modal large model is fine-tuned in the first stage using LoRA to reinforce its AU recognition capability. Subsequently, a second-stage fine-tuning is conducted on the FER dataset to guide the model in achieving emotion classification through AU-based reasoning. On this basis, decision-level joint prediction is further explored by combining the fine-tuned large model with the specialized small model OpenFace 3.0. Systematic experimental evaluations demonstrate that the proposed two-stage fine-tuned fusion model achieves significant superiority over both the non-fine-tuned Qwen3-VL and other mainstream deep models, and also substantially outperforms the standalone inference results of OpenFace 3.0, fully validating the effectiveness of the two-stage fine-tuning strategy in eliciting discriminative features for subtle facial geometric deformations and low-intensity expressions from large visual models. Meanwhile, the study also reveals that while pure large visual models possess formidable capabilities in high-level semantic understanding, they still exhibit inherent limitations in perceiving subtle facial deformations. The proposed method demonstrates clear efficacy in both improving prediction accuracy and enhancing interpretability, offering a novel technical pathway and empirical foundation for constructing highly robust and interpretable FER systems.

Related articles

Related articles are currently not available for this article.