Emotion Detection from Video and Audio and Text
DOI:
https://doi.org/10.64751/ajaccm.2026.v6.n3.736Abstract
The ability to recognize emotions in video, speech and text has become a critical area of study in the fields of artificial intelligence and human-computer interaction. With digital communication evolving into a multimodal experience, the ability to interpret human emotions in such contexts has become critical both for user experience and for better mental health diagnosis and for further development of affective computing technologies. This paper provides an overview of various approaches and architectures to emotion recognition from video, audio and textual information, and investigates the synergies and challenges of multimodal emotion recognition systems. The paper introduces and briefly reviews the importance of each modality in emotion detection first. Video analysis: Using computer vision to identify key facial features and gestures and body language as cues to the emotional state of a person. Audio Processing works on vocal characteristics like tone, pitch and speech patterns, using algorithms from signal processing and machine learning, to understand the emotions behind speech. Text analysis, on the other hand, utilizes natural language processing (NLP) methods to evaluate sentiment and emotional context in written words, taking both syntactic and semantic elements into account. The proposed systems can combine these three modalities to realize more accurate and robust emotion recognition, which is a reflection of human emotional expression.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.







