tech_surveillance8648 wordsRead on Arc Codex

Bridging the silos in affective AI: a critical perspective from data to society

Abstract Affective computing has grown into a field that spans emotion theory, multimodal data, computational modeling, human–AI interaction, social applications, and ethical evaluation, yet most existing surveys address only one or two of these layers at a time and fail to capture how the layers depend on one another. Rather than offering another comprehensive survey, this position paper reorganizes the field around a six-layer pipeline and reads the literature as a chain of dependencies in which upstream design choices propagate into downstream evaluation, deployment, and governance. Our central diagnosis is that four silo-bridge patterns recur across the field: disconnections at the theory/data (Pattern 1), modeling/interaction (Pattern 2), technology/ethics (Pattern 3), and data/social-application (Pattern 4) boundaries. These are not incidental implementation problems but recurring failure modes rooted in the pipeline structure itself, forming a cascade in which upstream patterns structurally enable downstream ones. We argue that this dependency chain is not merely a technical liability but a socio-philosophical one: as upstream inconsistencies propagate downstream, they threaten the individual’s authority to interpret their own emotions (emotional agency), pull affective expression toward machine-legible norms (algorithmic conformity), and relocate the power to define and act on emotion from persons to the institutions that operate these systems (institutional power). As a prescriptive program, we present five integrated design criteria: theory-explicit affective representation (DC1), affective reasoning with bounded intervention (DC2), longitudinal interaction-in-the-loop evaluation (DC3), deployment-specific affective accountability (DC4), and user-retained interpretive authority (DC5). The five criteria are designed as an interlocking response to the cascade rather than as independent recommendations. We situate affective AI not as a mere recognition or generation technology but as a sociotechnical system that can intervene in human emotional interpretation, interaction, and social decision-making, and call upon future research to move toward systems that are more coherent, more responsible, and more respectful of human emotional agency. Similar content being viewed by others Data availability No datasets were generated or analyzed during the current study. Materials availability Not applicable. Code availability Not applicable. References Ahir A, Gohokar V (2019) Driver inattention monitoring system: a review. In: Proceedings of the 2019 international conference on innovative trends and advances in engineering and technology (ICITAET), pp 188–194. https://doi.org/10.1109/ICITAET47105.2019.9170249 Andalibi N, Ingber AS (2025) Public perceptions about emotion ai use across contexts in the united states. In: Proceedings of the 2025 CHI conference on human factors in computing systems, pp 1–16. https://doi.org/10.1145/3706598.3713501 Averill FR (1980) A constructivist view of emotion. In: Theories of emotion, pp 305–339. https://doi.org/10.1016/B978-0-12-558701-3.50018-1 Barker D, Tippireddy MKR, Farhan A, Ahmed B (2025) Ethical considerations in emotion recognition research. Psychol Int 7(2):43. https://doi.org/10.3390/psycholint7020043 Barrett LF (2006) Are emotions natural kinds? Perspect Psychol Sci 1(1):28–58. https://doi.org/10.1111/j.1745-6916.2006.00003.x Barrett LF (2017) The theory of constructed emotion: an active inference account of interoception and categorization. Soc Cognit Affect Neurosci 12(1):1–23. https://doi.org/10.1093/scan/nsw154 Barrett LF, Adolphs R, Marsella S, Martinez AM, Pollak SD (2019) Emotional expressions reconsidered: challenges to inferring emotion from human facial movements. Psychol Sci Public Interest 20(1):1–68. https://doi.org/10.1177/1529100619832930 Bender EM, Friedman B (2018) Data statements for natural language processing: toward mitigating system bias and enabling better science. Trans Assoc Comput Linguist 6:587–604. https://doi.org/10.1162/tacl_a_00041 Benjamin R (2019) Race after technology: abolitionist tools for the new jim code. Polity Bickmore TW, Picard RW (2005) Establishing and maintaining long-term human-computer relationships. ACM Trans Comput Hum Interact 12(2):293–327. https://doi.org/10.1145/1067860.1067867 Bohus D, Horvitz E (2009) Models for multiparty engagement in open-world dialog. In: Proceedings of SIGDIAL 2009, pp 225–234 Bostan LAM, Klinger R (2018) An analysis of annotated corpora for emotion classification in text. In: Proceedings of the 27th international conference on computational linguistics, pp 2104–2119 Breazeal C (2003) Emotion and sociable humanoid robots. Int J Hum Comput Stud 59(1–2):119–155. https://doi.org/10.1016/S1071-5819(03)00018-1 Breazeal C (2004) Designing sociable robots. MIT Press. https://doi.org/10.7551/mitpress/2376.001.0001 Broekens J, Hilpert B, Verberne S, Baraka K, Gebhard P, Plaat A (2023) Fine-grained affective processing capabilities emerging from large language models. In: Proceedings of the international conference on affective computing and intelligent interaction, pp 1–8. https://doi.org/10.1109/ACII59096.2023.10388177 Buolamwini J, Gebru T (2018) Gender shades: intersectional accuracy disparities in commercial gender classification. In: Proceedings of the 1st conference on fairness, accountability and transparency. Proceedings of machine learning research, vol 81, pp 77–91 Burkhardt F, Paeschke A, Rolfes M, Sendlmeier WF, Weiss B (2005) A database of German emotional speech. In: Proceedings of the annual conference of the international speech communication association, vol 5. ISCA, Lisbon, pp 1517–1520. https://doi.org/10.21437/Interspeech.2005-446 Busso C, Bulut M, Lee C-C, Kazemzadeh A, Mower E, Kim S, Chang JN, Lee S, Narayanan SS (2008) IEMOCAP: interactive emotional dyadic motion capture database. Lang Resour Eval 42:335–359. https://doi.org/10.1007/S10579-008-9076-6 Calvo RA, D’Mello S (2010) Affect detection: an interdisciplinary review of models, methods, and their applications. IEEE Trans Affect Comput. https://doi.org/10.1109/T-AFFC.2010.1 Cambria E (2016) Affective computing and sentiment analysis. IEEE Intell Syst 31(2):102–107. https://doi.org/10.1109/MIS.2016.31 Cao H, Cooper DG, Keutmann MK, Gur RC, Nenkova A, Verma R (2014) CREMA-D: crowd-sourced emotional multimodal actors dataset. IEEE Trans Affect Comput 5(4):377–390. https://doi.org/10.1109/TAFFC.2014.2336244 Cassell J, Sullivan J, Prevost S, Churchill EF (2000) Embodied conversational agents. MIT Press. https://doi.org/10.7551/mitpress/2697.001.0001 Chalmers DJ (1995) Facing up to the problem of consciousness. J Conscious Stud 2(3):200–219. https://doi.org/10.1093/acprof:oso/9780195311105.003.0001 Chandra NA, Murtfeldt R, Qiu L, Karmakar A, Lee H, Tanumihardja E, Farhat K, Caffee B, Paik S, Lee C, Choi J, Kim A, Etzioni O (2025) Deepfake-Eval-2024: a multi-modal in-the-wild benchmark of deepfakes circulated in 2024. https://doi.org/10.48550/arXiv.2503.02857. arXiv: 2503.02857 Chavan V, Cenaj A, Shen S, Bar A, Binwani S, Del Becaro T, Funk M, Greschner L, Hung R, Klein S, Kleiner R, Krause S, Olbrych S, Parmar V, Sarafraz J, Soroko D, Withanage Don D, Zhou C, Vu HTD, Semnani P, Weinhardt D, Andre E, Kr ger J, Fresquet X (2025) Feeling machines: ethics, culture, and the rise of emotional ai. https://doi.org/10.48550/arXiv.2506.12437. arXiv:2506.12437 Cheng Z, Cheng Z-Q, He J-Y, Sun J, Wang K, Lin Y, Lian Z, Peng X, Hauptmann AG (2024) Emotion-LLaMA: multimodal emotion recognition and reasoning with instruction tuning. In: Proceedings of the 38th international conference on neural information processing systems, pp 110805–110853. https://doi.org/10.5555/3737916.3741434 Damasio A (1994) Descartes’ error: emotion, reason, and the human brain. Putnam Davidson RJ (2004) What does the prefrontal cortex“do’’in affect: perspectives on frontal EEG asymmetry research. Biol Psychol 67(1–2):219–234. https://doi.org/10.1016/j.biopsycho.2004.03.008 Davis MH (1983) Measuring individual differences in empathy: evidence for a multidimensional approach. J Pers Soc Psychol 44(1):113–126. https://doi.org/10.1037/0022-3514.44.1.113 Davison AK, Lansley C, Costen N, Tan K, Yap MH (2018) Samm: a spontaneous micro-facial movement dataset. IEEE Trans Affect Comput 9(1):116–129. https://doi.org/10.1109/TAFFC.2016.2573832 De Choudhury M, Gamon M, Counts S, Horvitz E (2013) Predicting depression via social media. In: Proceedings of the international AAAI conference on web and social media, vol 7, pp 128–137. https://doi.org/10.1609/icwsm.v7i1.14432 Decety J, Jackson PL (2004) The functional architecture of human empathy. Behav Cogn Neurosci Rev 3(2):71–100. https://doi.org/10.1177/1534582304267 Défossez A, Mazaré L, Orsini M, Royer A, Pérez P, Jégou H, Grave E, Zeghidour N (2024) Moshi: a speech-text foundation model for real-time dialogue. https://doi.org/10.48550/arXiv.2410.00037. arXiv:2410.00037 Demszky D, Movshovitz-Attias D, Ko J, Cowen A, Nemade G, Ravi S (2020) GoEmotions: a dataset of fine-grained emotions. In: Proceedings of the 58th annual meeting of the association for computational linguistics, pp 4040–4054. https://doi.org/10.18653/v1/2020.acl-main.372 DeVault D, Artstein R, Benn G, Dey T, Fast E, Gainer A, Georgila K, Gratch J, Hartholt A, Lhommet M, Lucas G, Marsella S, Morbini F, Nazarian A, Scherer S, Stratou G, Suri A, Traum D, Wood R, Xu Y, Rizzo A, Morency L-P (2014) Simsensei kiosk: a virtual human interviewer for healthcare decision support. In: Proceedings of the 2014 international conference on autonomous agents and multi-agent systems, pp 1061–1068. https://doi.org/10.5555/2615731.2617415 Devlin J, Chang M-W, Lee K, Toutanova K (2019) BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, pp 4171–4186. https://doi.org/10.18653/v1/N19-1423 Dinan E, Roller S, Shuster K, Fan A, Auli M, Weston J (2019) Wizard of wikipedia: knowledge-powered conversational agents. In: Proceedings of the international conference on learning representations. https://doi.org/10.48550/arXiv.1811.01241 D’Mello S (2013) A selective meta-analysis on the relative incidence of discrete affective states during learning with technology. J Educ Psychol 105(4):1082–1099. https://doi.org/10.1037/a0032674 D’Mello SK, Graesser A (2010) Multimodal semi-automated affect detection from conversational cues, gross body language, and facial features. User Model User-Adap Interact 20:147–187. https://doi.org/10.1007/s11257-010-9074-4 Duffy BR (2003) Anthropomorphism and the social robot. Robot Auton Syst 42(3–4):177–190. https://doi.org/10.1016/S0921-8890(02)00374-3 Ekman P (1992) An argument for basic emotions. Cognit Emot 6(3–4):169–200. https://doi.org/10.1080/02699939208411068 Ekman P, Friesen WV (1978) Facial action coding system: a technique for the measurement of facial movement. Consulting Psychologists Press. https://doi.org/10.1037/t27734-000 Elfenbein HA, Ambady N (2002) On the universality and cultural specificity of emotion recognition: a meta-analysis. Psychol Bull 128(2):203–235. https://doi.org/10.1037/0033-2909.128.2.203 Emmelkamp PMG, Meyerbrker K (2021) Virtual reality therapy in mental health. Annu Rev Clin Psychol 17:495–519. https://doi.org/10.1146/annurev-clinpsy-081219-115923 Fang CM, Liu AR, Danry V, Lee E, Chan SWT, Pataranutaporn P, Maes P, Phang J, Lampe M, Ahmad L, Agarwal S (2024) How AI and human behaviors shape psychosocial effects of chatbot use: a longitudinal randomized controlled study. https://doi.org/10.48550/arXiv.2503.17473. arXiv: 2503.17473 Fischer T, Biemann C (2024) Exploring large language models for qualitative data analysis. In: Proceedings of the 4th international conference on natural language processing for digital humanities, pp 423–437. https://aclanthology.org/2024.nlp4dh-1.41/ Fitzpatrick KK, Darcy A, Vierhile M (2017) Delivering cognitive behavior therapy to young adults with symptoms of depression and anxiety using a fully automated conversational agent (Woebot): a randomized controlled trial. JMIR Mental Health 4(2):19. https://doi.org/10.2196/mental.7785 Foucault M (1977) Discipline and punish: the birth of the prison. Vintage Books Gabriel S, Puri I, Xu X, Malgaroli M, Ghassemi M (2024) Can AI relate: testing large language model response for mental health support. In: Findings of the association for computational linguistics: EMNLP 2024, pp 2206–2221. https://doi.org/10.18653/v1/2024.findings-emnlp.120 Gebru T, Morgenstern J, Vecchione B, Vaughan JW, Wallach H, Daum H III, Crawford K (2021) Datasheets for datasets. Commun ACM 64(12):86–92 Geirhos R, Jacobsen J-H, Michaelis C, Zemel R, Brendel W, Bethge M, Wichmann FA (2020) Shortcut learning in deep neural networks. Nat Mach Intell 2:665–673. https://doi.org/10.1038/s42256-020-00257-z Gergen KJ (1985) The social constructionist movement in modern psychology. Am Psychol 40(3):266–275. https://doi.org/10.1037/0003-066X.40.3.266 Ghosal D, Majumder N, Poria S, Chhaya N, Gelbukh A (2019) DialogueGCN: a graph convolutional neural network for emotion recognition in conversation. In: Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing, pp 154–164. https://doi.org/10.18653/v1/D19-1015 Gilardi F, Alizadeh M, Kubli M (2023) ChatGPT outperforms crowd workers for text-annotation tasks. Proc Natl Acad Sci. https://doi.org/10.1073/pnas.2305016120 Goodfellow IJ, Erhan D, Carrier PL, Courville A, Mirza M, Hamner B, Cukierski W, Tang Y, Thaler D, Lee D-H, Zhou Y, Ramaiah C, Feng F, Li R, Wang X, Athanasakis D, Shawe-Taylor J, Milakov M, Park J, Ionescu R, Popescu M, Grozea C, Bergstra J, Xie J, Romaszko L, Xu B, Chuang Z, Bengio Y (2013) Challenges in representation learning: a report on three machine learning contests. In: Proceedings of the international conference on neural information processing, pp 117–124. https://doi.org/10.1007/978-3-642-42051-1_16 Gross JJ (1998) The emerging field of emotion regulation: an integrative review. Rev Gen Psychol 2(3):271–299. https://doi.org/10.1037/1089-2680.2.3.271 Gross JJ (2015) Emotion regulation: current status and future prospects. Psychol Inq 26(1):1–26. https://doi.org/10.1080/1047840X.2014.940781 Gross JJ, Levenson RW (1995) Emotion elicitation using films. Cognit Emot 9(1):87–108. https://doi.org/10.1080/02699939508408966 Guo Z, Lai A, Thygesen JH, Farrington J, Keen T, Li K (2024) Large language models for mental health applications: systematic review. JMIR Mental Health 11:57400. https://doi.org/10.2196/57400 Hagendorff T (2020) The ethics of AI ethics: an evaluation of guidelines. Mind Mach 30(1):99–120. https://doi.org/10.1007/s11023-020-09517-8 Halkiopoulos C, Gkintoni E, Aroutzidis A, Antonopoulou H (2025) Advances in neuroimaging and deep learning for emotion detection: a systematic review of cognitive neuroscience and algorithmic innovations. Diagnostics 15(4):456. https://doi.org/10.3390/diagnostics15040456 He H, Garcia EA (2009) Learning from imbalanced data. IEEE Trans Knowl Data Eng 21(9):1263–1284. https://doi.org/10.1109/TKDE.2008.239 Hegde K, Jayalath H (2025) Emotions in the loop: a survey of affective computing for emotional support. https://doi.org/10.48550/arXiv.2505.01542. arXiv:2505.01542 Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Comput 9(8):1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735 Hochschild AR (1979) The managed heart: commercialization of human feeling. University of California Press Huang C-ZA, Vaswani A, Uszkoreit J, Shazeer N, Simon I, Hawthorne C, Dai AM, Hoffman MD, Dinculescu M, Eck D (2018) Music transformer. https://doi.org/10.48550/arXiv.1809.04281. arXiv:1809.04281 Huang X, Hong X, Mao Q, Zheng W, Dhall A (2024) A survey on deep learning for group-level emotion recognition. IEEE Trans Comput Soc Syst 13(2):2475–2500. https://doi.org/10.1109/TCSS.2025.3638859 Hutto C, Gilbert E (2014) VADER: a parsimonious rule-based model for sentiment analysis of social media text. In: Proceedings of the international AAAI conference on web and social media, vol 8, pp 216–225 Imel ZE, Caperton DD, Tanana M, Atkins DC (2017) Technology-enhanced human interaction in psychotherapy. J Couns Psychol 64(4):385–393. https://doi.org/10.1037/cou0000213 Indrasiri PL, Kashyap B, Kolambahewage C, Nakisa B, Ijaz K, Pathirana PN (2024) VR based emotion recognition using deep multimodal fusion with biosignals across multiple anatomical domains. https://doi.org/10.48550/arXiv.2412.02283. arXiv:2412.02283 Inkster B, Sarda S, Subramanian V (2018) Delivering cognitive behavior therapy to young adults with symptoms of depression and anxiety using a fully automated conversational agent (Woebot): a randomized controlled trial. JMIR Mental Health 4(2):19. https://doi.org/10.2196/12106 Inoshita K (2024) Sentiment analysis of Japanese twitter users regarding the Ukraine-Russia War and its implications for security policy. In: 2024 11th international conference on information technology, computer, and electrical engineering, pp 338–343. https://doi.org/10.1109/ICITACEE62763.2024.10762783 Inoshita K, Harada R (2026) PersonaGen: persona-based synthetic data generation using multi-stage conditioning with large language models for emotion recognition. Int J Act Behav Comput 1:1–18. https://doi.org/10.60401/ijabc.133 Inoshita K, Mizuno S (2026) World model inspired sarcasm reasoning with large language model agents. https://doi.org/10.48550/arXiv.2512.24329 Inoshita K, Tomisu H, Zhou X, Kawai A, Yada K (2026) KDDA: a knowledge-driven domain and diversity alignment framework for emotion data generation with large language models. Int J Act Behav Comput 1:1–24 Inoshita K, Zhou X, Kawai A, Yada K (2026b) LLMs capture emotion labels, not emotion uncertainty: distributional analysis and calibration of human-LLM judgment gaps. https://doi.org/10.48550/arXiv.2604.27345 Irfan B, Kuoppamäki S, Skantze G (2024) Recommendations for designing conversational companion robots with older adults through foundation models. Front Robot AI. https://doi.org/10.3389/frobt.2024.1363713 Jiang W, Windl M, Tag B, Sarsenbayeva Z, Mayer S (2024) An immersive and interactive vr dataset to elicit emotions. IEEE Trans Visual Comput Graph 30(11):7343–7353. https://doi.org/10.1109/TVCG.2024.3456202 Jin Y, Liu J, Li P, Wang B, Yan Y, Zhang H, Ni C, Wang J, Li Y, Bu Y, Wang Y (2025) The applications of large language models in mental health: scoping review. J Med Internet Res 27:69284. https://doi.org/10.2196/69284 Jobin A, Ienca M, Vayena E (2019) The global landscape of AI ethics guidelines. Nat Mach Intell 1:389–399. https://doi.org/10.1038/s42256-019-0088-2 Johnson KT, Narain J, Thomas Q, Maes P, Picard RW (2023) Recanvo: a database of real-world communicative and affective nonverbal vocalizations. Sci Data. https://doi.org/10.1038/s41597-023-02405-7 Ju Z, Wang Y, Shen K, Tan X, Xin D, Yang D, Liu Y, Leng Y, Song K, Tang S, Wu Z, Qin T, Li X, Ye W, Zhang S, Bian J, He L, Li J, Zhao S (2024) NaturalSpeech 3: zero-shot speech synthesis with factorized codec and diffusion models. In: Proceedings of the 41st international conference on machine learning, pp 22605–22623. https://doi.org/10.5555/3692070.3692979 Kaplan AD, Kessler TT, Brill JC, Hancock PA (2021) Trust in artificial intelligence: meta-analytic findings. J Hum Factors Ergon Soc. https://doi.org/10.1177/00187208211013988 Khalil HA, Hammad SA, Munim HEAE, Maged SA (2023) Low-cost driver monitoring system using deep learning. IEEE Access 13:14151–14164. https://doi.org/10.1109/ACCESS.2025.3530296 Khan UA, Xu Q, Liu Y, Lagstedt A, Alamäki A, Kauttonen J (2024) Exploring contactless techniques in multimodal emotion recognition: insights into diverse applications, challenges, solutions, and prospects. Multimed Syst. https://doi.org/10.1007/s00530-024-01302-2 Kim Y (2014) Convolutional neural networks for sentence classification. In: Proceedings of the 2014 conference on empirical methods in natural language processing, pp 1746–1751. https://doi.org/10.3115/v1/D14-1181 Kim RS (2026) Formal and computational foundations for implementing affective sovereignty in emotion AI systems. Discov Artif Intell 6(235):501–507. https://doi.org/10.1007/s44163-026-01000-0 Koelstra S, Mühl C, Soleymani M, Lee J-S, Yazdani A, Ebrahimi T, Pun T, Nijholt A, Patras I (2012) DEAP: a database for emotion analysis using physiological signals. IEEE Trans Affect Comput 3(1):18–31. https://doi.org/10.1109/T-AFFC.2011.15 Kumar MJD, Rao MS, Narendra KC (2025) Multimodal emotion recognition: a comprehensive survey of datasets, methods, and applications. IEEE Access 13:201067–201097. https://doi.org/10.1109/ACCESS.2025.3636186 Lang PJ, Bradley MM, Cuthbert BN (2005) International affective picture system (iaps): affective ratings of pictures and instruction manual. Technical Report Technical Report A-6. University of Florida, Gainesville Lazarus RS (1991) Emotion and adaptation. Oxford University Press LeDoux JE (1998) The emotional brain: the mysterious underpinnings of emotional life. Weidenfeld & Nicolson Lei S, Dong G, Wang X, Wang K, Qiao R, Wang S (2024a) InstructERC: reforming emotion recognition in conversation with multi-task retrieval-augmented large language models. arXiv:2309.11911. https://doi.org/10.48550/arXiv.2309.11911 Lei S, Zhou Y, Tang B, Lam MWY, Liu F, Liu H, Wu J, Kang S, Wu Z, Meng H (2024b) SongCreator: lyrics-based Universal Song Generation. In: Proceedings of the 38th international conference on neural information processing systems, pp 80107–80140. https://doi.org/10.52202/079017-2546 Leite I, Castellano G, Pereira A, Martinho C, Paiva A (2014) Empathic robots for long-term interaction. Int J Soc Robot 6:329–341. https://doi.org/10.1007/s12369-014-0227-1 Li Y, Su H, Shen X, Li W, Cao Z, Niu S (2017) DailyDialog: a manually labelled multi-turn dialogue dataset. In: Proceedings of the 8th international joint conference on natural language processing, pp 986–995 Lian Z, Sun H, Chen L, Sun H, Sun L, Ren Y, Cheng Z, Liu B, Liu R, Peng X, Yi J, Tao J (2025) AffectGPT: a new dataset, model, and benchmark for emotion understanding with multimodal large language models. In: Proceedings of the 42nd international conference on machine learning, pp 36993–37014 Lin Z, Madotto A, Shin J, Xu P, Fung P (2019) MoEL: mixture of empathetic listeners. In: Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing, pp 121–132. https://doi.org/10.18653/v1/D19-1012 Lin H, Czarnek G, Lewis B, White JP, Berinsky AJ, Costello T, Pennycook G, Rand DG (2025) Persuading voters using human-artificial intelligence dialogues. Nature 648:394–401. https://doi.org/10.1038/s41586-025-09771-9 Liu Z, Shen Y, Lakshminarasimhan VB, Liang PP, Zadeh AB, Morency L-P (2018) Efficient low-rank multimodal fusion with modality-specific factors. In: Proceedings of the 56th annual meeting of the association for computational linguistics, pp 2247–2256. https://doi.org/10.18653/v1/P18-1209 Liu S, Zheng C, Demasi O, Sabour S, Li Y, Yu Z, Jiang Y, Huang M (2021) Towards emotional support dialog systems. In: Proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing, pp 3469–3483. https://doi.org/10.18653/v1/2021.acl-long.269 Liu Y, Dai W, Feng C, Wang W, Yin G, Zeng J, Shan S (2022) MAFW: a large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild. In: Proceedings of the 30th ACM international conference on multimedia. Association for Computing Machinery, pp 24–32. https://doi.org/10.1145/3503161.3548190 Livingstone SR, Russo FA (2018) The Ryerson audio-visual database of emotional speech and song (ravdess): a dynamic, multimodal set of facial and vocal expressions in north American English. PLoS ONE 13(5):0196391. https://doi.org/10.1371/journal.pone.0196391 Lotfian R, Busso C (2019) Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings. IEEE Trans Affect Comput 10(4):471–483. https://doi.org/10.1109/TAFFC.2017.2736999 Lubold N, Walker E, Pon-Barry H, Ogan A (2018) Automated pitch convergence improves learning in a social, teachable robot for middle school mathematics. In: Proceedings of the 19th international conference on artificial intelligence in education. Lecture notes in computer science, vol 10947, pp 282–296. https://doi.org/10.1007/978-3-319-93843-1_21 Lucey P, Cohn JF, Kanade T, Saragih J, Ambadar Z, Matthews I (2010) The extended Cohn-Kanade dataset (ck+): a complete dataset for action unit and emotion-specified expression. In: 2010 IEEE computer society conference on computer vision and pattern recognition workshops, pp 94–101. https://doi.org/10.1109/CVPRW.2010.5543262 Lutz CA (1988) Unnatural emotions: everyday sentiments on a micronesian atoll and their challenge to western theory. University of Chicago Press Ma Z, Zheng Z, Ye J, Li J, Gao Z, Zhang S, Chen X (2024) emotion2vec: self-supervised pre-training for speech emotion representation. In: Findings of the association for computational linguistics: ACL 2024, pp 15747–15760. https://doi.org/10.18653/v1/2024.findings-acl.931 Ma F, Yuan Y, Xie Y, Ren H, Liu I, He Y, Ren F, Yu FR, Ni S (2025) Generative technology for human emotion recognition: a scoping review. Inf Fusion 115:102753. https://doi.org/10.1016/j.inffus.2024.102753 Maas AL, Daly RE, Pham PT, Huang D, Ng AY, Potts C (2011) Learning word vectors for sentiment analysis. In: Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, pp 142–150 Madiega T (2024) Artificial intelligence act. Briefing PE 698.792. European Parliamentary Research Service Maeda T, Quan-Haase A (2024) When human-AI interactions become parasocial: agency and anthropomorphism in affective design. In: Proceedings of the 2024 ACM conference on fairness, accountability, and transparency. https://doi.org/10.1145/3630106.3658956 Majumder N, Poria S, Hazarika D, Mihalcea R, Gelbukh A, Cambria E (2019) DialogueRNN: an attentive RNN for emotion detection in conversations. In: Proceedings of the thirty-third AAAI conference on artificial intelligence, pp 6818–6825. https://doi.org/10.1609/aaai.v33i01.33016818 Mancuso V, Borghesi F, Chirico A, Bruni F, Sarcinella ED, Pedroli E, Cipresso P (2024) Iavrs international affective virtual reality system: psychometric assessment of 360 images by using psychophysiological data. Sensors 24(13):4204. https://doi.org/10.3390/s24134204 Maples B, Cerit M, Vishwanath A, Pea R (2024) Loneliness and suicide mitigation for students using GPT3-enabled chatbots. NPJ Ment Health Res. https://doi.org/10.1038/s44184-023-00047-6 McColl D, Hong A, Hatakeyama N, Nejat G, Benhabib B (2016) A survey of autonomous human affect detection methods for social robots engaged in natural HRI. J Intell Robot Syst 82:101–133. https://doi.org/10.1007/s10846-015-0259-2 McKeown G, Valstar M, Cowie R, Pantic M, Schroder M (2012) The SEMAINE database: annotated multimodal records of emotionally colored conversations between a person and a limited agent. IEEE Trans Affect Comput 3(1):5–17. https://doi.org/10.1109/T-AFFC.2011.25 McStay A (2018) Emotional AI: the rise of empathic media. SAGE Publications. https://doi.org/10.4135/9781526451293 Mehrabian A, Russell JA (1974) An approach to environmental psychology. MIT Press Miner AS, Milstein A, Schueller S, Hegde R, Mangurian C, Linos E (2016) Smartphone-based conversational agents and responses to questions about mental health, interpersonal violence, and physical health. JAMA Intern Med 176(5):619–625 Mitchell M, Wu S, Zaldivar A, Barnes P, Vasserman L, Hutchinson B, Spitzer E, Raji ID, Gebru T (2019) Model cards for model reporting. In: Proceedings of the conference on fairness, accountability, and transparency, pp 220–229. https://doi.org/10.1145/3287560.3287596 Mittelstadt B (2019) Principles alone cannot guarantee ethical AI. Nat Mach Intell 1:501–507. https://doi.org/10.1038/s42256-019-0114-4 Mohammad SM (2022) Ethics sheet for automatic emotion recognition and sentiment analysis. Comput Linguist 48(2):239–278. https://doi.org/10.1162/coli_a_00433 Mohammad SM, Turney PD (2012) Crowdsourcing a word-emotion association lexicon. Comput Intell. https://doi.org/10.1111/j.1467-8640.2012.00460.x Mohamed S, Png M-T, Isaac W (2020) Decolonial AI: decolonial theory as sociotechnical foresight in artificial intelligence. Philos Technol 33:659–684. https://doi.org/10.1007/s13347-020-00405-8 Mollahosseini A, Hasani B, Mahoor MH (2019) AffectNet: a database for facial expression, valence, and arousal computing in the wild. IEEE Trans Affect Comput 10(1):18–31. https://doi.org/10.1109/TAFFC.2017.2740923 Moon A-S, Kim H, Park Y-C, Lee J (2026) A survey on multimodal emotion recognition: methods, datasets, and future directions. Comput Mater Continua 1:1. https://doi.org/10.32604/cmc.2026.076411 Moors A, Ellsworth PC, Scherer KR, Frijda NH (2013) Appraisal theories of emotion: state of the art and future development. Emot Rev 5(2):119–124. https://doi.org/10.1177/1754073912468165 Mori H, Nishino H (2025) End-to-end conversational speech synthesis with controllable emotions in the dimensions of pleasantness and arousal. Acoust Sci Technol 46(1):70–77. https://doi.org/10.1250/ast.e24.13 Munezero MD, Montero CS, Sutinen E, Pajunen J (2014) Are they different? Affect, feeling, emotion, sentiment, and opinion detection in text. IEEE Trans Affect Comput 5(2):101–111. https://doi.org/10.1109/TAFFC.2014.2317187 Nabulsi J (2025) Affective sovereignty: a decolonising politics of emotion in palestine. Rev Int Stud. https://doi.org/10.1017/S0260210525100880 Niedenthal PM (2007) Embodying emotion. Science 316(5827):1002–1005. https://doi.org/10.1126/science.1136930 Norcross JC (2011) Psychotherapy relationships that work: evidence-based responsiveness. In: Psychotherapy relationships that work: evidence-based responsiveness, 2nd edn. Oxford University Press Northcutt CG, Athalye A, Mueller J (2021) Pervasive label errors in test sets destabilize machine learning benchmarks. In: Proceedings of the 35th international conference on neural information processing systems, track on datasets and benchmarks Ocumpaugh J, Baker RS, Gowda SM, Heffernan NT, Heffernan C (2014) Population validity for educational data mining models: a case study in affect detection. Br J Edu Technol 45:487–501. https://doi.org/10.1111/bjet.12156 Omarov B, Narynov S, Zhumanov Z (2022) Artificial intelligence-enabled chatbots in mental health: a systematic review. Comput Mater Continua 74(3):5105–5122. https://doi.org/10.32604/cmc.2023.034655 Ong DC (2021) An ethical framework for guiding the development of affectively-aware artificial intelligence. In: Proceedings of the 9th international conference on affective computing and intelligent interaction, pp 1–8. https://doi.org/10.1109/ACII52823.2021.9597441 Oorloff T, Koppisetti S, Bonettini N, Solanki D, Colman B, Yacoob Y, Shahriyari A, Bharaj G (2024) AVFF: audio-visual feature fusion for video deepfake detection. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition 2024, pp 27102–27112 Ortony A, Clore GL, Collins A (1988) The cognitive structure of emotions. Cambridge University Press. https://doi.org/10.1017/CBO9780511571299 Pang B, Lee L (2008) Opinion mining and sentiment analysis. Found Trends Inf Retr 2(1–2):1–135. https://doi.org/10.1561/1500000011 Park CY, Cha N, Kang S, Kim A, Khandoker AH, Hadjileontiadis L, Oh A, Jeong Y, Lee U (2020) K-EmoCon, a multimodal sensor dataset for continuous emotion recognition in naturalistic conversations. Sci Data. https://doi.org/10.1038/s41597-020-00630-y Parra-Gallego LF, Arias-Vergara T, Orozco-Arroyave JR (2025) Multimodal evaluation of customer satisfaction from voicemails using speech and language representations. Dig Signal Process 156(B):104820. https://doi.org/10.1016/j.dsp.2024.104820 Pei G, Haiying L, Lu Y, Wang Y, Hua S, Thihao L (2024) Affective computing: recent advances, challenges, and future trends. Intell Comput. https://doi.org/10.34133/icomputing.0076 Pelachaud C (2009) Studies on gesture expressivity for a virtual agent. Speech Commun 51(7):630–639. https://doi.org/10.1016/j.specom.2008.04.009 Picard RW (1997) Affective computing. MIT Press. https://doi.org/10.7551/mitpress/1140.001.0001 Picard RW, Vyzas E, Healey J (2001) Toward machine emotional intelligence: analysis of affective physiological state. IEEE Trans Pattern Anal Mach Intell 23(10):1175–1191. https://doi.org/10.1109/34.954607 Plutchik R (1980) A general psychoevolutionary theory of emotion. In: Plutchik R, Kellerman H (eds) Theories of emotion, vol 1. Academic Press, pp 3–33. https://doi.org/10.1016/B978-0-12-558701-3.50007-7 Poria S, Cambria E, Bajpai R, Hussain A (2017) A review of affective computing: from unimodal analysis to multimodal fusion. Inf Fusion 37:98–125. https://doi.org/10.1016/j.inffus.2017.02.003 Poria S, Majumder N, Mihalcea R, Hovy E (2019) Emotion recognition in conversation: research challenges, datasets, and recent advances. IEEE Access 7:100943–100953. https://doi.org/10.1109/ACCESS.2019.2929050 Poria S, Hazarika D, Majumder N, Naik G, Cambria E, Mihalcea R (2019b) MELD: a multimodal multi-party dataset for emotion recognition in conversations. In: Proceedings of the 57th annual meeting of the association for computational linguistics, pp 527–536. https://doi.org/10.18653/v1/P19-1050 Poria S, Majumder N, Hazarika D, Ghosal D, Bhardwaj R, Jian SYB, Hong P, Ghosh R, Roy A, Chhaya N, Gelbukh A, Mihalcea R (2021) Recognizing emotion cause in conversations. Cogn Comput 13:1317–1332. https://doi.org/10.1007/s12559-021-09925-7 Qu C, Che X, Yang Y, Zhang Z, Chang E, Zhang J, Zhu H, Yang L (2025) Enhancing emotion recognition in virtual reality: a multimodal dataset and a temporal emotion detector. Front Psychol. https://doi.org/10.3389/fpsyg.2025.1709943 Rahwan I, Cebrian M, Obradovich N, Bongard J, Bonnefon J-F, Breazeal C, Crandall JW, Christakis NA, Couzin ID, Jackson MO, Jennings NR, Kamar E, Kloumann IM, Larochelle H, Lazer D, McElreath R, Mislove A, Parkes DC, Pentland A, Roberts ME, Shariff A, Tenenbaum JB, Wellman M (2019) Machine behaviour. Nature 568:477–486. https://doi.org/10.1038/s41586-019-1138-y Rashkin H, Smith EM, Li M, Boureau Y-L (2019) Towards empathetic open-domain conversation models: a new benchmark and dataset. In: Proceedings of the 57th annual meeting of the association for computational linguistics, pp 5370–5381. https://doi.org/10.18653/v1/P19-1534 Rhue L (2018) Racial influence on automated perceptions of emotions. SSRN working paper. https://doi.org/10.2139/ssrn.3281765 Riek LD (2012) Wizard of Oz studies in HRI: a systematic review and new reporting guidelines. J Hum Robot Interact 1(1):119–136 Ringeval F, Sonderegger A, Sauer J, Lalanne D (2013) Introducing the recola multimodal corpus of remote collaborative and affective interactions. In: Proceedings of the 10th IEEE international conference and workshops on automatic face and gesture recognition, pp 1–8. https://doi.org/10.1109/FG.2013.6553805 Roller S, Dinan E, Goyal N, Ju D, Williamson M, Liu Y, Xu J, Ott M, Smith EM, Boureau Y-L, Weston J (2021) Recipes for building an open-domain chatbot. In: Proceedings of the 16th conference of the European chapter of the association for computational linguistics, pp 300–325. https://doi.org/10.18653/v1/2021.eacl-main.24 Russell JA (1980) A circumplex model of affect. J Pers Soc Psychol 39(6):1161–1178. https://doi.org/10.1037/h0077714 Sabour S, Liu S, Zhang Z, Liu J, Zhou J, Sunaryo A, Lee T, Mihalcea R, Huang M (2024) EmoBench: evaluating the emotional intelligence of large language models. In: Proceedings of the 62nd annual meeting of the association for computational linguistics, pp 5986–6004. https://doi.org/10.18653/v1/2024.acl-long.326 Sambasivan N, Kapania S, Highfill H, Akrong D, Paritosh P, Aroyo LM (2021) “Everyone wants to do the model work, not the data work”: data cascades in high-stakes AI. In: Proceedings of the 2021 CHI conference on human factors in computing systems, pp 1–15. https://doi.org/10.1145/3411764.3445518 Scassellati B, Admoni H, Matarić M (2012) Robots for use in autism research. Annu Rev Biomed Eng 14:275–294. https://doi.org/10.1146/annurev-bioeng-071811-150036 Scherer KR (2001) Appraisal considered as a process of multilevel sequential checking. In: Appraisal processes in emotion: theory, methods, research. Oxford University Press, pp 92–120 Scherer KR (2005) What are emotions? And how can they be measured? Soc Sci Inf 44(4):695–729. https://doi.org/10.1177/053901840505821 Scherer KR, Moors A (2019) The emotion process: event appraisal and component differentiation. Annu Rev Psychol 70:719–745. https://doi.org/10.1146/annurev-psych-122216-011854 Schlicher M, Li Y, Murthy SMK, Sun Q, Schuller BW (2025) Emotionally adaptive support: a narrative review of affective computing for mental health. Front Dig Health. https://doi.org/10.3389/fdgth.2025.1657031 Schmidt P, Reiss A, Duerichen R, Marberger C, Van Laerhoven K (2018) Introducing WESAD, a multimodal dataset for wearable stress and affect detection. In: Proceedings of the 20th ACM international conference on multimodal interaction, pp 400–408. https://doi.org/10.1145/3242969.3242985 Searle JR (1980) Minds, brains, and programs. Behav Brain Sci 3(3):417–424. https://doi.org/10.1017/S0140525X00005756 Sharma A, Miner A, Atkins D, Althoff T (2020) A computational approach to understanding empathy expressed in text-based mental health support. In: Proceedings of the 2020 conference on empirical methods in natural language processing, pp 5263–5276. https://doi.org/10.18653/v1/2020.emnlp-main.425 Sharma A, Lin IW, Miner AS, Atkins DC, Althoff T (2021) Towards facilitating empathic conversations in online mental health support: a reinforcement learning approach. In: Proceedings of the web conference 2021. https://doi.org/10.1145/3442381.3450097 Shi Y, Yu K, Dong Y, Chen F (2026) Large language models in education: a systematic review of empirical applications, benefits, and challenges. Comput Educ Artif Intell 10:100529. https://doi.org/10.1016/j.caeai.2025.100529 Shingjergji K, Iren D, Urlings C, Klemke R (2026) Affective computing in online higher education: a systematic literature review. Comput Educ Artif Intell 10:100499. https://doi.org/10.1016/j.caeai.2025.100499 Shum H-Y, He X-D, Li D (2018) From Eliza to Xiaoice: challenges and opportunities with social chatbots. Front Inf Technol Electron Eng 19(1):10–26. https://doi.org/10.1631/FITEE.1700826 Slater M (2009) Place illusion and plausibility can lead to realistic behaviour in an immersive virtual environment. Philos Trans R Soc B 364(1535):3549–3557. https://doi.org/10.1098/rstb.2009.0138 Socher R, Perelygin A, Wu J, Chuang J, Manning CD, Ng AY, Potts C (2013) Recursive deep models for semantic compositionality over a sentiment treebank. In: Proceedings of the 2013 conference on empirical methods in natural language processing, pp 1631–1642 Sofroniew N, Kauvar I, Saunders W, Chen R, Henighan T, Hydrie S, Citro C, Pearce A, Tarng J, Gurnee W, Batson J, Zimmerman S, Rivoire K, Fish K, Olah C, Lindsey J (2026) Emotion concepts and their function in a large language model. Anthoropic Somarathna R, Bednarz T, Mohammadi G (2023) Virtual reality for emotion elicitation—a review. IEEE Trans Affect Comput 14(4):2626–2645. https://doi.org/10.1109/TAFFC.2022.3181053 Spitale M, Axelsson M, Gunes H (2024) Appropriateness of LLM-equipped robotic well-being coach language in the workplace: a qualitative pilot study. https://doi.org/10.48550/arXiv.2401.14935. arXiv:2401.14935 Stark L (2018) Algorithmic psychometrics and the scalable subject. Soc Stud Sci. https://doi.org/10.1177/03063127187720 Stark L, Hoey J (2021) The ethics of emotion in artificial intelligence systems. In: Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp 782–793. https://doi.org/10.1145/3442188.3445939 Stieglitz S, Mirbabaie M, Ross B, Neuberger C (2018) Social media analytics: challenges in topic discovery, data collection, and data preparation. Int J Inf Manag 39:156–168. https://doi.org/10.1016/j.ijinfomgt.2017.12.002 Tang C, Yu W, Sun G, Chen X, Tan T, Li W, Lu L, Ma Z, Zhang C (2024) SALMONN: towards generic hearing abilities for large language models. In: Proceedings of the international conference on learning representations. https://doi.org/10.48550/arXiv.2310.13289 Tian Y, Huang T, Liu M, Jiang D, Spangher A, Chen M, May J, Peng N (2024a) Are large language models capable of generating human-level narratives? In: Proceedings of the 2024 conference on empirical methods in natural language processing, pp 17659–17681. https://doi.org/10.18653/v1/2024.emnlp-main.978 Tian L, Wang Q, Zhang B, Bo L (2024b) EMO: emote portrait alive generating expressive portrait videos with Audio2Video diffusion model under weak conditions. In: Computer vision—ECCV 2024, pp 244–260. https://doi.org/10.1007/978-3-031-73010-8_15 Troiano E, Oberländer L, Klinger R (2023) Dimensional modeling of emotions in text with appraisal theories: corpus creation, annotation reliability, and prediction. Comput Linguist 49(1):1–72. https://doi.org/10.1162/coli_a_00461 Tsai Y-HH, Bai S, Liang PP, Kolter JZ, Morency L-P, Salakhutdinov R (2019) Multimodal transformer for unaligned multimodal language sequences. In: Proceedings of the 57th annual meeting of the association for computational linguistics, pp 6558–6569. https://doi.org/10.18653/v1/P19-1656 Turing AM (1950) Computing machinery and intelligence. Mind 49:433–460 Tzirakis P, Trigeorgis G, Nicolaou MA, Schuller BW, Zafeiriou S (2017) End-to-end multimodal emotion recognition using deep neural networks. IEEE J Sel Top Signal Proces 11(8):1301–1309. https://doi.org/10.1109/JSTSP.2017.2764438 Vaidyam AN, Wisniewski H, Halamka JD, Kashavan MS, Torous JB (2019) Chatbots and conversational agents in mental health: a review of the psychiatric landscape. Can J Psychiatry. https://doi.org/10.1177/0706743719828977 Veale M, Borgesius ZF (2021) Demystifying the draft EU Artificial Intelligence Act. Comput Law Rev Int 22(4):97–112. https://doi.org/10.9785/cri-2021-220402 Wang Y, Guo J, Bai J, Yu R, He T, Tan X, Sun X, Bian J (2025) InstructAvatar: text-guided emotion and motion control for avatar generation. In: Proceedings of the AAAI conference on artificial intelligence, vol 39, pp 8132–8140. https://doi.org/10.1609/aaai.v39i8.32877 Wankhade M, Kulkarni C, Rao ACS (2025) A survey on aspect base sentiment analysis methods and challenges. Appl Soft Comput 167(A):112249. https://doi.org/10.1016/j.asoc.2024.112249 Winner L (1980) Do artifacts have politics? Daedalus 109(1):121–136 Witte BD, Reynaert V, Kieken D, Jabbour J, Demarey C, Dumoulin A, Possik J (2026) Immersive virtual reality learning and cognitive load: a multiple-day field study. Comput Hum Behav 176:108853. https://doi.org/10.1016/j.chb.2025.108853 Xia R, Ding Z (2019) Emotion-cause pair extraction: a new task to emotion analysis in texts. In: Proceedings of the 57th annual meeting of the association for computational linguistics, pp 1003–1012. https://doi.org/10.18653/v1/P19-1096 Xu S, Chen G, Guo Y-X, Yang J, Li C, Zang Z, Zhang Y, Tong X, Guo B (2024) VASA-1: lifelike audio-driven talking faces generated in real time. In: Proceedings of the 38th international conference on neural information processing systems, pp 660–684. https://doi.org/10.5555/3737916.3737937 Yan W-J, Li X, Wang S-J, Zhao G, Liu Y-J, Chen Y-H, Fu X (2014) Casme II: an improved spontaneous micro-expression database and the baseline evaluation. PLoS ONE 9(1):86041. https://doi.org/10.1371/journal.pone.0086041 Yan HY, Morrow G, Yang K-C, Wihbey J (2025) The origin of public concerns over AI supercharging misinformation in the 2024 U.S. presidential election. In: Harvard Kennedy School misinformation review. https://doi.org/10.37016/mr-2020-171 Yannakakis GN, Spronck P, Loiacono D, André E (2013) Player modeling. In: Artificial and computational intelligence in games, vol 6, pp 45–59. https://doi.org/10.4230/DFU.Vol6.12191.45 Zadeh A, Zellers R, Pincus E, Morency L-P (2016) Multimodal sentiment intensity analysis in videos: facial gestures and verbal messages (CMU-MOSI). IEEE Intell Syst 31:82–88. https://doi.org/10.1109/MIS.2016.94 Zadeh A, Chen M, Poria S, Cambria E, Morency L-P (2017) Tensor fusion network for multimodal sentiment analysis. In: Proceedings of the 2017 conference on empirical methods in natural language processing, pp 1103–1114. https://doi.org/10.18653/v1/D17-1115 Zadeh AB, Liang PP, Poria S, Cambria E, Morency L-P (2018) Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph. In: Proceedings of the 56th annual meeting of the association for computational linguistics. pp 2236–2246. https://doi.org/10.18653/v1/P18-1208 Zhang Y, Wang M, Wu Y, Tiwari P, Li Q, Wang B, Qin J (2024) DialogueLLM: context and emotion knowledge-tuned large language models for emotion recognition in conversations. Neural Netw 192:107901. https://doi.org/10.1016/j.neunet.2025.107901 Zhang W, Deng Y, Liu B, Pan S, Bing L (2024b) Sentiment analysis in the era of large language models: a reality check. In: Findings of the association for computational linguistics: NAACL 2024, pp 3881–3906. https://doi.org/10.18653/v1/2024.findings-naacl.246 Zhang Y, Zhao D, Hancock JT, Kraut R, Yang D (2025) The rise of AI companions: how human-chatbot relationships influence well-being. https://doi.org/10.48550/arXiv.2506.12605 Zhang Y, Yang X, Xu X, Gao Z, Huang Y, Mu S, Feng S, Wang D, Zhang Y, Song K, Yu G (2026) Affective computing in the era of large language models: a survey from the NLP perspective. Knowl-Based Syst 337:115411. https://doi.org/10.1016/j.knosys.2026.115411 Zhao S, Wang S, Soleymani M, Joshi D, Ji Q (2019) Affective computing for large-scale heterogeneous multimedia data: a survey. ACM Trans Multimed Comput Commun Appl 15(93):1–32. https://doi.org/10.1145/3363560 Zhao J, Zhang T, Hu J, Liu Y, Jin Q, Wang X, Li H (2022) M3ED: multi-modal multi-scene multi-label emotional dialogue database. In: Proceedings of the 60th annual meeting of the association for computational linguistics. Association for Computational Linguistics, pp 5699–5710. https://doi.org/10.18653/v1/2022.acl-long.391 Zheng W-L, Lu B-L (2015) Investigating critical frequency bands and channels for EEG-based emotion recognition with deep neural networks. IEEE Trans Auton Ment Dev 7(3):162–175 Zheng C, Sabour S, Wen J, Zhang Z, Huang M (2023) AugESC: dialogue augmentation with large language models for emotional support conversation. In: Findings of the association for computational linguistics: ACL 2023. Association for Computational Linguistics, pp 1552–1568. https://doi.org/10.18653/v1/2023.findings-acl.99 Zhou L, Gao J, Li D, Shum H-Y (2018) The design and implementation of XiaoIce, an empathetic social chatbot. Comput Linguist 46(1):53–93. https://doi.org/10.1162/coli_a_00368 Ziems C, Held W, Shaikh O, Chen J, Zhang Z, Yang D (2024) Can large language models transform computational social science? Comput Linguist 50:237–291. https://doi.org/10.1162/coli_a_00502 Zuboff S (2019) The age of surveillance capitalism: the fight for a human future at the new frontier of power. PublicAffairs Acknowledgements The author gratefully acknowledges the support of JST SPRING, the Nippon Foundation HUMAI Program, the GMO Internet Foundation, and the Telecommunications Advancement Foundation, which provided the funding and research environment that enabled this work. Funding This work was supported in part by JST SPRING under Grant Number JPMJSP2150, the Nippon Foundation HUMAI Program, the GMO Internet Foundation, and the Telecommunications Advancement Foundation. Author information Authors and Affiliations Contributions The author confirms sole responsibility for the following: study conception and design, methodology, analysis and interpretation of the results, and manuscript preparation. Corresponding author Ethics declarations Conflict of interest The authors declare no competing interests. Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Use of AI tools The author used AI-based language tools to support translation and language editing during manuscript preparation. All AI-assisted outputs were reviewed, revised, and verified by the author, who takes full responsibility for the final content of the manuscript. Additional information Publisher's Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Appendix A: Coding rubric for the prior-surveys comparison table Appendix A: Coding rubric for the prior-surveys comparison table This appendix documents the rubric used to score each prior survey listed in Table 1 along the two diagnostic axes used in this paper: cross-layer dependence treated as central and prescriptive stance. The purpose of the rubric is not to rank prior surveys but to make the differentiation claim auditable: a future reader who disagrees with the placement of any row should be able to identify which scoring anchor they would invoke instead and why. 1.1 A.1 Scoring axes Axis 1: Cross-layer dependence treated as central. This axis asks whether the survey treats inter-layer dependence—for instance theory–data, modeling–interaction, technology–ethics, or data–application—as a central analytic object rather than as a residual concern. The three-level scale is: - Yes: the survey thematically (at the section or chapter level) frames inter-layer disconnections as the object of analysis, and the structure of the survey is organized around such disconnections rather than around a single layer or a single modality. - Partial: the survey discusses two or more layers in interaction (e.g., couples ethics to method, or data to deployment), but the inter-layer disconnection itself is not the principal organizing object; cross-layer concern surfaces locally rather than systemically. - No: the survey is organized within a single layer (e.g., data, modeling, modality), reviews multiple layers in parallel without analyzing their dependence, or adopts a bibliometric stance in which layer-level analytic structure is not used. Axis 2: Prescriptive stance. This axis asks whether the survey advances a prescriptive program that goes beyond taxonomy or empirical taking-stock and is intended to constrain subsequent research practice. The two-level scale used in Table 1 is: - Yes: the survey advances a normative program, such as a foundational vision for the field, a prescriptive ethics agenda, or a set of design recommendations intended to constrain how subsequent work should proceed. The recommendation may be confined to one layer (e.g., ethics) or may be field-wide; what matters is that the survey takes a position rather than merely describing the literature. - No: the survey is descriptive, taxonomic, or bibliometric in stance. Local recommendations may appear (most surveys close with “future directions”), but the survey does not commit to a programmatic position whose acceptance would reshape research practice. A finer-grained four-level scale (Strong/Moderate/Weak/None) was considered during rubric development; it was collapsed to the binary Yes/No scale used in Table 1 because the disambiguation required to separate Strong from Moderate could not be applied uniformly across the eleven surveys without introducing reviewer-dependent variance. The collapse is conservative: surveys scored Yes satisfy at least the Moderate threshold, and surveys scored No do not reach the Moderate threshold. The richer scale is documented here for transparency but is not used in the table itself. 1.2 A.2 Inclusion criteria for the comparison set The comparison set in Table 1 consists of eleven prior surveys selected under the following inclusion rules. - 1. Surveys, not primary papers. Only works whose self-stated genre is review, survey, manifesto, or ethics sheet are included. Influential primary papers cited elsewhere in this paper (e.g., Ekman 1992, Russell 1980, Barrett 2017) are not eligible for the differentiation table because their genre is theoretical or empirical contribution rather than survey. - 2. Direct relevance to affective computing as a field. Surveys whose object is affective computing, emotion recognition, affect detection, emotion in dialog, or the ethics of emotion AI are included. Surveys of adjacent fields (general HCI, general dialog systems, general AI ethics) are excluded unless their scope explicitly subsumes affective computing. - 3. Temporal coverage. The set spans 1997–2025 and is intentionally chosen to include the foundational manifesto (Picard 1997), mature interdisciplinary reviews from the 2010s (Calvo and Mello 2010; Mccoll et al. 2016; Poria et al. 2017, 2019a; Zhao et al. 2019), the inflection point of ethics-explicit work (Mohammad 2022), and the 2024–2025 LLM- and foundation-model-era surveys (Pei et al. 2024; Kumar et al. 2025; Moon et al. 2026; Zhang et al. 2026). The intent is to cover the inflection points of the field rather than to be exhaustive. - 4. Representative scope. Where multiple surveys cover essentially the same scope in the same year, the comparison set retains the one most cited at the time of writing or, when citation counts are similar, the one whose framing is closest to the cross-layer concern of this paper. The two Dey et al. 2025 surveys are both retained because they target distinct slices (general multimodal MER vs. LLM-/foundation-model-based MER 2021–2025) and the differentiation between them is itself diagnostic of the field’s current trajectory. - 5. Exclusions. Workshop-only proceedings, vendor white papers, and surveys whose published version was not available at the time of writing are excluded. This rubric does not claim that no other survey could have been included; it claims that the eleven included surveys are representative of the comparison categories the differentiation table is designed to span. 1.3 A.3 Disambiguation procedures for borderline cases Three disambiguation procedures were applied during scoring; they are documented here so that the borderline judgments are auditable. D1: Foundational vision vs. cross-layer diagnosis. A survey whose mission is to define a new field (e.g., Picard 1997) is necessarily multi-topic: data, models, applications, and ethics are all touched. The rubric does not score such breadth as Yes on Axis 1. Yes on Axis 1 requires that the inter-layer dependence (one layer’s choice constraining another layer’s evaluation) be the analytic object. Foundational manifestos that span layers but do not analyze their dependence as such are scored No on Axis 1 even when they take a strong programmatic position on Axis 2. D2: Ethics-and-method coupling. A survey whose argument explicitly couples ethical evaluation to methodological design (e.g., Mohammad 2022, whose 50 considerations span task design, data, method, evaluation, and implications) is scored Partial on Axis 1. The coupling is real and intentional, but it does not reach the full cross-layer-as-central threshold: theory–data and modeling–interaction couplings are not the principal analytic objects. We note that the method layer (L3) and the ethics layer (L6) are not adjacent strata in the six-layer pipeline of Sect. 1; an ethics-method coupling is therefore substantively an L3–L6 inter-layer coupling, which is precisely why it justifies a Partial score rather than No, even though it falls short of the systemic cross-layer-as-central threshold required for Yes. Partial is also used when a survey integrates two adjacent layers without taking the overall layered structure as its frame. D3: Bibliometric and taxonomic surveys. A survey whose primary method is bibliometric (citation analysis, co-occurrence mapping; e.g., Pei et al. (2024)) is scored No on Axis 1 even when its descriptive coverage formally spans every layer. The rationale is that bibliometric methods describe the distribution of attention across topics rather than analyze inter-layer dependence; without an analytic claim about dependence, parallel coverage is not cross-layer dependence in the sense this paper requires. 1.4 A.4 Per-survey scoring rationales The following entries provide, for each row of Table 1, the scoring rationale on both axes. The scores reproduced below match Table 1 cell-for-cell; any divergence should be treated as an error. Where the rationale relies on disambiguation procedures D1–D3, the procedure is named. 1. Picard (1997). Axis 1: No. Foundational manifesto for the field; layers (data, models, applications, ethics) are surveyed but inter-layer dependence is not the analytic object (D1). Axis 2: Yes (foundational vision, not cross-layer diagnosis). The work issues a programmatic position by defining what affective computing should be. It is scored Yes on Axis 2 for that reason. The qualifier in Table 1 explicitly distinguishes this stance from cross-layer diagnosis. 2. Calvo and Mello (2010). Axis 1: No (layers reviewed in parallel). The survey covers face, voice, body, physiology, EEG, text, and multimodal channels in parallel, alongside emotion theory; coverage is broad, but the dependence among these strata is not made into the analytic object. Axis 2: No. The work is an interdisciplinary review with descriptive intent; it does not advance a prescriptive program in the sense Axis 2 requires. 3. Mccoll et al. (2016). Axis 1: No. The scope is autonomous human affect detection methods within HRI scenarios; the focus is methodological taxonomy within a single application context rather than inter-layer dependence. Axis 2: No. The survey is descriptive, oriented toward HRI practitioners selecting detection methods; no field-wide prescriptive position is advanced. 4. Poria et al. (2017). Axis 1: No (modality-centric). The survey is organized around unimodal (visual, audio, text) analysis and multimodal fusion; modality is the organizing principle, and inter-layer dependence is not foregrounded. Axis 2: No. Comprehensive review aimed at modeling-layer practitioners; the closing remarks discuss future directions but do not amount to a prescriptive program over the field. 5. Poria et al. (2019a). Axis 1: No (task-internal). The scope is the ERC task: research challenges, datasets, and recent neural approaches, such as CMN, ICON, and DialogueRNN. The analysis stays within one task family rather than treating cross-layer dependence as its central analytic object. Axis 2: No. Systematic review of one task; descriptive in stance. 6. Zhao et al. (2019). Axis 1: No. Focus on heterogeneous multimedia (images, music, videos, multimodal) at the modeling and data layers; cross-layer dependence is not the analytic object. Axis 2: No. Modality-and-content-centered survey without a prescriptive program. 7. Mohammad (2022). Axis 1: Partial (ethics–method coupling). The 50 considerations span task design, data, method, evaluation, and implications; this is a deliberate coupling of ethical reasoning to methodological choice (D2). The coupling does not extend to the full theory–data and modeling–interaction couplings that would warrant Yes, but it reaches further than No. Axis 2: Yes (prescriptive ethics agenda). The ethics sheet format itself is a programmatic intervention: it tells the field how ethical reasoning should be incorporated into emotion-recognition and sentiment-analysis research. This is a clear prescriptive program, even though its scope is bounded by the ethics layer. 8. Pei et al. (2024). Axis 1: No (bibliometric, not analytic). A bibliometric review of 33,448 papers from 1997–2023; the method is descriptive distribution analysis (D3). Even where every layer appears in the dataset, parallel coverage is not cross-layer dependence in the sense Axis 1 requires. Axis 2: No. The method does not yield a programmatic position; the contribution is a map of the literature, not a stance on it. 9. Kumar et al. (2025). Axis 1: No. Comprehensive survey of multimodal emotion recognition; modality and architectural lineage are the organizing axes. Axis 2: No. The survey is descriptive of recent technical advances; future-direction remarks do not constitute a prescriptive program. 10. Moon et al. (2026). Axis 1: No. Survey of multimodal emotion recognition 2021–2025 with focus on LLM- and foundation-model-based architectures; concentrates on the modeling layer with attention to architectural taxonomy of LLM-based MER. Axis 2: No. Technical review; descriptive rather than prescriptive in stance. 11. Zhang et al. (2026). Axis 1: No. LLM-era affective computing survey from an NLP perspective covering affect understanding and affect generation tasks, instruction tuning, prompt engineering, and benchmarks; the analysis stays at the modeling and benchmark layers. Axis 2: No. Technical survey oriented toward NLP practitioners; future-direction remarks are local rather than programmatic over the field. 1.5 A.5 Limitations of the rubric Three limitations of the rubric are recorded here in the interest of fair self-assessment. First, the scoring was performed by the authors of this paper without an independent second coder; inter-rater reliability (IRR) is therefore not measured. This is a real limitation, and the position-paper genre is not, in itself, an excuse for it. The decision to defer second-coder scoring rests on two grounds: (i) the binary collapse described in Sect. A.1 is intended precisely to reduce the surface on which IRR variance acts, and (ii) the per-survey rationales above expose every disambiguation step so that an independent reader can reproduce the scoring with public information. A future revision of this paper, or a follow-up, could readily fold in second-coder scoring; the rubric is designed to support that. Second, the inclusion criteria of Sect. A.2 are themselves selection choices, and any selection choice is in principle contestable. The rubric’s defense is not that the eleven surveys are uniquely correct but that the comparison categories they span (foundational manifesto, interdisciplinary review, modality-centric review, task-centric review, ethics sheet, bibliometric review, LLM/foundation-model survey) are the categories against which a position-paper claim of novelty must be checked. Third, the rubric does not score any survey on the further axis of cascade depth, that is, whether the survey, when it does treat inter-layer dependence, treats the dependence as a single coupling or as a chained cascade of couplings. Adding such an axis would be informative, but it would also collapse onto the position of this paper itself (since this paper is the explicit cascade claim), and the rubric’s purpose here is comparison, not self-promotion. The absence of this axis is therefore a deliberate scope decision rather than an oversight. The rubric, the per-survey rationales, and the table together support the position-paper claim of Sect. 8 and the prescriptive program of Sect. 9 at the level of evidence: any reader who disagrees with the placement of a row in Table 1 is invited to re-score that row using this rubric and to re-examine, against the new score, whether the cascade diagnosis and the DC1–DC5 program continue to follow. Rights and permissions Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law. About this article Cite this article Inoshita, K. Bridging the silos in affective AI: a critical perspective from data to society. AI & Soc (2026). https://doi.org/10.1007/s00146-026-03324-y Received: Accepted: Published: Version of record: DOI: https://doi.org/10.1007/s00146-026-03324-y

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached — you'll always get the same 5 for this article.