general9982 wordsRead on Arc Codex

A Deep Learning-Based Retrospective Evaluation Prediction System for Emotional Experiences: Temporal Dynamic Feature Extraction and ERP Neural Mechanisms of the Peak

Abstract Retrospective evaluations of emotional experiences are significantly influenced by the peak-end effect, yet existing electroencephalography (EEG)-based emotion recognition studies have focused on classifying immediate emotional states, leaving integration of the peak-end effect into predictive modeling of retrospective evaluations underexplored. This study proposes a theory-driven deep learning framework extracting peak-end neural features from event-related potential (ERP) signals to predict retrospective overall ratings. Thirty participants completed a sequential emotion induction paradigm based on the International Affective Picture System (IAPS), with peak position and endpoint intensity systematically manipulated. Three ERP components—Early Posterior Negativity (EPN), P300, Late Positive Potential (LPP)—were extracted from 64-channel EEG to construct the TCN-Attention-PeakEnd Gate (TAPE) model, which combines a temporal convolutional network (TCN) encoder with multi-head self-attention and a peak-end feature gating module that embeds the cognitive theory into the network as differentiable operations. Under leave-one-subject-out (LOSO) cross-validation, TAPE achieved MAE = 1.038, Pearson \(r\) = 0.654, and three-level classification accuracy of 70.4%, significantly outperforming eight baselines under Holm-Bonferroni-corrected paired-sample \(t\)-tests. Across 50 independent fivefold cross-validation evaluations, TAPE simultaneously attained the highest median \(r\) (0.682) and smallest interquartile range (0.054). External validation on SEED replicated TAPE's ranking above the strongest deep-learning baselines (84.7% vs 81.9% for Transformer). Ablation confirmed independent contributions of each module, and attention weight visualization revealed high consistency between learned temporal patterns and peak-end theoretical predictions. This study provides a "cognitive theory and data-driven" dual-track fusion modeling paradigm for EEG affective computing. Introduction After a sustained emotional event, individuals tend to form overall subjective evaluations from a few key memory fragments rather than the complete course. Cognitive psychology summarizes this as the “Peak-End Rule”: overall retrospective assessment is primarily determined by the moment of highest emotional intensity (the peak) and the ending state (the endpoint), while duration exerts minimal influence—a phenomenon known as “duration neglect” [1]. This effect has been validated across multiple contexts including pain perception, consumer experience, and everyday well-being, and has in recent years been progressively introduced into research on mental health assessment and clinical intervention [2]. However, neuroscientific research on the peak-end effect has largely stayed at the level of post-hoc comparison between behavioral retrospective ratings and ecological momentary assessment; systematic exploration of the EEG temporal dynamic features corresponding to this memory bias during emotional processing remains limited. Event-related potentials, with millisecond temporal resolution, are widely used to trace the time course of emotional processing. Studies show that the Late Positive Potential (LPP) peaks approximately 600 ms after stimulus onset, and its amplitude reliably reflects the motivational salience and arousal level of stimuli [3]. Deep learning has advanced EEG-based affective recognition: convolutional neural networks (CNN) combined with long short-term memory (LSTM) networks effectively capture EEG spatial patterns and temporal dependencies, with some models exceeding 90% accuracy on valence and arousal classification [4]. Researchers have also integrated differential entropy features with CNN-LSTM architectures, constructing two-dimensional grid matrices to preserve the spatial topological information among electrodes, thereby achieving superior performance in emotion classification tasks [5]. More notably, Transformer-based self-attention mechanisms have also been introduced into EEG-based emotion recognition, demonstrating clear advantages in capturing long-range temporal dependencies [6]. Despite continuous breakthroughs in immediate emotional state classification, systematic integration of deep learning with the ERP neural mechanisms of the peak-end effect for predicting retrospective overall evaluations of emotional experiences remains underexplored. This paper proposes a retrospective evaluation prediction system integrating ERP temporal dynamic features with deep learning. A sequential presentation paradigm of continuous emotional stimuli is used, with retrospective ratings after each experience. ERP components associated with the peak-end effect (LPP for arousal, EPN for early attention, P300 for cognitive evaluation) are extracted from EEG, with their temporal evolution as model inputs. A hybrid TCN–multi-head self-attention architecture automatically learns the neural representational weights of peak and endpoint moments. The innovations lie in three aspects: theoretically, extending the peak-end effect from behavioral description to explicit quantitative characterization of its ERP neural mechanisms, revealing differential LPP contributions at peak and endpoint; methodologically, proposing an ERP dynamic feature extraction strategy structured by the peak-end effect that captures retrospective-evaluation information directly from neural signals; and at the application level, providing an objective technical means for correcting symptom recall biases in clinical psychological assessment. The remainder is organized as follows. Section "Related Work" reviews the peak-end effect, ERP markers of emotional experience, EEG/ERP deep learning, and temporal-sequence feature extraction. Section "Methods" details the paradigm, preprocessing, TAPE architecture, and training/evaluation. Section "Results" reports ERP statistics, behavioral peak-end validation, regression and classification results, ablation, and attention weight visualization. Section "Discussion and Implications" discusses performance-theory associations, cognitive-neuroscientific significance of attention weights, and methodological implications. Section "Conclusion" concludes. Related Work Theoretical Foundations and Experimental Paradigms of the Peak-End Effect The peak-end effect was first proposed by Kahneman et al. in pain perception experiments. Its core hypothesis is that in retrospectively evaluating a continuous experience, individuals assign disproportionate weight to the moment of highest emotional intensity and the final moment rather than averaging across the process. Alaybek et al. conducted a large-scale meta-analysis across 174 effect sizes from 112 independent studies, reporting a composite effect size of r ≈ 0.581, robust across emotional valence, assessment tasks, and time spans, with predictive power significantly higher than alternative heuristics such as the initial moment, the moment of lowest intensity, and trend variability [7]. This meta-analysis also confirmed the “duration neglect” hypothesis—that is, the duration of an experience has virtually no effect on retrospective evaluation. Researchers have extended the peak-end effect to applied contexts. MĂŒller et al. used a threat-of-shock paradigm to compare retrospective anxiety ratings under different ending conditions, finding that participants ending under high-threat conditions rated their overall anxiety significantly higher; the result held under counterbalanced designs, evidencing applicability in negative emotional memory [8]. McCullough et al. focused on the competition between the peak-end and primacy effects in service experiences, finding that retrospective delay moderates their weights: longer delays favor the peak-end effect, while immediate evaluations favor the primacy effect [9]. These findings reveal how peak-end strength varies across evaluation time windows and inform the timing of retrospective ratings in the present paradigm. ERP Markers of Emotional Experience (LPP, EPN, P300, Etc.) Event-related potentials provide a unique window into the time course of emotional processing. The Late Positive Potential (LPP) is recognized as the core indicator of emotional arousal and motivational salience. Farkas and Sabatinelli compared the functional dissociation of the Early Posterior Negativity (EPN) and LPP using scene images with varying bodily exposure and emotional type, finding that EPN is highly sensitive to body-feature detection but weakly associated with subjective arousal, whereas LPP amplitude strongly and positively correlates with arousal ratings, indicating distinct functional roles at different stages [10]. Schupp et al. further advanced this research to the individual level, using a case-by-case analysis to verify that the enhancement effects of EPN and LPP for high-arousal stimuli can be stably replicated in the vast majority of participants, with EPN effects appearing within the 150–350 ms time window after stimulus onset, and LPP effects concentrated in the 350–750 ms interval, together constituting an electrophysiological chain of evidence for the prioritized processing of emotional stimuli [11]. P300 also plays an important role in emotion-cognition evaluation. Schindler and Bublatzky examined social contextual modulation of emotional face processing, observing selective P300 enhancement for attentional allocation to threatening faces, with significant differences between conditions with and without social background information [12]. P300 and LPP partially overlap in scalp distribution and time window, but their functional roles differ: P300 is more associated with stimulus categorization and cognitive evaluation, while LPP reflects sustained elaborative processing of emotional meaning [13]. At the individual level, the case-by-case analysis by Schupp et al. mentioned above has already demonstrated that the emotion-modulation effects of EPN and LPP can be stably replicated in the vast majority of individuals [11], providing reliability guarantees for their use as input features in predictive models. These findings provide a sufficient theoretical basis for the selection of LPP, EPN, and P300 as neural feature markers of the peak-end effect in this study. Deep Learning Methods for EEG/ERP-Based Affective Recognition Deep learning has advanced EEG-based affective recognition. Li et al. proposed a Transformer neural architecture search method for EEG emotion recognition, achieving superior classification performance on SEED and DEAP by automatically optimizing attention-module structural parameters, demonstrating the potential of data-driven architecture search [14]. Convolutional networks capture EEG spatial-spectral features while long short-term memory networks model temporal dependencies. Huang et al. proposed an emotion recognition model integrating CNN with bidirectional LSTM and attention, applying attention weighting along the temporal dimension to enhance key time-step contributions and achieving superior valence and arousal classification on DEAP over single-network baselines [15]. Li et al. integrated spatial, temporal, and inter-region connectivity features into a unified multi-scale convolutional framework (STC-CNN), demonstrating competitive recognition performance in both valence and arousal binary classification tasks [16]. Following Transformer's success in natural language processing, self-attention has been introduced into EEG analysis. Liu et al. proposed the interpretable Transformer ERTNet, in which temporal convolutional layers automatically learn frequency-selective features resembling bandpass filters, and spatial-convolutional weight distributions map to scalp topographies, substantially enhancing interpretability while maintaining classification accuracy [17]. Xu et al. developed AMDET, an EEG Transformer with multi-dimensional attention that applies weights separately to channel and temporal dimensions, exceeding prior methods on SEED cross-subject classification [18]. Peng et al.'s proposed temporal relative Transformer encoding method introduced a temporal relative position encoding strategy on the basis of channel attention, enabling the model to better capture the relative relationships between different time steps in EEG signals [19]. These advances indicate that attention mechanisms have unique advantages for the extraction and selection of EEG emotional features, and have laid a strong methodological foundation for the broader adoption of deep learning across EEG-based affective and clinical tasks. Recent work has extended deep learning to a range of EEG-based affective and clinical applications. Jadhav proposed an EEG stress detection framework based on Convolutional Spiking Neural Networks (CSNNs), combining discrete wavelet transform features with spike-based temporal dynamics and achieving competitive accuracy on the PhysioNet EEG dataset with advantages suited to edge deployment [20]. Malviya et al. examined the discriminative power of canonical EEG frequency-band features (delta, theta, alpha, beta, gamma) for stress detection during arithmetic tasks, showing that conventional classifiers combined with neurophysiological priors retain practical value under modest sample sizes [21]. Building on such pipelines, hybrid architectures combining convolutional and recurrent modules have emerged as more expressive alternatives: Malviya and Mal proposed a DWT (Discrete wavelet transform)-based CNN-BLSTM model that attained robust stress-level classification on multi-channel EEG recordings [22]. The effectiveness of CNN-LSTM hybridization with principled hyperparameter search has been further corroborated by Jain et al., whose Grey-Wolf-optimized CNN-LSTM for multilingual sentiment classification substantially outperformed both traditional and standalone deep-learning baselines [23]. These works establish CNN-LSTM hybrids and data-driven hyperparameter search as robust design patterns for affective computing, and directly motivate the Bayesian-optimization tuning strategy adopted for TAPE. In parallel, federated learning frameworks such as the FLWCO scheme by Dash et al. for diabetes and heart-disease prediction demonstrate collaborative training without exposing raw biosignals [24], providing a paradigm transferable to multi-site EEG monitoring. Notwithstanding this progress, prior efforts have primarily relied on interpretability constraints from filter-bank or topographic mappings and on multi-dimensional attention reshaping [17, 18], whereas explicit, differentiable gating mechanisms tied to retrospective memory-bias theories such as the peak-end rule have received limited attention. Moreover, existing attention-based EEG models largely target immediate emotional state classification, leaving retrospective evaluation prediction of multi-stimulus emotional sequences as an underexplored direction, which motivates the present study. Review of Temporal Series Dynamic Feature Extraction Models The Temporal Convolutional Network (TCN), designed for time-series modeling, achieves long-range dependency coverage through causal dilated convolutions while maintaining parameter efficiency. Qin et al. combined TCN with an efficient channel attention module in ETCNet, demonstrating advantages in capturing EEG temporal features for motor imagery, with a lightweight design suitable for real-time applications [25]. Xie et al. further proposed a bidirectional feature pyramid attention TCN that introduces multi-head attention for adaptive fusion of TCN outputs at different temporal scales, addressing limitations of standard TCN in integrating multi-scale temporal patterns [26]. In EEG affective computing, combining TCN with attention is emerging as a new trend. Altaheri et al. comprehensively reviewed deep learning for motor imagery EEG classification, noting that TCN–Transformer hybrids simultaneously address local temporal patterns and global contextual dependencies [27]. Li et al. proposed GMSS, a graph-based multi-task self-supervised model integrating spatial jigsaw, frequency jigsaw, and contrastive learning, achieving strong cross-subject emotion recognition on SEED [28]. Chen et al. proposed DAMGCN, a dual-attention graph convolutional network that uses Transformer self-attention to weight electrode channels and frequency bands separately, capturing spatial associations across brain regions through graph structures [29]. Qiu et al. proposed the semi-supervised fine-tuning self-supervised graph attention network SFT-SGAT, which captures functional connectivity dynamics among brain regions via graph attention and reduces reliance on labeled data through self-supervised pre-training, providing insights for attention-based temporal feature modeling in emotion recognition [30]. Dar et al. approached from the perspective of memory recall, using deep learning to decode EEG patterns related to emotion-evoked memory recall, achieving effective extraction of neural features of the recall process in an emotion charting task [31]. Synthesizing the above, EEG affective computing has accumulated rich foundations in feature extraction and model architecture; however, nearly all such work targets recognition of immediate emotional states, lacking systematic integration of temporal dynamic features with memory bias theories—particularly the peak-end effect. On this gap, this paper proposes an ERP dynamic feature extraction strategy oriented toward the peak-end effect structure and achieves predictive modeling of retrospective evaluation via a TCN–multi-head self-attention hybrid architecture. Methods Experimental Design and ERP Data Collection Paradigm This study extracts peak-end effect neural features from EEG ERPs and uses deep learning to predict retrospective evaluations of emotional experiences. Thirty right-handed participants (aged 18–28 years, equal males and females) were recruited; all had no neurological or psychiatric history and normal or corrected-to-normal vision. The protocol was approved by the institutional ethics committee and participants provided informed consent. Emotional induction materials were selected from the International Affective Picture System (IAPS) [32]. Sixty images were selected from each of three categories—high-arousal positive, high-arousal negative, and low-arousal neutral—based on IAPS standardized valence and arousal ratings, with valence and arousal balanced within each category. To simulate the temporal dynamics of real emotional experiences, a sequential emotion induction paradigm was adopted: each trial contained six consecutively presented images forming an “emotional sequence,” arranged according to a preset arousal curve to yield clear peak and endpoint moments. Peak position (2nd, 3rd, or 4th image) and endpoint intensity (high, medium, low) were fully crossed via a Latin square design to systematically manipulate the two key variables of the peak-end effect [33]. Each image was presented for 2000 ms with a random 800–1200 ms inter-stimulus interval to avoid ERP overlap. After a 3-s blank buffer, participants provided a retrospective overall evaluation of the sequence on a 9-point Likert scale; this rating served as the continuous label for model training [3]. The experiment comprised 90 trials divided into 3 blocks, with participants free to rest between blocks. Figure 1 illustrates the complete temporal procedure of a single trial. EEG data were collected using a 64-channel Ag/AgCl electrode cap following the extended international 10–20 system, with reference electrodes at bilateral mastoids and ground at the anterior midline. Sampling rate was 1000 Hz and electrode impedance was maintained below 5 kΩ. Horizontal and vertical electrooculogram (HEOG/VEOG) were recorded for artifact correction [34]. Figure 2 illustrates the electrode distribution and annotation of regions of interest. ERP Signal Preprocessing and Peak-End Feature Definition Preprocessing of raw EEG signals was performed in the EEGLAB toolbox [35]. A 0.1–30 Hz bandpass filter removed slow drifts and high-frequency electromyographic noise, followed by 50 Hz notch filtering for power-line interference. Bad channels were repaired via spherical spline interpolation, and Extended Infomax Independent Component Analysis (ICA) decomposed EEG into independent components [36]. The ICLabel automatic classifier was used to determine the source probability of each component; artifact components including ocular, myogenic, and cardiac sources were identified and removed, and the clean signal was reconstructed by retaining brain-source components [37]. Preprocessed data were segmented with image onset as the anchor; each epoch spanned −200 to + 1000 ms with baseline correction using the pre-stimulus 200-ms interval. After rejecting epochs exceeding \(\pm\) 100 ÎŒV, retained epochs were averaged to extract three ERP component feature parameters: EPN as the mean amplitude within 150–350 ms at occipital-temporal electrodes (O1/O2/P7/P8) [38]; the P300 component was characterized by peak amplitude and peak latency within the 300–500 ms window at central-parietal electrodes (Cz/CPz/Pz) [39]; and the LPP component was measured as the mean amplitude within the 400–800 ms window at midline parietal electrodes (Pz/POz/Oz) [40]. The ERP feature parameters are organized into structured representations reflecting the peak-end effect. Peak and endpoint moments serve as anchoring features (rather than global mean or minimum intensity) based on three converging lines of evidence. Behaviorally, Alaybek et al.'s meta-analysis established that peak and endpoint variables jointly account for substantially greater variance in retrospective evaluations than alternative summary statistics [7]. Neurophysiologically, LPP amplitude within 400–800 ms exhibits the strongest positive correlation with subjective arousal ratings, while EPN indexes early attentional capture and P300 indexes stimulus categorization [10, 11]. The maximum-LPP moment provides a physiologically interpretable anchor for peak intensity, and the endpoint feature is the ERP vector of the final stimulus, capturing the most recent neural state at retrospective rating. The peak is defined on absolute LPP amplitude because both high-arousal positive and negative stimuli elicit comparable positive LPP deflections, yielding a valence-invariant operationalization consistent with the arousal-driven character of the peak-end rule. Formally, let an emotional sequence contain \(N\) images, with the ERP feature vector for the \(i\)-th image denoted as \({\mathbf{e}}_{i} \in {\mathbb{R}}^{d}\), where \(d = 7\) is the feature dimensionality (comprising the amplitudes and latencies of the three ERP components described above); the ERP feature matrix for the entire sequence is \(\mathbf{E}={[{\mathbf{e}}_{1},{\mathbf{e}}_{2},\dots ,{\mathbf{e}}_{\text{N}}]}^{{\top}}\in {\mathbb{R}}^{N\times d}\). The peak feature is defined as the feature vector corresponding to the moment of maximum absolute LPP mean amplitude within the sequence: where \(\overline{{{\text{LPP}}}}_{i}\) denotes the mean LPP amplitude of the \(i\)-th image within the 400–800 ms window, and \(k^{*}\) is the index of the image with the maximum absolute LPP amplitude within that sequence. The endpoint feature is directly taken as the feature vector of the last image in the sequence, \({\mathbf{e}}_{{{\text{end}}}} = {\mathbf{e}}_{N}\). The aforementioned peak feature \({\mathbf{e}}_{{{\text{peak}}}}\) and endpoint feature \({\mathbf{e}}_{{{\text{end}}}}\) will serve as inputs to the peak-end attention module specifically designed within the model. Deep Learning Model Architecture for Temporal Dynamic Feature Extraction The predictive model proposed in this study uses a temporal convolutional network (TCN) as its backbone, integrating a multi-head self-attention mechanism and a peak-end feature gating module to construct an end-to-end retrospective evaluation prediction framework. Figure 3 presents the overall model architecture. - (1) Temporal Convolutional Encoder The feature encoding component of the model adopts a TCN structure composed of stacked residual blocks below. The core operation of TCN is dilated causal convolution, which expands the receptive field through exponentially increasing dilation rates while preserving causality, enabling the network to capture long-range temporal dependencies without introducing recurrent structures [25]. The dilated causal convolution operation for input sequence \(x\) at time step \(t\) is defined as: where \(f\) is a convolutional kernel of size \(K\), \(d\) is the dilation factor, and \(j\) is the index within the kernel. By incrementing the dilation rate according to \(d = 2^{l}\) (where \(l\) is the layer index) across successive residual blocks, TCN can achieve an effective receptive field of \((K - 1) \cdot 2^{L} + 1\) time steps within \(L\) residual blocks. Each residual block contains two dilated causal convolutional layers, each followed sequentially by batch normalization, ReLU activation, and spatial Dropout operations, with residual connections performing dimensionality alignment via \(1 \times 1\) convolution [42]. Figure 4 illustrates the internal structure of the TCN residual block and the dilation rate incrementation scheme. - (2) Multi-Head Self-Attention Module The output sequence of the TCN encoder is fed into the multi-head self-attention module to model global dependencies among all time steps within the sequence [43]. For the hidden representation \({\mathbf{H}} \in {\mathbb{R}}^{{N \times d_{h} }}\) output by the TCN, three sets of learnable linear mappings are used to generate the Query, Key, and Value matrices respectively; the attention computation follows the scaled dot-product formula: where \({\mathbf{Q}} = {\mathbf{HW}}^{Q}\), \({\mathbf{K}} = {\mathbf{HW}}^{K}\), \({\mathbf{V}} = {\mathbf{HW}}^{V}\) are the query, key, and value matrices respectively, and \(d_{k}\) is the key vector dimension, preventing large dot products from causing softmax gradient saturation. The model uses 8 parallel attention heads whose outputs are concatenated and projected via a linear transformation to yield the final attention output \({\mathbf{A}} \in {\mathbb{R}}^{{N \times d_{h} }}\) [44]. Figure 5 illustrates the computation flow of the multi-head self-attention mechanism. - (3) Peak-End Feature Gating Fusion Module To explicitly embed the cognitive mechanisms of the peak-end effect into the model, this study designs a Peak-End Gate fusion module. This module receives the peak feature \({\mathbf{e}}_{{{\text{peak}}}}\) and endpoint feature \({\mathbf{e}}_{{{\text{end}}}}\) defined above, and adaptively modulates the contribution weights of each time step in the attention output through learnable gating signals. The gating signal is computed as: where \([{\mkern 1mu} \cdot {\mkern 1mu} ,{\mkern 1mu} \cdot {\mkern 1mu} ]\) denotes vector concatenation along the feature dimension, \({\mathbf{W}}_{g} \in {\mathbb{R}}^{{d_{h} \times 2d}}\) and \({\mathbf{b}}_{g} \in {\mathbb{R}}^{{d_{h} }}\) are learnable parameters, and \(\sigma\) is the Sigmoid activation function. The gating signal \({\mathbf{g}} \in (0,1)^{{d_{h} }}\) applies element-wise weighted modulation to the attention output; after global average pooling, the result is fed into a two-layer fully connected regression head to output the final retrospective evaluation prediction score \(\hat{y}\). This design embeds the theoretical assumptions of the peak-end effect into the network structure in the form of differentiable operations, endowing the model with theoretical constraints on top of data-driven learning [45]. Model Training Strategy and Evaluation Metrics The prediction target is the retrospective overall rating (1–9 continuous value) provided after each trial, so training uses a regression framework. The loss combines mean squared error (MSE) and Huber loss to balance sensitivity to small errors with robustness to large ones [46]: where \(\hat{y}_{i}\) and \(y_{i}\) are the predicted value and ground-truth label for the \(i\)-th sample respectively, \(M\) is the mini-batch size, \(\delta_{\varepsilon } ( \cdot )\) is the Huber loss function with threshold \(\varepsilon = 1.0\), and \(\lambda\) is the balancing coefficient (set to 0.5 in this study). The optimizer is AdamW [47], with an initial learning rate of \(1 \times 10^{ - 3}\) and a cosine annealing schedule for gradual decay during training. The total number of training epochs is 200, the batch size is 32, and the Dropout rate is 0.3. To prevent overfitting, \(L_{2}\) regularization (weight decay coefficient \(5 \times 10^{ - 4}\)) is applied to the fully connected layers, and an early stopping strategy is adopted (patience = 20) [48]. To rule out sensitivity to coarse hyperparameter choices, all key hyperparameters were tuned via Bayesian-optimization-based search on an inner validation split before the main experiments. The search space comprised: number of TCN residual blocks \(L \in \{ 2,3,4,5\}\); convolutional kernel size \(K \in \{ 3,5,7\}\); hidden dimensionality \(d_{h} \in \{ 64,128,256\}\); number of attention heads \(\in \{ 4,8,16\}\); learning rate \(\in [1 \times 10^{ - 4} ,1 \times 10^{ - 2} ]\) (log-uniform); Dropout rate \(\in \{ 0.1,0.2,0.3,0.4,0.5\}\); and loss balancing coefficient \(\lambda \in [0.1,0.9]\). A total of 50 Bayesian-optimization iterations were executed using a Gaussian-process surrogate model with the expected-improvement acquisition function, taking validation-set Pearson r as the optimization target. The final selected hyperparameter combination, summarized in Table 1, was then frozen for all subsequent benchmark experiments to prevent test-set leakage. The complete configuration of model hyperparameters is shown in Table 1. Two additional baseline families were introduced to quantify the gain of deep temporal modeling over conventional approaches: (i) shallow regression models—Linear and Ridge—trained directly on the concatenated ERP feature vector across the six-image sequence, with the Ridge regularization coefficient selected via fivefold cross-validation from \(\{ 10^{ - 3} ,10^{ - 2} ,10^{ - 1} ,1,10,10^{2} \}\); and (ii) a classical EEG feature-engineering baseline (denoted SVR-PSD (Power spectral density)-DE (Differential entropy)), in which Power Spectral Density (PSD) and Differential Entropy (DE) features are extracted from the five canonical frequency bands (\(\delta\), \(\theta\), \(\alpha\), \(\beta\), \(\gamma\)) per image, concatenated across the six-image sequence, and fed into a Support Vector Regressor with RBF kernel. Combined with the SVR (RBF), Random Forest, LSTM, CNN-LSTM, and Transformer baselines, these form a multi-tier benchmark spanning linear models, classical EEG feature engineering, and advanced deep architectures. Considering the high inter-individual variability of EEG signals, model evaluation adopts a Leave-One-Subject-Out (LOSO) cross-validation strategy [49]. In each fold, one participant's data is the test set, with the remaining data split 9:1 for training/validation; the cycle continues until each participant has served as the test subject once. LOSO rigorously evaluates generalization to completely unseen participants and avoids performance overestimation from within-subject data leakage [50]. To further verify that the advantage is not contingent on a particular partition, ten independent runs of stratified fivefold cross-validation (each with a different random seed) were conducted across all methods, yielding a total of \(10 \times 5 = 50\) independent performance evaluations per model, the distribution of which is summarized via box-and-whisker plots in Section "Regression Prediction Performance of the Model". Evaluation metrics span regression and classification. For regression, mean absolute error (MAE) and root mean squared error (RMSE) measure the deviation between predicted and true ratings, while Pearson correlation coefficient r is used to assess the linear correlation between predicted and true values [51]. In the classification dimension, retrospective ratings are divided into three levels—“low,” “medium,” and “high”—according to tertiles, and classification accuracy and weighted \(F_{1}\)-score are computed to enable cross-comparison with existing emotion recognition studies [52]. The weighted \(F_{1}\)-score is formally defined as the support-weighted average of per-class \(F_{1}\) scores: where \(C = 3\) is the number of classes, \(n_{c}\) is the number of samples belonging to class \(c\), \(N\) is the total number of samples, and \(P_{c}\) and \(R_{c}\) are respectively the precision and recall of class \(c\). The weighted formulation aggregates per-class \(F_{1}\) scores using their support as weights, thereby providing a more faithful summary of overall model performance under the moderately class-imbalanced tertile partitioning. Beyond converting continuous regression outputs to three-level labels via tertile cuts (the default reporting pipeline), a directly-trained classification variant of TAPE was developed to verify multi-class predictive capability under an end-to-end classification objective. This variant preserves the TCN encoder, multi-head self-attention module, and peak-end feature gating module, replacing only the regression head with a three-way softmax head. The training loss is the class-weighted cross-entropy: where \(\hat{p}_{i,c}\) is the predicted probability that sample \(i\) belongs to class \(c\), \(y_{i,c} \in \{ 0,1\}\) is the one-hot label, and \(w_{c} = N/(C \cdot n_{c} )\) is the inverse-frequency class weight compensating for tertile-induced class imbalance. All other hyperparameters (TCN depth, attention heads, learning rate, Dropout, weight decay, batch size, epochs) follow Table 1, with the same LOSO protocol. Performance of this directly-trained variant is reported alongside the tertile-derived results in Section "Classification Performance and Comparative Analysis of the Model". The statistical inference settings are specified as follows. (1) Repeated-measures ANOVA design. Repeated-measures ANOVAs are used in Section "Descriptive Statistics of ERP Features and Behavioral Validation of the Peak-End Effect" to examine differences in ERP component parameters (EPN mean amplitude, P300 peak amplitude, P300 peak latency, LPP mean amplitude) across the three emotional conditions (high-arousal positive, high-arousal negative, low-arousal neutral). Emotional condition is the sole within-subject factor with \(k\) = 3 levels and \(n\) = 30 participants, so the omnibus \(F\) statistic is reported as \(F\)(2, 58) with numerator df = \(k\) − 1 = 2 and denominator df = (\(n\) − 1)(\(k\) − 1) = 58. Mauchly's test of sphericity is performed prior to each ANOVA; when violated (\(p\) < 0.05), Greenhouse–Geisser correction is applied and the corrected \(p\)-value is reported. Effect sizes are reported as partial \(\eta^{2}\) (Cohen's benchmarks: small 0.01, medium 0.06, large 0.14). All omnibus ANOVA tests are two-sided at \(\alpha\) = 0.05. Post-hoc pairwise comparisons use paired-sample \(t\)-tests with Bonferroni correction across the three contrasts (\(\alpha_{corrected}\) = 0.017). (2) Paired-sample \(t\)-test design for inter-model comparisons. Paired-sample \(t\)-tests in Sections "Regression Prediction Performance of the Model", "Classification Performance and Comparative Analysis of the Model", "Ablation Study Results", and "External Dataset Validation" compare TAPE against each baseline or ablation variant on cross-validation metrics. Pairing is defined by fold identity: under LOSO, each of the 30 folds yields one metric per model, and the paired difference is \(d_{i} = m_{TAPE}^{(i)} - m_{baseline}^{(i)}\), where \(m^{(i)}\) denotes the metric value (MAE, RMSE, Pearson \(r\), or accuracy) on the \(i\)-th fold. The null hypothesis \(H_{0}\): \(\mu_{d}\) = 0 is tested two-sided at \(\alpha\) = 0.05 with \(t\)(df = 29). For the SEED external validation (Section "External Dataset Validation"), pairing is across the 15 SEED participants, yielding \(t\)(df = 14). Prior to each test, the Shapiro–Wilk test is applied to the paired-difference distribution; when normality is violated (\(p\) < 0.05), the Wilcoxon signed-rank test substitutes and its \(z\) statistic is reported. (3) Effect sizes, confidence intervals, and multiple-comparison correction. Effect sizes are Cohen's \(d_{z} = M_{d} /SD_{d}\) (small 0.20, medium 0.50, large 0.80). The two-sided 95% CI of the mean paired difference is \(M_{d} \pm t_{0.975,df} \times SD_{d} /\sqrt n\), reported in Sections "Regression Prediction Performance of the Model", "Ablation Study Results", and "External Dataset Validation". Across the seven inter-baseline comparisons (Section "Regression Prediction Performance of the Model"), the four ablation comparisons (Section "Ablation Study Results"), and the three SEED comparisons (Section "External Dataset Validation"), Holm-Bonferroni sequential correction controls the family-wise error rate at \(\alpha_{family}\) = 0.05: raw \(p\)-values are sorted ascending and the \(k\)-th smallest is compared against \(\alpha /(m - k + 1)\), where \(m\) is the family size. All statistical analyses use SPSS 27.0 and Python 3.10 with SciPy 1.11.3 (scipy.stats). All metrics report the mean and standard deviation across the 30 participants over the 30 LOSO folds. Figure 6 illustrates the execution procedure of LOSO cross-validation. To further verify the independent contribution of the peak-end feature gating module, an ablation protocol was designed: separately removing the peak-end gating module (retaining only TCN-Attention), removing attention (retaining only TCN-PeakEnd Gate), and replacing TCN with a standard LSTM as the temporal encoder, validating each module's contribution by comparing each variant under the same LOSO evaluation framework [27]. Results Descriptive Statistics of ERP Features and Behavioral Validation of the Peak-End Effect The 30 participants completed 2,700 trials (90 per participant); after removing artifact-contaminated trials, 2,574 valid trials were retained (95.3% retention). Descriptive statistics of ERP components across the three emotional conditions are shown in Table 2. Repeated-measures ANOVAs were performed on each ERP component with emotional condition as the sole within-subject factor (\(k\) = 3 levels, \(n\) = 30), yielding omnibus \(F\)(2, 58). Mauchly's sphericity test was non-significant for all four dependent variables (all \(p\) > 0.10), so no Greenhouse–Geisser correction was applied. Effect sizes are partial \(\eta^{2}\) (small 0.01, medium 0.06, large 0.14). All three ERP components differed significantly across conditions: high-arousal negative elicited the largest LPP amplitude (10.41 \(\pm\) 3.05 \(\mu\) V), significantly greater than positive and neutral, consistent with motivational salience theory; EPN showed pronounced negative deflection under both high-arousal conditions, indicating preferential attentional capture within 150–350 ms; P300 latency was longest under neutral, suggesting slower cognitive evaluation of low-arousal stimuli. Post-hoc pairwise comparisons used paired-sample \(t\)-tests with Bonferroni correction (three contrasts per component, \(\alpha_{corrected}\) = 0.017); results in Table 3. For LPP: high-arousal negative exceeded positive (\(t\)(29) = 4.05, \(p_{corrected}\) = 0.001, \(d_{z}\) = 0.74, 95% CI [0.73, 2.23] \(\mu\) V); both high-arousal conditions substantially exceeded neutral (Negative vs Neutral: \(t\)(29) = 13.05, \(d_{z}\) = 2.38; Positive vs Neutral: \(t\)(29) = 11.50, \(d_{z}\) = 2.10; both \(p_{corrected}\) < 0.001). For EPN: both high-arousal conditions showed larger negative deflections than neutral (Negative vs Neutral: \(t\)(29) = −11.59, \(d_{z}\) = 2.12; Positive vs Neutral: \(t\)(29) = −11.97, \(d_{z}\) = 2.19; both \(p_{corrected}\) < 0.001); the two high-arousal valences did not differ (\(t\)(29) = −2.26, \(p_{corrected}\) = 0.094, \(d_{z}\) = 0.41). For P300 amplitude: both high-arousal conditions were significantly larger than neutral (both \(p_{corrected}\) < 0.001, \(d_{z}\) = 1.16 and 1.40) with no significant difference between valences (\(t\)(29) = 2.45, \(p_{corrected}\) = 0.062, \(d_{z}\) = 0.45). For P300 latency: neutral was delayed relative to both high-arousal conditions (Neutral vs Negative: \(t\)(29) = 3.20, \(p_{corrected}\) = 0.009, \(d_{z}\) = 0.58; Neutral vs Positive: \(t\)(29) = 4.54, \(p_{corrected}\) < 0.001, \(d_{z}\) = 0.83); high-arousal conditions did not differ (\(t\)(29) = 1.25, \(p_{corrected}\) = 0.663, \(d_{z}\) = 0.23). These results confirm that ERP modulation is driven primarily by arousal intensity, agreeing with the rationale for these three components as peak-end neural markers. To validate the peak-end effect behaviorally, hierarchical regression on the behavioral data used retrospective overall ratings as the dependent variable, with three predictors entered separately: the sequence mean arousal rating (global mean), the peak-image arousal rating, and the endpoint-image arousal rating. The global-mean model accounted for 42.7% of the variance (\(R^{2} = .427\)), and after additionally entering the peak and endpoint variables, the \(R^{2}\) increment was 0.138 (\(\Delta R^{2} = .138\), \(p < .001\)), bringing the total explained variance to 56.5%. The standardized regression coefficient for the peak variable was \(\beta = .314\) (\(p < .001\)), and for the endpoint variable was \(\beta = .267\) (\(p < .001\)), with both being independently significant predictors. Figure 7 presents the partial regression scatter plots of peak-end features on retrospective ratings. Regression Prediction Performance of the Model Under LOSO cross-validation, regression performance of the proposed TCN-Attention-PeakEnd Gate (TAPE) model against multiple baselines is presented in Table 4. Baselines span three tiers: shallow regression (Linear, Ridge), traditional machine learning and classical EEG-feature methods (SVR-RBF, Random Forest, SVR-PSD-DE), and deep learning (standard LSTM, CNN-LSTM, Transformer). TAPE achieved optimal performance across all three regression metrics. Compared with the second-best Transformer, TAPE reduced MAE by 13.3% (1.197 → 1.038), RMSE by 12.3%, and improved Pearson \(r\) by 10.5%. MAE was reduced ≈42.1% versus the strongest shallow baseline (Ridge) and 40.9% versus the classical-EEG-feature baseline SVR-PSD-DE, showing that gain derives jointly from deep temporal modeling and ERP-feature representation. Paired-sample \(t\)-tests with Holm-Bonferroni correction confirmed TAPE outperformed every comparison on MAE (all corrected \(p\) < 0.01); full statistics (test values, corrected \(p\)-values, Cohen's \(d_{z}\), and 95% CIs of the mean paired difference in MAE, baseline minus TAPE) are in Table 5. Under LOSO, each paired comparison uses \(n\) = 30 folds, df = 29, with CIs computed as \(M_{d} \pm t_{0.975,29} \times SD_{d} /\sqrt n\) (\(t_{0.975,29}\) = 2.045). Cohen's \(d_{z}\) ranges from 0.71 (vs Transformer) to 1.83 (vs Linear), indicating medium-to-very-large effect sizes. To evaluate stability under different random partitions, stratified fivefold cross-validation was repeated ten times with distinct random seeds, yielding 50 independent evaluations per model. The distribution of Pearson r across these 50 evaluations is shown in Fig. 8 (box-and-whisker plot), with median and IQR of each model summarized in Table 6. TAPE consistently achieves the highest median Pearson r (0.682) and smallest IQR (0.054), indicating its advantage is not driven by favorable data splits. TAPE's IQR is smaller than every baseline (Transformer 0.062, CNN-LSTM 0.067, LSTM 0.072, all shallow baselines exceeding 0.080), suggesting the peak-end gating inductive bias also reduces cross-partition variance. Figure 9 presents the relationship between predicted and true values of the TAPE model across all participants. To more intuitively present the distribution of prediction accuracy at the individual level, Fig. 10 presents violin plots of Pearson correlation coefficients for each of the 30 participants. As can be seen from Fig. 10, the Pearson r distribution across 30 participants ranged 0.51–0.78, with median 0.67 and interquartile range 0.09, indicating relatively stable predictive capability across individuals. In contrast, the LSTM model showed markedly greater inter-individual variability (IQR 0.16), with 5 participants having r values below 0.40, reflecting the limitations of recurrent networks in cross-subject generalization. Classification Performance and Comparative Analysis of the Model After converting retrospective ratings to three-level labels via tertile cuts, classification performance is presented in Table 7. For completeness, the directly-trained variant TAPE-Cls (using the class-weighted cross-entropy loss in Eq. (7)) is reported alongside the regression-derived results. TAPE (regression-then-tertile pipeline) achieved 70.4% classification accuracy, 4.7 pp above Transformer and ≈20 pp above traditional SVR. The weighted \(F_{1}\) was 0.698, highest among all comparison models excluding the direct-classification variant. TAPE-Cls (directly trained) attained slightly higher accuracy (71.2%), showing the architecture is compatible with both regression and classification objectives; the 0.8 pp gap suggests the regression-then-tertile pipeline is near-optimal when continuous ratings are available. Paired \(t\)-tests with Holm-Bonferroni correction confirmed TAPE outperformed every baseline at \(p\) < 0.01; against Transformer: \(t\)(29) = 3.94, \(p_{corrected}\) < 0.001, \(d_{z}\) = 0.72, 95% CI of mean paired difference [2.26, 7.14] pp. Figure 11 presents the normalized confusion matrix on the three-level classification task. The confusion matrix shows the “high evaluation” category had the highest recall (74.6%), followed by “low evaluation” (71.2%), and “medium evaluation” lower (65.3%). This pattern is common in emotion recognition: extreme emotional states have more discriminable neural representations, while moderate-intensity experiences overlap more with the extremes in ERP feature space. Ablation Study Results To quantify each module's contribution, systematic ablation was conducted per Section "Model Training Strategy and Evaluation Metrics". Table 8 presents the comparison across four model variants and the complete TAPE model. The ablation results reveal hierarchical component contributions. Removing peak-end gating increased MAE by 11.0% and decreased accuracy by 3.5 pp, demonstrating the effectiveness of gating-based peak-end embedding. Removing multi-head self-attention decreased Pearson \(r\) from 0.654 to 0.589 (− 9.9%), showing global temporal dependency modeling is indispensable. Replacing TCN with LSTM further increased MAE by 17.1%, validating dilated causal convolution over recurrent structures for ERP time series. The TCN-backbone-only variant was lowest; its gap from the complete model (MAE 0.236, accuracy 8.6 pp) reflects the joint gain from attention and gating. Paired \(t\)-tests with Holm-Bonferroni correction confirmed all MAE differences are significant (Table 9), with \(N_{folds}\) = 30, df = 29; 95% CIs use \(t_{0.975,29}\) = 2.045. Progressively increasing effect sizes indicate each ablated module makes a non-trivial independent contribution. Figure 12 presents a radar chart comparison of the four core metrics across each ablation variant. Attention Weight Visualization and Neural Mechanism Interpretation To test whether the model's temporal attention patterns match peak-end theoretical expectations, the multi-head self-attention weight matrices of the trained TAPE model were analyzed. Attention weights from all test trials were grouped and averaged by image position (1st–6th) and emotional condition. Table 10 presents the average attention weights by position. Under high-arousal conditions (especially negative), positions 3 and 4 received the highest attention weights (0.208 and 0.197)—precisely where peak stimuli most frequently occurred by design. The endpoint (position 6) weight was significantly higher under positive (0.189) than negative (0.162), an asymmetry suggesting a stronger endpoint-impression effect on retrospective evaluation for positive experiences. Under neutral, the six-position distribution was relatively uniform (0.159–0.173), indicating the model assigns approximately equal attention when sequences lack salient peak stimuli. Figure 13 presents an interaction heatmap of attention weights across sequence positions and emotional conditions. Correlation analysis between the gating signals and ERP features was conducted across all 2,574 valid trials. A moderate positive correlation was found between gating activation values and LPP amplitude (\(r\) = 0.483, 95% CI [0.451, 0.513] via Fisher \(z\) transformation, \(p\) < 0.001), indicating that the gating mechanism captured LPP signal features associated with emotional arousal salience during learning. This provides computational-level evidence for the ERP neural basis of the peak-end effect: the model did not merely memorize input–output mappings but learned a feature-weighting strategy consistent with cognitive theory. Table 11 presents the correlation analysis results between gating signal activation values and each ERP component. The correlation strength between the peak gating signal and LPP amplitude (\(r = 0.483\)) was higher than its correlations with EPN (\(r = 0.312\)) and P300 (\(r = 0.358\)), indicating that gating relies primarily on sustained emotional processing signals carried by LPP when encoding peak information. The endpoint gating signal showed a similar but weaker correlation pattern; its correlation with P300 latency did not reach significance (\(r = - 0.142\), \(p = .061\)), suggesting that endpoint cognitive evaluation speed has limited influence on endpoint gating. These correlations support the model design: using LPP amplitude as the peak-feature anchor is physiologically meaningful. External Dataset Validation To further evaluate whether the proposed TAPE architecture generalizes beyond the laboratory dataset collected in this study, external validation was conducted on the publicly available SEED dataset [53]. SEED comprises 62-channel EEG from 15 right-handed Chinese participants viewing 15 emotional film clips (5 positive, 5 negative, 5 neutral), each ≈4 min. Since SEED is a single-stimulus emotion classification task without retrospective evaluation labels on multi-stimulus sequences, the external validation focused on the transferability of TAPE's feature-extraction and temporal-modeling capabilities rather than the full sequential-rating pipeline. Each 4-min clip was uniformly segmented into twelve 20-s sub-segments, with ERP-style temporal features (mean amplitudes within EPN-, P300-, and LPP-equivalent windows after baseline-aligned re-referencing) extracted as a 12-step feature sequence per clip. The clip-level emotion label (positive/neutral/negative) was the prediction target. The TCN encoder, multi-head self-attention, and peak-end gating modules were retained with hyperparameters as in Table 1; the regression head was replaced with a 3-way softmax head trained with the class-weighted cross-entropy loss in Eq. (7). Peak position was determined by the maximum-absolute-LPP rule and endpoint by the final sub-segment. LOSO was applied across the 15 SEED participants. Classification performance of TAPE against the three strongest baselines on SEED is in Table 12. TAPE attained 84.7% three-class accuracy on SEED, surpassing Transformer (81.9%), CNN-LSTM (78.6%), and LSTM (75.4%). Paired-sample \(t\)-tests confirmed the TAPE-vs-Transformer gain is significant (\(t\)(14) = 3.21, \(p\) = 0.006, Cohen's \(d_{z}\) = 0.83, 95% CI [0.93, 4.67] pp, with df = 14 from pairing across 15 SEED participants). The absolute TAPE–Transformer gap on SEED (2.8 pp) is narrower than on the in-house dataset (4.7 pp), consistent with the theoretical expectation that peak-end gating attains its largest benefit when the task involves retrospective evaluation of multi-stimulus sequences. Per-subject accuracies on SEED are shown in Fig. 14. The consistent ranking of TAPE above Transformer, CNN-LSTM, and LSTM across two independent datasets supports that its architectural advantage is not confined to the specific paradigm or sample, while the narrower gap on SEED reflects that the peak-end gating mechanism is tailored to retrospective evaluation tasks and yields its largest benefit there. Discussion and Implications Association Between Model Prediction Performance and the Peak-End Effect Theory TAPE achieved MAE = 1.038, Pearson \(r\) = 0.654, and classification accuracy of 70.4% under LOSO cross-validation, ranking first among all comparison methods. Table 13 summarizes the contribution hierarchy of the peak-end theory-guided modular design: removing the peak-end gating module increased MAE by 0.114 (≈11.0%), removing attention decreased Pearson \(r\) by 0.065 (≈9.9%), and replacing TCN with LSTM caused a 6.3-pp accuracy loss. The three ablations each caused the most prominent degradation in a different dimension, showing that the three core modules serve irreplaceable and synergistic roles. All three differences reached statistical significance with medium-to-large effect sizes (Cohen's \(d_{z}\) = 0.64–0.95) under Holm-Bonferroni-corrected paired-sample \(t\)-tests (Section "Ablation Study Results"), ruling out random variation. The 4.7-pp accuracy improvement of TAPE over the standard Transformer (65.7%) derives not from increased parameter count or depth but from the structural constraints imposed by peak-end theory. The gating module compels the network to attend to peak and endpoint moments, an inductive bias mirroring human retrospective memory biases. Statistical robustness is supported by the paired \(t\)-test against Transformer (\(t\)(29) = 3.87, \(p\) < 0.001 after Holm-Bonferroni correction, \(d_{z}\) = 0.71; Table 5) and by 50 independent fivefold cross-validation evaluations, where TAPE attained the highest median Pearson \(r\) (0.682) and smallest IQR (0.054) (Fig. 8, Table 6). Concurrent gains in central tendency and reduction in dispersion indicate that peak-end constraints raise accuracy and stabilize behavior across partitions, consistent with well-motivated inductive biases mitigating variance under limited supervision. This supports Kahneman's "snapshot model": retrospective evaluation relies on neural imprints of a few key moments rather than moment-by-moment integration. Nonetheless, TAPE's \(r\) = 0.654 remains below the behavioral peak-end ceiling (hierarchical \(R^{2}\) = 0.565), suggesting that higher-level processes—prefrontal-mediated emotion regulation and metacognitive judgment—may not be fully captured by ERP channels alone. Cognitive Neuroscientific Significance of Attention Weight Distributions The attention weight patterns in Section "Results" provide a window into how deep learning models “implicitly” learn the peak-end cognitive structure. Under high-arousal negative, positions 3 and 4 received weights 0.208 and 0.197—where peak stimuli were most likely by design—while non-peak positions 1 and 5 received only 0.131 and 0.149. Under neutral, the distribution was nearly uniform (0.159–0.173), consistent with a further corollary of peak-end theory: when arousal variability is low, retrospective evaluation reverts toward a mean strategy across the sequence. This condition-dependent pattern was not preset but emerged automatically from end-to-end training; its consistency with cognitive theory provides empirical evidence for interpretability research on deep learning in affective computing. Another notable finding is that the endpoint weight (position 6) under positive (0.189) was significantly higher than under negative (0.162); this valence-dependent asymmetry has been underexplored in prior behavioral peak-end research. One explanation is that the higher endpoint weight in positive experiences reflects the neural basis of the “ending-on-a-high-note effect”: when a positive sequence ends pleasantly, sustained positive LPP processing reinforces the endpoint contribution in retrospective integration, whereas in negative sequences the endpoint weight is diluted by the higher peak weight. The 0.483 correlation between gating and LPP amplitude, substantially higher than EPN (0.312) or P300 (0.358), corroborates this: sustained emotional elaboration in LPP is the primary signal source for peak-end gating. Methodological Implications and Application Prospects The core methodological insight is that embedding mature cognitive-psychology theories into deep learning architectures as differentiable operations is a feasible and productive strategy. Traditional EEG emotion recognition often pursues purely data-driven end-to-end mappings, neglecting the theoretical foundations of emotional cognition. TAPE demonstrates that this "theory-data" dual-driven approach improves predictive accuracy and interpretability; as shown in Table 11, the ordered correlations between gating signals and ERP components enable computational validation of neural mechanism hypotheses. The approach can be extended to other EEG tasks involving cognitive biases, e.g., embedding anchoring or framing effects into decoding models for decision-related ERP signals. The framework can extend to multiple domains. In clinical psychology, retrospective symptom assessment (e.g., PHQ-9) is subject to peak-end bias; TAPE's core ideas can transplant into wearable-EEG dynamic symptom monitoring by capturing ERP temporal features. In user experience assessment, combining EEG with interactive processes and peak-end gating to extract key experience nodes offers more objective, refined tools than traditional questionnaires. Cross-context transferability is evidenced by SEED (Sect. "External Dataset Validation"), where TAPE maintained its ranking above the strongest baselines (84.7% vs 81.9% Transformer, 78.6% CNN-LSTM, 75.4% LSTM under LOSO). The narrower TAPE–Transformer gap on SEED (2.8 pp vs 4.7 pp on the in-house dataset) is consistent with peak-end gating attaining its largest benefit for multi-stimulus retrospective evaluation while still providing gains in single-stimulus classification. Transfer from laboratory to real-world scenarios still faces engineering challenges regarding signal quality, individual differences, and real-time efficiency. Four future directions extend this work. First, evaluate peak-end gating on larger multi-center EEG datasets spanning heterogeneous sites, hardware, and cultural populations to test generalization beyond the current single-laboratory sample. Second, develop a real-time TAPE implementation via structured pruning, knowledge distillation, and low-latency inference kernels for wearable EEG headsets during naturalistic tasks; edge deployment will further allow integration with low-power spiking-network EEG pipelines already used for stress detection [20, 22]. Third, extend peak-end gating to multimodal integration jointly modeling EEG, ECG (Electrocardiography), GSR (Galvanic skin response), respiration, and eye-tracking, so that peak and endpoint anchors are identified from cross-modal arousal indicators rather than EEG alone, improving robustness and ecological validity. Fourth, pursue clinical translation in longitudinal symptom monitoring, where TAPE-derived estimates are compared against validated scales (PHQ-9, GAD-7) to quantify and correct peak-end recall bias in depression, anxiety, and post-traumatic stress disorder cohorts. Conclusion This study proposes an ERP signal decoding framework integrating the peak-end effect theory with deep learning for retrospective evaluation prediction. A sequential emotion induction paradigm based on IAPS images was designed, systematically manipulating peak position and endpoint intensity, and feature parameters of three ERP components—EPN, P300, LPP—were extracted from 64-channel EEG as neural markers of the peak-end effect. The TCN-Attention-PeakEnd Gate (TAPE) model, combining dilated causal convolution, multi-head self-attention, and peak-end feature gating, achieved MAE = 1.038, Pearson r = 0.654, and three-level classification accuracy of 70.4% under leave-one-subject-out cross-validation across 30 participants, significantly outperforming eight baselines spanning shallow regression, classical EEG-feature engineering, and deep learning architectures under Holm–Bonferroni-corrected paired-sample t-tests. The performance advantage of TAPE was further confirmed to be stable across 50 independent fivefold cross-validation evaluations, in which the model consistently attained both the highest median Pearson r and the smallest interquartile range among all methods; external validation on SEED replicated TAPE's ranking above the strongest deep-learning baselines (84.7% versus 81.9% for Transformer), supporting cross-dataset generalizability. Ablation confirmed independent contributions of the peak-end gating module (MAE + 11.0% upon removal), the multi-head self-attention module (Pearson r − 9.9% upon removal), and the TCN encoder (accuracy − 6.3 pp when replaced by LSTM); attention weight visualization revealed high consistency between temporal patterns learned and theoretical expectations of the peak-end effect. Hierarchical regression validated at the behavioral level the independent predictive efficacy of peak and endpoint variables for retrospective ratings (\(\Delta R^{2} = 0.138\)), and the significant positive correlation of 0.483 between gating signals and LPP amplitude provides neurophysiological evidence at the computational level for this theoretical hypothesis. This study has several limitations. Although the SEED replication provides initial evidence of generalizability, the sample size (\(N\) = 30) of the in-house experiment may constrain conclusions about broader inter-individual transfer. Emotional induction was limited to static images, leaving ecological validity to be improved. The gating weighting for peak and endpoint features is fixed and does not account for potentially heterogeneous peak-end bias tendencies across individuals. The four directions in Section "Methodological Implications and Application Prospects"—large-scale multi-center validation, real-time on-device implementation, multimodal integration, and clinical translation to longitudinal symptom monitoring—constitute a coherent roadmap. Along these, dynamic video or virtual reality emotion induction in larger cohorts should enhance ecological validity, while adaptive gating and personalized transfer learning will improve robustness under cross-subject and cross-scenario conditions. Data Availability All data generated in this study are included in the manuscript, and raw data can be obtained from the corresponding author. References Horwitz AG, McCarthy K, Sen S. A review of the peak-end rule in mental health contexts. Curr Opin Psychol. 2024;58:101845. https://doi.org/10.1016/j.copsyc.2024.101845. Scharbert J, Utesch K, Reiter TF, et al. If you were happy and you know it, clap your hands! Testing the peak-end rule for retrospective judgments of well-being in everyday life. Eur J Pers. 2025;39(1):55–69. https://doi.org/10.1177/08902070241235969. Horwitz AG, Zhao Z, Sen S. Peak-end bias in retrospective recall of depressive symptoms on the PHQ-9. Psychol Assess. 2023;35(4):378–81. https://doi.org/10.1037/pas0001219. Cheng Z, Bu X, Wang Q, et al. EEG-based emotion recognition using multi-scale dynamic CNN and gated transformer. Sci Rep. 2024;14:31319. https://doi.org/10.1038/s41598-024-82705-z. Wang T, Huang X, Xiao Z, et al. EEG emotion recognition based on differential entropy feature matrix through 2D-CNN-LSTM network. EURASIP J Adv Signal Process. 2024. https://doi.org/10.1186/s13634-024-01146-y. Qiao Y, Mu J, Xie J, et al. Music emotion recognition based on temporal convolutional attention network using EEG. Front Hum Neurosci. 2024;18:1324897. https://doi.org/10.3389/fnhum.2024.1324897. Alaybek B, Dalal RS, Fyffe S, et al. All’s well that ends (and peaks) well? A meta-analysis of the peak-end rule and duration neglect. Organ Behav Hum Decis Process. 2022;170:104149. https://doi.org/10.1016/j.obhdp.2022.104149. MĂŒller UWD, Gerdes ABM, Alpers GW. Time is a great healer: peak-end memory bias in anxiety – induced by threat of shock. Behav Res Ther. 2022;159:104206. https://doi.org/10.1016/j.brat.2022.104206. McCullough H, Padgett D, Han S, et al. First impressions vs. the peak-end rule: episodic evaluations in a service experience and the moderating effect of retrospective delay. J Bus Res. 2024;185:114899. https://doi.org/10.1016/j.jbusres.2024.114899. Farkas AH, Sabatinelli D. Emotional perception: divergence of early and late event-related potential modulation. J Cogn Neurosci. 2023;35(6):941–56. https://doi.org/10.1162/jocn_a_01988. Schupp HT, Kirmse U, SchmĂ€lzle R. Replication at the level of the individual: EPN and LPP components of affective stimulus processing. Psychophysiology. 2023;60(9):e14318. https://doi.org/10.1111/psyp.14318. Schindler S, Bublatzky F. Attention and emotion: an integrative review of emotional face processing as a function of attention. Cortex. 2024;172:82–107. https://doi.org/10.1016/j.cortex.2023.11.013. Hajcak G, Foti D. Significance?... Significance! Empirical, methodological, and theoretical connections between the late positive potential and P300 as neural responses to stimulus significance: an integrative review. Psychophysiology. 2020;57(7):e13570. https://doi.org/10.1111/psyp.13570. Li C, Chen J, Fan H, et al. EEG-based emotion recognition via transformer neural architecture search. IEEE Trans Industr Inf. 2023;19(4):6016–25. https://doi.org/10.1109/TII.2022.3169146. Huang Z, Ma Y, Wang R, et al. A model for EEG-based emotion recognition: CNN-Bi-LSTM with attention mechanism. Electronics (Basel). 2023;12(14):3188. https://doi.org/10.3390/electronics12143188. Li T, Fu B, Wu Z, et al. EEG-based emotion recognition using spatial-temporal-connective features via multi-scale CNN. IEEE Access. 2023;11:41859–67. https://doi.org/10.1109/ACCESS.2023.3269467. Liu R, Chao Y, Ma X, et al. ERTNet: an interpretable transformer-based framework for EEG emotion recognition. Front Neurosci. 2024;18:1320645. https://doi.org/10.3389/fnins.2024.1320645. Xu Y, Zhong C, Quan S, et al. AMDET: Attention based multiple dimensions EEG transformer for emotion recognition. IEEE Trans Affect Comput. 2024;15(3):1067–77. https://doi.org/10.1109/TAFFC.2023.3328408. Peng G, Zhao K, Zhang H, et al. Temporal relative transformer encoding cooperating with channel attention for EEG emotion analysis. Comput Biol Med. 2023;154:106537. https://doi.org/10.1016/j.compbiomed.2023.106537. Jadhav A. Advancing EEG based stress detection using spiking neural networks and convolutional spiking neural networks. Sci Rep. 2025;15(1):26267. https://doi.org/10.1038/s41598-025-10270-0. Malviya L, Khandelwal S, Mal S. Mental stress detection using EEG extracted frequency bands. In: Kumar R, Ahn CW, Sharma TK, editors. Soft computing: theories and applications, vol. Vol. 425. Singapore: Springer; 2022. p. 311–21. https://doi.org/10.1007/978-981-19-0707-4_28. Malviya L, Mal S. A novel technique for stress detection from EEG signal using hybrid deep learning model. Neural Comput Appl. 2022;34(22):19819–30. https://doi.org/10.1007/s00521-022-07540-7. Jain V, Malviya L, Anjana S. Optimized hybrid deep learning for cross-linguistic sentiment analysis: a novel approach. J Cloud Comput. 2025;14(1):30. https://doi.org/10.1186/s13677-025-00753-w. Dash S, et al. Privacy-preserving diabetes and heart disease prediction via federated learning and WCO. Int J Comput Intell Syst. 2025;18(1):217. https://doi.org/10.1007/s44196-025-00956-8. Qin Y, Li B, Wang W, et al. ETCNet: an EEG-based motor imagery classification model combining efficient channel attention and temporal convolutional network. Brain Res. 2024;1823:148673. https://doi.org/10.1016/j.brainres.2023.148673. Xie X, Chen L, Qin S, et al. Bidirectional feature pyramid attention-based temporal convolutional network model for motor imagery electroencephalogram classification. Front Neurorobot. 2024;18:1343249. https://doi.org/10.3389/fnbot.2024.1343249. Altaheri H, Muhammad G, Alsulaiman M, et al. Deep learning techniques for classification of electroencephalogram (EEG) motor imagery (MI) signals: a review. Neural Comput Appl. 2023;35:14681–722. https://doi.org/10.1007/s00521-021-06352-5. Li Y, Chen J, Li F, et al. GMSS: graph-based multi-task self-supervised learning for EEG emotion recognition. IEEE Trans Affect Comput. 2023;14(3):2512–25. https://doi.org/10.1109/TAFFC.2022.3170428. Chen W, Liao Y, Dai R, et al. EEG-based emotion recognition using graph convolutional neural network with dual attention mechanism. Front Comput Neurosci. 2024;18:1416494. https://doi.org/10.3389/fncom.2024.1416494. Qiu L, Zhong L, Li J, et al. SFT-SGAT: a semi-supervised fine-tuning self-supervised graph attention network for emotion recognition and consciousness detection. Neural Netw. 2024;180:106643. https://doi.org/10.1016/j.neunet.2024.106643. Dar MN, Akram MU, Subhani AR, Khawaja SG, Reyes-Aldasoro CC, Gul S. Insights from EEG analysis of evoked memory recalls using deep learning for emotion charting. Sci Rep. 2024;14(1):17080. https://doi.org/10.1038/s41598-024-61832-7. Branco D, Gonçalves ÓF, Badia SBI. A systemas(IAPS) around the world. Sensors. 2023;23(8):3866. https://doi.org/10.3390/s23083866. Scharbert J, Utesch K, Reiter T, et al. If you were happy and you know it, clap your hands! Testing the peak-end rule for retrospective judgments of well-being in everyday life. Eur J Pers. 2025;39(1):1085–104. https://doi.org/10.1177/08902070241235969. Klug M, Berg T, Gramann K. Optimizing EEG ICA decomposition with data cleaning in stationary and mobile experiments. Sci Rep. 2024;14:14119. https://doi.org/10.1038/s41598-024-64919-3. Delorme A. EEG is better left alone. Sci Rep. 2023;13:2372. https://doi.org/10.1038/s41598-023-27528-0. Atti I, Belardinelli P, Ilmoniemi RJ, et al. Measuring the accuracy of ICA-based artifact removal from TMS-evoked potentials. Brain Stimul. 2024;17(1):10–8. https://doi.org/10.1016/j.brs.2023.11.005. Pion-Tonachini L, Kreutz-Delgado K, Makeig S. ICLabel: an automated electroencephalographic independent component classifier, dataset, and website. Neuroimage. 2019;198:181–97. https://doi.org/10.1016/j.neuroimage.2019.05.026. Schindler S, Bublatzky F. Attention and emotion: an integrative review of emotional face processing as a function of attention. Cortex. 2020;130:362–86. https://doi.org/10.1016/j.cortex.2020.06.010. Polich J. Updating P300: an integrative theory of P3a and P3b. Clin Neurophysiol. 2007;118(10):2128–48. https://doi.org/10.1016/j.clinph.2007.04.019. Hajcak G, Foti D. Significance?... Significance! Empirical, methodological, and theoretical connections between the late positive potential and P300 as neural responses to stimulus significance: an integrative review. Psychophysiology. 2020;57(6):e13570. https://doi.org/10.1111/psyp.13570. Bai S, Kolter JZ, Koltun V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271. 2018 arXiv:1803.01271. He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. InProceedings of the IEEE conference on computer vision and pattern recognition 2016 (pp. 770-778). https://doi.org/10.1109/CVPR.2016.90 Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ɓ, Polosukhin I. Attention is all you need. Adv Neural Informat Process Syst 2017;30. https://doi.org/10.48550/arXiv.1706.03762 Song Y, Zheng Q, Liu B, et al. EEG Conformer: Convolutional transformer for EEG decoding and visualization. IEEE Trans Neural Syst Rehabil Eng. 2023;31:710–9. https://doi.org/10.1109/TNSRE.2022.3231133. Wang X, Ren Y, Luo Z, et al. Deep learning-based EEG emotion recognition: current trends and future perspectives. Front Psychol. 2023;14:1126994. https://doi.org/10.3389/fpsyg.2023.1126994. Rakhmatulin I, Dao MS, Nassibi A, et al. Exploring convolutional neural network architectures for EEG feature extraction. Sensors (Basel). 2024;24(3):877. https://doi.org/10.3390/s24030877. Loshchilov I, & Hutter F (2019). Decoupled weight decay regularization. In International Conference on Learning Representations. https://doi.org/10.48550/arXiv.1711.05101 Prechelt L. Early stopping — but when? In: Orr GB, MĂŒller KR, editors. Neural networks: tricks of the trade, vol. Vol. 7700. 2nd ed. Berlin, Heidelberg: Springer; 2012. p. 53–67. https://doi.org/10.1007/978-3-642-35289-8_5. Saha S, Baumert M. Intra- and inter-subject variability in EEG-based sensorimotor brain computer interface: a review. Front Comput Neurosci. 2020;13:87. https://doi.org/10.3389/fncom.2019.00087. Roy Y, Banville H, Albuquerque I, et al. Deep learning-based electroencephalography analysis: a systematic review. J Neural Eng. 2019;16(5):051001. https://doi.org/10.1088/1741-2552/ab260c. Schober P, Boer C, Schwarte LA. Correlation coefficients: appropriate use and interpretation. Anesth Analg. 2018;126(5):1763–8. https://doi.org/10.1213/ANE.0000000000002864. Li X, Song D, Zhang P, et al. Exploring EEG features in cross-subject emotion recognition. Front Neurosci. 2018;12:162. https://doi.org/10.3389/fnins.2018.00162. Zheng WL, Lu BL. Investigating critical frequency bands and channels for EEG-based emotion recognition with deep neural networks. IEEE Trans Auton Ment Dev. 2015;7(3):162–75. https://doi.org/10.1109/TAMD.2015.2431497. Funding This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors. Author information Authors and Affiliations Contributions Zhongtang Guo: Conceptualization, Methodology, Formal analysis, Data curation, Writing—original draft. Corresponding author Ethics declarations Ethical Approval Not applicable. Conflict of interest The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this research article. Competing interests The authors declare no competing interests. Additional information Publisher's Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Rights and permissions Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/. About this article Cite this article Guo, Z. A Deep Learning-Based Retrospective Evaluation Prediction System for Emotional Experiences: Temporal Dynamic Feature Extraction and ERP Neural Mechanisms of the Peak-End Effect. Cogn Comput 18, 111 (2026). https://doi.org/10.1007/s12559-026-10657-9 Received: Accepted: Published: Version of record: DOI: https://doi.org/10.1007/s12559-026-10657-9

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached — you'll always get the same 5 for this article.