Self-Organized Context Dependent Processing in Neural Networks
Abstract
Context dependent processing (CDP) enables flexible responses to stimuli based on varying circumstances. It is a crucial aspect of higher cognition that supports primates’ adaptive behavior in dynamic environments. However, the structural requirements for its emergence remain poorly understood. The current study bridges this gap by investigating what network architectures enable the spontaneous development of CDP in artificial neural networks (ANNs). We demonstrate that networks with two parallel pathways that merge via cross-product self-organize towards CDP (we termed SOCDP), with spontaneously emerging functional modules for processing contextual and sensory information, respectively, mirroring the key aspects of CDP in the primate brain. Further analysis revealed that the pressure to efficiently navigate high-dimensional information space—a challenge akin to overcoming the curse of dimensionality—is the driver behind the SOCDP. Furthermore, we identified crucial factors such as network size, structural asymmetry, and task complexity in influencing the emergence of functional specialization in the system. These results suggest that CDP can arise spontaneously given proper network structural predisposition, providing insights into the computational principles underlying flexible cognition in biological systems and how structural constraints shape functional organization in neural architectures.
Similar content being viewed by others
Introduction
Context dependent processing (CDP) is a fundamental aspect of primate cognition that enables flexible responses to similar or even the same stimuli based on varying circumstances. For example, reaction to seeing a snake could vary dramatically based on whether you are at a zoo, hiking in the wilderness, or watching a nature documentary. At its core, CDP involves effectively decomposing the sensory information into various contextual dimensions and allowing them to modulate the behavioral response in an adaptive way (Fig. 1a) [1, 2]. In the primate brain, this capability stems from the interplay between the sensory-motor pathways and the prefrontal cortex (PFC). While sensory-motor pathways process the stimulus-response mapping, the PFC can enhance or suppress certain mappings based on contextual information (Fig. 1b) [3,4,5,6], such as the environment, goals, expectations, and previous experiences. This framework supports flexible, context-appropriate response without rote memorization of myriads of possible scenarios [7].
Recurrent neural networks were used to replicate the complex dynamics of the PFC during context-dependent tasks, showing that specific population-level dynamics can capture the selection and integration of sensory inputs based on context [8]. Additionally, some studies have explored the dimensionality and geometry of neural representations for context-dependent decision-making tasks, highlighting how neural circuits dynamically reconfigure their connectivity and activation patterns based on the context [9]. Moreover, artificially designed modules for modulating sensory pathways based on externally provided contextual signals can support the development of distinct context-specific mappings within the same network architecture [10]. However, previous research has demonstrated that neural networks can implement CDP through explicitly designed modules for sensory and contextual processing [11, 12]. The origin of CDP—whether it emerges spontaneously or requires explicit, supervised learning—remains poorly understood [13, 14]. We still lack an understanding of what network architectures enable the spontaneous emergence of CDP.
To investigate the structural requirements underlying the emergence of CDP, here we employed artificial neural networks (ANNs) as a computational modeling framework to understand how structural organization gives rise to functional specialization in cognitive systems. ANNs have proven valuable for modeling complex neural representation and processing in the brain [15, 16]. Compared to traditional computational neuroscience models that focus primarily on neural dynamics, ANNs offer a unique advantage: they perform actual classification and recognition tasks, allowing us to examine how task performance pressures drive functional organization. Specifically, through control over network architecture and task design [17, 18], we examined how input information can be decomposed into different contextual domains and then used them to modulate the internal processing in ANNs to achieve contextually appropriate behavior [19]. Our findings provide insights into how the interplay between structural architecture and task demands can give rise to context-dependent processing capabilities, shedding light on the organizational principles that may underlie flexible cognition in biological neural systems.
Method
Contextual MNIST Dataset
We used the MNIST dataset as the source dataset. The original MNIST dataset contains 60,000 training images across 10 classes and 10,000 images for testing, each depicting a single handwritten digit from 0 to 9. All images are grayscale images with a size of 28\(\times \)28\(\times \)1. To introduce CDP, we modified the MNIST dataset by adding contextual information to the image backgrounds. Each image’s border was divided into 32 blocks (3\(\times \)3 each). From these blocks, 10 non-overlapping blocks were randomly selected and illuminated to indicate a specific context, resulting in \(\textrm{C}_{10}^{32}\) possible contexts. Prior to any training run, distinct block combinations were drawn without replacement from this combinatorial pool. Each sampled combination was assigned a unique context index and served as the fixed visual pattern for all images belonging to that context. This set of combinations was fixed once at the start of the experiment and held constant across all training runs under that experimental condition, ensuring that comparisons across runs reflect differences in training dynamics rather than differences in the context set. Different contexts may share individual illuminated blocks, but their overall 10-block combinations are guaranteed to be distinct. Each selected block lights up to indicate its context and the labels were modified to reflect the original digit plus the context index times 10. This modification resulted in multiple distinct contexts being embedded within the image backgrounds.
To construct the dataset, the indices of the full training and test sets were independently shuffled and divided into context_num equally-sized splits using numpy.array_split. Each split was assigned one of the context_num pre-defined context patterns, ensuring that each context is associated with an approximately equal number of samples and that the digit class distribution within each context mirrors that of the original MNIST dataset. The same assignment procedure was applied independently to both the training and test sets using the same set of context patterns, ensuring that the context distribution is matched across train and test splits. The context visual patterns were generated once prior to all experiments using a fixed seed (random.seed(12)) in generate_location_list, guaranteeing that the set of context patterns is identical across all experimental conditions and runs. We defined the number of unique contexts as “context_num”, varying this parameter to create a range of experimental conditions. Initially, contexts were increased from 2 to 20 in single increments. Beyond 20 contexts, we increased the number by increments of 5, up to a total of 50 contexts. Consequently, this approach resulted in 25 different variations of context, ensuring a comprehensive examination of how contextual complexity affects network performance.
Network Structures
We constructed five distinct neural network structures: Simple, DotProduct, CrossProduct, Sparse DotProduct and Sparse CrossProduct networks. All networks originated from a common foundation of two convolutional layers dedicated to feature extraction, followed by layers with varying architectures. The initial convolutional layer of this network is equipped with eight 3x3 convolutional kernels, while the subsequent layer contains sixteen 3Ă—3 kernels. Following each convolutional operation, we applied a Rectified Linear Unit (ReLU) activation function and a Max-pooling layer.
Simple network served as the baseline and did not include bifurcated pathways. It contained a fully connected layer following the convolutional layers, with the number of neurons specified by the “middle_size” parameter. DotProduct and CrossProduct networks featured bifurcated pathways, with information processed in parallel streams. The DotProduct network recombined information using a dot product operation, while the CrossProduct network used a cross product operation for combining the outputs of the two pathways. Sparse DotProduct and Sparse CrossProduct networks were variants of the DotProduct and CrossProduct networks, with sparse connectivity enforced to examine the impact of reduced inter-neuronal connections on learning and context-dependent processing. Specifically, given pathway outputs \(f_1 \in \mathbb {R}^m\) and \(f_2 \in \mathbb {R}^n\), the cross product fusion operation computes the outer product:
where \(\otimes \) denotes the outer product. The resulting matrix F is subsequently flattened into a vector of dimension \(m \times n\) and passed to the subsequent fully connected layers. This operation is fixed and non-learnable, introducing no additional parameters beyond those in the two pathways themselves. In this sense, the CrossProduct fusion is a special case of bilinear interaction, in which the joint feature distribution of the two pathways is explicitly represented through multiplicative coupling rather than additive superposition.
Sparse coding is a prominent organizational principle in sensory cortices, where neural populations represent stimuli using relatively few active units at any given time [20]. To investigate whether such within-pathway sparsity facilitates or otherwise interacts with the emergence of functional specialization, we introduced sparse variants of the DotProduct and CrossProduct networks by incorporating Iterative Shrinkage-Thresholding Algorithm (ISTA) layers [21], which are designed to enforce sparsity on the activations. ISTA layers are implemented to enforce sparsity within each pathway of the Sparse DotProduct and Sparse CrossProduct Networks after bifurcation, which is achieved by transforming the intermediate activations through a learned dictionary, promoting sparse coding of these activations. The sparse representation \(Z^{Sparse}\) is computed as:
where: Z represents the activations before applying the ISTA step. D is the dictionary matrix used for sparse coding, which is learned during training. \(\eta \) is the step size parameter with a default value of 0.1. \(\lambda 1\) is the sparsity regularization parameter with a default value of 0.01. ReLU is the Rectified Linear Unit, ensuring non-negativity of the activations.
For clarity, we denote the neuron count in the first fully connected layer of the Simple network as “middle_size”. To maintain computational and parameter comparability, adjustments were made to the neuron counts in the fully connected layers. Specifically, the neuron count in the second fully connected layer of the Simple network was calibrated to equal the square of the neuron count in the feature layer of the CrossProduct network. For the bifurcated networks (DotProduct, CrossProduct, and their sparse variants), neuron counts in each pathway were set to 50% of those in the first fully connected layer in the Simple network.
The network, trained on various randomly initialized configurations, underwent 100 iterations per configuration using torch.optim.SGD (LR: 0.001, Momentum: 0.9) over 10 epochs without a fixed initial seed, with a batch size of 64. During the training process, we saved all the models’ weights at each epoch, along with training loss, and both training and testing accuracy rates, for further analysis. We used Cross-Entropy Loss as our primary loss function, which is suitable for our multi-class classification task.
Quantifying Functional Specialization
In the field of generative modeling, particularly in the generation of images, evaluating the quality and diversity of generated samples is crucial. The Fréchet Inception Distance (FID) has emerged as a prominent metric for assessing these aspects. It has been widely adopted due to its effectiveness in capturing the visual quality of generated images. While FID traditionally measures the similarity between distributions of generated and real images, our study adapts this metric to evaluate the functional specialization within the bifurcated neural network structures. We leverage FID to quantify the distinctiveness in feature representation between the specialized pathways of the network. Assuming there are M types of contexts and N types of digits, the calculation involves comparing the feature distributions for each context/digit category against others. The FID value between the \(i^{th}\) and \(j^{th}\) context/digit types is calculated using the following formula:
where \(\mu _{i}\) and \(\mu _{j}\) are the mean feature vectors of responses to the \(i^{th}\) and \(j^{th}\) context/digit types of the specialized pathways, \(\sum _{i}\) and \(\sum _{j}\) are the covariance matrices for responses to the \(i^{th}\) and \(j^{th}\) context/digit types of the specialized pathways, and Tr denotes the trace of a matrix, capturing the sum of the diagonal elements. Unlike traditional usage, where FID compares generated images to a real dataset, our study employs FID to contrast the feature distributions between the two specialized pathways within the network. This novel application serves to quantify the degree of functional differentiation that the network achieves post bifurcation. We further compute the FID Score:
where \(FID_{\text {context}}^{(1)}(k)\) and \(FID_{\text {digit}}^{(1)}(k)\) are \(k^{th}\) Retraining experiment FID scores for pathway 1, \(FID_{\text {context}}^{(2)}(k)\) and \(FID_{\text {digit}}^{(2)}(k)\) are for pathway 2. This customized FID scores calculation provides a quantifiable measure of context and digit functional specialization within the network, reflecting the network’s ability to self-organize and differentiate tasks between its bifurcated pathways.
Neuronal Response Visualization
In the realm of machine learning, interpretability has become a crucial aspect, especially for complex models like deep neural networks and ensemble methods. SHapley Additive exPlanations (SHAP) has gained prominence as a powerful tool for explaining the output of these models. SHAP is a method that explains the prediction of any machine learning model by computing the contribution of each aspect or component of the input data to the prediction. It is grounded in the principles of fairness and consistency, aiming to provide transparent and understandable explanations. SHAP is based on the Shapley value, a concept from cooperative game theory. The Shapley value assigns a value to each player (or feature in the context of SHAP) that represents their contribution to the total payout (or prediction). The formula for the Shapley value for a feature i is:
where N is the set of all features. S is a subset of features excluding i, v(S) is the prediction model’s output when only the features in S are used, and \(\phi _{i}\) is the Shapley value for feature i, representing its contribution. In our study, we employed SHAP to unravel the distinct roles of the bifurcated pathways in the Cross Product network. The SHAP values facilitated an in-depth analysis of how each pathway contributes to the task of context and number recognition, thereby affirming their functional specialization. Specifically, through the SHAP method, we established a mapping relationship between each pixel in the input image and the neurons in each pathway of the bifurcated pathways. This approach allowed us to clearly discern which parts of the input image each pathway’s neurons are sensitive to, whether it be the context or the number. For the analyses underlying Figs. 2c, S2, and S4, we used DeepSHAP (via shap.DeepExplainer), with a background set consisting of the first batch of 64 test images; SHAP values were computed for 4 representative input images per pathway configuration. For the supplementary asymmetry analyses (Fig. S9), GradientSHAP (via shap.GradientExplainer) was used with a background set of 64 test images and 8 explained images per configuration. To provide a pathway-level quantitative summary of functional specialization, we additionally computed the mean signed SHAP value for each pathway separately over the digit region and the background region (Fig. S3), allowing suppression of non-preferred regions to be distinguished from mere indifference.
Statistical Analyses
All statistical tests were performed across 100 independent training runs per experimental condition. For performance comparisons (Tables S1 and S2), Mann-Whitney U tests were used to compare final-epoch metrics between CrossProduct and Simple/DotProduct networks; this non-parametric test makes no normality assumption and is appropriate for comparing distributions across repeated independent runs. Effect size was quantified as rank-biserial correlation r, where \(r < 0\) indicates CrossProduct outperforms the comparison network and \(r > 0\) indicates the reverse. All p-values were corrected for multiple comparisons using the Benjamini-Hochberg false discovery rate procedure. For functional specialization (Table S3), a one-sided binomial test (\(H_0:p = 0.5\), \(H_1:p > 0.5\)) was applied to the FID Score of the CrossProduct network to assess whether the rate of spontaneous functional specialization significantly exceeds chance. A chi-square test of independence was additionally used to compare specialization rates between CrossProduct and DotProduct networks.
Results
Emergence of Functional Specialization and CDP in Networks
In the primate brain, the sensory-motor pathway are modulated by another context-sensitive pathway involving the PFC. Inspired by such an architecture, here we examine networks with two parallel pathways (Fig. 1c), named bifurcated network, to see if they can develop the ability to extract contextual information and carry out CDP through end-to-end training without predefined functional specialization. Specifically, we tested five distinct network structures (Fig. 1d). Among them, Simple is the baseline with non-bifurcated structure, DotProduct/CrossProduct indicates the information processed by the two pathways are recombined by the process of dot product/cross product (the latter implemented as the outer product of the two pathway feature vectors; see Methods for details), and Sparse indicate the corresponding sparse variants (see methods for details). All networks were trained to perform a modified version of the MNIST task [22], which we named the “Contextual MNIST”. Besides the standard MNIST digits, we added a series of contextual patterns to the border of the images (Fig. 1e), and the same digits with different contexts should be recognized as different classes. We found all five networks could eventually be trained to perform the task well, but with various learning speeds (Fig. 1f). CrossProduct and its sparse variant exhibited the fastest learning speed, followed by the Simple network, which in turn outperformed DotProduct and its sparse variant. As CrossProduct turned out to be the more promising bifurcated architecture, and the sparsity played a relatively minor role, next we focused on the analyses of the CrossProduct network. The negligible effect of sparsity regularization constitutes an informative negative result: it suggests that within-pathway coding density is not a determining factor for SOCDP, and that the cross-product fusion operation is the dominant architectural condition for the spontaneous emergence of functional specialization. For completeness, the DotProduct network counterparts of Figs. 2 and S2 are provided in Fig. S4. As shown there, neither pathway in the DotProduct network develops a clear representational asymmetry between digit and background information.
Surprisingly, we found that in the CrossProduct network trained to perform the Contextual MNIST task, the initial functional symmetry between the two pathways was spontaneously broken in the majority of cases. Figure 2a shows the results of one example network [23]. The representation in one pathway (F1) is better clustered according to the background pattern compared to the digit, while that in the other pathway (F2) is the opposite, indicating that F1 is more dedicated for processing information about background patterns and F2 is more dedicated for processing digits. After the fusion of these two pathway, the activities can represent both digits and background patterns, similar to that found in the simple network, suggesting that the information processed separately in different pathways could be effectively recombined by the cross product operation. Instances where the two pathways did not exhibit clear specialization are presented in the supplementary materials (Fig. S1).
To better quantify the emergence of functional specialization, we employed the Fréchet Inception Distance (FID) method [24] (see Methods for details), which measures representational distance among different classes. As depicted in Fig. 2b, F1 activity exhibited larger FID among different background patterns, thus better suited for classifying those patterns. Conversely, F2 activity exhibited larger FID among different digits.
To examine the representation of each pathway more closely, we utilized the SHapley Additive exPlanations (SHAP) method [25] (see Methods for details), which probes the relationship between each pixel in the input image and each neuron’s activation in the network. Consistent with the overall functional separation demonstrated above, we found that the neurons in Pathway F1 showed a significantly higher response to the area of background patterns and almost no response to the central part of the image where digits are presented. Conversely, neurons in Pathway F2 responded more vigorously to the digits area and minimally to the background patterns (Fig. 2c). The complete set of SHAP visualizations for all neurons is provided in the supplementary materials (Fig. S2). To quantify this dissociation at the pathway level, we computed the mean signed SHAP value for each pathway over the digit and background regions separately (Fig. S3). F1 showed a strongly positive mean SHAP for the background region and near-zero for the digit region, while F2 showed the opposite pattern. This sign asymmetry provides quantitative confirmation of the functional dissociation shown in Fig. 2c.
Taken together, these results demonstrate that the functional symmetry in a bifurcated network can be spontaneously broken, accompanied by the gain of ability to decompose the inputs into different contextual domains. Furthermore, these separately processed information can be recombined later to support effective CDP. Importantly, these can be achieved spontaneously through end-to-end training, suggesting the emergence of CDP is a consistent tendency observed in networks with proper structure and when they are faced with contextual sensitive tasks. We termed this phenomenon as the self-organized CDP (SOCDP).
The observed functional specialization differs from classical dual-pathway models of visual processing (e.g., dorsal "where" vs. ventral "what" streams) [26]in that our pathways self-organize based on task demands rather than being pre-specified for spatial versus object processing. This emergent specialization aligns more closely with flexible context-dependent representations observed in prefrontal cortex [8], where the same neural population can dynamically reconfigure to process different task-relevant dimensions. Importantly, our cross-product integration mechanism provides a computational account for how separately processed information streams can be flexibly combined—a question not fully addressed by previous models of CDP.
The Influence of Task Design on Functional Specialization
Next we investigate how the task design may contribute to spontaneous functional specialization in the network. Firstly, we examine if the functional specialization can occur in other context sensitive task settings. To this end, we modified the Contextual MNIST task used above. Specifically, instead of using different background patterns as contextual cues, we used different colors in this version of the task. We found that the results, in terms of functional specialization, were largely the same between the two versions of the task (Fig. 3a), indicating that the emergence of functional specialization and CDP is not dependent on specific choice of contextual cues.
Secondly, we examine what aspects of those tasks actually drive the functional specialization. Specifically, we tested two scenarios: 1) the task requires the network to discriminate in total NĂ—M classes depending on the combination of context and sensory input, where N is the number of different contexts and M is the number of different sensory inputs (this is the scenario we tested in the previous section) and 2) although the classification depends on the combination of sensory inputs and contexts, the task requires the network to discriminate only M classes (Fig. 3a). We call the former the non-degenerated task, which mimics the situation where the network needs to deal with the curse of dimensionality due to combinatorial explosion, and the latter the degenerated task, which imposes less combinatorial pressure on the network. Interestingly, we found that, unlike the non-degenerated task (Fig. 2b), the degenerate task did not drive the functional specialization (Fig. 3b), suggesting that the pressure of dealing with the curse of dimensionality is an important factor that leads to functional specialization and CDP. This result parallels the concept of degenerate coding in systems neuroscience [27], where multiple distinct network states producing equivalent outputs constrain the emergence of functionally differentiated representations; our degenerate mapping condition instantiates this principle at the task level.
The above result suggests that in the non-degenerated task, the network overcame the curse of dimensionality [28,29,30] by decoupling the sensory and context cues. In this way, if the specific input-output mapping is changed, the network does not need to relearn the entire task. Instead, it can still rely on the acquired capability of discriminating different sensory and context cues, and only need to relearn their new combination. To verify such functional advantage, we first trained the networks to discriminate N×M classes, and then randomly shuffled the class labels—mimicking a set of new tasks encountered in dynamical environments, and restrained the network. Importantly, only the last fully connected layer was subjected to retraining, while the other parts of the network were fixed. As a control, a non-bifurcated network was tested using the same procedure.
Figure 3c shows that the bifurcated network with functional specialization exhibited significant advantages in both the learning speed and final accuracy when only the final layer was retrained, indicating the strategy of decoupling sensory and context cues indeed makes the network more efficient in adapting to dynamical environments. It is worth noting that when the Simple network is retrained on all layers, it achieves final performance comparable to that of the CrossProduct network retrained on only the last layer. Rather than undermining the advantage of CrossProduct, this observation provides a complementary perspective: to match the performance of a CrossProduct network that updates only its final layer, the Simple network must retrain its entire parameter set. This disparity directly reflects the degree to which the learned representations are reusable. In the CrossProduct network, functional specialization factorizes context and sensory representations into separate pathways; when the task remapping occurs, only the final integration layer requires updating because the upstream representations remain valid. In the Simple network, context and sensory features are entangled throughout, so meaningful adaptation requires adjusting representations at all levels. The similar final performance achieved by the two strategies therefore quantifies, rather than undermines, the efficiency advantage conferred by SOCDP. Full results are presented in the supplementary materials (Fig. S5).
The Influence of Task Difficulty and Network Size on Functional Specialization
In this section, we investigate how the difficulty of the task and the network size can affect the functional specialization in bifurcated networks. Specifically, we examined five network sizes: 16, 32, 64, 128, and 256 (number of neurons in the bifurcated layer). Task difficulty was modulated by varying the number of different contexts, ranging from 2 to 50. We found that the DotProduct network exhibited a consistently low level of functional specialization across various network sizes and task difficulties (Fig. 4a). In contrast, the CrossProduct network demonstrated fluctuating functional specialization depending on these two factors (Fig. 4b). We note that the influence of network size is not monotonic, with middle sized networks with 32 and 64 neurons in the bifurcated layer leading to more pronounced functional specialization. The influence of task difficulty on functional specialization is overall much weaker than the network size. As shown in Fig. 4b, for smaller network sizes, difficult tasks tended to inhibit functional specialization, but for larger network sizes, difficult tasks tended to strengthen functional specialization.
We further examined how the training affect functional specialization. We found that functional specialization exhibited a dynamic pattern of change along the training process (Fig. 4c). Specifically, larger networks reached the peak of functional specialization earlier in the process, then exhibited a decline of it. Conversely, smaller networks exhibited a steady increase in specialization throughout the training process. Notably, examination of the training loss curves revealed that all networks learned the task well enough by the second training epoch. This rapid task acquisition, coupled with the divergent trajectories of functional specialization, suggests that the networks dynamically adjust their internal representations and division of labor even after achieving high performance. The observed dynamics indicate that, given a proper structure architecture, functional specialization is frequently observed across the network configurations tested, though its manifestation is network-size dependent. To provide a performance baseline against which these specialization patterns can be interpreted, Fig. S6 presents the training dynamics of the Simple CNN network across the same range of network sizes and task complexities.
Asymmetry in Network Structures Leading to Stabilized Functional Specialization
Next, we investigated the impact of structural asymmetry of the bifurcated networks on functional specialization. We hypothesized that asymmetry in representational capacity would lead to more stable functional division [31, 32], with the pathway having larger capacity dedicated for more difficult processing, and vice versa. To test this, we compared symmetrical networks, where both pathways were equally sized at half the total layer size, with asymmetrical networks, where one pathway was scaled to three-quarters and the other to one-quarter of the total layer size (Fig. 5a).
The training losses of both symmetrical and asymmetrical networks were nearly identical, as depicted in Fig. 5b(1). However, a notable distinction was observed in the FID Score, which was higher in the asymmetrical network (Fig. 5b(2)). Analysis of 100 training sessions revealed that symmetrical networks displayed dynamic role alternation between pathways for context and digit recognition with nearly equal probability (Fig. 5b(4)). The color encoding in Fig. 5b(4-5) represents four possible outcomes across repeated runs: (i) both F1 and F2 are more sensitive to digits; (ii) F1 specializes for digits and F2 for context; (iii) F2 specializes for digits and F1 for context; and (iv) both F1 and F2 are more sensitive to context. Functional specialization corresponds to outcomes (ii) and (iii), which together account for the rising curve analogous to the FID Score in Fig. 5b(2). The random alternation we describe refers to the fact that outcomes (ii) and (iii) occur with approximately equal probability across runs of the symmetric network, meaning that the assignment of roles to pathways is not determined by the architecture but varies randomly across initializations. In contrast, asymmetrical networks consistently allocated the larger pathway to digit recognition and the smaller to context recognition (Fig. 5b(5)). This specialization pattern correlates with the relative complexity of context and digit recognition, as evidenced by the loss curves (Fig. 5b(3)). The context loss converges more rapidly and to a lower value compared to the digit loss, indicating that context recognition is a simpler task. Consequently, in asymmetrical networks, the smaller pathway consistently handles the less complex context recognition, while the larger pathway tackles the more demanding digit recognition task.
These findings indicate that structural asymmetry in neural networks promotes stable functional specialization. In asymmetrical networks, the pathway with greater representational capacity consistently undertakes more complex tasks, leading to a solidification of roles. This contrast with symmetrical networks, where specialization alternated randomly between pathways. This structural bias towards stable specialization may serve as a fundamental mechanism for robust CDP, potentially explaining the specialized pathways observed in biological neural systems for sensory processing and contextual modulation. Importantly, supplementary analyses demonstrate that this qualitative difference between symmetric and asymmetric networks is preserved across a wide range of task complexities (Fig. S8), suggesting that structural asymmetry operates as a dominant and robust determinant of specialization stability, independent of the combinatorial pressure imposed by the task. Further visual evidence is provided in Fig. S9, which presents SHAP value maps for both symmetric and asymmetric networks at middle_size = 32 and 64.
The contrast between symmetric and asymmetric networks can be interpreted through the lens of integration and segregation principles [33]. The symmetric network, where pathway roles alternate randomly across runs, reflects a high-integration regime in which neither pathway establishes a stable independent representational identity. Structural asymmetry drives a transition toward functional segregation, with each pathway developing a dedicated role commensurate with its capacity: the larger pathway consistently handles the more demanding digit recognition task and the smaller pathway the simpler context recognition task (Fig. 5(3)). This capacity-to-complexity matching is consistent with the network economy principle, whereby resources are allocated in proportion to task demands rather than distributed symmetrically.
Discussion and Conclusion
The ability to process information in a context-dependent manner enables systems to adapt their responses to similar stimuli based on varying contextual cues, allowing for more flexible and appropriate reactions to a wide range of situations [34, 35]. In biological systems, CDP is crucial for adaptive behavior, allowing organisms to respond flexibly to environmental cues and internal states [36,37,38]. Understanding CDP in computational models provides a window into the principles governing adaptive behavior in biological systems, offering mechanistic insights into how neural architectures support context-sensitive cognition. At its core, CDP represents a powerful solution to the curse of dimensionality. By decoupling context from sensory input, CDP effectively reduces the dimensionality of the input-output mapping, allowing for more efficient and generalizable learning [39, 40]. This decoupling is crucial because it enables the system to treat context and sensory information as separate, interacting variables rather than a single, monolithic input. Consequently, the system can learn general principles about how context modulates sensory processing, rather than having to learn every possible combination of context and sensory input independently [41, 42]. This not only reduces the required learning capacity but also enhances the system’s ability to generalize to novel situations.
The CDP involves the ability to distinguish contextual information and to exploit it to achieve flexible, situationally appropriate behavior. It seems to be a sophisticated information processing strategy that may require a specific design. However, we know that genetic predisposition is only capable of shaping the structural feature of neural networks, but cannot dictate subtle functional specializations required for CDP. The main finding of the current study is that, given the appropriate structural foundations [43, 44] and task pressures, CDP can emerge in a self-organized way, at least within the task families and architectural configurations examined here, thus offering a potential explanation for its prevalence in primate cognition and pointing towards more adaptable AI systems. This work illustrates the deep interdependence between structural architecture and functional specialization. This bidirectional influence between structure and function is a fundamental principle in neuroscience, known as structure-function coupling [45,46,47]. Our work provides a computational demonstration of this principle, showing how it can arise in artificial systems and potentially shedding light on similar processes in biological neural networks.
Importantly, our approach differs from previous computational models of CDP in fundamental ways. Unlike models that rely on temporal dynamics for context switching [8] or pre-trained orthogonal representations [9], our work demonstrates that functional specialization can emerge from architectural constraints alone, without explicit supervision for separate pathway functions. Furthermore, while our cross-product operation may seem biologically abstract, it captures the essential computational principle of interactions between neural populations, which have been observed in various brain regions including prefrontal and parietal cortices. This multiplicative mechanism provides a concrete implementation of how structure can give rise to function, offering a mechanistic account that complements the more abstract structure-function coupling principle. More precisely, relative to Mante et al. (2013) [8], we identify a feedforward structural route to CDP that does not require recurrent temporal dynamics. Relative to Flesch et al. (2022) [9], our results concern pathway-level dissociation rather than within-pathway representational geometry. Relative to Zeng et al. (2019) [10], who take modularity as a design input, we treat it as an emergent output. And relative to Dobs et al. (2022) and Bakhtiari et al. (2021) [48, 49], who document emergent specialization for low-level sensory features, our work addresses the higher-level decoupling of contextual and sensory information underlying flexible stimulus-response mappings.
It is also worth contrasting this mechanism with attention-based approaches to context-dependent processing. Attention mechanisms address a distinct, engineering-oriented question—how to improve task performance by dynamically weighting input features—rather than explaining how functional specialization emerges from structural constraints. Critically, the additive nature of attention outputs (weighted sums with softmax normalization) is less suited to expressing the conditional remapping that defines CDP, compared to the multiplicative interactions implemented by cross-product fusion, which more closely reflect the gain modulation observed in cortical circuits [50]. Furthermore, because attention mixes information within a shared representational space, pathway-level functional dissociation analogous to SOCDP is unlikely to arise naturally in attention-based architectures, making direct comparison between the two frameworks less straightforward than it might initially appear.
While it might seem intuitive that separate pathways would naturally lead to functional specialization [49, 51], our results reveal that the development of distinct roles for context and digit processing is not guaranteed by the mere presence of bifurcation. Instead, it arises from a complex interplay of factors including network architecture, task structure, and training dynamics [48, 52]. The superiority of cross-product fusion over dot-product in fostering specialization highlights that not all bifurcated structures are equally conducive to functional division, mirroring the importance of specific connectivity patterns in biological neural networks for enabling complex cognitive functions [33, 53]. Non-degenerate output mappings, which allow for clear differentiation between context and object categories, promote specialization, whereas degenerate mappings do not, highlighting the critical role of task structure in guiding functional differentiation. Our findings reveal a complex, non-linear relationship between network size, task complexity, and functional specialization. Moderate-sized networks exhibited the most pronounced specialization, while task difficulty’s impact varied with network size, which indicates that this process is finely balanced and not a simple consequence of pathway separation. This finding echoes the principle of "neural efficiency" observed in human brain development, where optimal cognitive performance is associated with a balance between neural specialization and integration [54, 55]. This capacity-pressure balance can be understood as follows: when network capacity is insufficient, neither pathway accumulates enough representational resources to develop a dedicated functional role; conversely, when capacity is excessive, redundancy reduces the pressure for functional division of labor, as either pathway alone could in principle accommodate the full task. The optimal size range identified here (32–64 neurons in the bifurcated layer) is specific to the Contextual MNIST setting; more complex tasks would likely shift this range upward, a prediction that warrants systematic empirical investigation in future work.
An important interpretive question concerns whether SOCDP is driven by the specific structure of the cross-product interaction or merely by the increased representational capacity that the outer-product expansion introduces. First, the DotProduct network applies element-wise multiplicative interaction while preserving the original pathway dimensionality—yet consistently fails to produce SOCDP, establishing that multiplicative interaction per se is not sufficient. Second, the Simple network’s second fully connected layer was calibrated to match the \((M/2)^2\)-dimensional fusion output of the CrossProduct, yet shows no spontaneous specialization, indicating that high-dimensional representational capacity without structural bifurcation and factorization is likewise insufficient. A gradient-based analysis offers a relatively intuitive explanation for the effectiveness of the cross-product operation: for the outer product \(F = f_1 \otimes f_2\), where \(F[i,j] = f_1[i] \times f_2[j]\), the gradient with respect to pathway F1’s output is \(\partial L/\partial f_1[i] = \sum _j (\partial L/\partial F[i,j]) \cdot f_2[j]\)—each pathway’s learning signal is continuously mediated by the partner pathway’s current representation. This bidirectional, pathway-mediated gradient coupling creates a persistent inductive bias toward functional complementarity that neither a dimensionality-matched single-pathway MLP nor an element-wise multiplicative fusion can replicate, because neither preserves the factorized, pathway-separated representational structure of the outer product.
While our study draws inspiration from the biological architecture of the PFC and its role in modulating sensory-motor pathways, it is essential to emphasize that the neural network model we employed represents an abstract computational framework [56]. Our bifurcated network architecture is designed to investigate the emergence of CDP at a conceptual level, rather than to directly mimic the specific neural circuitry of the PFC [57, 58]. The PFC and our bifurcated networks are instances of broader principles rather than direct analogs. Throughout this work, references to biological phenomena—including PFC modulation, gain modulation, and structure-function coupling—should be understood as computational analogies and mechanistic parallels, not as claims of direct neural plausibility, and the generalizability of our findings to natural neural systems remains to be established through direct comparison with neurophysiological data. By focusing on abstract computational principles, our study contributes to understanding the fundamental organizational principles underlying CDP in cognitive systems. This computational approach complements neurobiological findings by revealing how architectural constraints can spontaneously give rise to functional specialization observed in cortical processing. This abstraction allows us to explore the fundamental mechanisms underlying CDP without being constrained by the specific anatomical and physiological details of any particular brain region [59, 60]. We also note that the Contextual MNIST task operationalizes CDP as context-sensitive classification with combinatorial structure, rather than as flexible remapping of identical sensory inputs under contextual control, as in Stroop-like paradigms. Our results should therefore be interpreted as evidence for spontaneous functional specialization under combinatorial task pressure, with the generalization to stricter formulations of CDP remaining an important direction for future work.
In conclusion, our study not only demonstrates the spontaneous emergence of CDP in ANNs but also provides a framework for understanding this crucial cognitive function [61, 62]. By revealing the complex interplay between structure, function, and task demands, our work contributes to a deeper understanding of adaptive information processing in both biological and artificial systems. As we continue to unravel the principles underlying SOCDP, we advance our understanding of the fundamental organizing principles of higher cognition and the structural-functional relationships that enable adaptive behavior in complex cognitive systems.
Limitations of the Study
Our study has several limitations that should inform interpretation of the results. First, our experimental paradigm uses artificial contextual manipulations (border patterns) rather than naturalistic contexts, limiting direct comparison with biological CDP studies that often involve temporal or goal-based contexts. Additionally, the task operationalizes CDP as context-sensitive classification with combinatorial structure rather than as flexible remapping under identical sensory inputs; whether SOCDP generalizes to the stricter neuroscientific formulation of CDP remains to be established. Second, while cross-product integration proved superior in our tasks, this may be specific to our experimental design; other integration mechanisms might excel under different conditions. Third, our analysis focused on relatively simple networks and tasks—scaling to more complex scenarios with multiple contexts or content types remains unexplored. Relatedly, the optimal network size range for functional specialization identified in our experiments (32–64 neurons in the bifurcated layer) is specific to the Contextual MNIST setting. This range reflects a capacity-pressure balance: insufficient capacity prevents either pathway from developing a dedicated functional role, while excessive capacity reduces the pressure for functional division of labor. Whether this balance scales predictably with task complexity, and whether the underlying dynamics generalize to larger and deeper architectures, remains an open question for future investigation. Fourth, the training regime (backpropagation) differs fundamentally from biological learning, potentially affecting the generalizability to natural neural systems. Finally, we did not investigate how our model handles scenarios with hierarchical or interdependent contexts, which are common in real-world CDP. The present dual-pathway structure, while biologically motivated by the two-stream organization of sensory-motor and PFC-mediated pathways, imposes a fixed architectural constraint that limits the model’s ability to simultaneously handle tasks involving a large number of independent contextual dimensions. Extending the framework to multi-pathway or hierarchically organized architectures— where deeper levels of functional specialization might emerge analogously to hierarchical cortical representations—represents a natural direction for future work, and may also offer a more principled account of how the primate brain manages the combinatorial complexity of naturalistic contexts. Future work should address these limitations through more diverse task designs, biologically plausible learning rules, and systematic comparison with neurophysiological data.
Data Availability
The analysis codes are available at: https://doi.org/10.6084/m9.figshare.29144234. The MNIST dataset is available at https://opendatalab.com/OpenDataLab/MNIST/tree/main/raw.
References
Okayasu M, Inukai T, Tanaka D, Tsumura K, Shintaki R, Takeda M, et al. The stroop effect involves an excitatory-inhibitory fronto-cerebellar loop. Nat Commun. 2023;14(1):27.
Costa TL, Orsten-Hooge K, Gaudêncio Rêgo G, Wagemans J, Pomerantz JR, Sérgio Boggio P. Neural signatures of the configural superiority effect and fundamental emergent features in human vision. Scientif Rep. 2018;8(1):13954.
Miller EK. The prefontral cortex and cognitive control. Nat Rev Neurosci. 2000;1(1):59–65.
Miller EK, Cohen JD. An integrative theory of prefrontal cortex function. Annual Rev Neurosci. 2001;24(1):167–202.
Koechlin E, Summerfield C. An information theoretical approach to prefrontal executive function. Trends Cogn Sci. 2007;11(6):229–35.
Fuster J. The prefrontal cortex. Academic press; 2015.
Sun W, Advani M, Spruston N, Saxe A, Fitzgerald JE. Organizing memories for generalization in complementary learning systems. Nat Neurosci. 2023;26(8):1438–48.
Mante V, Sussillo D, Shenoy KV, Newsome WT. Context-dependent computation by recurrent dynamics in prefrontal cortex. Nature. 2013;503(7474):78–84.
Flesch T, Juechems K, Dumbalska T, Saxe A, Summerfield C. Orthogonal representations for robust context-dependent task performance in brains and neural networks. Neuron. 2022;110(7):1258–70.
Zeng G, Chen Y, Cui B, Yu S. Continual learning of context-dependent processing in neural networks. Nat Mach Intell. 2019;1(8):364–72.
Meunier D, Lambiotte R, Bullmore ET. Modular and hierarchically modular organization of brain networks. Front Neurosci. 2010;4:200.
Kirsch L, Kunze J, Barber D. Modular networks: Learning to decompose neural computation. Adv Neural Inf Process Syst. 2018;31.
Lee I, Lee C-H. Contextual behavior and neural circuits. Front Neural Circ. 2013;7:84.
Zamboni G et al (2016) Functional specialization and network connectivity in brain function. Oxford textbook of cognitive neurology and dementia. 2016. p. 32–41.
Yamins DL, DiCarlo JJ. Using goal-driven deep learning models to understand sensory cortex. Nat Neurosci. 2016;19(3):356–65.
Giordano BL, Esposito M, Valente G, Formisano E. Intermediate acoustic-to-semantic representations link behavioral and neural responses to natural sounds. Nat Neurosci. 2023;26(4):664–72.
Vogel AC, Power JD, Petersen SE, Schlaggar BL. Development of the brain’s functional network architecture. Neuropsychol Rev. 2010;20:362–75.
Béna G, Goodman DF. Dynamics of specialization in neural modules under resource constraints. 2021. arXiv preprint arXiv:2106.02626
Rueckl JG, Cave KR, Kosslyn SM. Why are and where processed by separate cortical visual systems? a computational investigation. J Cogn Neurosci. 1989;1(2):171–86.
Olshausen BA, Field DJ. Emergence of simple-cell receptive field properties by learning a sparse code for natural images. Nature. 1996;381(6583):607–9.
Yu Y, Buchanan S, Pai D, Chu T, Wu Z, Tong S, et al. White-box transformers via sparse rate reduction. Adv Neural Inf Process Syst. 2023;36:9422–57.
Deng L. The mnist database of handwritten digit images for machine learning research [best of the web]. IEEE Signal Process Magaz. 2012;29(6):141–2.
Van der Maaten L, Hinton G. Visualizing data using t-sne. J Mach Learn Res. 2008;9(11).
Obukhov A, Krasnyanskiy M. Quality assessment method for gan based on modified metrics inception score and fréchet inception distance. In: Software engineering perspectives in intelligent systems: Proceedings of 4th computational methods in systems and software 2020. Springer; 2020. vol. 14. p. 102–114.
Lundberg S. A unified approach to interpreting model predictions. 2017. arXiv preprint arXiv:1705.07874
Thompson JA, Sheahan H, Dumbalska T, Sandbrink JD, Piazza M, Summerfield C. Zero-shot counting with a dual-stream neural network model. Neuron. 2024;112(24):4147–58.
Di Plinio S, Northoff G, Ebisch S. The degenerate coding of psychometric profiles through functional connectivity archetypes. Front Human Neurosci. 2024;18:1455776.
Köppen M. The curse of dimensionality. In: 5th online world conference on soft computing in industrial applications (WSC5). 2000. vol. 1. p. 4–8.
Kuo FY, Sloan IH. Lifting the curse of dimensionality. Notic AMS. 2005;52(11):1320–8.
Altman N, Krzywinski M. The curse (s) of dimensionality. Nat Methods. 2018;15(6):399–400.
Palmer AR. Symmetry breaking and the evolution of development. Science. 2004;306(5697):828–33.
Li R, Bowerman B. Symmetry breaking in biology. Cold Spring Harbor Perspect Biol. 2010;2(3):a003475.
Bullmore E, Sporns O. The economy of brain network organization. Nat Rev Neurosci. 2012;13(5):336–49.
Siqi-Liu A. Context-specific adjustments of cognitive flexibility. Ph.D. dissertation, Duke University; 2023.
Rikhye RV, Gilra A, Halassa MM. Thalamic regulation of switching between cortical representations enables cognitive flexibility. Nat Neurosci. 2018;21(12):1753–63.
Cohen JD, Servan-Schreiber D. Context, cortex, and dopamine: a connectionist approach to behavior and biology in schizophrenia. Psychol Rev. 1992;99(1):45.
Kuchibhotla KV, Gill JV, Lindsay GW, Papadoyannis ES, Field RE, Sten TAH, et al. Parallel processing by cortical inhibition enables context-dependent behavior. Nat Neurosci. 2017;20(1):62–71.
Xu D, Dong M, Chen Y, Delgado AM, Hughes NC, Zhang L, et al. Cortical processing of flexible and context-dependent sensorimotor sequences. Nature. 2022;603(7901):464–9.
Ganguli S, Sompolinsky H. Compressed sensing, sparsity, and dimensionality in neuronal information processing and data analysis. Annual Rev Neurosci. 2012;35(1):485–508.
Bayones L, Zainos A, Alvarez M, Romo R, Franci A, Rossi-Pool R. Orthogonality of sensory and contextual categorical dynamics embedded in a continuum of responses from the second somatosensory cortex. Proceed Nation Acad Sci. 2024;121(29):e2316765121.
Hoke KL, Pitts NL. Modulation of sensory-motor integration as a general mechanism for context dependence of behavior. Gener Comparat Endocrinol. 2012;176(3):465–71.
Taylor JA, Ivry RB. Context-dependent generalization. Front Human Neurosci. 2013;7:171.
Bertolero MA, Yeo BT, D’Esposito M. The modular and integrative functional architecture of the human brain. Proceed Nation Acad Sci. 2015;112(49):E6798–807.
Lambiotte R, Schaub MT. Modularity and dynamics on complex networks. Cambridge University Press; 2021.
Park H-J, Friston K. Structural and functional brain networks: from connections to cognition. Science. 2013;342(6158):1238411.
Preti MG, Van De Ville D. Decoupling of brain function from structure reveals regional behavioral specialization in humans. Nat Commun. 2019;10(1):4747.
Sarwar T, Tian Y, Yeo BT, Ramamohanarao K, Zalesky A. Structure-function coupling in the human connectome: A machine learning approach. NeuroImage. 2021;226:117609.
Dobs K, Martinez J, Kell AJ, Kanwisher N. Brain-like functional specialization emerges spontaneously in deep neural networks’’. Sci Adv. 2022;8(11):eabl8913.
Bakhtiari S, Mineault P, Lillicrap T, Pack C, Richards B. The functional specialization of visual cortex emerges from training parallel pathways with self-supervised predictive learning. Adv Neural Inf Process Syst. 2021;34:25 164-25 178.
Salinas E, Thier P. Gain modulation: a major computational principle of the central nervous system. Neuron. 2000;27(1):15–21.
Rueffler C, Hermisson J, Wagner GP. Evolution of functional specialization and division of labor. Proceed Nation Acad Sci. 2012;109(6):E326–35.
Praczyk T. Emerging modularity during the evolution of neural networks. J Artif Intell Soft Comput Res. 2023;13(2):107–26.
Pulvermüller F, Tomasello R, Henningsen-Schomers MR, Wennekers T. Biological constraints on neural network models of cognitive function. Nat Rev Neurosci. 2021;22(8):488–502.
Neubauer AC, Fink A. Intelligence and neural efficiency. Neurosci Biobehav Rev. 2009;33(7):1004–23.
Dunst B, Benedek M, Jauk E, Bergner S, Koschutnig K, Sommer M, et al. Neural efficiency as a function of task demands. Intelligence. 2014;42:22–30.
Cohen JD, Braver TS, O’Reilly R. A computational approach to prefrontal cortex, cognitive control and schizophrenia: recent developments and current challenges, Philosophical transactions of the royal society of london. Series B: Biol Sci. 1996;351(1346):1515–27.
Soltani A, Koechlin E. Computational models of adaptive behavior and prefrontal cortex. Neuropsychopharmacology. 2022;47(1):58–71.
Heald JB, Wolpert DM, Lengyel M. The computational and neural bases of context-dependent learning. Annual Rev Neurosci. 2023;46(1):233–58.
Haber SN, Liu H, Seidlitz J, Bullmore E. Prefrontal connectomics: from anatomy to human imaging. Neuropsychopharmacology. 2022;47(1):20–40.
Hass J, Hertäg L, Durstewitz D. A detailed data-driven network model of prefrontal cortex reproduces key features of in vivo activity. PLoS Computat Biol. 2016;12(5):e1004930.
Richards BA, Lillicrap TP, Beaudoin P, Bengio Y, Bogacz R, Christensen A, et al. A deep learning framework for neuroscience. Nat Neurosci. 2019;22(11):1761–70.
Hassabis D, Kumaran D, Summerfield C, Botvinick M. Neuroscience-inspired artificial intelligence. Neuron. 2017;95(2):245–58.
Acknowledgements
The authors thank Prof. Frederic Alexandre for valuable discussions and support in methodological design. The authors also thank all members of the lab for their support.
Funding
This work was funded by the Strategic Priority Research Program of the Chinese Academy of Sciences (CAS) via grants XDB1010301 and XDB1010302, CAS Project for Young Scientists in Basic Research via grant YSBR-041, and the International Partnership Program of the Chinese Academy of Sciences (CAS) via grant 173211KYSB20200021.
Author information
Authors and Affiliations
Contributions
G.H., S.Y., and F.A. designed research; G.H. and Y.C. performed research; G.H. and Y.C. analyzed data; G.H. and S.Y. wrote the paper; S.Q. contributed initial observations; and S.Y. and F.A. provided methodological guidance.
Corresponding authors
Ethics declarations
Competing Interests
The authors declare no competing interests.
Additional information
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary Information
Below is the link to the electronic supplementary material.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
About this article
Cite this article
Hao, G., Chen, Y., Qin, S. et al. Self-Organized Context Dependent Processing in Neural Networks. Cogn Comput 18, 112 (2026). https://doi.org/10.1007/s12559-026-10624-4
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s12559-026-10624-4
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.