Top-down feedback in deep neural networks leads to functional differences during audiovisual integration

  1. Mashbayar Tugsbayar  Is a corresponding author
  2. Mingze Li
  3. Eilif B Muller
  4. Blake Richards
  1. McGill University, Canada
  2. Mila Quebec AI Institute, Canada
  3. Department of Neurosciences, Faculty of Medicine, Université de Montréal, Canada
  4. Centre de Recherche Azrieli du CHU Sainte-Justine, Canada
6 figures, 6 tables and 1 additional file

Figures

Figure 1 with 1 supplement
Description of models.

(a) Each area of the model receives driving feedforward input and modulatory feedback input. Feedback input alters the gain of a neuron, but it does not affect its threshold of activation (multiplicative feedback). In later experiments, we explore an alternative mechanism, where it weakly affects the threshold of activation (composite feedback). (b) Modeled regions and their externopyramidization values (i.e. thickness and relative differentiation of supragranular layers, used as a proxy measure for sensory-associational hierarchical position). Note higher overall externopyramidization values in the occipital lobe compared to the temporal lobe. (c) Using the above hierarchical measures, we constructed models where each connection has a direction (i.e. regions send either feedforward or feedback connections to other regions). In the brainlike model based on human cytoarchitectural data, where all visual regions send feedforward to and receive feedback from the auditory regions. In the reverse model, all auditory regions send feedforward to and receive feedback from visual regions, while connections within a modality remain the same. (d) The resulting artificial neural network (ANN). Outputs of image identification tasks are read out taken from IT, outputs of audio identification tasks are read out from A4, an auditory associational area. Connections between modules are simplified for illustration.

Figure 1—figure supplement 1
Connectivity of the three random models.

The connectivity of the random models was consistent throughout all experiments.

Multimodal visual tasks.

(a) Training conditions. Models must identify the visual stimulus given an ambiguous image and a matching audio clue (VS1) or unambiguous image and distracting audio (VS2). (b, c) Accuracy across epochs for tasks VS1 and VS2 on holdout datasets, (d) Trained models were given an ambiguous visual stimulus and a nonmatching audio stimulus (VS3) to assess which modality they align most closely with. (e) Alignment of trained models across epochs based on task VS3. Scenario was shown to model at the end of each training epoch, but never during training. (f) Models were additionally trained and tested on image stimuli only to assess their baseline performance (VS4). (g) Accuracy of models across epochs on task VS4. N=10 for all experiments.

Figure 3 with 2 supplements
Multimodal auditory tasks.

(a) Training conditions. Models must identify the auditory stimulus given an ambiguous audio and a matching visual clue (AS1) or unambiguous audio and distracting image (AS2). (b, c) Accuracy across epochs for tasks AS1 and AS2 on holdout datasets, (d) Trained models were given an ambiguous audio stimulus and a nonmatching visual stimulus (AS3) to assess which modality they align most heavily with. (e) Alignment of trained models across epochs based on task AS3. Scenario was shown to model at the end of each training epoch, but never during training. (f) Models were additionally trained and tested on audio stimuli only to assess their baseline performance (AS4), (g) Accuracy of models across epochs on task AS4. N=10 for all experiments.

Figure 3—figure supplement 1
Performance of larger model (hidden state size 32) on AS tasks.

Performance is increased across the board, but difference in stimuli preference is still visible.

Figure 3—figure supplement 2
Effect of process time on AS task performance.

Final epoch accuracy in model given short (3P), standard (7P, equal to number of model layers), and long (10P) processing times when given (a) ambiguous audio, helpful visual clue, (b) clean audio and distracting image, and (c, d) ambiguous audio and nonmatching visual input.

Composite versus multiplicative feedback in multimodal tasks.

(a, c, e) Test performances of models with composite feedback and drive-only models trained on visual tasks (VS1 and VS2). The final epoch accuracy of models with composite feedback (C) is compared to that of models with multiplicative feedback (M) shown in previous figures (b, d, f) Test performances of models trained on auditory tasks (AS1 and AS2).

Figure 5 with 1 supplement
Audiovisual switching task.

(a) All models were given a new audiovisual output area (AV) connecting to IT and A4. (b) The output area receives an attention flag telling which stream of information to attend to (visual or auditory). The models with feedback use composite feedback. (c, d) Test performance of models with composite feedback and drive-only models on all tasks. The models were trained simultaneously on all tasks. (e) Alignment of models given ambiguous visual and ambiguous audio input with differing labels. N=10 for all experiments.

Figure 5—figure supplement 1
Models with multiplicative feedback on the audiovisual switching task.

The task and the drive-only (blue) models are the same between this figure and Figure 5.

Model activity during multimodal tasks.

(a) Information flows through the model from area to area across time. At the first time step, only the primary visual and auditory areas will process information. The areas they feed forward to are activated at the next time step, incorporating top-down information if there is any. (b) Comparison of t-SNE reduced latent space and clustering metric in three areas of the brainlike model at different time stages on task VS2 (ignore audio stimulus). (c) Neighborhood Hit scores in all areas of the trained models across time. Trained models were taken from experiments in Figure 4.

Tables

Table 1
Parameters for the modified ConvGRU equations.
ParameterDefinition
rlReset gate for layerl
ulUpdate gate for layerl
hlHidden state of layerl
hl1Feedforward input from the previous layer
hl+1Feedback input from the next layer
mModulatory top-down signal
mmodModulatory component of top-down signal (composite feedback)
mdDriving component of top-down signal (composite feedback)
cCandidate activation
WrLearnable weight matrix for reset gate
WuLearnable weight matrix for update gate
WcLearnable weight matrix for candidate activation
WmLearnable weight matrix for modulatory signal
brBias vector for reset gate
buBias vector for update gate
bcBias vector for candidate activation
bmBias vector for modulatory signal
σSigmoid activation function
tanhHyperbolic tangent activation function
ReLURectified Linear Unit activation function
[x;y]Concatenation along the channel dimension
Convolution operation
Element-wise multiplication
Appendix 1—table 1
Mean and standard deviation of final epoch accuracy across VS tasks for small models with multiplicative feedback (Figure 2).
BrainlikeReverseRand1Rand2Rand3
VS10.868±0.0090.914±0.0110.885±0.0130.880±0.0300.836±0.011
VS20.933±0.0060.945±0.0130.909±0.0210.884±0.0580.800±0.052
VS3 (image align)0.446±0.0070.440±0.0060.436±0.0100.424±0.0160.406±0.027
VS3 (audio align)0.216±0.0080.242±0.0110.237±0.0070.247±0.0080.255±0.029
VS4 (image recognition)0.964±0.0030.939±0.0070.967±0.0020.907±0.0250.959±0.009
Appendix 1—table 2
Mean and standard deviation of final epoch accuracy across AS tasks for small models with multiplicative feedback (Figure 3).
BrainlikeReverseRand1Rand2Rand3
AS10.972±0.0060.880±0.0460.968±0.0390.879±0.0470.887±0.054
AS20.944±0.0160.937±0.0090.945±0.0050.945±0.0150.946±0.007
AS3 (image alignment)0.509±0.0660.164±0.1020.452±0.1810.165±0.0950.204±0.137
AS3 (audio alignment)0.326±0.0530.775±0.1550.428±0.2070.764±0.1570.729±0.187
AS4 (audio recognition)0.928±0.0120.938±0.0100.936±0.0160.919±0.0220.922±0.015
Appendix 1—table 3
Mean and standard deviation across VS tasks for small models with composite feedback (Figure 4).
BrainlikeReverseRand1Rand2Rand3
VS10.928±0.0140.907±0.0070.946±0.0070.945±0.0080.941±0.012
VS20.949±0.0100.941±0.0120.962±0.0070.671±0.4020.811±0.317
VS3 (image align)0.442±0.0070.445±0.0120.450±0.0060.329±0.1630.384±0.128
VS3 (audio align)0.247±0.0120.235±0.0110.247±0.0080.483±0.3260.368±0.260
AS10.967±0.0070.897±0.0540.964±0.0380.892±0.0530.877±0.059
AS20.928±0.0130.928±0.0080.940±0.0110.935±0.0100.933±0.014
AS3 (image align)0.453±0.0730.200±0.1120.444±0.1370.193±0.1280.192±0.130
AS3 (audio align)0.368±0.0650.722±0.1670.423±0.1640.741±0.1730.741±0.189
Appendix 1—table 4
Mean and standard deviation of final epoch accuracy across audiovisual switching tasks (Figure 5).
BrainlikeReverseRand1Rand2Rand3
VS10.909±0.0200.869±0.0230.912±0.0220.883±0.0270.875±0.028
VS20.864±0.0420.285±0.2540.874±0.0550.399±0.3060.430±0.319
AS10.982±0.0080.900±0.0580.983±0.0140.917±0.0620.913±0.065
AS20.897±0.0290.925±0.0230.922±0.0180.937±0.0170.937±0.019
Image align0.619±0.0440.228±0.1380.570±0.0730.273±0.1550.286±0.163
Audio align0.228±0.0500.665±0.1930.273±0.0830.611±0.2060.600±0.242
Appendix 1—table 5
Mean and standard deviation of final epoch accuracy for big models on AS tasks (Figure 3—figure supplement 1).
BrainlikeReverseRand1Rand2Rand3
AS10.987±0.0040.949±0.0360.984±0.0080.979±0.0080.977±0.010
AS20.952±0.0110.949±0.0090.953±0.0110.962±0.0120.971±0.002
AS3 (image align)0.556±0.0360.201±0.1280.358±0.1020.366±0.0450.353±0.105
AS3 (audio align)0.280±0.0150.689±0.2030.477±0.1240.431±0.0450.484±0.135

Additional files

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Mashbayar Tugsbayar
  2. Mingze Li
  3. Eilif B Muller
  4. Blake Richards
(2026)
Top-down feedback in deep neural networks leads to functional differences during audiovisual integration
eLife 14:RP105953.
https://doi.org/10.7554/eLife.105953.3