Peer review process
Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.
Read more about eLife’s peer review process.Editors
- Reviewing EditorMing MengUniversity of Alabama at Birmingham, Birmingham, United States of America
- Senior EditorHuan LuoPeking University, Beijing, China
Reviewer #1 (Public review):
Summary:
This manuscript investigates whether the human brain contains a shared category-general representation of gender across faces, bodies, and gender-associated objects. The authors acquired fMRI data while participants viewed male and female stimuli from three categories in a one-back task. They then used searchlight MVPA, cross-category decoding, regression-based RSA, CNN vs. brain representational comparisons, and PPI analyses. Their main finding is that gender information could be decoded from distributed occipitotemporal regions within each category, whereas a cluster in the rMTG showed convergence across cross-category decoding and RSA. The authors concluded that this rMTG representation resembles intermediate layers of fine-tuned CNNs and that face and body gender processing share similar functional connectivity patterns.
Strengths:
The question is potentially important, particularly for social cognition, object recognition, and the use of neural network models to interpret high-level visual representations. Previous behavioral studies have shown cross-category adaptation between bodies and faces, and even between gender-associated objects and faces, so the attempt to test for a neural counterpart using fMRI is well motivated. The use of multiple complementary analyses including within-category decoding, cross-category decoding, regression RSA, CNN comparisons, and effective connectivity analyses is also a strength. The convergence of cross-category MVPA and RSA in a right MTG cluster is potentially interesting and deserves attention.
Weaknesses:
The largest problem is conceptual. The term gender is used as if it refers to the same construct across faces, bodies, and objects. This is not self-evident. In faces and bodies, the stimuli seem to contain visual cues from which observers infer binary gender categories. In objects, however, the relevant information is almost gender stereotype, cultural association, or learned semantic association. These are not equivalent constructs. The manuscript therefore needs to distinguish much more carefully between perceived gender, biological sex cues, gender-associated visual features, and gender stereotypes. Without this distinction, the title and main conclusion are too broad. The object condition is particularly problematic. Javadi & Wee (2012) showed that gender-associated objects can bias subsequent judgments of ambiguous face gender, and they discussed two possible mechanisms, including shared neural substrates or top-down modulation induced by the gender concept. However, their behavioral adaptation study does not directly demonstrate that objects, faces, and bodies are encoded in the same neural representational format. The present manuscript treats these object stimuli as if they provide evidence about the same kind of gender representation as faces and bodies, but that step requires additional empirical support. Independent ratings of object gender association, cultural familiarity, visual similarity, and semantic category are essential here.
A second major concern is stimulus control. The face images were taken from Chinese male and female actors, the body images were headless bodies in underwear, and the object images were selected because of prior gender associations. This design introduces many possible confounds: hairstyle, makeup, skin texture, body shape, clothing, color, luminance, object category, object function, curvature, spatial frequency, and cultural familiarity. Cross-category decoding can be significant even when a classifier relies on shared visual statistics rather than an abstract gender code. For example, female-associated stimuli may differ from male-associated stimuli in color, shape, brightness, texture, or semantic category in ways that are consistent across faces, bodies, and objects. The present analyses do not adequately rule out these alternatives. Foster et al. (2019) are especially relevant in this respect. They reported that body sex could be decoded from both body- and face-responsive regions. However, the sex of well-controlled faces, for example faces excluding hairstyle cues, could not be decoded from face- or body-responsive regions. This finding should make the authors more cautious. The fact that the present study used more ecological face stimuli may increase sensitivity to gender-related cues, but it also increases the possibilities that decoding is driven by uncontrolled external features rather than by an abstract gender representation. Accordingly, because no additional visual, semantic, or stereotype-based model RDMs were included in the RSA analysis, this result alone cannot establish an abstract, category-independent gender representation. Any systematic difference between male- and female-associated images will load onto the gender RDM. At least, the authors should include additional model RDMs for low-level visual features. In addition, the current RSA analysis has another limitation. The neural RDMs are based on only six condition-level patterns, producing a 6 × 6 matrix. The theoretical model includes only binary gender and category RDMs. This is too coarse to support the claim of category-independent gender representation. Ideally, all the RSA analysis should be performed at the item level rather than at the condition level.
The cross-category decoding result in rMTG is promising but not yet conclusive. The authors identify a right MTG cluster by overlapping thresholded maps from three cross-category decoding analyses. This is useful descriptively, but it does not by itself establish a common representational code. The overlap of thresholded maps depends on the chosen threshold. If the authors want to make a formal conjunction claim, they should use a valid conjunction-null approach such as a minimum-statistic conjunction evaluated under the appropriate conjunction null, rather than simply displaying the intersection of thresholded maps. Even if this approach cannot be adopted in this study, the issue should be included as a limitation.
In the PPI analysis, the reported similarity between face and body connectivity matrices is a little bit small (r = 0.08). The claim of a shared functional network should therefore be softened unless the authors test whether this correlation is significantly larger than the face-object and body-object correlations, correct for multiple comparisons, account for the non-independence of matrix elements, and report participant-level distributions and confidence intervals.
Reviewer #2 (Public review):
Summary:
The study tests whether male/female-related information is represented in a form that generalizes across faces, bodies, and gender-associated objects. Using within- and cross-category MVPA, regression RSA, comparisons with fine-tuned CNNs, and connectivity analyses, the authors identify a right middle temporal gyrus region whose patterns generalize across the three stimulus classes. They conclude that this region provides a category-general, mid-level representation of gender and acts as a neural hub.
Strengths:
The question is novel and important, while the logic of the study is straightforward. Examining faces, bodies, and objects within the same participants provides a useful extension beyond the predominantly face-based literature. Cross-category decoding is also a stronger test of shared information than simple anatomical overlap between within-category maps. The combination of MVPA, RSA, computational modelling, and connectivity analysis is ambitious, and the replication of the CNN layer profile with both AlexNet and VGG16 is a useful characterization of relevant information.
Weaknesses:
(1) The construct labelled "gender" is not equivalent across stimulus classes. For faces and bodies, the male/female label is intended to track a property of the depicted person, albeit one inferred imperfectly from appearance; for objects, masculinity or femininity is not an intrinsic property of the object but a culturally contingent association that may vary across observers and contexts. Treating both as levels of a single binary factor risks conflating person-category information with gender-stereotypic object associations and interpreting their common neural discriminability as evidence for one abstract concept of gender. The term "object gender" could also be confused with grammatical gender in some languages (e.g., French or German).
(2) The CNN analysis does not isolate the shared male/female component. The authors correlate the complete six-condition neural RDM with the complete CNN RDM. However, rMTG also carries substantial information about whether an image is a face, body, or object. Consequently, the peak correspondence with Conv4 may reflect category structure rather than the representation that supports cross-category male/female decoding. The current analysis does not establish that shared gender-related information specifically depends on mid-level features.
(3) The connectivity interpretation is overstated. PPI measures task-dependent covariance; it does not establish information transmission, directionality, or an upstream-to-downstream processing sequence. The reported face-body connectivity similarity is also small (r=.08). Also, describing rMTG as a "hub" is not justified without network-centrality measures, lesion evidence, or causal perturbation.
The authors partly achieve their aims. The results provide credible evidence that patterns in rMTG contain information that generalizes across binary male/female-labelled faces and bodies and masculine/feminine-associated objects. They do not yet establish a genuinely abstract representation of gender, a specifically gender-related correspondence with intermediate CNN layers, or a neural hub that transmits information through a directed network. With more precise framing and targeted reanalysis, the study could make a useful contribution to research on social vision and cross-category representation.
Reviewer #3 (Public review):
Summary:
In this work, the authors investigate whether gender information is encoded in the brain in a way that is invariant to the object being perceived. They design an fMRI experiment in which 22 participants perform a one-back repetition detection task in a block design. Images shown are of three types (faces, objects, and bodies) and of two perceived genders, male and female. They perform MVPA, RSA, and functional connectivity analyses to determine whether gender information is invariant to the type of image being perceived. They report an area in the posterior right middle temporal gyrus (rMTG) that is found in their gender decoding analysis across categories. To confirm that this area encodes gender information, they perform a regression-based RSA with category and gender model RDMs, and report that the gender model RDM is significantly correlated with brain representations in that area. Finally, to further investigate the representations in this area, they perform a model-based RSA in which they first fine-tune a deep neural network for gender classification, and then study the correlation between model RDMs and brain RDMs. Consistent with a previous report in face processing (Jiahui et al., 2023), they find that gender information is more consistent with representations in middle-to-late layers of the networks. Additional functional connectivity and PPI analyses are reported to reveal differences in co-fluctuation of brain activity within occipital and parietal nodes when perceiving different types of male/female images. Based on these results, the authors conclude that rMTG represents gender information invariant of the category perceived, although rMTG also afforded decoding of category information.
Strengths:
Whether perceived gender is represented in a manner invariant to the category of the stimulus is a legitimate and interesting question, and one of relevance particularly to the face and person perception literature.
The model-based RSA, in which RDMs from networks fine-tuned for gender classification are compared against brain RDMs, is an interesting approach, and the layer-wise profile the authors obtain converges with a previous report in the face processing literature (Jiahui et al., 2023).
Weaknesses:
A substantial number of inferences are drawn on the basis of weak statistical methods and a suboptimal design. My concerns are set out below, ordered by severity.
(1) The statistical tests are not appropriate for classification and RSA, and are prone to false positives. Classification accuracies and RSA correlations may be positively biased, and the true null distribution may therefore be centered above the nominal chance level, or above zero in the case of RSA. Testing against a theoretical value with a one-sample t-test under these conditions inflates the false positive rate, especially with few test samples per classification, and does not afford valid population inference for information-like measures (Combrisson & Jerbi, 2015; Allefeld et al., 2016). The concern applies to every inferential claim in the manuscript, including the identification of the rMTG cluster on which the paper's central conclusion rests. The established remedy is permutation testing, in which the labels are randomly permuted and the full analysis, including cross-validation, is re-computed so that any bias is captured in the empirical null distribution (Stelzer et al., 2013; Etzel & Braver, 2013). This approach has been applied in comparable face-decoding studies using both classification and RSA (Guntupalli et al., 2017). I raise this methodological concern here because it is the clearest way to convey why the reported statistics cannot be safely interpreted at face value.
(2) The decoding analyses do not appear to test generalization to left-out stimuli. From my reading of the design, each run contained all six conditions presented three times in random order, with each block containing 12 images (10 unique plus two repetitions serving as catch trials). If all images were presented in every run, the same images would be present in both the training and test sets of the cross-validation. Under these conditions, the interpretation of a general "gender" code is difficult to justify: the classifier may be exploiting low-level image features specific to the particular exemplars rather than gender per se. This bears directly on the paper's central claim, which concerns an abstract, category-invariant representation of gender, a claim that requires decoding to generalize to stimuli the classifier has not encountered.
(3) There is no evidence that participants perceived the stimuli's gender as the authors assumed. Perceived gender may be subject-specific, yet no norming is reported establishing that participants actually rated or processed the stimuli according to the gender the authors assigned to each image. Some images are likely to be more ambiguous than others. This is a construct validity issue rather than an analysis issue: the class labels used throughout the decoding analyses, and the gender model RDM used in the RSA, both rest on an assumption about the participants' percepts that is never tested against the participants themselves.
(4) The rMTG ROI reported in Figure 2c appears to overlap almost perfectly with the motion-sensitive area hMT+. The reported effects may therefore be driven, at least in part, by low-level motion signals arising from the rapid on/off changes of the stimuli and the associated optic flow. I am not claiming that the results are fully driven by this, but no control reported in the manuscript rules it out, and this region is the centerpiece of the paper's conclusion.
(5) Stimulus size is confounded with category in the functional connectivity analyses. The authors report that functional connectivity differed between faces and objects, and between bodies and objects. However, faces and bodies were shown with the same visual extent, while objects were larger. Given that the nodes being investigated are in visual areas, it is unclear how these differences can be attributed to category rather than to the low-level difference in stimulus size. The same confound bears on the behavioral task performed within the scanner: participants can perform the one-back task more easily, simply by detecting size differences, since two images of different sizes are clearly not the same image, rather than by processing the image content. This affects what can be assumed about participants' attention to the stimulus category or gender.
(6) No motion quality control is reported for the functional connectivity analyses. Functional connectivity is well known to be highly susceptible to head motion, yet the manuscript reports no summary of how much subject motion there was, no indication of whether volumes with excessive motion were removed or censored, and no account of quality control on the measured data more generally.
(7) The use of famous faces introduces an avoidable confound. The face stimuli were famous faces. Famous and familiar faces are known to recruit substantially more widespread activity than unfamiliar faces, extending well beyond the core visual system (Gobbini & Haxby, 2007; Natu & O'Toole, 2011; Visconti di Oleggio Castello et al., 2017; Kovacs, 2020). For a study focused specifically on gender, this introduces a source of variance that unfamiliar faces would have avoided, and it complicates the comparison of the face conditions against the body and object conditions.
(8) The rationale and benefit of fine-tuning the deep neural networks are not established. The manuscript does not report the original, non-fine-tuned accuracy of the models that required fine-tuning, so the benefit of the procedure cannot be assessed; given that the final validation accuracy is low, it is unclear that fine-tuning actually helped. AlexNet and VGG are trained for object classification on large datasets, and fine-tuning with 2,000 training images may not be sufficient to genuinely shift the objective. Whether the activation patterns and RDMs changed in any significant manner after fine-tuning is not reported, and the rationale for selecting the specific layers used is not stated.
(9) Taken together, the analyses as presented do not establish the paper's central claim. My concern is not that the reported effects are necessarily absent, but that the combination of statistical tests that do not account for possible positive bias, a cross-validation scheme that may not guarantee generalization across stimuli, a key region that coincides with a motion-sensitive area, and gender labels that were never validated against participants' own perception leaves too many open questions for the results to be evaluated as they stand.
(10) I would add one broader consideration. Perceived gender is likely to depend on culture and to vary across individuals. A binary male/female contrast in 22 participants, without evidence that those participants perceived the stimuli as the authors intended, is a narrow operationalization of a construct that is unlikely to be so simple. Even if the analyses were fully sound, caution would be warranted in generalizing from this design to claims about how the brain universally represents gender.