Larger language models better align with neural representations of natural language

  1. Zhuoqiao Hong
  2. Haocheng Wang  Is a corresponding author
  3. Zaid Zada
  4. Harshvardhan Gazula
  5. David Turner
  6. Bobbi Aubrey
  7. Leonard Niekerken
  8. Werner Doyle
  9. Sasha Devore
  10. Patricia Dugan
  11. Daniel Friedman
  12. Orrin Devinsky
  13. Adeen Flinker
  14. Uri Hasson
  15. Samuel Nastase
  16. Ariel Y Goldstein
  1. Department of Psychology and the Neuroscience Institute, Princeton University, United States
  2. McGovern Institute for Brain Research, Massachusetts Institute of Technology, United States
  3. New York University Grossman School of Medicine, United States
  4. Business School, Data Science Department and Cognitive Science Department, Hebrew University, Israel
5 figures, 4 tables and 1 additional file

Figures

Naturalistic language comprehension model comparison framework.

(A) Participants listened to a 30 min story while undergoing electrocorticography (ECoG) recording. A word-level aligned transcript was obtained and served as input to four language models of varying size from the same GPT-Neo family. (B) For every layer of each model, a separate linear regression encoding model was fitted on a training portion of the story to obtain regression weights that can predict each electrode separately. Then, the encoding models were tested on a held-out portion of the story and evaluated by measuring the Pearson correlation of their predicted signal with the actual signal. (C) Encoding model performance (correlations) was measured as the average over electrodes and compared between the different language models.

Figure 2 with 6 supplements
Model performance improves with increasing model size.

(A) The relationship between model size (measured as the number of parameters, shown on a log scale) and perplexity: as the model size increases, perplexity decreases. Each data point corresponds to a model. (B) The relationship between model size (shown on a log scale) and brain encoding performance: correlations for each model are calculated by averaging the maximum correlations across all lags and layers across electrodes. As the model size increases, the encoding performance increases. Each data point corresponds to a model. The error bars represent standard error. (C) For the GPT-Neo model family, the relationship between encoding performance and layer number. Encoding performance is best for intermediate layers. The shaded colors represent standard error. (D) Same as C, but the layer number was transformed to a layer percentage for better model comparison.

Figure 2—figure supplement 1
Heatmap comparing encoding performance between models.

Encoding performance for a model is represented by the maximum correlation across lags and layers per electrode. Paired t-tests are performed across electrodes (df = 159). The result is blue if t<0, meaning the larger model (the model with a higher number of parameters, represented on the x-axis) outperforms the smaller model (the model with a smaller number of parameters, represented on the y-axis). The result is red if t>0, meaning the smaller model outperforms the larger model. The shades of the colors represent significance (p<0.001 or p<0.01 or not significant, two-sided, false discovery rate [FDR]-corrected). There is a positive relationship between model size and encoding performance when models are smaller than 3 billion parameters. When models are larger than 7 billion parameters, we observed a plateau in the maximal encoding performance.

Figure 2—figure supplement 2
Model performance improves with increasing model size.

To control for the different embedding dimensionality across models, we standardized all embeddings to the same size using principal component analysis (PCA) and trained linear encoding models using ordinary least-squares (OLS) regression (Figure 2). (A) Replication of Figure 2B, the relationship between model size (shown on a log scale) and brain encoding performance: encoding performance increases as model size increases. Each data point corresponds to a model. (B) Ridge regression encoding outperforms PCA+OLS regression encoding for all 20 transformer-based language models (paired two-sided t-test across electrodes, df = 159, p<0.001, Bonferroni-corrected). Each data point corresponds to a model.

Figure 2—figure supplement 3
Lag-wise encoding for the GPT-Neo Family.

Top: Lag-wise encoding for all four models of the GPT-Neo family, averaged across electrodes. Bottom: Lag-wise encoding difference for the three bigger models compared to SMALL, averaged across electrodes. The dots represent lags where XL significantly outperformed SMALL (paired two-sided t-test across electrodes, df = 159, p<0.001, Bonferroni-corrected). XL significantly outperformed SMALL in encoding models for most lags from 2000 ms before word onset to 575 ms after word onset.

Figure 2—figure supplement 4
The relationship between encoding performance and layer number for pretrained and untrained SMALL.

For pretrained SMALL, encoding performance is best for intermediate layers. For untrained SMALL with randomly initialized weights, encoding performance is best for the 0th layer. Encoding performance is significantly higher for pretrained SMALL than for untrained SMALL for every layer (paired two-sided t-test across electrodes, df = 159, p<0.001, Bonferroni-corrected). The shaded colors represent standard error.

Figure 2—figure supplement 5
Contextual embeddings from large language models (LLMs) outperform classical speech features and GloVe embeddings.

(A) Maximum encoding performance (across lags) for acoustic features (60 dimensions), linguistic features (191 dimensions), word frequency (2 dimensions), acoustic features+linguistic features+word frequency (253 dimensions), GloVe (50 dimensions), SMALL (768 dimensions), and XL (6144 dimensions) using ridge regression. SMALL and XL showed significantly better performance than other features (paired two-sided t-test across electrodes, df = 159, p<0.001). (B) Replication of A using principal component analysis (PCA) and ordinary least-squares (OLS) regression. To control for the different dimensions of the different feature spaces, we standardized all feature sets to 50 dimensions (2 dimensions for word frequency) using PCA and trained OLS regression encoding models. SMALL and XL showed significantly better performance than other embeddings (paired two-sided t-test across electrodes, df = 159, p<0.001). (C) Encoding performance across lags for all features. Shaded colors represent standard error across electrodes. (D) Replication of C using PCA and OLS regression.

Figure 2—figure supplement 6
Model performance improves with increasing dataset size.

(A) For SMALL, the relationship between the percentage of training data and brain encoding performance across layers. Encoding performance increases as the training dataset size increases. The shaded colors represent standard error across electrodes. (B) Same as A, but for model XL.

Figure 3 with 1 supplement
Model performance improves with increasing model size across electrodes and regions of interest (ROIs).

(A) Maximum correlation per electrode for SMALL. The encoding model achieves the highest correlations in superior temporal gyrus (STG) and inferior frontal gyrus (IFG). (B) For MEDIUM, LARGE, and XL, the percentage difference in correlation relative to SMALL for all electrodes with significant encoding differences. The encoding performance is significantly higher for the bigger models for almost all electrodes across the brain (pairwise t-test across cross-validation folds). (C) Maximum encoding correlations for SMALL and XL for each ROI (middle STG [mSTG], anterior STG [aSTG], Brodmann area 44 [BA44], Brodmann area 45 [BA45], and temporal pole [TP] area). The encoding performance is significantly higher for XL for all ROIs except TP. Each data point corresponds to an electrode in the corresponding ROI. (D) Percent difference in correlation relative to SMALL for all ROIs. As model size increases, the percent change in encoding performance also increases for mSTG, aSTG, and BA44. After the MEDIUM model, the percent change in encoding performance plateaus for BA45 and TP. The shaded colors represent standard error.

Figure 3—figure supplement 1
Brain map of electrodes in five regions of interest (ROIs) across the cortical language network: middle superior temporal gyrus (mSTG, n=28 electrodes), anterior superior temporal gyrus (aSTG, n=13 electrodes), Brodmann area 44 (BA44, n=19 electrodes), Brodmann area 45 (BA45, n=26 electrodes), and temporal pole (TP, n=6 electrodes).
Relative layer preference varies with model size.

(A) Relative layer (in percentage of total number of layers) with peak encoding performance for all four GPT-Neo models: the larger the model size, the earlier relative layer where the encoding performance peaks. (B) The relationship between model size (shown on a log scale) and best encoding layer (in percentage) for all four model families: as the model size increases, the best encoding layer (in percentage) decreases, although the rate of decrease is different between model families. We estimate a linear regression model per model family of the form: best percent layer~log(model size). The slopes (β) indicate the decrease in the relative best-performing layer at increasing log model size; p-values are obtained from a Wald test against the null hypothesis that the slope is 0. Each data point corresponds to a model. (C) Best relative encoding layer (in percentage) for all four GPT-Neo models. (D) Best encoding layer for XL with electrodes that peak in the first half of the model (layers 0–22). (E) Best encoding layer (in percentage) for SMALL and XL for each region of interest [ROI] (middle superior temporal gyrus [mSTG], anterior superior temporal gyrus [aSTG], Brodmann area 44 [BA44], Brodmann area 45 [BA45], and temporal pole [TP]). Each data point corresponds to an electrode in the corresponding ROI.

Figure 5 with 1 supplement
Encoding performance across lags does not vary with model size.

(A) Average region of interest (ROI) encoding performance for SMALL and XL models. Middle superior temporal gyrus (mSTG) encoding peaks first before word onset, then anterior superior temporal gyrus (aSTG) peaks after word onset, followed by Brodmann area 44 (BA44), Brodmann area 45 (BA45), and temporal pole (TP) encoding peaks at around 400 ms after onset. The dots represent the peak lag for each ROI. (B) Lag with best encoding performance correlation for each electrode, using SMALL and XL model embeddings. Only electrodes with the best lags that fall within 600 ms before and after word onset are plotted.

Figure 5—figure supplement 1
The optimal lags for each electrode do not exhibit significant variation when transitioning between SMALL and XL models.

(A) Scatter plot of best-performing lag for SMALL and XL models, colored by max correlation. Each data point corresponds to an electrode. (B) Scatter plot of best-performing lag for SMALL and XL models, colored by regions of interest (ROIs). Each data point corresponds to an electrode. Only the electrodes in Figure 3—figure supplement 1 are included.

Tables

Appendix 1—table 1
Summary of four families of open large language models: GPT-2, GPT-Neo, OPT, and Llama-2.

Context length is the maximum context length for the model, ranging from 1024 to 4096 tokens. The model name is the model’s name as it appears in the transformers package from Hugging Face (Wolf et al., 2019). Model size is the total number of parameters; M represents million, and B represents billion. The number of layers is the depth of the model, and the hidden embedding size is the internal width.

Model familyContext lengthModel nameModel sizeNumber of layersHidden embedding size
GPT-21024Distilgpt282 M6768
gpt2124 M12768
gpt2-medium355 M241024
gpt2-large774 M321280
gpt2-xl1.5 B481600
GPT-Neo2048EleutherAI/gpt-neo-125M125 M12768
EleutherAI/gpt-neo-1.3B1.3 B242048
EleutherAI/gpt-neo-2.7B2.7 B322560
EleutherAI/gpt-neox-20b20 B446144
OPT2048facebook/opt-125m125 M12768
facebook/opt-350m350 M241024
facebook/opt-1.3b1.3 b242048
facebook/opt-2.7b2.7 B322560
facebook/opt-6.7b6.7 B324096
facebook/opt-13b13 B405120
facebook/opt-30b30 B487168
facebook/opt-66b66 B649216
Llama-24096meta-llama/Llama-2-7b-hf7 B324096
meta-llama/Llama-2-13b-hf13 B405120
meta-llama/Llama-2-70b-hf70 B808192
Appendix 1—table 2
Classic speech features.

Lower-level acoustic features include phoneme (39 classes), place of articulation (9 classes), manner of articulation (9 classes), and voiced or voiceless (3 classes). Linguistic features include part of speech (17 classes), tag (50 classes), function or content (3 classes), prefix (30 classes), suffix (44 classes), dependency (45 classes), whether the word is an alpha character (binary), and whether the word is a stop word (binary). Word frequency includes the frequency of the word in the Google Web Trillion Word Corpus and in our dataset.

Feature familyFeatureDimension
Acoustic featuresPhonemes39
Place of articulation9
Manner of articulation9
Voiced or voiceless3
Linguistic featuresPart of speech (POS)17
Tag (detailed POS)50
Function or content3
Prefix30
Suffix44
Dependency45
Is alpha word1
Is stop word1
Word frequencyGoogle Web Trillion Word Corpus1
In our dataset1
Appendix 1—table 3
Summary statistics and paired t-test results for maximum correlations between SMALL and XL models across five regions of interest (ROIs).

Encoding performance for the XL model significantly surpassed that of the SMALL model in whole brain, middle superior temporal gyrus (mSTG), anterior superior temporal gyrus (aSTG), Brodmann area 44 (BA44), and Brodmann area 45 (BA45).

ROIElectrode numSMALL corr meanSMALL corr stdXL corr meanXL corr stdCorr diff t-valueCorr diff p-value
All1600.2550.0970.2840.10916.7914.770e–37
mSTG280.3350.0870.3770.09811.5845.538e–12
aSTG130.2740.0960.3110.1086.1115.247e–5
BA44190.2240.0680.2570.0826.3405.663e–6
BA45260.2890.1050.3160.1136.1072.205e–6
TP60.2430.0860.2630.0861.9910.103
Appendix 1—table 4
Summary statistics and paired t-test results for best-performing layers (in percentage) for the SMALL model across five regions of interest (ROIs).

The best-performing layer (in percentage) occurred earlier for electrodes in middle superior temporal gyrus (mSTG) and anterior superior temporal gyrus (aSTG) and later for electrodes in Brodmann area 44 (BA44), Brodmann area 45 (BA45), and temporal pole (TP).

ROIsmSTG
M=49.107
SD = 9.168
aSTG
M=50.000
SD = 22.82
BA44
M=60.965
SD = 11.471
BA45
M=60.897
SD = 14.485
TP
M=59.722
SD = 15.290
aSTGt(39)=–0.180
p=0.429
BA44t(45)=–3.930
p=1.450e–5
t(30)=–1.797
p=0.041
BA45t(52)=–3.601
p=3.538e–5
t(37)=–1.820
p=0.038
t(43)=0.017
p=0.507
TPt(32)=–2.276
p=0.0148
t(17)=–0.943
p=0.179
t(23)=0.214
p=0.584
t(30)=0.177
p=0.570

Additional files

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Zhuoqiao Hong
  2. Haocheng Wang
  3. Zaid Zada
  4. Harshvardhan Gazula
  5. David Turner
  6. Bobbi Aubrey
  7. Leonard Niekerken
  8. Werner Doyle
  9. Sasha Devore
  10. Patricia Dugan
  11. Daniel Friedman
  12. Orrin Devinsky
  13. Adeen Flinker
  14. Uri Hasson
  15. Samuel Nastase
  16. Ariel Y Goldstein
(2026)
Larger language models better align with neural representations of natural language
eLife 13:RP101204.
https://doi.org/10.7554/eLife.101204.3