Single unit data.

a. Participants had to learn the order of 4 pictures presented one after another. After a 3-second delay, we queried the sequence position of one of the pictures. b. On each trial, we used the same four pictures but changed the order in which they were displayed. c. The averaged population trajectory throughout a full trial of the task in neural activity space shows that different task stages occupy different regions in neural activity space. d. Decoding task stage from the population firing rates. Left: average decoded task stage (columns) given true task stage firing rates (rows) across trials, with chance level (1/8) indicated on the color scale. Right: decoding error. e. Raster plot across trials for a hippocampal cell showing significant task-stage modulation. The red line is the average firing rate. f. Raster plot across trials for an entorhinal cell showing significant task-stage modulation. The red line is the average firing rate. g. Proportion of cells showing significant task-stage coding (initiation, stim1–4, delay, question, feedback; one-way ANOVA). h-l: Example picture or coordinate cells. We averaged the firing rate of each cell across all trials of a particular picture (column) in a particular position (row). h. Hippocampal cell coding for picture 1. i. Hippocampal coordinate cell that increases its firing from the first to the last sequence position. j. Entorhinal cell coding for picture 3. k. Entorhinal coordinate cell that decreases its firing from the first to the last sequence position. l. Other examples of picture (top) and position (bottom) specific cells in hippocampus (left) and entorhinal cortex (right) averaged across time within trial stage as a compact variant of the plots in h-k. The maximum average firing rate is indicated by the white number for each cell. m. Proportions of picture and position specific cells in hippocampus (left) and entorhinal cortex (right). n. Variance explained through time from stimulus presentation by picture and position factors on hippocampus (left) and entorhinal cortex (right).

fMRI experimental design and behaviour.

a-d. Schematic of logic for anatomical hypothesis: coordinates systems organised anatomically by scale in EC. a-b. Grid cells are a coordinate system for space. a. Simulated example grid rate maps. Grid cells in rodent entorhinal cortex group together in anatomically separable modules of different spatial scales (rows). Cells within a module exhibit different shifts or ‘phases’ (columns, in differently dashed boxes). Adapted from Stemmler et al., 20153. b. Together, these grid cell modules implement a factorised hierarchical coordinate system, here illustrated in 1D for clarity. Given the tuning curves of each grid cell in a module (left, different dashes), a particular firing pattern across cells (e.g. cell 3 in module 1 is active while cells 1 and 2 are silent) specifies a particular (module scale dependent) distribution over locations (right) - the readout of the coordinate. Combining the multi-scale coordinates across modules yields a hierarchical code for one specific location (grey, bottom). c. We hypothesise that similar factorised hierarchical codes in the human brain provide a scaffold for memories. For example, the hierarchical representation ‘16 July 1969’ corresponds to the memory of the first man on the moon. d. For physical space, this gradient of coordinate representations from small to large scales is organised anatomically in rodent entorhinal cortex along the dorsal (small scale grid cells) to ventral (large scale grid cells) axis. In humans, this axis has rotated into the posterior-to-anterior direction. e. We test this hypothesis in a four-day neuroimaging experiment. Participants learn auditory sequences (day 1-2) with specific locations tagged by pictures (day 3). We show the same pictures in random order in the MRI scanner on day 4, without any auditory stimuli. After the scan, we ask participants to draw the sequences and place the pictures at their associated locations (an example drawing in Fig S2.2) f. Hierarchical structure of the auditory sequences (full sequences in Fig S2.1). g. Participants learned two sequences with identical structures but different sounds. h. Example pictures and their specific sequence locations that participants learned on the third day. The 1D position of every picture in the auditory sequence can be pinpointed using four coordinates. In the fMRI scanner we showed pictures in a random order without sounds, with the goal of reactivating the representations of the corresponding auditory coordinates. i. Distance between incorrectly placed pictures in the post-scan drawings of day 4 and their correct locations. The three strongest peaks in the histogram correspond to incorrectly identifying the coordinate in one level while correctly identifying all other coordinates. j. The average accuracy of the post-scan drawings across all pictures for each hierarchical coordinate.

Coordinate representations are anatomically organised.

a. We present pictures associated to specific coordinates in random order in the MRI scanner. b. To find coordinate representations, we specify hypothesised similarities between pictures with respect to each level of the hierarchy. For example, the dog and the book have high similarity for coordinate C2 (blue: both 1), but low similarity for coordinate C4 (red: 4 versus 1). c. We use RSA to detect within entorhinal cortex (left) where neural similarities match the hypothesised similarities (example participant, middle), along the anterior-posterior axis (right). d. To quantify the A-P order of coordinate representations, we calculate an “orderness metric” from differences between projections (Yi) of each coordinate representation’s centre of mass (CoMi) on the anterior-posterior axis. If the gradient is present, the differences are expected to be positive, because we always subtract the lower from the higher level, and the projection value increases in anterior direction. e-g. Visualisation of group averaged coordinate representations in dataset 1 (e), dataset 2 (f), and both combined (g). h-j. Orderness metric for neighbouring levels’ pairs only (left), including all pairs of levels (middle), and averaged across participants against a null distribution from shuffled level labels (right) in dataset 1 (h), dataset 2 (i), and both combined (j).

Abstract generalisable coordinate system.

a. We presented pictures associated to specific coordinates in random order in the MRI scanner. Notice that some were paired with a sequence 1 coordinate during training (yellow dashed line) and others with sequence 2 (purple dashed line). b. Across-sequence similarity matrices. To find coordinate representations, we specify hypothesised similarities between pictures with respect to each level of the hierarchy. This time we include only similarities between pictures from different auditory sequences for further analysis. c-e. Visualisation of group averaged coordinate representations in dataset 1 (c), dataset 2 (d), and both combined (e). f-h. Orderness metric for neighbouring levels pairs only (left), including all pairs of levels (middle), and averaged across participants against a null distribution from shuffled levels’ labels (right) in dataset 1 (f), dataset 2 (g), and both combined (h).

a. Combinatorial electrode. Right: schematic showing macro- (blue) and microelectrode (green) contacts along the depth electrode. Left: example MRI from a patient illustrates electrode placement, with visible macro- and microcontacts, with an overlaid schematic of the electrode positioned slightly below its actual location in the MRI. b. Electrode localization in MNI space, showing electrode positions in the Freesurfer average brain model, with the entorhinal cortex (EC, green) and hippocampus (HC, orange) highlighted. Electrodes are color-coded according to their respective anatomical regions. c. Spike sorting example, showing raw recordings (top) and recordings filtered in 300-3000 Hz range (middle) from a single microelectrode channel. Detected spikes are marked and color-coded corresponding to the three identified neurons (bottom). Individual spike waveforms (lighter color) and averaged waveforms (darker color) are displayed for three identified neurons.

Spike Sorting Quality Assessment.

Summary of key metrics used to evaluate the reliability of spike sorting: a. the number of neuronal units identified per electrode contact, b. the percentage of inter-spike intervals (ISI) shorter than 3 ms, c. the average firing rate in Hz d. the peak amplitude (µV) of detected spikes, calculated as the mean across waveform peaks for identified neurons, and e. the signal-to-noise ratio (SNR) of the peak of the mean waveform, defined as the ratio of the peak amplitude of the neuron’s mean spike waveform to the background noise level (calculated as the peak of the averaged waveform divided by the mean of the standard deviations of individual spike waveforms).

P-values table for Fig 3.

P-values for orderness metric for neighbouring pairs of levels (left column), including all pairs of levels (middle column), and averaged across participants against a null distribution from shuffled level labels (right column) in dataset 1 (first row), dataset 2 (second row), and both combined (third row). The rows correspond to figure 3h-j: first row to 3h, second row to 3i, third row to 3j. The columns also correspond to figure 3h-j: first column to 3h-j left, the second column to 3h-j middle and the third column to 3h-j right. For Dataset 1, the significance threshold was corrected for multiple comparisons across both hemispheres.

Summary firing rate plots for all hippocampal cells.

For each cell we average firing rates across time and across trial stages with a particular picture (columns) in a particular position (rows), then plot the 4×4 firing rate map with the maximum firing rate in yellow and the minimum firing rate in blue. The number above each plot indicates the maximum averaged firing rate, i.e. the firing rate for the yellow square (if that number is 0, it means the cell didn’t fire during the stimulus presentation stages of the task, but it was active in different task stages). The box outline denotes whether the cell is significantly tuned to picture and position (blue), position only (orange), picture only (yellow), or neither (grey).

Summary firing rate plots for all entorhinal cells.

For each cell we average firing rates across time and across trial stages with a particular picture (columns) in a particular position (rows), then plot the 4×4 firing rate map with the maximum firing rate in yellow and the minimum firing rate in blue. The number above each plot indicates the maximum averaged firing rate, i.e. the firing rate for the yellow square (if that number is 0, it means the cell didn’t fire during the stimulus presentation stages of the task, but it was active in different task stages). The box outline denotes whether the cell is significantly tuned to picture and position (blue), position only (orange), picture only (yellow), or neither (grey).

Single unit tuning to only position 2 and 3.

Cells that significantly code for stimulus position, picture, or both (analogous to Fig 1m) when excluding all data from position 1 and 4.

Single unit position preference.

Number of cells within position coding cells that have a maximum firing rate for each position (left), with the average z-scored firing rates of each group plotted throughout the stimulus presentations stages of the trial (right).

Auditory sequences.

Participants learnt two sequences with identical structures but different sounds.

An example participant drawing of the auditory sequence and pictures in it.

The letters are indicating the sounds in the sequence, and the numbers are referring to the pictures. We did not explicitly instruct participants about the hierarchical nature of the sequence; the division along blocks and rows is their own interpretation.

a. T-statistic for each level’s coordinate representation (C1 is pink, C2 is blue, C3 is green, C4 is red) averaged across participants within A-P slices as a function of slice A-P position in MNI space. Note that the peaks are ordered according to the hypothesised order (both datasets together). b. To test for the anatomical consistency of EC coordinate representations across datasets, we compare a level’s t-statistic at that level’s centre of mass from the other dataset (c: congruent) against its t-statistics at other levels centre of mass from the other dataset (i: incongruent). See Methods for more details. c. Congruent effects were stronger than non-congruent effects, p=0.0039.

Cells and sessions per patient.

Number of sessions completed for each patient, with the result number of hippocampal (hpc) and entorhinal (ec) units and the corresponding neurons significantly tuned to position (hpc_tbls1.

The orderness metrics computed on centres of mass defined in the right entorhinal mask.

Orderness metric for neighbouring pairs of levels (left column), including all pairs of levels (middle column), and averaged across participants against a null distribution from shuffled level labels (right column) in dataset 1 (first row), dataset 2 (second row), and both combined (third row). See corresponding p-values in the table below.

P-values table, corresponding to the Fig S3.2, for the orderness metrics computed on centres of mass defined in the right entorhinal mask.

For Dataset 1, the significance threshold was corrected for multiple comparisons across both hemispheres.

The orderness metrics computed on centres of mass defined in the stricter entorhinal mask (atlas probability 75% instead of the 50% probability mask used in all other analyses, see Methods).

Orderness metric for neighbouring pairs of levels (left column), including all pairs of levels (middle column), and averaged across participants against a null distribution from shuffled level labels (right column) in dataset 1 (first row), dataset 2 (second row), and both combined (third row). See corresponding p-values in the table below.

P-values table, corresponding to the Fig S3.3, for the orderness metrics computed on centres of mass in the stricter entorhinal mask.

For Dataset 1, the significance threshold was corrected for multiple comparisons across both hemispheres.

Discrete orderness metric

(see Methods) for neighbouring pairs of levels (left column), including all pairs of levels (middle column), and averaged across participants against a null distribution from shuffled level labels (right column) in dataset 1 (first row), dataset 2 (second row), and both combined (third row). See corresponding p-values in the table below.

P-values table, corresponding to the Fig S3.4, for the discrete orderness metric.

For Dataset 1, the significance threshold was corrected for multiple comparisons across both hemispheres.

The orderness metrics computed on centres of mass defined on statistical maps that result from separate regressions (fitting 4 similarity regressors separately (one for each level), Methods).

Orderness metric for neighbouring pairs of levels (left column), including all pairs of levels (middle column), and averaged across participants against a null distribution from shuffled level labels (right column) in dataset 1 (first row), dataset 2 (second row), and both combined (third row). See corresponding p-values in the table below.

P-values table, corresponding to the Fig S3.5, for the orderness metrics computed on statistical maps that result from a multiple regression.

For Dataset 1, the significance threshold was corrected for multiple comparisons across both hemispheres.

P-values table for Fig 4: across-sequence similarities matrices.

P-values for orderness metric for neighbouring pairs of levels (left column), including all pairs of levels (middle column), and averaged across participants against a null distribution from shuffled level labels (right column) in dataset 1 (first row), dataset 2 (second row), and both combined (third row). The rows correspond to Fig 4f-h: first row to 3f, second row to 3g, third row to 3h. The columns also correspond to Fig 3f-h: first column to 3f-h left, the second column to 3f-h middle and the third column to 3f-h right. For Dataset 1, the significance threshold was corrected for multiple comparisons across both hemispheres.

The null result for orderness metrics computed on centres of mass defined in the hippocampal mask.

Orderness metric for neighbouring pairs of levels (left column), including all pairs of levels (middle column), and averaged across participants against a null distribution from shuffled level labels (right column) in dataset 1 (first row), dataset 2 (second row), and both combined (third row).

Whole-brain statistical maps for the representational similarity analysis (RSA) of each hierarchical layer coordinate.

Maps are shown at a low statistical threshold for visualization purposes. Note that these group-level maps were not used for any statistical analysis. Instead, we use subject-level statistical maps to define a within-subject measure of representational order (the “orderness metric”), which is the reported statistic in all analyses. This within-subject order measure is more sensitive because it avoids signal loss due to averaging in a small brain region that is difficult to align across participants. Because widespread activation is typical in whole-brain fMRI analyses, our study used a strictly hypothesis-driven approach, focusing specifically on representational order within subjects within the entorhinal cortex.

Grid cell and place cell hierarchies.

a. Grid cell tuning curves, exhibiting periodic responses at different phases within a module (dotted/dashed/solid lines) at different scales across modules (colours). b. Decoded location from each module independently (colours) and jointly (grey, bottom row). These grid responses implement a factorised hierarchy, where each module encodes an independent hierarchical coordinate, and decoding a single location is only possible by combining the representations across modules – just like a calendar event is only fully determined by time, date, and month and not by any one of those. c. Place cell tuning curves that show individual peaks at different locations (dotted/dashed/solid lines) at different spatial scales (colours). d. Decoded location from each module independently (colours) and jointly (grey, bottom row). The location can be decoded separately from every module, at different levels of granularity – just like a street address that specifies street, town, and country.