A membrane insertion code for intrinsically disordered proteins

  1. Fidha Nazreen Kunnath Muhammedkutty
  2. Huan-Xiang Zhou  Is a corresponding author
  1. Department of Chemistry, University of Illinois Chicago, United States
  2. Department of Physics, University of Illinois Chicago, United States
6 figures, 4 tables and 1 additional file

Figures

Classification of aromatic-centered motifs into membrane inserters and non-inserters by MD simulations.

(A) Initial dataset containing 10 sequences (sequence 1 from TrkA, 2 from Drp1, 3 and 4 from CD3ε, and 5–10 from CD3ζ). (B) Initial snapshot, time trace of aromatic Ztip, and snapshot at 980.7 ns for sequence 1. (C) Initial snapshot, time trace of aromatic Ztip, and snapshot at 964.3 ns for sequence 6. (B) and (C) represent simulations where the peptides are inserted or not inserted, respectively, into the membrane. Insertion means at least 80% of the frames in the last 100 ns (shaded region) have aromatic Ztip at least 3.1 Å (indicated by a red dashed line) below the phosphate plane, which is set to Z=0. Time traces are smoothed using the moving average in a 2.5-ns window. The membrane is displayed as surface, with phosphates and all atoms above in orange while those below in gray; the peptide is displayed as cartoon in cyan, with the central aromatic residue as sticks in magenta and red. (D) The number of replicate simulations, out of a total of 20, where the peptide is inserted into the membrane.

Figure 2 with 5 supplements
Differences and similarities in membrane insertion and cation-π interactions among F-, W-, and Y-centered peptides.

(A) Aromatic Ztip distribution for an F-centered peptide (sequence 1). Insets: snapshot with F deeply inserted to interact with lipid acyl chains (bottom), or above the membrane to form cation-π interaction with a choline group (top). (B) Aromatic Ztip distribution for a W-centered peptide (sequence 2). Insets: snapshot with W deeply inserted to interact with lipid acyl chains and form a hydrogen bond with a glycerol oxygen atom (bottom), or in the headgroup region to form cation-π interaction with a choline group (top). (C) Aromatic Ztip distribution for a Y-centered peptide (sequence 3). Insets: snapshot with Y in the headgroup region to form a hydrogen bond with a glycerol oxygen atom (bottom), or near the membrane to form cation-π interaction with a choline group (top). In A–C, results were accumulated from the last 100 ns of 20 simulations of each peptide; a red dashed line is drawn at Ztip = –3.1 Å. (D) Illustration of stable positions of F, W, and Y with respect to the membrane. (E) Extremely strong correlation between the percentage of membrane-inserting frames in MD simulations and the percentage of membrane-inserting conformations in PPM runs (correlation coefficient = 0.995). The line represents the fitted regression line.

Figure 2—figure supplement 1
Aromatic tilt angle distributions in membrane-inserted frames of MD simulations.

(A) Definition of tilt angle. (B–D) Tilt angle distributions for F-, W-, and Y-centered sequences 1, 2, and 3, respectively, along with snapshots illustrating aromatic poses at major and minor peak angles. (E) Time traces of the aromatic Ztip and Z coordinate of the glycerol C2 carbon atom in a partner lipid in the simulation of Y-centered sequence 3 with 57.0% inserting frames. Lines at Z=5 Å indicate hydrogen bond breakup; hydrogen bond reformation may be with a different lipid molecule. A red line is drawn at Z=–3.1 Å. (F) Correlation between the aromatic Ztip and Z coordinate of the glycerol C2 carbon atom (correlation coefficient = 0.781). The red line represents the fitted regression line.

Figure 2—figure supplement 2
Sampling of intermediate and (partially) inserted states and formation of cation-π interactions by aromatic-centered peptides.

(A, B) Aromatic Ztip distributions for F- and W-centered sequences 1 and 2 in all simulations, in membrane-inserting simulations, and among membrane-interacting or cation-π forming frames. Insets: aromatic Ztip distributions in partly inserting simulations. (C) Corresponding results for Y-centered sequence 3, except that partly inserting simulations are presented in the main panel. In all panels, a red line is drawn at Ztip = –3.1 Å.

Figure 2—figure supplement 3
Ztip profiles of aromatic-centered peptides calculated from the last 100 ns of (partly) membrane-inserting simulations.

(A, B) Results from 3 and 8 membrane-inserting simulations, respectively, of F- and W-centered sequences 1 and 2. Results from individual simulations are in color; averages among the simulations are in black. (C) Corresponding results for Y-centered sequence 3, except the two simulations are partly inserting. In all panels, a red line is drawn at Ztip = –3.1 Å.

Figure 2—figure supplement 4
Comparison in membrane insertion profile between our simulations with a 9-residue peptide and Wang et al.’s simulations with a 26-residue IDR from TrkA.

The insertion distance was calculated following the protocol of Wang et al., 2019 as the difference between the Z coordinate of the all-atom center of mass of a residue and the average Z coordinate of the lipid phosphorus atoms in each frame; the values in the last 50 ns of the three membrane-inserting simulations were pooled and the median is plotted.

Figure 2—figure supplement 5
Partial helix formation by the WRGML motif of Drp1 in the membrane-inserting simulations.

(A) Snapshot of W-centered sequence 2 in a membrane-inserting simulation, with WRGML forming an amphipathic helix. (B, C) Snapshots in two other membrane-inserting simulations with no helix formation. (D) Cα secondary chemical shifts, averaged over the three membrane-inserting simulations of (A–C).

Figure 3 with 1 supplement
Three-step pathways to membrane insertion.

(A) F-centered sequence 1. Basic side chains at positions −4, –3, and –1 first tether one end of the peptide to the membrane surface. R[+4] then tethers the peptide at the other end, situating the central F for cation-π interaction. Finally, the aromatic side chain is inserted into the acyl chain region. (B) W-centered sequence 2. The R[+1] side chain initiates membrane contact. With possible stabilization by the membrane penetration of an aliphatic residue at position +3 or +4, the central W forms cation-π interaction with a choline and then enters the acyl chain region. (C) Y-centered sequences. The initial membrane contact can be formed by either a flanking basic side chain (R[–3]; via electrostatic interaction) or the central Y itself (via cation-π interaction). After stabilizing at the membrane surface by R[–3] and L[+3], the central Y lowers into the headgroup region to form a hydrogen bond with a glycerol oxygen atom. The aromatic ring occasionally dips into the acyl chain region, bringing the lipid glycerol with it and producing local curvature.

Figure 3—figure supplement 1
Time traces of Ztip coordinates depicting pathways to membrane insertion.

(A–C) Results for F-, W-, and Y-centered sequences in two simulations. In B and C, the peptides initially moved away from membranes and then came back; these initial portions of the simulations are skipped. In addition, in (B) the peptide stayed in the intermediate state for over 200 ns; the initial period in the intermediate state is skipped. Time traces are smoothed using the moving average in a 20.1-ns window. Residues are labeled as X[i], where X is the one-letter code for an amino acid and i is the position relative to the central aromatic residue. In all panels, a red line is drawn at Ztip = –3.1 Å.

Figure 4 with 1 supplement
Classification of aromatic-centered motifs into membrane inserters and non-inserters by PPM runs.

(A) Disordered conformations of 9-residue motifs generated by TraDES. (B) The pose of an F-centered peptide (sequence 1) from a PPM run. The phosphate plane is represented by arrays of small salmon spheres. (C) A scatter plot of ΔG and Ztip data for a batch of 100 conformations of sequence 1. Dashed lines represent cutoffs for defining inserters (ΔG≤–4.1 kcal/mol and Ztip <–1.5 Å). 15 conformations satisfy the cutoffs. (D) Extremely strong correlation between the percentage of MD simulations classified as membrane-inserting and the percentage of inserting conformations in PPM runs (correlation coefficient = 0.997). The line represents the fitted regression line.

Figure 4—figure supplement 1
Clear separation of membrane inserters and non-inserters by PPM.

For each of the 10 sequences in our initial dataset, the number of conformations in a batch of 100 that satisfy the cutoffs for both aromatic Ztip and ΔG, when averaged over 10 batches, is displayed as a blue circle. Error bars represent standard deviations across the 10 batches. The gray block illustrates the separation between inserters and non-inserters.

Figure 5 with 2 supplements
Flanking residues as determinants of membrane insertion.

(A) Illustration of the library of 3×58 = 1.2×106 sequences. (B) Percentages of membrane inserters and prediction accuracy of AroMIP. (C) L and R are enriched in all flanking positions of membrane inserters. (D) N, E, and G are enriched in all flanking positions of non-inserters. In C and D, an orange dashed line is drawn at a frequency of 0.2 expected of random distributions. (E) q parameters for F-, W-, and Y-centered sequences.

Figure 5—figure supplement 1
Distributions of amino acids in flanking positions of aromatic-centered membrane inserters and non-inserters.

(A, B) F-centered inserters and non-inserters. (C, D) W-centered inserters and non-inserters. (E, F) Y-centered inserters and non-inserters. An orange dashed line is drawn at a frequency of 0.2 expected of random distributions.

Figure 5—figure supplement 2
Convergence of q parameters at increasing training data size.

(A–C) F-, W-, and Y-centered sequences.

Figure 6 with 2 supplements
Implementation of AroMIP on 2.4×104 aromatic-centered motifs in IDRs of the human proteome.

(A) Illustration of 2.4×104 aromatic-centered motifs in IDRs of the human proteome. This dataset was split 80:20 for training and testing, respectively. (B) q parameters for F-, W-, and Y-centered motifs. (C) Performance of AroMIP on the training and test sets.

Figure 6—figure supplement 1
Correlations of AroMIP insertion scores and membrane-binding free energies for aromatic-centered motifs in the human proteome.

(A–C) Correlations of insertion scores and binding free energies calculated using the octanol scale, for F-, W-, and Y-centered motifs, respectively. Each point represents a unique 9-residue motif; the number of points is 10,992, 5700, and 7616 in A–C, respectively. Red lines represent the fitted regression lines. (D) Correlation coefficients for two scales: the octanol scale and the membrane interfacial scale.

Figure 6—figure supplement 2
Correlations of AroMIP-predicted membrane-insertion propensities and MD and PPM results.

(A) Correlation of membrane-insertion propensity and percentage of membrane-inserting MD simulations (correlation coefficient = 0.918). (B) Corresponding correlation percentage of membrane-inserting conformations in PPM (correlation coefficient = 0.889). Lines represent the fitted regression lines.

Tables

Table 1
MD, PPM, and AroMIP results on the initial dataset.
Seq. #MD
% inserting frames
PPM
% inserting conformations
AroMIP
QPInserter?
120.614.82.3190.699Y
256.647.63.5910.782Y
34.50.50.0110.011N
40.01.70.0380.037N
51.20.10.0040.004N
60.00.00.0120.012N
70.00.10.0060.006N
80.11.00.0100.010N
90.00.70.0020.002N
100.00.10.0020.002N
Table 2
Top scoring motifs in IDRs of the human proteome.
Uniprot IDProtein nameSequence# of aromatic# of aliphatic*# of basicMembrane-insertion propensity
F-centered
K7EN89MidnolinRFILFKRPW3331.000
Q70EL4UBP43LFSRFLLAL2511.000
H3BN95INTS14FPLPFPFPS3501.000
Q70EL4UBP43LRRLFSRFL2330.999
Q96JE7Sec16BGFGWFSWFR5010.999
K7EQI6ARHGAP33LLPFFPHMP2600.999
E9PJ49CCKBRLMPVFLIPR1710.999
H0Y542BCAS1LGLAFRKFF3320.999
Q9HCE9Anoctamin-8AFLSFKFLK3320.999
O15063GARRE1RTWPFPEFF4210.999
W-centered
Q9UHK0NUFIP1SWMFWAMLP3501.000
Q96JE7Sec16BSGFGWFSWF5001.000
H3BUN7SH2B1PWLSWSPWL3401.000
O75807GADD34FLKAWVYWP4411.000
H0YFY4HMGA2RPRKWWLLM2430.999
KAI2556058COL13A1SWASWFTWT4100.999
Q9BYE0hHes7PPAFWRPWP3510.999
H7C1I9MAST4LLEPWFLPP2600.999
Q9Y4K4MEKKK 5FMLQWNPFV3400.999
H0YFY4HMGA2PRKWWLLMK2430.999
Y-centered
A0PJX2TLDC2LRWRYTRLP2330.778
A0A7P0T883C2CD3SLLLYPLAF2600.765
P48634PRRC2APPFMYPPYL3600.762
P48634PRRC2AMYPPYLPFP3600.762
A0A1B0GV45Myosin XVIIIALMMRYLYRP2520.760
  1. *

    Aliphatic residues are: L, I, M, V, P, and A.

  2. GenBank ID.

Table 3
AroMIP predictions on 12 IDPs or IDRs.
IDP namePrevious studySequence*,AroMemIn*,
MARCKS
(Uniprot P29966; residues 152–176)
Solid-state NMR (Zhang et al., 2003)KKKKKRFSFKKSFKLSGFSFKKNKKF158 0.953
F160 0.992
F164 0.956
F169 0.956
F171 0.742
Tat
(Uniprot P04610; residues 1–86)
Solution NMR (Ghanam et al., 2023)MEPVDPRLEPWKHPGSQPKTACTTCYCKKCC
FHCQVCFTTKALGISYGRKKRRQRRRPPQGSQTHQVSLSKQPTSQPRGDPTGPKE
W11 0.727
Y26 0.000
F32 0.288
F38 0.509
Y47 0.050
Translocated intimin receptor
(Uniprot B7UM99; residues 388–550)
Solution NMR (Vieira et al., 2024)RRNQPAEQTTTTTTHTVVQQQTGGNTPAQGGTDAT
RAEDASLNRRDSQGSVASTHWSDSSSEVVNPYAEVGGARNSL
SAHQPEEHIYDEVAADPGYSVIQNFSGSGPVTGRLIGTPGQGIQSTYALL
ANSGGLRLGMGGLTSGGESAVSSVNAAPTPGPVRFV
W443 0.179
Y454 0.024
Y474 0.002
Y483 0.014
F489 0.347
Y511 0.140
mGluR3
(Uniprot Q14832; residues 839–879)
Solution NMR; all-atom MD simulations (Mancinelli et al., 2024)LHLNRFSVSGTGTTYSQSSASTYVPTVCNGREVLDSTTSSLF844 0.716
Y853 0.002
Y861 0.036
Bap1
(Uniprot A0A7Z7YFH0; residues 415–471)
All-atom MD simulations; tryptophan fluorescence (Huang et al., 2026)YLGLEWKTKTVPYLGVEWRTKTVSYWFFGW
HTKQVAYLAPVWKEKTIPYAVPVTLSK
W420 0.838
Y427 0.061
W432 0.787
Y439 0.289
W440 1.000
F441 0.998
F442 0.997
W444 0.998
Y451 0.139
W456 0.863
Y463 0.263
KtrB
(Uniprot O87953; residues 1–24)
All-atom MD simulations (Stautz et al., 2025)MTQFHQRGVFYVPDGKRDKAKGGEF10 0.746
Y11 0.065
ABHD5
(Uniprot Q9DBL9; residues 1–33)
All-atom Gaussian accelerated MD simulations; hydrogen-deuterium exchange (Kumar et al., 2025)MKAMAAEEEVDSADAGGGSGWLTGWLPTWCPTSW21 0.858
W25 1.000
W29 0.989
Prolactin Receptor
(Uniprot P16471; residues 260–280)
Coarse-grained MD simulations; solution NMR (Araya-Secchi et al., 2023)GYSMVTCIFPPVPGPKIKGFDF268 0.985
OPA1
(Uniprot O60313; residues 768–782)
Cryo-EM (Nyenhuis et al., 2023; von der Malsburg et al., 2023)GPDWKKRWLYWKNRTQW775 0.999
Y777 0.398
W778 0.990
Gasdermin D
(Uniprot P57764; residues 41–53)
Cryo-EM; (Xia et al., 2021) all-atom MD simulations (Schaefer and Hummer, 2022)VRKPSSSWFWKPRYKW48 0.996
F49 0.989
W50 0.998
Gasdermin B
(Uniprot Q8TAX9-1; residues 41–51)
Cryo-EM (Wang et al., 2023)GEKRTFFGCRHF46 0.854
F47 0.920
PLAT1
(Uniprot O65660; residues 37–49)
All-atom MD simulations (Kulke et al., 2024)TGSIWKAGTDSIIW41 0.551
  1. *

    Inserters are indicated by bold letters, whereas non-inserters are indicated by an underline.

  2. The value after the residue number is the membrane-insertion propensity.

Table 4
Details of MD simulations in membranes.
Seq. #Box dimensions (Å3)# of atoms
182 x 82 x 10364,818
281 x 81 x 10364,647
381 x 81 x 10666,106
482 x 82 x 10163,196
582 x 82 x 10464,865
682 x 82 x 10465,384
782 x 82 x 10365,315
881 x 81 x 10564,809
982 x 82 x 10364,150
1082 x 82 x 10465,265

Additional files

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Fidha Nazreen Kunnath Muhammedkutty
  2. Huan-Xiang Zhou
(2026)
A membrane insertion code for intrinsically disordered proteins
eLife 15:RP111515.
https://doi.org/10.7554/eLife.111515.3