Cluster size determines internal structure of transcription factories in human cells

  1. Massimiliano Semeraro
  2. Giuseppe Negro  Is a corresponding author
  3. Giada Forte  Is a corresponding author
  4. Antonio Suma
  5. Giuseppe Gonnella
  6. Peter Cook
  7. Davide Marenduzzo
  1. Dipartimento Interateneo di Fisica, Università degli Studi di Bari and INFN, Sezione di Bari, Italy
  2. SUPA School of Physics and Astronomy, University of Edinburgh, United Kingdom
  3. Sir William Dunn School of Pathology, University of Oxford, United Kingdom
10 figures and 1 additional file

Figures

Figure 1 with 2 supplements
Toy model, with transcription units (TUs) coloured randomly (the random string).

(A) Overview. (i) Yellow, red, and green TFs (25 of each colour) bind strongly (when in an on state) to 100 TUs beads of the same colour in a string of 300 beads (representing 3 Mb), and weakly to blue beads. TU beads are positioned regularly and coloured randomly. TFs switch between off and on states at rates αoff=105τB1 and αon=αoff/4 (τB Brownian time, see SI). (ii) The sequence of bars reflects the random sequence of yellow, red, and green TUs (blue beads not shown). (B) Snapshot of a typical conformation obtained after a simulation (transcription factors, TFs not shown). Inset: enlargement of boxed area. TU beads of the same colour tend to cluster and organise blue beads into loops. (C) Bridging-induced phase separation drives clustering and looping. Local concentrations of red, yellow, and green TUs and TFs might appear early during the simulation. Red TF 1 – which is multivalent – has bound to two red TUs and so forms a molecular bridge that stabilises a loop; when it dissociates it is likely to re-bind to one of the nearby red TUs. As red TU 2 diffuses through the local concentration, it is also likely to be caught. Consequently, positive feedback drives growth of the red cluster (until limited by molecular crowding). Similarly, the yellow and green clusters grow as yellow TF 3 and green TF 4 are captured. (D) Bar heights give transcriptional activities of each TU in the string (average of 100 runs each lasting 8105τB). A TU bead is considered to be active whilst within 2.24σ6.7×108m of a TF:pol complex of similar colour. Dashed boxes: regions giving the three clusters in the inset in (B). (E) Pearson correlation matrix for the activity of all TUs in the string.

Figure 1—figure supplement 1
Varying transcriptional activity threshold does not alter transcriptional profiles.

(A and B) Bar heights give transcriptional activities of each transcription unit (TU) in the toy string (average of 100 runs each lasting 8×105τB). A TU bead is considered to be active, respectively, whilst within 1.8σ5.4×108m and 1.1σ3.3×108m of a TF:pol complex of similar colour. (C and D) Difference between the transcriptional profile obtained with threshold 2.25σ (see Figure 1D) and those obtained with threshold 1.8σ and, 1.1σ, respectively.

Figure 1—figure supplement 2
Random string with just one colour.

Transcription units (TUs) in the random string considered in the main text are coloured randomly yellow, red, or green; here, instead, every TU has the same colour. Averages are obtained from 100 runs, each lasting 8×105τB, (A) The transcriptional activity profile is much flatter than that obtained with the 3-colour version. (B) The Pearson correlation matrix also lacks red blocks along the diagonal.

Simulating effects of mutations.

Yellow transcription unit (TU) beads 1920, 1950, 1980, 2010, 2040, and 2070 in the random string have the highest transcriptional activity. One to four of these beads are now mutated by recolouring them red. (A) The sequence of bars reflects the sequence of yellow, red, and green TUs in random strings with 1, 2, and 4 mutations (blue beads not shown). Black boxes highlight mutant locations. (B) Typical snapshots of conformations with (i) one, and (ii) four mutations. (C) Transcriptional-activity profiles of mutants (averages over 100 runs, each lasting 8×105τB). Bars are coloured according to TU colour. Black boxes: activities of mutated TUs. (D) Activities (+/- SDs) of wild-type (yellow) and different mutants. Three mutations: TUs 1950, 1980, and 2010 mutated from yellow to red. (E) Typical kymographs for (i) wild-type, corresponding to the same original string presented in Figure 1, and (ii) 4-mutant cases, in which four yellow TUs have been mutated to red. Each row reports the transcriptional state of a TU during one simulation. Black pixels denote inactivity, and others activity; pixels colour reflects TU colour. Blue boxes: region containing mutations. (F) Pearson correlation matrices for wild-type and 4-mutant cases. Black boxes: regions containing mutations (mutations also change patterns far from boxes).

Reducing the concentration of yellow transcription factors (TFs) reduces the transcriptional activity of most yellow transcription units (TUs) while enhancing the activities of some red TUs.

(A) Overview. Simulations are run using the random string with the concentration of yellow TFs reduced by 30%, and activities determined (means from 100 runs each lasting 8×105τB). (B) Activity profile. Dashed boxes: activities fall in the region containing the biggest cluster of yellow TUs seen with 100% TFs, as those of an adjacent red cluster increase. (C) Differences in activity induced by reducing the concentration of yellow TFs. This plot is obtained by subtracting the transcriptional activity of the wild-type, Figure 1D, from that of the current system in panel B. (D) Pearson correlation difference matrix. This plot is obtained by subtracting the Pearson correlation matrix of the wild-type, Figure 1E, from that of the current system. Boxes: regions giving the 3three clusters from Figure 1B, inset.

Clustering similar transcription units (TUs) in 1D genomic space increases transcriptional activity.

(A) Simulations involve toy strings with patterns (dashed boxes) repeated one or six times. Activity profiles plus Pearson correlation matrices are determined (100 runs, each lasting 8×105τB). (B) The 6-pattern yields a higher mean transcriptional activity (arrow highlights difference between the two means). (C) The 6-pattern yields higher positive correlations between TUs within each pattern and higher negative correlations between each repeat.

Figure 5 with 2 supplements
Transcription unit (TU) transcriptional networks and demixing.

Simulations are run using the toy models indicated, and complete correlation networks (qualitatively reminiscent of gene regulatory networks) constructed from Pearson correlation matrices. (A) Simplified network given by the random string. TUs from first (bead 30) to last (bead 3000) are shown as peripheral nodes (coloured according to TU); black and dashed grey edges denote statistically-significant positive and negative correlations, respectively (above a threshold of 0.2, corresponding to a p-value 5×102). The complete network consists of n=100 individual TUs, so that there are nc=(1002)=4950 pairs of TUs couples; we find 990 black and 742 gray edges. Since p-value.nc=223, most interactions (edges) are statistically significant. Networks shown here only correlations (i) between red TUs, and (ii) between red and green TUs. (ii) (B) Average Pearson correlation (shading shows +/- SD, and is usually less than line/spot thickness) as a function of genomic separation for the (i) random, (ii) 6-, and (iii) 1-pattern cases. Correlation values at fixed genomic distance are taken from super-/sub-diagonals of Pearson matrices. Red dots give mean correlation between TUs of the same colour (three possible combinations), and blue dots those between TUs of different colours (4four possible combinations). Cartoons depict contents of typical clusters to give a pictorial representation of mixing degree (as this determines correlation patterns); see SI for exact values of θdem.

Figure 5—figure supplement 1
Complete regulatory networks, and demixing coefficients for toy strings.

(A) Networks involving positive (black) and negative (dashed grey) correlations between all coloured beads are presented, for the models considered in the main text in Figure 5. (i) The random string. The network is complex and entails both positive and negative correlations. (ii) The 6-pattern repeat. There are many black edges within repeats and virtually none between repeats, many grey edges between repeats, and few edges between transcription units (TUs) with different colours. (iii) The 1-pattern repeat. There are few long black edges. (B) Average of θdem (error bars: weighted SD) for various toy strings. Red lines: values for random string. Compared to the random case, 6-pattern is more demixed, and the 1-pattern less demixed. As discussed in the main text, same-colour and different-colour correlations diverge from each other in the 6-pattern case, and roughly overlap in the 1-pattern.

Figure 5—figure supplement 2
Activity correlations found in toy chromosomes without weakly-binding beads.

Correlation values (shading shows +/- SD, and is usually less than line/spot thickness) at fixed genomic distance are taken from super-/sub-diagonals of Pearson matrices. Red dots give mean correlations between transcription units (TUs) of same colour, and blue dots those between TUs of different colours. (i) Random string, (ii) 6-pattern string, and (iii) 1-pattern string with non-binding beads instead of weakly-binding ones.

Figure 6 with 2 supplements
Comparison of transcriptional activities of transcription units (TUs) on different human chromosomes determined from simulations and GRO-seq.

(A) Overview of panels (A–C). The 35784 beads on a string representing HSA14 in HUVECs are of four types: TUs active only in HUVECs (red), ‘house-keeping’ TUs – ones active in both HUVECs and ESCs (green), ‘euchromatic’ ones (blue), and ‘heterochromatic’ ones (gray). Red and green transcription factors (TFs) bind to TUs of the same colour with strong (specific) interactions, while they experience a weak (non-specific) attraction to euchromatin. Interactions between both red and green TFs and heterochromatin are purely repulsive. (B) Snapshot of a typical conformation, showing both specialised and mixed clusters. (C) TU activities seen in simulations and GRO-seq are ranked from high to low, binned into quintiles, and activities compared. (D) Spearman’s rank correlation coefficients for the comparison between activity data obtained from analogous simulations and GRO-seq for the chromosomes and cell types indicated.

Figure 6—figure supplement 1
Spearman’s rank-correlation coefficients determined by comparing activity data obtained from simulations (one- and two-colour) and GRO-seq for chromosomes and cell types indicated.

Grey bars: one-colour cases. Coloured bars: two-colour cases where values are for ‘all’ TUs, just housekeeping (‘hk’), and HUVEC- or GM12878-specific ones. (i) HSA 14 in HUVEC. (ii) HSA 18 in HUVEC. (iii) HSA 19 in HUVEC. (iv) HSA 14 in GM12878.

Figure 6—figure supplement 2
Comparing experimental Hi-C and numerical contact maps for HUVEC HSA 14.

After running 100 simulations for HUVEC HSA 14, contact maps for two chromosome fragments are generated and compared to an experimental Hi-C map provided by The ENCODE Project Consortium, 2012. (i) fragment 67.5−70.5 MB (beads 22,500−23,500). (ii) fragment 100.8−106.8 MB (beads 33,600−35,500). In both cases the binning is 30 kbp, each pixel is coloured according to the number of contacts and fair data agreement is found (Pearson coefficient r0.7, p-value <106 two-sided t-test). The diagonal is made white for graphical clarity.

Figure 7 with 6 supplements
Small clusters tend to be unmixed, large ones mixed.

After running one simulation for HSA 14 in HUVECs, clusters are identified. (A) Snapshot of a typical final conformation (transcription units, TUs, non-binding beads, and transcription factors (TFs) in off state not shown). Insets: a large mixed cluster and a small demixed one. (B) Example clusters with different numbers of TFs/cluster (2, 10, 20, 30, 40) chosen to represent the range seen from all-red to all-green (with 3three intervening bins). Black numbers: observed number of clusters of that type seen in the simulation. (C) Average of the demixing coefficient θdem (error bars: SD), showing a crossover between demixed and mixed clusters with increasing cluster size. Values of 1 and 0 are completely demixed and completely mixed, respectively. Grey area: demixed regime where θdem is >0.5.

Figure 7—figure supplement 1
Small clusters tend to be less mixed than large ones in simulations of various human chromosomes in two cell types.

After running 100 simulations for each chromosome in each cell type, all clusters are identified, numbers of housekeeping and cell-type-specific transcription units (TUs) in them counted, and demixing coefficients determined. Each dot in a plot is positioned according to TF content and value of θ~dem=(2xred1) with xred the fraction of red beads in a cluster (so that θ~dem=+1 θ~dem=1 correspond to fully demixed red and green clusters, respectively). Each dot is coloured according to its frequency of occurrence (right-hand bars). (i) HSA 14 in HUVEC. (ii) HSA 18 in HUVEC. (iii) HSA 19 in HUVEC. (iv) HSA 14 in GM12878. All panels have similar patterns; each indicates there are high numbers of demixed small clusters (top and bottom left of each panel), and lower numbers of mixed larger clusters (at middle right in each panel).

Figure 7—figure supplement 2
Cluster analysis for HSA 14 in GM12878.

After running 100 simulations, clusters are identified. (A) Snapshot of a typical final conformation (transcription units, TUs, non-binding beads, and transcription factors TFs, in off state not shown). Insets: a large mixed cluster and a small demixed one. (B) Cartoons showing representative contents of clusters. We consider five types of clusters (ranging from all green to all red, y axis), and 5 values for the total numbers of TFs/cluster (2, 10, 20, 30, 40 x axis). Cartoons: cluster composition. Black numbers: observed number of clusters in the simulation. (C) Average of the demixing coefficient θdem as a function of cluster size (error bars: SD). Gray area: demixed regime (corresponding to specialised factories) where θdem is >0.5. White area: mixed regime (corresponding to mixed factories associated with HOTs) where θdem is <0.5.

Figure 7—figure supplement 3
Cluster analysis for confined HSA 14 in HUVEC.

After running 100 simulations, clusters are identified for HSA 14 in HUVEC. The system is confined into an ellipsoidal territory with aspect ratio chosen according to typical experimental values (Wang et al., 2017) (semiaxes were 22.24σ:34.24σ:41.80σ, or 0.67μm:1.03μm:1.25μm). Confinement is enforced by modifying the source code in LAMMPS to describe an ellipsoidal indenter. This introduces a soft force towards the centre of the ellipsoid, only if beads exit the confining ellipsoid. (A) Snapshot of a typical confined conformation. Insets: a large mixed cluster and a small demixed one (non-binding beads and transcription factors, TFs, in off state not shown). (B) All clusters are identified, numbers of housekeeping and cell-type-specific transcription units (TUs) in them counted, and demixing coefficients determined. Each dot in the plot is positioned according to TF content and value of θ~dem and coloured according to its frequency of occurrence (right-hand bar). (C) Average of the demixing coefficient θdem as a function of cluster size (error bars: SD). Gray area: demixed regime where θdem>0.5. White area: mixed regime where θdem<0.5.

Figure 7—figure supplement 4
Cluster analysis for HSA 14 in HUVEC with smaller transcription factors (TFs).

Clusters are identified for HSA 14 in HUVEC after running 80 simulations with TFs size 0.5σ (A) and 100 simulations with TFs size 0.16σ (B). Panels report average of the demixing coefficient θdem as a function of cluster size (error bars: SD). Grey area: demixed regime (corresponding to specialised factories) where θdem>0.5. White area: mixed regime (corresponding to mixed factories associated with HOTs) where θdem<0.5.

Figure 7—figure supplement 5
Effect of genomic separation on activity correlations found in whole human chromosomes.

Correlation values (shading shows +/- SD, and is usually less than line/spot thickness) at fixed genomic distance are taken from super-/sub-diagonals of Pearson matrices. Red dots give mean correlations between transcription units (TUs) of same colour, and blue dots those between TUs of different colours. (i) HSA 14 in HUVEC. (ii) HSA 14 in GM 12878.

Figure 7—figure supplement 6
Left: Kymograph showing, at each time step, the transcriptionally active transcription units (TUs), colour-coded according to their assigned type.

This kymograph is identical to that reported in Figure 2E (i) of the main text. Right: Corresponding kymograph displaying the number of green transcription factors (TFs) within a sphere of radius 10σ centred on each green TU.

Highly predictive hetero-morphic polymer (HiP-HoP) model simulations: small clusters tend to be unmixed, large ones mixed.

(A) Snapshot of a configuration adopted by HSA14 in HUVECs, within the HiP-HoP model. Grey regions represent less accessible chromatin regions poor in H3K27ac, while cyan regions represent those enriched in H3K27ac. In addition, H3K27me3 and H3K9me3 peaks determine the chromatin binding sites for polycomb-like and heterochromatin proteins, and are represented in yellow and blue, respectively. As in the previous DNase hypersensitive site (DHS) multicolour model, transcription units (TUs) only present in HUVEC are represented in red, while the house-keeping ones in green. (B–C) Example of clusters of proteins: large mixed cluster (B) and a small demixed one (C). (D) Average of the demixing coefficient θdem (error bars: SD). Values of 1 and 0 correspond to completely demixed and completely mixed clusters, respectively. Grey area: demixed regime where θdem is >0.5.

Author response image 1
Heatmap comparing activity correlations of TUs in the random string under normal conditions (top half) and with reduced yellow-TF concentration (bottom half).
Author response image 2
Kymograph showing the TU activity during a typical run in the 6-pattern case.

Each row reports the transcriptional state of a TU during one simulation. Black pixels correspond to inactive TUs, red (yellow, green) pixels correspond to active red (yellow, green) TUs.

Additional files

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Massimiliano Semeraro
  2. Giuseppe Negro
  3. Giada Forte
  4. Antonio Suma
  5. Giuseppe Gonnella
  6. Peter Cook
  7. Davide Marenduzzo
(2026)
Cluster size determines internal structure of transcription factories in human cells
eLife 14:RP103955.
https://doi.org/10.7554/eLife.103955.3