Figures and data

Visual description of the dilution and plate counting process and the need for its analysis.
a) A sample of volume V is mixed with sterile medium in a larger vial (e.g., 19V, for a total dilution of 20). A portion of volume V of this diluted sample is then spread on a plate and incubated. Further tenfold dilutions are performed by transferring V into a vial with 9V of medium before plating. The process continues until a reasonable number of counts is obtained. Panel created using BioRender (https://BioRender.com/espzw5y). b) The histogram on the left shows simulated counts across 1000 plates, obtained by diluting samples from the ground truth distribution by a factor of 200. The usual method of multiplying the counts by the total dilution factor produces the orange histogram (in the right), which is significantly broader than the ground truth (black line, right) due to additional stochasticity introduced by dilution and plating. The ground truth distribution is a Gaussian with a mean of 8000 and a standard deviation of 500. REPOP’s reconstruction (blue) corrects the broadening. c) A similar simulation to (b) where the ground truth population distribution is trimodal (parameters in Fig. 3). Here, the stochasticity introduced by dilution and plating obscures the three peaks, making them difficult to distinguish in the observed counts. Nevertheless, REPOP successfully reconstructs the trimodal distribution.

REPOP workflow from plate counts to reconstructed population statistics.
The input consists of observed colony counts and their corresponding dilution factors, organized as pairs (ki, ϕi), each arising from a single biological sample. From these data, REPOP evaluates the probability of each observed dilution count pair conditioned on the underlying bacterial number, p(ki, ϕi ∣ ni), thereby capturing the stochasticity introduced by dilution and plating. These likelihood contributions are then combined to reconstruct the distribution of the original bacterial population across samples. Icons in the top-left are from BioRender (https://BioRender.com/espzw5y).

For ease of reference, we summarize the differences among the three models taken into account in the present article.

REPOP can reconstruct multimodal populations even when they are not directly observable by naively multiplying colony counts by dilution.
Similar to Fig. 1b and c, we show above the observed colony counts, the naïve estimate obtained by multiplying counts by the dilution factor, and the reconstruction produced by REPOP, compared to the ground truth. In both cases, the population size n is drawn from a mixture of three Gaussians, shown as the black curves. In a), the mixture has means (4000, 8000, 14000), standard deviations (200, 1500, 1000), and weights (0.25, 0.4, 0.35). In b), the mixture has means (8000, 16000, 24000), standard deviations (1000, 1000, 1000), and weights (0.3, 0.2, 0.5). Observed counts are sampled from (1), with dilution factor 200, and are shown as histograms on the left. In (a), a trimodal structure is visible in the observed count histogram. However, in b), the trimodal structure is strongly obscured by stochasticity from dilution and plating, making it difficult to discern even in large datasets. Despite this, applying REPOP to datasets of increasing size shows that, although small datasets of 25 plates may miss some modes, larger datasets allow REPOP to accurately recover the underlying multimodal structure in both cases. We quantify this behavior further in Appendix 2, where we analyze the error between the reconstructed and ground-truth means, as well as the relative entropy between the reconstructed and ground-truth distributions, as a function of the number of plates. Together, these results demonstrate REPOP’s ability to infer the true population despite stochastic noise introduced by dilution and plating.

Not accounting for the cutoff can lead to incorrectly attributed multimodality.
This figure compares population reconstructions with and without accounting for the cutoff effect, (using Model 3 and Model 2 respectively). a) simulated data generated with a cutoff kCO = 50. The ground truth population consists of three Gaussian components with means 





Reconstructing, from real experimental plate counting, the population from vials with different optical densities.
a) Distribution of observed colony counts obtained as described in text. The data is passed to the system without vial identification, aiming to reconstruct the underlying structure. b) Results from a simulated dataset with 750 datapoints (plates). The increased sample size improves the estimation of the four main components, (which appear as approximately three peaks because two components substantially overlap), yielding a more accurate reconstruction.

Results for the population of E.coli-fed C. elegans’ gut at days 3, 5, 7, and 9 of adulthood.
Here we are able to observe the shift in the population distribution obtained with plate counting shifts to higher population for longer-living C. elegans. To better visualize the distribution and model fits, we use a split-log x-axis: bacterial counts from 0 to 100 are shown on a linear scale, and counts above 100 on a logarithmic scale. The population distribution shifts toward higher values each day, as expected from colonization and replication within the gut. If biological significance, such as different C. elegans types, as proposed in e.g. (Boddu et al., 2024) is to be inferred from this multimodality, a rigorous separation of population heterogeneity and plating stochasticity, as implemented by REPOP, is essential.

REPOP reconstructs marginal population distributions when individual plates distinguish two species.
A synthetic dataset of 1500 samples contains two correlated bacterial populations. Black points denote the true underlying populations (nA, nB ), which are drawn from a joint distribution with three well-separated modes, as described in the text. Orange points denote the corresponding plate counts after a dilution factor of 200, naively reconstructed as counts multiplied by the dilution factor. Although the true population (black) modes are well separated, stochasticity introduced by dilution and plating causes the observed measurements(orange) to overlap. We assume that the two species can be distinguished on each plate, so that separate counts are obtained for species A and B. Applying REPOP separately to the counts from each species recovers the corresponding marginal population distributions, as shown in the side panels.

Decay of error metrics as the dataset size increases for the distributions shown in Fig. 3.
We show the relative error defined in (23) (left) and the KL divergence defined in (24) (right) for the distributions in Fig. 3a and Fig. 3b. Both metrics decrease as the number of plates increases, indicating improved agreement between the REPOP reconstruction and the ground-truth.