Pathogen-phage geomapping to overcome resistance

  1. Camilla Do  Is a corresponding author
  2. Keiko Christine Salazar
  3. James D Chang
  4. Justin R Clark
  5. Austen Lee Terwilliger
  6. Paul Ruchhoeft
  7. Paul Nicholls
  8. Anthony W Maresso
  1. Department of Molecular Virology and Microbiology, Baylor College of Medicine, United States
  2. TAILΦR LABS, Baylor College of Medicine, United States
  3. Department of Electrical & Computer Engineering, University of Houston, United States
  4. Section of Infection Diseases, Department of Medicine, Baylor College of Medicine, United States

eLife Assessment

This important study establishes an environmental sampling workflow for the discovery of bacteriophages capable of infecting antibiotic-resistant pathogens. The authors convincingly demonstrate the effectiveness of the approach, even with the limited sampling scheme and the current challenges in viral taxonomy. This study will interest researchers working on bacterial infections, environmental microbiology, and phage-based alternatives for addressing antimicrobial resistance.

https://doi.org/10.7554/eLife.109259.3.sa0

Abstract

The rise of antibiotic resistance has renewed interest in bacteriophages as therapeutic alternatives. However, coevolution of phage and bacteria will naturally give rise to phage-resistant pathogens, complicating phage therapy efforts. A critical bottleneck in the production of phage therapeutics is the discovery of virulent phages against resistant pathogens. Conventional methods for discovery are time-consuming, biased, and laborious, limiting the potential for identifying suitable phage candidates. To overcome these limitations, we combined small-volume environmental sampling with 16 S rRNA sequencing to identify reservoirs where bacterial hosts co-exist with their phage predators. This strategy, which we term geographical phage mapping (geΦmapping), pinpoints ecological ‘hotspots’ for targeted phage hunting. We further developed a portable phage hunting device (ΦHD) that generates highly enriched phage concentrates directly from these reservoirs. By integrating geΦmapping with high-throughput enrichment, we constructed the RΦ library, a diverse collection of novel phages. We captured and isolated 36 new phages targeting extremely resistant organisms across various ESKAPE pathogens when conventional phage hunting and experimental evolution approaches failed.

Introduction

The rise of antibiotic resistance has brought global healthcare to the precipice of a post-antibiotic era, driving interest in bacteriophage (phage) therapy as an alternative strategy to combat resistant bacterial infections (Naghavi et al., 2024; Pirnay et al., 2024; Green et al., 2023). Unlike antibiotics, phages evolve with their host; both locked in a relentless evolutionary arms race. Bacteria and phage continuously expand their repertoires of resistance and counter-resistance strategies. Among many examples, phages counter bacterial defenses by exposing receptors with enzymes, evading restriction enzymes with DNA base modification, and evolving anti-CRISPR proteins (Hutinet et al., 2019; Katz et al., 2024). In response, bacteria resist phage infection by blocking adsorption and nucleic acid injection and using restriction enzymes and CRISPR-Cas9 to target phage nucleic acids (Garneau et al., 2010; He et al., 2024). This dynamic interplay of phage resistance poses a barrier to effective phage therapy (Oechslin, 2018). TAILΦR is a phage center dedicated to providing lytic phages to physicians to treat patients battling multidrug-resistant infections (Terwilliger et al., 2020). Across 263 distinct cases encompassing 462 bacterial isolates, 22% of isolates given to TAILOR remain untreated despite an extensive library of over 325 phages targeting 15 bacterial species (Figure 1A). Information regarding TAILΦR’s phage and bacterial library can be found in Supplementary files 1–3. When no phages in the library can target a pathogen, TAILΦR turns to directed evolution (to expand the host range of a given phage) or environmental sampling (to find natural phages). Such approaches often fail to yield a suitable phage, and so discovering new, virulent phages that can effectively target these pathogens becomes a critical and rate-limiting step (Figure 1B). As phage therapy spreads, resistant isolates will inevitably emerge under increased selection pressure, analogous to antibiotic resistance after the introduction of penicillin (Rammelkamp and Maxon, 1942). This puts the onus onto phage discovery where improvements are critically needed (Olsen et al., 2020b).

Overview and schematic of ΦHD.

(A) Pie charts of patient isolate status and percentage of bacterial strain with no phages for in TAILΦR’s library. Unfortunately, 24% of isolates that TAILΦR receives have no phages; at 35%, most of the isolates are Pseudomonas aeruginosa. (B) Diagram from phage request to phage discovery and limitation. Clinicians and their patients can seek phage therapy for antibiotic-resistant infections when there are no other approved treatments available. After approval, clinicians send patient isolates to TAILΦR. TAILΦR’s phage library and wastewater concentrates are screened against the patient isolate. If phages are discovered, they are moved onto the next stage for preparation where phages are ultimately purified and tested for safety and efficacy before being handed to the clinician. When there are no phages for that isolate in the library, TAILΦR searches for phages through environmental sampling. Alternatively, they can train related phages to infect the isolate through directed phage evolution. (C) Diagram of strategies employed in our study. Geographical phage (Φ) mapping (geΦmapping) allowed us to pinpoint target sites rich in our pathogen of interest from our environment (left). Once a location is chosen, we used our capture device, the phage hunting device (ΦHD), to filter, concentrate, and enrich our sample with phage and bacteria (middle). Bar graph shown here was derived from Figure 2F as an example for phage concentration. With the samples, we further analyzed them with phage metagenomics (ΦMICS) and screened, isolated, and characterized phage present in the samples (right). This led to the curation of the resistant phage (Φ) library (RΦ-Library). (D) Schematic of ΦHD. The components on the right of the reservoir draw, filter, and concentrate water through ΦHD. On the left of the reservoir, the remaining component concentrates any material in the reservoir. Created with BioRender.com.

Phage engineering is one method of constructing phages for improved therapeutic potential. Examples include generating lytic derivatives of lysogenic phages and altering phage tail fibers to adjust their host range (Dedrick et al., 2019; Yehl et al., 2019). In contrast, phage-host competition has naturally addressed phage resistance over billions of years of evolution. The challenge lies in isolating these phages from nature. Conventional methods for phage discovery can be time-consuming or introduce bias (Olsen et al., 2020b). Phage hunters are limited to sampling small volumes of water, a constraint that typically necessitates the use of enrichment techniques to increase the likelihood of phage detection (Sada and Tessema, 2024; Van Twest and Kropinski, 2009; Aghaee et al., 2021; Kenney and Gómez-Duarte, 2024). Enrichments can favor fast-growing or highly virulent phages, potentially overlooking those that may be therapeutically relevant but less abundant or slower to propagate (Muniesa et al., 2005). To sample larger volumes, dead-end filtration (limited by low filtration speed) and iron chloride flocculation (limited by phage inactivation) are typically used (Supplementary file 4). Roux et al., 2016; Gios et al., 2024; Xiong et al., 2023; Chen et al., 2023; Labbé et al., 2020; Malki et al., 2020; Angly et al., 2006; Schoenfeld et al., 2008; Gu Liu et al., 2024; Heffron et al., 2019; Langenfeld et al., 2021 The speed and phage bias limitations are acceptable for metagenomic approaches but unhelpful for phage therapy, leaving a critical gap for patients that urgently need phages for phage-resistant bacterial infections.

Viral metagenomics involves identification of largely uncultivated viral genomes from the environment. Unfortunately, analysis is hindered by the vast diversity of viruses, lack of consensus sequences, and limited reference genomes available (Roux et al., 2019). Some 70% of all viral genomes are unclassified, often referred to as ‘viral dark matter’ (Roux et al., 2015). Studying unknown viral genomes can advance viral ecology and may contribute useful gene products for research and therapeutics (Huss et al., 2025). Viral metagenomics are frequently limited by low biomass, sampling biases, and limitations in sampling remote sites (Thurber et al., 2009; Gowers et al., 2019).

In this study, we sought to improve current methods for phage discovery to tackle phage-resistant pathogens and enhance gene discovery. We used a bioprospecting approach, collecting samples in East Texas, USA, and applied 16 S rRNA analysis to identify pathogenic reservoirs and hence their phages. We dubbed this process, ‘Geographical Phage (Φ) Mapping,’ or ‘GeΦmapping.’ After identifying pathogen-rich sources, we deployed a high-throughput phage hunting device, ΦHD, for sample processing (Figure 1C and D). We used this approach of combining molecular bioprospecting and high-volume sampling to address problematic ‘extremely phage-resistant organisms’ (XΦROs) in clinical cases.

Results

GEΦMAPPING of water sources from Houston and Galveston, Texas

Phages are found where their hosts thrive (Clokie et al., 2011). In marine ecosystems, for example, phage-to-bacterium ratios often reach 10:1, reflecting close associations between bacterial density and phage populations (Chibani-Chennoufi et al., 2004). We identified sites of high pathogen abundance to determine where to hunt for phages. We sampled 15 waterways around Houston and Galveston, Texas, USA thrice and performed 16 S rRNA analysis to examine the bacterial constituents (Figure 2A). We performed rarefaction analysis to estimate species richness relative to sampling effort. The logarithmic shape of the rarefaction curve indicates that most of the abundant taxa present have been detected in each sample, although sample saturation was not fully achieved (Figure 2—figure supplement 1A, top). To further quantify species richness, we compared α-diversity indices (inverse Simpson, Shannon, and Chao) and noticed that White Oak Bayou was the richest (Figure 2—figure supplement 1B). To compare the membership of each site’s microbiome, we analyzed β-diversity using principal coordinate analysis (PCoA) based on Bray–Curtis dissimilarity (Figure 2B and C). Samples of seawater, sewage, and freshwater clustered distinctively, whereas samples from brackish water mingled between freshwater and seawater sites (Figure 2B and C). We performed an analysis of molecular variance (AMOVA) to compare samples derived from brackish, sea, fresh, and sewage. AMOVA revealed significant structuring across all habitat comparisons, with among-group variance exceeding within-group variance in every case. Differentiation was strongest between fresh and sewage samples (Fₛ=28.71, p<0.001), followed by sea-sewage (Fₛ=20.21, p<0.001) and brackish-sewage (Fₛ=15.92, p<0.001), indicating pronounced divergence associated with wastewater. Comparisons involving brackish with seawater (Fₛ=4.65–13.16, p≤0.006) and brackish-fresh (Fₛ=5.96–13.16, p≤0.001) were also significant (Fₛ=4.65–13.16, p≤0.006). Interestingly, a freshwater sample from Vince Bayou, where Acinetobacter was prominent, clustered with sewage samples suggesting its microbial profile resembled that of sewage. Of note, a contaminated industrial facility, formerly an oil processor and wastewater treatment plant, is situated upstream of this location (Environmental Protection Agency, 2024). After heavy storms, toxic wastes have reportedly contaminated the area with reports of hundreds of dead fish (Kelly, 2024). During sampling, toxic chemical contamination may have shifted the microbiome to resemble sewage-like communities, causing resilient genera like Acinetobacter to dominate.

Figure 2 with 2 supplements see all
GEΦMAPPING highlights Bray’s Bayou as a Pseudomonas-rich target site, and performance of ΦHD sampling of several environmental sites.

(A) Geographical map of freshwater samples obtained in Houston, TX, USA. (B) Principal Coordinate Analysis (PCoA) analysis showing β-diversity differences between sewage, sea, and freshwater sources. The black arrow indicates the Vince Bayou sample. (C) PCoA analysis (Bray–Curtis dissimilarity) of environmental water bodies with Bray’s Bayou target sites highlighted (purple). (D) Heat map with z-score normalized by site shows distinct patterns of microbial prevalence at various sites. (E) Heat map, z-score normalized by pathogen, highlights the potential of Bray’s Bayou as a target site for Pseudomonas and enteric phages. (F) Quantification of anti-Pseudomonas (left) and anti-Vibrio (right) plaques (PFU/mL) from different stages of processing (input, first stage, and second stage processing). (G) Heatmap showing anti-pathogen phages (log10[PFU/mL]) from freshwater sites. We display the mean and SEM of three technical repeat measurements. (H) Plaque assay on P. aeruginosa PAO1 of unprocessed, first stage concentrate, and second stage concentrate of a freshwater sample. (I) TEM of unprocessed water from freshwater. (J) Plaque assay on Vibrio parahaemolyticus 17802 of unprocessed, first stage concentrate, and second stage concentrate seawater sample. (K) TEM of unprocessed water from seawater, respectively. Representative images shown (H–K). Statistical analysis. (F, left panel) Phage titers were compared using a two-way ANOVA with Dunnett’s multiple comparisons test (α=0.05; n=3 technical replicates per condition), revealing significant effects of sampling site (F(1,12) = 419.7, p<0.0001), processing stage (F(2,12) = 670.2, p<0.0001), and their interaction (F(2,12) = 338.3, p<0.0001), indicating that the effect of processing stage differed between sites. Second-stage processing significantly increased phage yield relative to unprocessed input at both Bray’s Bayou (mean difference = 25,744 PFU/mL, p<0.0001) and Clear Creek (mean difference = 4,325 PFU/mL, p<0.0001); first-stage processing did not reach significance at either site (p=0.075 and p=0.945, respectively). (F, right panel) Phage titers across processing stages were compared using a one-way ANOVA with Dunnett’s multiple comparisons test (α=0.05; n=3 technical replicates per condition), revealing a significant effect of processing stage (F(2,6) = 48.00, p=0.0002). Second-stage processing yielded significantly higher titers relative to unprocessed input (p=0.0003), while first-stage processing did not differ significantly from input (p>0.9999). We display the mean and SD of three replicates for both graphs.

To focus on pathogens, we extracted the counts of pathogen-related taxa and generated heatmaps normalized by site (Figure 2D) and by taxa (Figure 2E). By highlighting selected pathogenic taxa, we can compare patterns across the environments relative to the other taxa; this does not imply absolute abundance or general enrichment of these taxa in the environment. Genus-level classification includes both pathogen and non-pathogenic species. Normalizing by site, we determined the microbial composition of each location. We detected Vibrio in Gulf of Mexico and West Bay; both expected in marine environments (Figure 2D). We noted Enterobacteriaceae was most common in waterways draining rural and suburban areas, while Pseudomonas was more common in larger urban waterways (Brays and Buffalo Bayous; Figure 2D, black rectangles). Again, Vince Bayou was an outlier, highly enriched with Acinetobacter.

Normalizing by taxa allowed us to find sites of interest for ΦHD sampling (Figure 2E, black rectangles). The data showed predominance of Vibrio at marine sites, while Pseudomonas, Enterobacteriaceae, and Enterococcus were best represented at Brays Bayou. Vince Bayou remained enriched with Acinetobacter. Of the 22% of TAILΦR isolates without a phage, 35% of these were Pseudomonas (Figure 1A, Supplementary file 3). This made Brays Bayou, rich in Pseudomonas, an ideal target for high volume sampling.

Environmental sampling freshwater and seawater with ΦHD

Although phages are abundant in nature, harvesting phages against pathogens can be difficult (Sada and Tessema, 2024; Kenney and Gómez-Duarte, 2024). Typically, phage hunters are restricted to small volumes and sometimes rely on enrichment techniques to compensate (Sada and Tessema, 2024; Kenney and Gómez-Duarte, 2024). These techniques vary by sample and host but take days with no guaranteed success. After selecting our sampling location, the next challenge was obtaining a highly concentrated sample through high-volume sampling to ensure sufficient phage abundance for effective downstream screening.

We approached this using ΦHD, which was inspired by the various research groups that used tangential flow filtration (TFF) for viral metagenomics (Angly et al., 2006; Schoenfeld et al., 2008; Fulton et al., 2009). We designed ΦHD to have a robust filtration train combined with a high-flux concentration component, allowing for low shear fluid path and high portability (Figure 1D). We used a 297 μm hose filter to screen out larger debris, followed by sequential 25 and 5 µm filter cartridges to remove larger microorganisms such as plankton and protists. We then used two hollow fiber TFF units, operating in parallel for concentration with flow rates in the low shear region. The first unit performs ultrafiltration and fills the reservoir with retentate, whereas the second takes from that reservoir and continuously concentrates the retentate in a closed loop. These are powered by a dual-headed peristaltic pump equipped with a lithium-ion battery and solar panels to allow prolonged operation in the field. During spike-in studies, we spiked pond water (60 L) with phage (1×103 PFU/mL of ΦJB10; Figure 2—figure supplement 2A). We found minimal losses of phage and bacteria during passage through the filtration train and obtained an average 20-fold increase in phage yield (Figure 2—figure supplement 2B). After the spike-in study, we evaluated our cleaning-in-place protocol (Figure 2—figure supplement 2C, Post CIP) for the TFF units and found it effectively removed phages and bacteria, confirming ΦHD is reusable.

In the field, we routinely achieved flow rates of 200 L/hr and gathered samples ranging from 400 to 1200 L (Supplementary file 5). Initial testing on-site proved the filtration/concentration train was robust enough to manage the highly polluted waterways of Houston (Figure 2—figure supplement 2D–E, Santa Ana Capture Site [SACS]). Although not all Pseudomonas phages will be detected, we selected Pseudomonas aeruginosa PA01 as the indicator strain because of its broad susceptibility profile and its role as a permissive host. After first stage processing with ΦHD, phage yield increased by 95-fold in comparison to unprocessed sample (Figure 2—figure supplement 2E, right panel). We also sampled the more pristine waters of Pedernales Falls (PF) and Hamilton Pool Preserve (HP), and incorporated a secondary in-lab concentration stage to further increase retentate concentrations (Figure 2—figure supplement 2F). Phage yields from inputs from PF and HP were near zero but increased to 102-3 PFU/mL with secondary processing. Similar treatment of SACS retentates yielded an average 3501-fold increase in phage yield (Figure 2—figure supplement 2G). This shows ΦHD, with second stage concentration, greatly increased the detection of viable phages from a given site (Figure 2—figure supplement 2G and H). Screening with other bacteria showed more phages against other pathogens from polluted SACS water compared to pristine sites (Figure 2—figure supplement 2I).

To test the combination of geΦmapping and ΦHD, we selected one site that was highlighted by our map as Pseudomonas-, Enterobacteriaceae-, and Enterococcus-rich (Brays Bayou, Figure 2E) and one low-richness site (Clear Creek). We processed 400 L/site and tracked phage yields by endogenous anti-Pseudomonas (PAO1) phages. ΦHD significantly increased phage yields at each site (166-fold at Brays Bayou and 520-fold at Clear Creek, Figure 2F and H). Using second stage concentrates, we screened a panel of patient isolates of different genera to compare hits and found more lytic phages at Brays Bayou (7/12 species) than Clear Creek (3/12 species), phenotypically validating our 16S-based mapping strategy (Figure 2G). We employed an identical filtration train at Galveston Bay sampling the marine, saltwater biome. We tracked concentrations with Vibrio parahaemolyticus 17,802 and could only identify plaques after second stage concentration (Figure 2F-right panel, J), showing utility in salt water. To test whether phage concentration increased without host-selection bias, we used transmission electron microscopy (TEM). We observed a wide range of phage morphologies with intact tails in concentrates, while the input was mostly clear of particles (Figure 2I and K). This provides further objective evidence of unbiased concentration of intact phages through the ΦHD process.

GEΦMAPPING of wastewater treatment plants in Houston, Texas

Pathogens belonging to Enterococcus and Enterobacteriaceae, including Escherichia and Klebsiella, are known inhabitants of wastewater, so we reasoned wastewater would be prime choice for finding phages against these pathogens (Brumfield et al., 2022; Drulis-Kawa et al., 2011). In a similar fashion to freshwater mapping, we sampled influents of seven different wastewater treatment plants (WWTPs) twice to generate geΦmaps of the pathogenic landscape of wastewater plants. Figure 3A shows the catchment area for each WWTP in Houston. The logarithmic shape of rarefaction curves showed we adequately sampled from each site (Figure 2—figure supplement 1A, bottom). PCoA of β-diversity (Bray–Curtis dissimilarity) demonstrated close clustering of sewage samples despite substantial diversity in both human populations and localization of industrial and commercial facilities. Unexpectedly, influent samples collected from Almeda Sims WWTP showed unusual abundance of Chloroflexi, known to be important in biodegradation in wastewater bioreactors (Figure 3B, yellow; Bovio-Winkler et al., 2023).

Figure 3 with 1 supplement see all
GEΦMAPPING identified Enterococcus and Enterobacteriaceae was rich at West University WWTP, and performance of ΦHD sampling in wastewater.

(A) Schematic of geographical areas drained by WWTPs in Houston, TX, USA. (B) β-diversity analysis shows variation in OTUs present in WWTPs by PCoA analysis based on Bray–Curtis dissimilarity (AB, colors represent differing WWTPs). (C) Heatmap analysis of pathogenic taxa highlights (black box) location of enteric pathogens (z-score normalized by taxa). (D) 16 S analysis further localizes pathogenic taxa to wastewater influent (z-score normalized by taxa). (E) Quantitation of plaques on index strains confirms 16 S sequencing data. (F) Titration of VLPs from different WWTP processing areas on PAO1 (P. aeruginosa) and K-12 (Escherichia coli) index strains (representative data shown). (G) Outline of prefiltration and sedimentation step for wastewater sites. (H) Quantification of anti-Pseudomonas plaques and anti-Escherichia plaques from different stages of processing shows phage retention through the system and concentration of the end products. (I) Titration of VLPs from different WWTP processing areas on PAO1 (P. aeruginosa). (J) TEM of unprocessed ΦHD input from wastewater. (K) TEM of second-stage concentrated wastewater. Representative images shown (F, I–K). Statistical analysis. (E) Phage titers across wastewater treatment stages and indicator strains were compared using a two-way ANOVA with Tukey’s multiple comparisons test (α=0.05; n=3 technical replicates per condition), revealing significant effects of treatment stage (F(2,12) = 399.3, p<0.0001), indicator strain (F(1,12) = 35.31, p<0.0001), and their interaction (F(2,12) = 16.87, p=0.0003), indicating that the effect of treatment stage differed between indicator strains. Influent yielded significantly higher phage titers than both contact and aeration tanks for PAO1 (p<0.0001 for both) and K-12 (p<0.0001 for both), while no significant difference was detected between contact and aeration tanks for either strain (PAO1: p=0.093; K-12: p>0.9999). (H) Phage titers across wastewater processing stages and indicator strains were compared using a two-way ANOVA with Dunnett’s multiple comparisons test (α=0.05; n=3 technical replicates per condition). A significant effect of processing stage was detected (F(4,20) = 31.68, p<0.0001), while neither indicator strain (F(1,20) = 0.188, p=0.669) nor the interaction between strain and processing stage (F(4,20) = 0.262, p=0.899) reached significance, indicating that processing stage drove phage yield increases independent of indicator strain. Second-stage processing yielded significantly higher titers relative to unprocessed influent for both PAO1 (p<0.0001) and K-12 (p<0.0001), while sedimentation, pre-filtration, and first-stage processing did not differ significantly from unprocessed influent for either strain (all p>0.6). We display the mean and SD of three replicates for both graphs.

We normalized by taxa and generated heatmaps to make our wastewater geΦmap (Figure 3C). The Park Ten WWTP influent was rich in Pseudomonas, Acinetobacter, and Stenotrophomonas, but it was being decommissioned during our study. West University WWTP had higher levels of Enterobacteriaceae and Enterococcus, so we selected it for further investigation. The wastewater treatment process alters the microbiome of wastewater, so we compared the pathogenic taxa throughout the process. We sampled raw wastewater filtered only through large steel bars that remove large debris (influent). Similarly, we sampled wastewater mixed with activated sludge (contact tank) and wastewater exposed to activated sludge with air bubbling (aeration tank; Burton, 2013). α-diversity indices showed that microbial diversity was richest in the contact tank and aeration tank of West University WWTP (Figure 2—figure supplement 1B), but the influent tank had the highest abundance of the selected pathogenic taxa (Figure 3D). We confirmed this 16 S genomic data by comparing titers of phages against indicator strains (PAO1 for Pseudomonas and K-12 for Escherichia coli) which showed the majority of these phages were found in the influent (Figure 3E–F).

Wastewater sampling with ΦHD

We used a modified ΦHD method to process 400 L of influent wastewater from West University WWTP (Figure 3G, Figure 3—figure supplement 1). To avoid pump issues caused by gravity, we lifted the wastewater 1–2 m onto an elevated deck into large tubs. We allowed this material to settle for 30 min to sediment large debris and then we pre-filtered the sample using 297 µm hose filters and two spin-down sediment traps (150 and 63 µm screens) to produce a clarified product. This pre-filtered material was then used as the ΦHD input. To ensure these additional filtration steps did not capture significant quantities of phage, we titered the material at each stage against indicator strains (PAO1 and K12) and found negligible losses (Figure 3H). We titered input and second stage concentrates (Figure 3H and I) on Pseudomonas and Escherichia indicators with final phage yields increasing by 118-fold and 108-fold, respectively. We attempted to look for non-selective concentration of virus-like particles (VLPs) with TEM; however, this was complicated by residual particulate matter. Nonetheless, we were able to appreciate some phages in retentates but not the unconcentrated input material (Figure 3J and K). We also screened the second stage concentrate against the previous panel of clinical and laboratory strains. This yielded phages for 9 of 12 bacterial strains screened, and phage titers were higher than at other sites (darker red; Figure 2G).

Collective shallow shotgun metagenomic sequencing

Phage diversity varies significantly across different biomes due to factors such as host availability, ecological niches, and environmental conditions. We sought to compare the viromes of all our ΦHD concentrate (Figure 4). Before that, to test the utility of ΦHD for metagenomic applications, we compared viral metagenomes extracted from an unprocessed sample (5 L) and a ΦHD concentrate (60 mL of 6667 X retentate, corresponding to 400 L of processed sample). The majority of α-diversity indices were similar, suggesting comparable within-sample diversity; however, β-diversity based on Bray–Curtis dissimilarity (0.6062) indicates highly moderate differences in viral community composition. (Supplementary file 6). Heatmap of viral population abundances showed no clear clustering or distinct patterns between the two samples (Figure 4—figure supplement 1A). Nevertheless, we detected substantial increases in unique viral contigs (Supplementary file 7, Figure 4—figure supplement 1B and C). We compared putative Pseudomonas-infecting viruses from the input and ΦHD samples, and the ΦHD sample contained both different members from genetically distinct groups and a greater absolute number of unique anti-Pseudomonas viruses (Figure 4—figure supplement 1D–F).

Figure 4 with 2 supplements see all
Taxonomic classification of viral contigs from metagenomic datasets reveals distinct viral signatures as well as biased presence of viruses of pathogenic hosts between different locations.

(A) vConTACT2 gene-sharing network of 16,198 viral operational taxonomic units (vOTUs; ≥5 kb) from West University WWTP, Brays Bayou, Clear Creek, and Galveston. Each dot (node) represents a vOTU, and each line (edge) represents the similarity between each genome. Additionally, 3508 reference genomes from the Prokaryotic Viral RefSeq v201 database are shown in grey. (B) Pie-chart of viral cluster status of vOTUs from vConTACT2. (C) Pie-chart of classified (family-level) vs unclassified vOTU from PhaGCN. (D) Bar graph of family-level classification (PhaGCN counts) of topmost abundant genera across the samples. Remaining genera are grouped together as ‘others’ in grey. Unclassified vOTUs are in black. (E) Heatmap with z-score normalized by viral taxa to show distinct viral signatures between sites (pheatmap separated each cluster). (F) Venn-diagram of shared vOTUs (cluster mode: amino acid identity [AAI] 45%, protein coverage [PC] 80%) depicts the number of unique vOTUs present at each site and shared between sites. (G) Heatmap with z-score normalized by host reveals the relevance of sampling specific sites for phage hunting.

Given the apparent utility of ΦHD in concentrating samples for metagenomics, we looked at the viromes of fresh (Brays), brackish (Clear Creek), waste (West University WWTP), and seawater (Galveston). Based on α-diversity indices, either freshwater or wastewater samples represent the richest sources of viral diversity (Supplementary file 8). Heatmap of viral abundance and β-diversity analysis with PCoA confirm distinct viral community compositions between biomes (Figure 4—figure supplement 2). We clustered our putative, non-redundant viral sequences into vOTUs, or virus operational taxonomic units. We generated a gene-sharing network with vConTACT2 with each node representing a vOTU, colored by source and clustered them with reference genomes from the Prokaryotic Viral RefSeq v201 database (Figure 4A). We counted 16,198 vOTUs ≥5 kb and 5095 vOTUs ≥10 kb.

39.1% of vOTUs clustered with a reference genome at the genus level, while 20.3% were outliers sharing only one to two genes with other genomes. 30.3% were singletons with no genes related to any other genomes in the data set (Figure 4A and B). We also used PhaGCN to classify each virus, but only 33.5% of viruses could be classified to the family-level (Figure 4C). Of the classified genomes, Kyanoviridae (green) and Autographiviridae (orange) made up the largest identifiable taxa in Clear Creek and Galveston (Figure 4D). This agrees closely with previous findings that Kyanoviridae, containing T4-like viruses that infect cyanobacteria, and Autographiviridae, T7-like viruses, have been found in deep-sea viromes (Zheng et al., 2025). Due to current limitations in viromics, a large portion of our metagenomic dataset remains unknown. Collectively, between 60.9% (outlier/singles from vConTACT2) and 66.5% (unclassified from PhaGCN) of vOTUs are in this category. A substantial fraction is considered viral dark matter, but we used the remaining identifiable sequences to provide some biological context. This enables us to validate our sampling strategies and compare samples and biomes from another. We used counts from PhaGCN classifications to generate heatmaps normalized by viral taxa to see distinct viral clusters from each biome type. There are four viral signatures present, separating wastewater, seawater, freshwater, and brackish water (Figure 4E). We further clustered viruses from each biome together to identify overlaps and visualized them with Venn diagrams (Figure 4F). We saw that overlaps of viral genomes reflected the physical environment; for example, the brackish water of Clear Creek overlapped between seawater and freshwater – representing its status as an admixture of the two. Next, we used CHERRY for host prediction and made heatmaps normalized by hosts. A large majority of phages with pathogenic hosts were found in the wastewater virome (Figure 4G, black box).

To increase phage diversity and minimize the emergence of phage resistance, curating a large, genetically diverse phage library is advantageous. Such a resource allows researchers and clinicians to select phages tailored to their patient needs, including targeting specific phage receptors, maximizing host coverage, designing phage cocktails, or reducing dependence on any single phage (Markwitz et al., 2022). We compared Pseudomonas- and Escherichia-infecting viruses from freshwater and wastewater. Phylogenetic analysis of Pseudomonas-infecting viruses showed significant increases in diversity when combining vOTUs from the two biomes (Figure 5A). Counting and clustering the unique genomes, there were no sequences that were in common between sewage and freshwater (Figure 5B). Analyzing the Escherichia-infecting viruses, phylogenetic diversity was greatly increased by including freshwater phages, and the two biomes had no sequences in common (Figure 5C and D). This shows the benefit of creating a phage library using ΦHD from multiple biomes.

Relevance of multi-site sampling for increased diversity and collection of unique Pseudomonas and Escherichia phages.

(A) Phylogenetic tree of all vOTUs predicted to have Pseudomonas host from West University WWTP wastewater and Brays Bayou freshwater metagenomic dataset. Branch-length is non-scaled. (B) Venn-diagram of vOTUs-grouping (cluster mode: AAI 45%, 15 shared protein (SP), PC 80%) shows no common vOTUs shared between sites. (C) Phylogenetic tree of all vOTUs predicted to have Escherichia host from wastewater and Brays Bayou metagenomic dataset. Branch-length is non-scaled. (D) Venn-diagram of vOTUs-grouping (cluster mode: AAI 45%, 15 shared protein [SP], PC 80%) shows no common vOTUs shared between sites. Bar graphs (A, B) represent the genomic similarities between a contig and its closest neighbor based on genome-wide sequence similarities computed by tBLASTx from VipTree.

Resistant phage library (RΦ-Library)

Last-line antibiotics serve as a critical safety net for treating life-threatening infections by multi-drug-resistant bacteria. Their use is restricted to preserve efficacy and slow the development of drug resistance. In a similar manner, the RΦ-library was conceptualized as a last-line phage library to combat extremely phage-resistant organisms (XΦROs). To test our approach of geΦmapping and ΦHD and generate a therapeutic phage library for XΦROs that posed clinical challenges, we used 17 XΦROs from the TAILΦR collection. These failed all conventional phage discovery methods (Figure 6A), including sewage spotting and Appelmans protocol (for host range expansion). As screening ΦHD concentrates allows searching far larger volumes (Figure 6B), we immediately found therapeutic phage candidates for all 17 XΦROs using freshwater (Brays Bayou) and sewage (West University) sources (Figure 6C). We purified (Figure 6D) and characterized 36 new phages by genomic analysis and TEM (Figure 6D and E, Supplementary file 9). This collection of phages is what we termed the RΦ-library. Compared to TAILOR’s sequenced phages, proteomic dendrograms (Figure 6E) showed phages with both subtle and high genetic variation, including new genera exemplified by MYC30C1 and MYC30C2 against Enterococcus (Figure 6—figure supplements 1 and 2). Our technique also isolated the jumbo phage UCS29C1 against E. coli, which was closely related to the Klebsiella jumbo phage Miami, showing that even large phages can survive the process intact (Figure 6—figure supplements 1 and 2; Mora et al., 2021). These successes make GEΦMAPPING and ΦHD sampling excellent additions to TAILΦR’s phage discovery pipeline, allowing for isolation of novel phages with greater precision and efficiency (Figure 6F).

Figure 6 with 2 supplements see all
ΦHD increases search volumes and allowed efficient discovery of phages for phage-resistant bacterial pathogens.

(A) Schematic of ΦHD compared to standard identification and discovery of therapeutic phages. (B) Effective volumes that can be searched comparing standard spotting of concentrates with high-concentration, high-volume ΦHD retentates (n=2 retentates with differing final concentrations). We display the mean and SD of two replicates. (C) Heatmap showing success/failure of ΦHD compared to the TAILΦR library in finding therapeutic phage candidates. (D) Plaque assays (left column) on phage-resistant strains EUA02 (K. pneumoniae) and HPC3.1 (P. aeruginosa). Individual plaques (middle columns) were streaked and isolated from plaque assays. TEM images (right columns) of select isolated phages infecting EUA02 and HPC3.1. Representative images shown. (E) Dendrograms of all isolated phages for Pseudomonas, Escherichia, Klebsiella, Enterococcus. Phages from ΦHD are in red, and phages from the TAILΦR library in blue. (F) Diagram of the implementation of ΦHD and geΦmapping to the phage discovery pipeline. Created with BioRender.com.

Discussion

Our study leverages the billion-year-long evolutionary war between phages and their hosts, focusing on discovering and harnessing natural phages. We report: (i) the design, construction, and successful use of a high-throughput phage capture device (ΦHD); (ii) the use of this device to identify regional sources of both bacteria and phages (geΦmapping); (iii) the use of geΦmapping coupled with ΦHD sampling to identify novel phages that target phage-resistant strains when all other phage matching solutions failed (RΦ-Library); and (iv) the characterization of new metagenomes. Collectively, this work provides a rich resource for therapeutic phages against resistant ESKAPEs (Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, and Enterobacter spp.) and a rich gene library for new biological discovery (Figure 6F).

Wastewater, areas of high human activity, and dense urban development are particularly rich in phages capable of targeting antibiotic-resistant pathogens (Sada and Tessema, 2024; Aghaee et al., 2021; Blanch et al., 2006; Piracha et al., 2014; Chan et al., 2016). This is largely due to the elevated microbial loads in these settings, which naturally support a higher abundance and diversity of phages (Gulino et al., 2020; Olsen et al., 2020a). Human activity imposes evolutionary pressure on bacteria, driving the emergence of antibiotic-resistant strains—and, by extension, their phage counterparts (Pruden et al., 2012). To strategically identify optimal sampling locations, we employed geΦmapping to pinpoint environments with high pathogen loads (Figures 2E and 3C). We identified Brays Bayou as an excellent locale for anti-Pseudomonas and other phages, with agreement between genomic mapping and phenotypic phage surveys. We extended our approach to sewage, not only localizing to a single WWTP for phage hunting but to a specific section of that plant (Figure 3A–F).

This targeting enabled the use of ΦHD, which produced phage-rich retentates with phage yields increasing by 100- to 3500-fold each sample in comparison to unprocessed input (Figures 2F and 3H, Figure 2—figure supplement 2G). These retentates were used to target XΦROs in clinical cases resulting in not only the generation of the RΦ-library (Figure 6C) and discovery of new genera of phages (Figure 6—figure supplements 1 and 2) but also treatment in patient cases. Currently, one patient has received two phages that necessitated the ΦHD approach, and six more patient cases are in progress that required ΦHD-derived phages (not published). This illustrates the critical need for phage discovery being addressed by the geΦmapping and ΦHD approach.

Our study is limited by multiple factors. Our current geΦmap is limited to the Gulf Coast area around Houston, TX, USA, and thus is of limited utility to others. However, the overall approach of localizing phage reservoirs is generalizable. With regards to the geomap-guided phage hunt, we only compared one low host-availability control site (Clear Creek) with two high host-availability sites (wastewater and Brays Bayou). More comparisons, especially incorporating even more sites with varied host availability and including additional hosts to screen, would strengthen the claim. Beyond our control, we were limited by weather conditions, as temperature and rainfall can drastically alter both microbial and viral composition. In turn, this limited our sampling time to days with consistent weather (i.e. dry, warm days) for accurate comparisons. In the field, it is technically challenging to remove all phages and bacteria from the hollow fiber units at the end of a sampling run. We employ simple backflushing through the retentate port, which increases phage yield by 25–50%. More phages could be stuck in the hollow fibers. Given the amount of particulate removed later during cleaning, we are certain recovery percentages can be further increased with additional backflushing. While the ΦHD system is portable, the total assembly is large and requires at least two people to carry and operate. We are currently developing a smaller, portable version for single-person use to address this. This would allow for researchers to sample in more remote locations and require less manpower. Finally, our approach is limited to aquatic biomes, but it is well known that the soils of forests and other areas (animal waste, sand, and so on) harbor huge phage biodiversity. We are currently working on ways to address this for both viral ecology and clinical phage discovery.

Collectively, we used a bioprospecting approach to localize therapeutic phages and harvest them en masse with ΦHD. This resulted in large metagenomic data sets informing the viral ecology of the Gulf Coast, critical therapeutic phages for challenging clinical cases, and tools for others to use in phage discovery.

Methods

Sample acquisition and preparation for 16S mapping

4 L freshwater and seawater samples from 15 locations were collected in containers cleaned with 10% bleach and copiously pre-rinsed in ddH2O. Locations and weather were precisely recorded (Supplementary file 10). Maps were made with EasyMapMaker, 2024. Samples were collected in triplicates with the exception of Brays Bayou @ HEB. Based on 16 S mapping and data not shown here, this location is biologically equivalent to Brays Bayou @ Medical Center due to its short distance and lack of inflow between them. For sampling purposes, we used this location due to its easier access to the bayou and safety reasons. Samples were transported back to the laboratory within hours of gathering samples and stored at 4 °C until further processing the next day. We centrifuged approximately 400 mL of each sample at 10,000 × g for 20 min to pellet bacteria and debris. 250 mg was measured out in pre-weighed tubes. In some samples, where there were insufficient pellets (e.g. Buffalo @ SACS, Galveston, West Bay, Sims Bayou, Halls @ Keith Weiss Park), we centrifuged an additional 400 mL of water until 250 mg was met. Sewage samples were obtained from the influent tank of various wastewater treatment plants by plant workers and transported to BCM and handled in a similar fashion. We centrifuged 50 mL wastewater influent at 10,000 × g for 20 minutes. 250 mg of the pellet was measured in pre-weighed tubes.

DNA preparation and 16S processing

DNA was extracted using the DNeasy PowerSoil Kit (Qiagen, Germany) according to the manufacturer’s instructions. Quality was evaluated by A260/280 ratio using a NanoDrop ND-1000 spectrophotometer (Thermo Fisher, USA). Samples with an A260/280 ratio outside of the 1.8–2.0 range were further processed using a Monarch Genomic DNA Purification kit (NEB, USA) using the ‘clean-up’ protocol in the manufacturer’s instructions. Samples were reprocessed from the start if DNA concentrations were insufficient. Samples were sent for sequencing to Novogene. The V4 region of the 16 S gene was amplified using the 515 F (5′-GTGCCAGCMGCCGCGGTAA-3′) and 806 R (5′-GGACTACHVGGGTWTCTAAT-3′) primers (Apprill et al., 2015; Parada et al., 2016), and then sequenced on a Novaseq6000 and 100 K raw reads collected (Novogene, China). We included negative controls with each sequencing batch, consisting of the PowerSoil kit without sample to generate a ‘kitome’ to assess the laboratory and kit background. For these control samples, 16 S rRNA amplification and sequencing was completed despite lack of high-quality DNA inputs. Raw sequences were uploaded to SRA (BioProject #: PRJNA1309115).

16S data analysis

Raw reads were processed in mothur (v1.48.0) guided by the MiSeq SOP using SILVA (v132) reference files and seed (Schloss et al., 2009; Kozich et al., 2013; Yilmaz et al., 2014; Quast et al., 2013). Owing to the size, number of samples, and relative lack of computational resources, we used the cluster.split command with taxlevel = 4 to cluster the sequences into OTUs. We generated rarefaction curves using default parameters and calculated alpha-diversity statistics in mothur, subsampling 50,000 random sequences from each group. Beta diversity analysis and PCoA was likewise performed in mothur using Bray–Curtis distance matrices and visualized in Graphpad Prism (v.10.1.2, GraphPad Software, USA). Using R, we performed AMOVA to confirm our β-diversity comparisons. To analyze and visualize the identified taxa in each sample, we generated Krona charts on a Galaxy web server (Ondov et al., 2011; Abueg et al., 2024). We generated heatmaps of the relative amounts of specified taxa by calculating z-scores in R (R Development Core Team, 2010). We calculated z-scores both by site and by taxa, and displayed both using Kolde, 2018.

ΦHD environmental sampling

A filtration/concentration device, ΦHD, was designed to filter microbes and phages from water. The filtration train starts with a 297 µm stainless steel mesh hose filter (3.6’’ in height and 2’’ in width) attached to silicone peristaltic tubing and goes into a water source (the feed line). The hose filter excludes trash, plant matter, and fish from the system. The tube connects to a Countertop Filtration System equipped sequentially with a 25 µm blown polypropylene cartridge and a 5 µm spun polypropylene cartridge. The filters remove more debris (soil, sand, etc.) and larger microorganisms like bacteria. Tubing then connects to the feed port of a 100 kDa hollow fiber unit (modified polyether sulfone (mPES), 65 cm, KrosFlo Hollow Fiber Filter Module). Anything smaller than 100 kDa (viruses, smaller bacteria, etc.) is filtered by TFF and filtered through the permeate ports. The retentate port ejects filtered water into a (pre-cleaned with 10% bleach and diH20-rinsed) 4 L reservoir. Stop valves are connected to the retentate line and bottom permeate line to control flow rates and fill the jacket of the unit. We use one head of a dual-headed peristaltic pump to propel water through the filtration train. The pump is powered with a travel-size battery station, Jackery Explorer 300, that was charged with solar panels.

Water was filtered and concentrated into a single reservoir, and a second 100 kDa hollow fiber unit was used to further concentrate the water into smaller volumes. The second head of the pump continuously cycled retentates from reservoir to unit. We used a custom-built aluminum stand that allows easy field use of the hollow fiber units. The concentrated product was expected to be 500 mL to 1 L in volume. ΦHD sampling in-field is the first stage of processing (Figure 3—figure supplement 1). Samples were brought back to the lab in 1–2 hr. We centrifuged the samples, generally 1–2 L, at 10,000 × g for 20 minutes and filtered the samples with 0.22 µm vacuum-filters (Steritop filters (PES)). Samples were stored at 4 °C until second stage processing in lab (next day). A smaller version of the 100 kDa hollow fiber unit was used to achieve volumes of 50–200 mL per sample (mPES, 20 cm, MidiKros Hollow Fiber Filter Module). Retentate flow rates are controlled by Luer lock stoppers on the retentate line and bottom permeate line. Concentrations of each sample were recorded accordingly after secondary concentration (Supplementary file 5).

Cleaning-in-place (CIP) protocol of hollow fiber units (TFF units)

After each ΦHD run, we cleaned the TFF units to regenerate them. For 5–10 min, water was flushed through the permeate ports on both ends of the TFF unit. This reverse flow removes the majority of biological build-up in the units. The outer jacket and hollow fibers were then filled with 0.5 M HCl and left overnight to incubate at room temperature. The next day, we rinsed the units with water then soaked outer jacket and hollow fibers in 0.5 M NaOH overnight. Once rinsed with diH2O, the units are ready to be used again. For more polluted waters (such as wastewater), we repeated the CIP protocol. We’ve successfully used the units over 10 times for this study (Supplementary file 5).

ΦHD wastewater sampling

Wastewater has a high burden of flotsam (toilet paper, wet wipes, and feminine hygiene production, etc.) as well as smaller particulates (Figure 3—figure supplement 1B). We pre-filtered wastewater and used the pre-filtered input as the feed to ΦHD. Two silicone peristaltic tubes drew wastewater 1–2 m from the influent tank of West University WWTP and into two large tubs, pre-cleaned with 70% ethanol and Clorox wipes. This helped avoid gravitation issues such as slower flow rates and battery tripping. Wastewater was allowed to settle for 30 min. This allowed larger, denser material to sediment to the bottom. Two tubes with styrofoam floaters and barbed hose filters propelled water through two spin down filters – containing 150 and 63 µm mesh cartridges. The floaters were essential in avoiding the settled debris at the bottom, and the larger filters helped remove more debris. Wastewater was stored in two pre-cleaned tubs; this became the input for ΦHD where we processed the water as described above. For wastewater processing, an additional valve on the feed line is needed to slow down flow rate. Similarly, we centrifuged the first stage concentrate (1–2 L) at 10,000 × g for 20 min and filtered the samples with 0.22 µm vacuum filters. Second stage processing lowered the volume to 150–200 mL.

Bacterial growth and plaque assays

A table of strains used in the study can be found in Supplementary file 11. From a single colony, bacterial cultures were inoculated in Luria Broth (LB, Sigma-Aldrich) and grown overnight at 37 °C with shaking at 180 rpm. For plaque assays, overnight cultures were suspended in 0.75% LB top agar with 10–1000 µL of ΦHD concentrate and poured onto 1.5% LB plates to solidify. Plates were incubated overnight at 37 °C. Plaques are counted next day. Phage titers were visualized with Graphpad Prism (v.10.6.1, GraphPad Software, USA). For phage streaks, a single plaque on agar plates was picked and resuspended in 25 µL of phosphate buffer saline. Bacterial cultures were suspended on 0.75% LB top agar and poured onto LB plates. Before the top agar solidifies, the resuspended plaque was streaked with a pipette. To isolate individual phages from a plaque, the phage was continuously streaked out until plaque morphologies are consistent (at least twice).

DNA preparation and shotgun shallow metagenomic sequencing

11.5 mL of second stage concentrates from ΦHD (Supplementary file 5) were centrifuged with a SW-41Ti rotor at 37.5 k rpm for 2 hr to pellet VLPs. Supernatant was removed, and the pelleted viruses are used for DNA extraction using the DNeasy PowerSoil kit (Qiagen, Germany) according to manufacturer’s protocols. The extracted DNA was sent for PCR-free library preparation followed by shallow shotgun metagenomic sequencing on the NovaSeq X Plus system platform using the PE150 strategy (Novogene, Beijing). This resulted in >2 G of raw data per sample with Q30>85. Raw sequences were uploaded to SRA (BioProject #: PRJNA1308632).

Viral metagenomic analysis

Raw data was uploaded into KBASE (Arkin et al., 2018) and examined with FastQC (0.12.1) (Andrews, 2010) then trimmed with Trimmomatic (v0.36) (Bolger et al., 2014) with default settings (sliding window size 4, sliding minimum quality 15, post tail crop length null, head crop length 0, leading minimum quality 3, trailing minimum quality 3, minimum read length 36; Bolger et al., 2014). Post-trimmed reads were assessed with FastQC (0.12.1). Reads were assembled into contigs with metaSPAdes (v3.15.3); minimum contig length was 300≤2000 bp and k-mer sizes were 21, 33, 55, 77, 99, and 127 (Nurk et al., 2017). Contigs were uploaded into CyVerse for further analysis with VIBRANT (Swetnam et al., 2024; Kieft et al., 2020). VIBRANT was used to identify and extract viral contigs from the dataset; default settings were used (length of bp was 1000, number of ORFs per scaffold was 4). Viral contigs from individual samples were clustered into viral operational taxonomic units (vOTUs) using average nucleotide identity (ANI)-based clustering (95% ANI, 85% alignment coverage) on PhaBOX (Shang et al., 2023). We used CheckV (Nayfach et al., 2021) to assess quality and completeness of vOTUs. vOTUs were annotated with PROKKA (v1.14.5) (Seemann, 2014) on KBASE, and gene-sharing networks were made with vConTACT2 (v0.9.19; Bolduc et al., 2017). Settings were set to default, and the reference database used for cluster taxonomy was NCBI Bacterial and Archaeal Viral RefSeq V201. We visualized the gene-networks with Cytoscape (Shannon et al., 2003). vOTUs were further taxonomically classified with PhaGCN (Shang et al., 2021; minimum length ≥5 kb, 75% amino acid identity [AAI], shared protein 15, and 80% protein coverage) and predicted host of each vOTUs with CHERRY (Shang and Sun, 2022; minimum length ≥5 kb, 75% AAI, shared protein 15, and 80% protein coverage, 90% CRISPRs identity). Again, we generated heatmaps of the relative amounts of identified taxa or predicted hosts by calculating z-scores and displayed them using pheatmap in RStudio (R Development Core Team, 2010; Kolde, 2018).

vOTUs with predicted hosts in the genera Pseudomonas and Escherichia were manually extracted from the datasets of wastewater (West University WWTP) and freshwater (Brays Bayou). We clustered vOTUs with AAI-based clustering mode on PhaBOX (45% AAI, 15 shared protein, and 80% protein coverage). We used VIPTree (Nishimura et al., 2017) phylogenetic analysis then visualized the resulting Newick trees using ggtree (Yu et al., 2017) in RStudio.

Statistical analysis of metagenomic sequences

Raw data was uploaded into KBASE and examined with FastQC (0.12.1) then trimmed with Trimmomatic (v0.36) with default settings (sliding window size 4, sliding minimum quality 15, post tail crop length null, head crop length 0, leading minimum quality 3, trailing minimum quality 3, minimum read length 36). Post-trimmed reads were assessed with FastQC (0.12.1).

We used MetaPop to analyze viral metagenomic sequence data at the interpopulation (macrodiversity) level (Gregory et al., 2022). We aligned and mapped filtered, trimmed reads of each biome to our reference genome with bowtie2 (v2.5.4+galaxy0). The 16,198 vOTUs (95% ANI, 85% alignment coverage) from wastewater, seawater, brackish, and freshwater were combined and used as our reference genome. Generated BAM output files from bowtie2 and the reference genome were used as input for MetaPop. MetaPop’s macrodiversity analysis includes raw population abundance, normalized population abundance, and α-diversity calculations. The following analyses were conducted in RStudio using tidyverse, vegan, and pheatmap. From normalized population abundance tables, we calculated z-scores and derived heatmaps to display feature-level patterns. We calculated β-diversity using Bray–Curtis dissimilarity, then we performed PCoA with the ordination of the Bray–Curtis matrix. To assess robustness, 2000 features were randomly subsampled and analysis repeated across 1000 bootstrap iterations. Resulting ordinations were aligned to a reference with Procrustes alignment. Mean coordinates and standard deviations were calculated for each sample, and scatter plots were generated.

Generation of RΦ collection

We obtained clinical isolates from the collection of TAILΦR Labs at Baylor College of Medicine (BCM). We selected isolates for which no phage had been isolated despite exhaustive searches in sewage spotting and testing of the TAILΦR phage library. Study of these clinical isolates was approved by BCM’s Institutional Review Board (IRB). We grew each isolate overnight in LB media, mixed 100 μL of this overnight culture with 3 mL of 0.7% top agar and either 100 μL or 1 mL of environmental or sewage ΦHD concentrate as indicated. Plates were incubated overnight at 37 °C. Next day, plaques that we considered as lytic in appearance (non-cloudy, cleared zones on bacterial lawns) were picked with a pipette tip and streaked in top agar with the same clinical strain prepared in an identical manner. Plaques were streaked three times to ensure purity. We endeavored to pick two to three different plaque morphologies per clinical isolate for each source (environment or sewage). The purified plaques were picked and mixed in 100 μL of phage buffer (10 mM Tris pH 8.0, 100 mM NaCl, 10 mM MgSO4). We then made plate lysates by mixing 50 μL of this phage-laden material with 100 μL bacterial overnight of the indicated strain and 3 mL 0.7% top agar and incubating overnight. These were harvested by scraping the top agar from the plates, rinsing the plate surface with 3 mL phage buffer and centrifuging this mixture at 2200 × g for 10 minutes and removing residual bacteria with a 0.22 μm PVDF filter (Merck Millipore, Ireland). These phages were titered using standard methods on 0.75% top agar on the indicated strain. Anti-HPC3.1 phages underwent plate lysate production in a laboratory strain PAO1 owing to the amount of mucus produced, and were subsequently titered on the indicated clinical strain. Any other phages found in this study were noted in Supplementary file 12.

Genetic analysis of the RΦ collection

We employed various methods of DNA preparation for the phages. For samples with adequate titers, we used the EZNA Universal Pathogen Kit (Omega Bio-Tek, USA) according to the manufacturer’s instructions. Phage strains with lower titers or that failed to produce adequate DNA using the EZNA kit were processed by generation of 9 mL of plate lysate (as above) and centrifugation with a Sw-41Ti rotor at 37.5 k rpm for 2 hr to pellet VLPs. This pellet was then processed using a Qiagen PowerSoil Kit using the manufacturer’s instructions. For DNA samples with ratios outside of 1.8–2.0, we further purified them with the NEB Genomic DNA Purification Kit using the manufacturer’s ‘clean-up’ protocol. Whole-genome sequencing using standard protocols was done by BCM’s CMMR core. We processed short reads data using FastQC (v0.12.1; Yehl et al., 2019) to assess data quality, removed low-quality sequences (Q>30 cut-off), and removed the adapters using Trimmomatic (v0.39; Sada and Tessema, 2024) with default settings. We verified the removal of adapters and quality cutoffs with another run of FastQC. These clean reads were then assembled using SPAdes (v3.15.3; DNA source standard, minimum contig length 200≤500; Bankevich et al., 2012) and annotated using RASTtk (v1.073; default setting). We identified candidate viral sequences from these assemblies using (Brettin et al., 2015) VirSorter2 (discarding sequences with <2 hallmark genes, requiring hallmark genes on all sequences, only output high confidence viral sequences, and requiring viral genes to be annotated; Guo et al., 2021). Sequences with high dsDNA phage scores, >four to six hallmark genes, relatively low cellular gene fractions, and coverage suggestive of high copy numbers above the bacterial host were examined manually, and BLASTn (Altschul et al., 1990) was used to determine phage or not. Confirmed phage sequences were then re-annotated with RASTtk (v1.073) to provide final phage genomes for the RΦ library. We compared TAILOR sequences (pre-existing data) with the sequences of purified phages from the RΦ library. We used VIPTree (Nishimura et al., 2017) for a proteomic dendrogram and then visualized the resulting Newick trees using ggtree (Yu et al., 2017) in RStudio. Several Klebsiella phages that were isolated on related host strains were identical, and so we removed these duplicates prior to visualization. All phage genomes were re-annotated with Pharokka v 1.9.0 (Bouras et al., 2023). Coding sequences were predicted with PHANOTATE (McNair et al., 2019). tRNAs were predicted with rRNAscan-SE 2.0 (Chan et al., 2021). tmRNAs were predicted with Aragorn (Laslett and Canback, 2004). CRISPRs were predicted with CRT (Bland et al., 2007). Functional annotations were generated by matching each CDS to the PHROGs (Terzian et al., 2021), VFDB (Chen et al., 2005), and CARD (Alcock et al., 2020) databases using MMseqs2 (Steinegger and Söding, 2017; Larralde and Camargo, 2023; Larralde, 2022) and PyHMMER (Larralde and Zeller, 2023). Phage genomes were rotated with Dnaapler (Bouras et al., 2024). Contigs were matched to their closest hit in the INPHARED (Cook et al., 2021) database using mash (Ondov et al., 2016). GenBank accession numbers for phage genomes are on Supplementary files 9 and 12.

Electron microscopy and figures

Various environmental, sewage, and purified phage samples were imaged at BCM’s Cryo-EM Advanced Technology Core Facility. Samples were deposited on grids (Quantifoil 2/2 200 Cu +2 nm ThinC; Quantifoil Micro Tools GmbH, Jena, Germany) and negatively stained with 2% uranyl acetate using standard techniques. Images were collected with either a JEOL 1230 electron microscope (JEOL, Japan) at 80 kV equipped with a 4000 by 4000 Gatan Ultrascan charge-coupled device camera (Gatan, Ametek) or a JEOL 2100 (200 kV) electron microscope equipped with a Gatan 4k x 4k CCD. Images were collected at numerous magnifications, and representative examples were selected and viewed in NIH ImageJ.

Appendix 1

Appendix 1—key resources table
Reagent type (species) or resourceDesignationSource or referenceIdentifiersAdditional information
Strain, strain background (Acinetobacter baumannii)17978ATCC 17978 Bouvet and Grimont, 1986Received from TAILФR LABS; bacterial strain originally isolated from fatal meningitis; Used in the multidrug-resistant isolate panel for phage screening - Figure 2G
Strain, strain background (Achromobacter xylosoxidans)UTS1Gift-TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from cystic fibrosis (CF) lung infection; Used in the multidrug-resistant isolate panel for phage screening - Figure 2G
Strain, strain background (Enterobacter cloacae)SLC1Gift-TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from; Used in the multidrug-resistant isolate panel for phage screening - Figure 2G
Strain, strain background (Enterococcus faecalis)VRE001Gift - TAILOR Labs. El Haddad et al., 2022Received from TAILФR LABS; Vancomycin-resistant isolate; bacterial strain originally isolated from patient stool; Used in the multidrug-resistant isolate panel for phage screening - Figure 2G
Strain, strain background (Escherichia coli)K12Gift - TAILOR Labs. Blattner et al., 1997Received from TAILФR LABS; Wild-type, common laboratory strain; Used in the multidrug-resistant isolate panel for phage screening - Figure 2G; also used as our phage indicator strain (Figure 3E, F, H and I)
Strain, strain background
(E. coli)
JJ2528Gift - TAILOR Labs. Price et al., 2013Received from TAILФR LABS; ExPEC ST131 H30-R; Used in the multidrug-resistant isolate panel for phage screening - Figure 2G
Strain, strain background (Klebsiella aerogenes)UCF1Gift - TAILOR Labs. Terwilliger et al., 2021Received from TAILФR LABS; bacterial strain originally isolated from hip prosthesis wound; Used in the multidrug-resistant isolate panel for phage screening - Figure 2G
Strain, strain background (Klebsiella pneumoniae)EUA2Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from urinary tract infection (UTI); Used in the multidrug-resistant isolate panel for phage screening - Figure 2G
Strain, strain background (Pseudomonas aeruginosa)DSA497Gift - TAILOR LabsReceived from TAILФR LABS; Ramig LAB; bacterial strain originally isolated from UTI; Used in the multidrug-resistant isolate panel for phage screening - Figure 2G
Strain, strain background
(P. aeruginosa)
PA01Gift - TAILOR Labs. Holloway, 1955Received from TAILФR LABS; Common laboratory strain; bacterial strain originally isolated from wound; Used in the multidrug-resistant isolate panel for phage screening - Figure 2G; also used as our indicator strain (Figure 3E, F and H)
Strain, strain background (Staphylococcus aureus)AZM28Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from dog skin swab #13–1 on MSA; Used in the multidrug-resistant isolate panel for phage screening - Figure 2G
Strain, strain background (Staphylococcus pseudintermedius)AZM69Gift - TAILOR LabsReceived from TAILФR LABS; Coagulase positive; bacterial strain originally isolated from canine ear infection; Used in the multidrug-resistant isolate panel for phage screening - Figure 2G
Strain, strain background
(P. aeruginosa)
AR0246CDC and FDA Antibiotic Resistance Isolate Bank; Pseudomonas aeruginosa (PSA) PanelReceived from TAILФR LABS;
MDR strain; CDC and FDA Antibiotic Resistance Isolate Bank; Used in the multidrug-resistant isolate panel for phage screening - Figure 2G
Strain, strain background
(P. aeruginosa)
AR0266CDC and FDA Antibiotic Resistance Isolate Bank; Pseudomonas aeruginosa (PSA)Received from TAILФR LABS MDR strain; CDC & FDA Antibiotic Resistance Isolate Bank; Used in the multidrug-resistant isolate panel for phage screening - Figure 2G
Strain, strain background
(E. coli)
BSL02Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from UTI; Used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(E. coli)
UCS29Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from UTI; Used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(E. coli)
UCS37Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from urine (pyelonephritis, sepsis); Used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(E. coli)
UCF06Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from bone (infected amputation); Used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(K. aerogenes)
MYC16Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from shoulder prosthetic joint infection; Used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(K. pneumoniae)
EUA02Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from UTI; Used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(K. pneumoniae)
EUA03Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from UTI; Used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(K. pneumoniae)
UCS26Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from UTI, bacteremia; Used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(K. pneumoniae)
UCS26.1Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from UTI, bacteremia; Used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(S. aureus)
MYC20Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from sputum (bronchiectasis); Used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(E. coli)
UCS17.2Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from UTI; Used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(P. aeruginosa)
CSM01Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from sputum (lung abscess, adjacent sternal osteomyelitis); Used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(P. aeruginosa)
UCS18Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from bile duct (biliary infection, sepsis); used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(P. aeruginosa)
UCS20Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from chest drain (lung, chest wall infection); used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(P. aeruginosa)
UCS23Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from UTI; used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(P. aeruginosa)
MGB1.4Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from sputum (CF lung); used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background
(P. aeruginosa)
HPC3.1Gift - TAILOR LabsReceived from TAILФR LABS; bacterial strain originally isolated from sputum (lung); used in the XΦRO screening panel for phage isolation - Figure 6C;
Strain, strain background (Vibrio parahaemolyticus)17802ATCC 17802 Fujino et al., 1974Purchased from ATCC; isolated from Shirasu food poisoning; used to produce Figure 2F and J (right panel); used to isolate Vibrio phages
Strain, strain background (Vibrio vulnificus)27562ATCC 27562 Reichelt et al., 1976Purchased from ATCC; isolated from human blood; used as part of our panel to screen concentrated seawater for Vibrio phages; did not result in phage
Strain, strain background (Vibrio parahaemolytic phage)VP1This paperNovel bacteriophage isolated in this study; isolated from seawater concentrate; strain used to isolate was 17802
Strain, strain background
(V. parahaemolytic phage)
VP2This paperNovel bacteriophage isolated in this study; isolated from seawater concentrate; strain used to isolate was 17802
Strain, strain background
(V. parahaemolytic phage)
VP3This paperNovel bacteriophage isolated in this study; isolated from seawater concentrate; strain used to isolate was 17802
Strain, strain background
(V. parahaemolytic phage)
VP4This paperNovel bacteriophage isolated in this study; isolated from seawater concentrate; strain used to isolate was 17802
Strain, strain background (Escherichia phage)UCS29C1This paperNovel bacteriophage isolated in this study; isolated from UCS29; used to produce Figure 6E
Strain, strain background
(E. phage)
UCS29C2This paperNovel bacteriophage isolated in this study; isolated from UCS29; used to produce Figure 6E
Strain, strain background
(E. phage)
UCS37C1This paperNovel bacteriophage isolated in this study; isolated from UCS37; used to produce Figure 6E
Strain, strain background
(E. phage)
UCS37C2_1This paperNovel bacteriophage isolated in this study; isolated from UCS37; used to produce Figure 6E
Strain, strain background
(E. phage)
UCS37C2_2This paperNovel bacteriophage isolated in this study; isolated from UCS37; used to produce Figure 6E
Strain, strain background
(E. phage)
BSL02C1_1This paperNovel bacteriophage isolated in this study; isolated from BSL02; used to produce Figure 6E
Strain, strain background
(E. phage)
BSL02C1_2This paperNovel bacteriophage isolated in this study; isolated from BSL02; used to produce Figure 6E
Strain, strain background
(E. phage)
BSL02C2This paperNovel bacteriophage isolated in this study; isolated from BSL02; used to produce 6E
Strain, strain background (Enterococcus phage)MYC30C1_1This paperNovel bacteriophage isolated in this study; isolated from MYC30; used to produce Figure 6E
Strain, strain background
(E. phage)
MYC30C1_2This paperNovel bacteriophage isolated in this study; isolated from MYC30; used to produce Figure 6E
Strain, strain background
(E. phage)
UCS17X2C1This paperNovel bacteriophage isolated in this study; isolated from UCS17; used to produce Figure 6E
Strain, strain background
(E. phage)
UCS17X2C2This paperNovel bacteriophage isolated in this study; isolated from UCS17; used to produce Figure 6E
Strain, strain background
(Klebsiella phage)
UCS26C1_1This paperNovel bacteriophage isolated in this study; isolated from UCS26; used to produce Figure 6E
Strain, strain background
(K. phage)
UCS26C1_2This paperNovel bacteriophage isolated in this study; isolated from UCS26; used to produce Figure 6E
Strain, strain background
(K. phage)
EUA02C1This paperNovel bacteriophage isolated in this study; isolated from EUA02; used to produce Figure 6E
Strain, strain background
(K. phage)
EUA3C1This paperNovel bacteriophage isolated in this study; isolated from EUA03; used to produce Figure 6E
Strain, strain background
(K. phage)
MYC16C2This paperNovel bacteriophage isolated in this study; isolated from MYC16; used to produce Figure 6E
Strain, strain background
(K. phage)
UCS26X1C1_1This paperNovel bacteriophage isolated in this study; isolated from UCS26; used to produce Figure 6E
Strain, strain background
(K. phage)
UCS26X1C2_2This paperNovel bacteriophage isolated in this study; isolated from UCS26; used to produce Figure 6E
Strain, strain background
(K. phage)
UCS26X1C2_1This paperNovel bacteriophage isolated in this study; isolated from UCS26; used to produce Figure 6E
Strain, strain background
(K. phage)
UCS26X1C2_2This paperNovel bacteriophage isolated in this study; isolated from UCS26; used to produce Figure 6E
Strain, strain background
(K. phage)
UCS26X1C2_3This paperNovel bacteriophage isolated in this study; isolated from UCS26; used to produce Figure 6E
Strain, strain background
(K. phage)
UCS26X1C2_4This paperNovel bacteriophage isolated in this study; isolated from UCS26; used to produce Figure 6E
Strain, strain background
(K. phage)
UCS26X1C2_5This paperNovel bacteriophage isolated in this study; isolated from UCS26; used to produce Figure 6E
Strain, strain background
(K. phage)
UCS26X1C2_6This paperNovel bacteriophage isolated in this study; isolated from UCS26; used to produce Figure 6E
Strain, strain background
(K. phage)
UCS26C2_1This paperNovel bacteriophage isolated in this study; isolated from UCS26; used to produce Figure 6E
Strain, strain background
(K. phage)
UCS26C2_2This paperNovel bacteriophage isolated in this study; isolated from UCS26; used to produce Figure 6E
Strain, strain background
(K. phage)
UCS26C2_3Novel bacteriophage isolated in this study; isolated from UCS26; used to produce Figure 6E
Strain, strain background
(K. phage)
MYC16C1This paperNovel bacteriophage isolated in this study; isolated from MYC16; used to produce Figure 6E
Strain, strain background (Pseudomonas phage)CSM01C1This paperNovel bacteriophage isolated in this study; isolated from CSM01; used to produce Figure 6E
Strain, strain background
(P. phage)
UCS20C3This paperNovel bacteriophage isolated in this study; isolated from UCS20; used to produce Figure 6E
Strain, strain background
(P. phage)
UCS18C3_1This paperNovel bacteriophage isolated in this study; isolated from UCS18; used to produce Figure 6E
Strain, strain background
(P. phage)
UCS18C3_2This paperNovel bacteriophage isolated in this study; isolated from UCS18; used to produce Figure 6E
Strain, strain background
(P. phage)
UCS20C5This paperNovel bacteriophage isolated in this study; isolated from UCS20; used to produce Figure 6E
Strain, strain background
(P. phage)
UCS20C4This paperNovel bacteriophage isolated in this study; isolated from UCS20; used to produce Figure 6E
Strain, strain background
(P. phage)
6X1C4This paperNovel bacteriophage isolated in this study; isolated from 6.1; used to produce Figure 6E
Strain, strain background
(P. phage)
UCS20C1This paperNovel bacteriophage isolated in this study; isolated from UCS20; used to produce Figure 6E
Strain, strain background
(P. phage)
MGB1X4C4This paperNovel bacteriophage isolated in this study; isolated from MGB1.4; used to produce Figure 6E
Strain, strain background
(P. phage)
MGB1X4C1This paperNovel bacteriophage isolated in this study; isolated from MGB1.4; used to produce Figure 6E
Strain, strain background
(P. phage)
UCS18C4This paperNovel bacteriophage isolated in this study; isolated from UCS18; used to produce Figure 6E
Strain, strain background
(P. phage)
MGB1X4C3This paperNovel bacteriophage isolated in this study; isolated from MGB1.4; used to produce Figure 6E

Data availability

Raw shotgun metagenomic sequences were uploaded to NCBI SRA (Sequence Read Archive) under BioProject PRJNA1308632. 16S rRNA amplicon sequences were uploaded to NCBI SRA under BioProject PRJNA1309115. Genome sequences generated in this study have been deposited in the National Center for Biotechnology Information (NCBI) GenBank database, with accession numbers provided in the Supplementary files 9 and 12. We have uploaded our raw data, annotated genomes, generated data/output, analyzed data, and source data to Dataverse (https://doi.org/10.7910/DVN/ZREY8T).

The following data sets were generated
    1. Nicholls P
    (2026) Harvard Dataverse
    Phage Genomes - Geomapping and PhiHD Project.
    https://doi.org/10.7910/DVN/ZREY8T

References

    1. Abueg LAL
    2. Afgan E
    3. Allart O
    4. Awan AH
    5. Bacon WA
    6. Baker D
    7. Bassetti M
    8. Batut B
    9. Bernt M
    10. Blankenberg D
    11. Bombarely A
    12. Bretaudeau A
    13. Bromhead CJ
    14. Burke ML
    15. Capon PK
    16. Čech M
    17. Chavero-Díez M
    18. Chilton JM
    19. Collins TJ
    20. Coppens F
    21. Coraor N
    22. Cuccuru G
    23. Cumbo F
    24. Davis J
    25. De Geest PF
    26. de Koning W
    27. Demko M
    28. DeSanto A
    29. Begines JMD
    30. Doyle MA
    31. Droesbeke B
    32. Erxleben-Eggenhofer A
    33. Föll MC
    34. Formenti G
    35. Fouilloux A
    36. Gangazhe R
    37. Genthon T
    38. Goecks J
    39. Beltran ANG
    40. Goonasekera NA
    41. Goué N
    42. Griffin TJ
    43. Grüning BA
    44. Guerler A
    45. Gundersen S
    46. Gustafsson OJR
    47. Hall C
    48. Harrop TW
    49. Hecht H
    50. Heidari A
    51. Heisner T
    52. Heyl F
    53. Hiltemann S
    54. Hotz H-R
    55. Hyde CJ
    56. Jagtap PD
    57. Jakiela J
    58. Johnson JE
    59. Joshi J
    60. Jossé M
    61. Jum’ah K
    62. Kalaš M
    63. Kamieniecka K
    64. Kayikcioglu T
    65. Konkol M
    66. Kostrykin L
    67. Kucher N
    68. Kumar A
    69. Kuntz M
    70. Lariviere D
    71. Lazarus R
    72. Bras YL
    73. Corguillé GL
    74. Lee J
    75. Leo S
    76. Liborio L
    77. Libouban R
    78. Tabernero DL
    79. Lopez-Delisle L
    80. Los LS
    81. Mahmoud A
    82. Makunin I
    83. Marin P
    84. Mehta S
    85. Mok W
    86. Moreno PA
    87. Morier-Genoud F
    88. Mosher S
    89. Müller T
    90. Nasr E
    91. Nekrutenko A
    92. Nelson TM
    93. Oba AJ
    94. Ostrovsky A
    95. Polunina PV
    96. Poterlowicz K
    97. Price EJ
    98. Price GR
    99. Rasche H
    100. Raubenolt B
    101. Royaux C
    102. Sargent L
    103. Savage MT
    104. Savchenko V
    105. Savchenko D
    106. Schatz MC
    107. Seguineau P
    108. Serrano-Solano B
    109. Soranzo N
    110. Srikakulam SK
    111. Suderman K
    112. Syme AE
    113. Tangaro MA
    114. Tedds JA
    115. Tekman M
    116. Cheng (Mike) Thang W
    117. Thanki AS
    118. Uhl M
    119. van den Beek M
    120. Varshney D
    121. Vessio J
    122. Videm P
    123. Von Kuster G
    124. Watson GR
    125. Whitaker-Allen N
    126. Winter U
    127. Wolstencroft M
    128. Zambelli F
    129. Zierep P
    130. Zoabi R
    131. The Galaxy Community
    (2024) The Galaxy platform for accessible, reproducible, and collaborative data analyses: 2024 update
    Nucleic Acids Research 52:W83–W94.
    https://doi.org/10.1093/nar/gkae410
  1. Book
    1. Burton FL
    (2013)
    Wastewater Engineering: Treatment and Resource Recovery
    McGraw-Hill Education.
  2. Book
    1. Fulton J
    2. Douglas T
    3. Young AM
    (2009) Isolation of viruses from high temperature environments
    In: Clokie MRJ, Kropinski AM, editors. Bacteriophages. Humana Press. pp. 43–54.
    https://doi.org/10.1007/978-1-60327-164-6
    1. Naghavi M
    2. Vollset SE
    3. Ikuta KS
    4. Swetschinski LR
    5. Gray AP
    6. Wool EE
    7. Robles Aguilar G
    8. Mestrovic T
    9. Smith G
    10. Han C
    11. Hsu RL
    12. Chalek J
    13. Araki DT
    14. Chung E
    15. Raggi C
    16. Gershberg Hayoon A
    17. Davis Weaver N
    18. Lindstedt PA
    19. Smith AE
    20. Altay U
    21. Bhattacharjee NV
    22. Giannakis K
    23. Fell F
    24. McManigal B
    25. Ekapirat N
    26. Mendes JA
    27. Runghien T
    28. Srimokla O
    29. Abdelkader A
    30. Abd-Elsalam S
    31. Aboagye RG
    32. Abolhassani H
    33. Abualruz H
    34. Abubakar U
    35. Abukhadijah HJ
    36. Aburuz S
    37. Abu-Zaid A
    38. Achalapong S
    39. Addo IY
    40. Adekanmbi V
    41. Adeyeoluwa TE
    42. Adnani QES
    43. Adzigbli LA
    44. Afzal MS
    45. Afzal S
    46. Agodi A
    47. Ahlstrom AJ
    48. Ahmad A
    49. Ahmad S
    50. Ahmad T
    51. Ahmadi A
    52. Ahmed A
    53. Ahmed H
    54. Ahmed I
    55. Ahmed M
    56. Ahmed S
    57. Ahmed SA
    58. Akkaif MA
    59. Al Awaidy S
    60. Al Thaher Y
    61. Alalalmeh SO
    62. AlBataineh MT
    63. Aldhaleei WA
    64. Al-Gheethi AAS
    65. Alhaji NB
    66. Ali A
    67. Ali L
    68. Ali SS
    69. Ali W
    70. Allel K
    71. Al-Marwani S
    72. Alrawashdeh A
    73. Altaf A
    74. Al-Tammemi AB
    75. Al-Tawfiq JA
    76. Alzoubi KH
    77. Al-Zyoud WA
    78. Amos B
    79. Amuasi JH
    80. Ancuceanu R
    81. Andrews JR
    82. Anil A
    83. Anuoluwa IA
    84. Anvari S
    85. Anyasodor AE
    86. Apostol GLC
    87. Arabloo J
    88. Arafat M
    89. Aravkin AY
    90. Areda D
    91. Aremu A
    92. Artamonov AA
    93. Ashley EA
    94. Asika MO
    95. Athari SS
    96. Atout MMW
    97. Awoke T
    98. Azadnajafabad S
    99. Azam JM
    100. Aziz S
    101. Azzam AY
    102. Babaei M
    103. Babin F-X
    104. Badar M
    105. Baig AA
    106. Bajcetic M
    107. Baker S
    108. Bardhan M
    109. Barqawi HJ
    110. Basharat Z
    111. Basiru A
    112. Bastard M
    113. Basu S
    114. Bayleyegn NS
    115. Belete MA
    116. Bello OO
    117. Beloukas A
    118. Berkley JA
    119. Bhagavathula AS
    120. Bhaskar S
    121. Bhuyan SS
    122. Bielicki JA
    123. Briko NI
    124. Brown CS
    125. Browne AJ
    126. Buonsenso D
    127. Bustanji Y
    128. Carvalheiro CG
    129. Castañeda-Orjuela CA
    130. Cenderadewi M
    131. Chadwick J
    132. Chakraborty S
    133. Chandika RM
    134. Chandy S
    135. Chansamouth V
    136. Chattu VK
    137. Chaudhary AA
    138. Ching PR
    139. Chopra H
    140. Chowdhury FR
    141. Chu D-T
    142. Chutiyami M
    143. Cruz-Martins N
    144. da Silva AG
    145. Dadras O
    146. Dai X
    147. Darcho SD
    148. Das S
    149. De la Hoz FP
    150. Dekker DM
    151. Dhama K
    152. Diaz D
    153. Dickson BFR
    154. Djorie SG
    155. Dodangeh M
    156. Dohare S
    157. Dokova KG
    158. Doshi OP
    159. Dowou RK
    160. Dsouza HL
    161. Dunachie SJ
    162. Dziedzic AM
    163. Eckmanns T
    164. Ed-Dra A
    165. Eftekharimehrabad A
    166. Ekundayo TC
    167. El Sayed I
    168. Elhadi M
    169. El-Huneidi W
    170. Elias C
    171. Ellis SJ
    172. Elsheikh R
    173. Elsohaby I
    174. Eltaha C
    175. Eshrati B
    176. Eslami M
    177. Eyre DW
    178. Fadaka AO
    179. Fagbamigbe AF
    180. Fahim A
    181. Fakhri-Demeshghieh A
    182. Fasina FO
    183. Fasina MM
    184. Fatehizadeh A
    185. Feasey NA
    186. Feizkhah A
    187. Fekadu G
    188. Fischer F
    189. Fitriana I
    190. Forrest KM
    191. Fortuna Rodrigues C
    192. Fuller JE
    193. Gadanya MA
    194. Gajdács M
    195. Gandhi AP
    196. Garcia-Gallo EE
    197. Garrett DO
    198. Gautam RK
    199. Gebregergis MW
    200. Gebrehiwot M
    201. Gebremeskel TG
    202. Geffers C
    203. Georgalis L
    204. Ghazy RM
    205. Golechha M
    206. Golinelli D
    207. Gordon M
    208. Gulati S
    209. Gupta RD
    210. Gupta S
    211. Gupta VK
    212. Habteyohannes AD
    213. Haller S
    214. Harapan H
    215. Harrison ML
    216. Hasaballah AI
    217. Hasan I
    218. Hasan RS
    219. Hasani H
    220. Haselbeck AH
    221. Hasnain MS
    222. Hassan II
    223. Hassan S
    224. Hassan Zadeh Tabatabaei MS
    225. Hayat K
    226. He J
    227. Hegazi OE
    228. Heidari M
    229. Hezam K
    230. Holla R
    231. Holm M
    232. Hopkins H
    233. Hossain MM
    234. Hosseinzadeh M
    235. Hostiuc S
    236. Hussein NR
    237. Huy LD
    238. Ibáñez-Prada ED
    239. Ikiroma A
    240. Ilic IM
    241. Islam SMS
    242. Ismail F
    243. Ismail NE
    244. Iwu CD
    245. Iwu-Jaja CJ
    246. Jafarzadeh A
    247. Jaiteh F
    248. Jalilzadeh Yengejeh R
    249. Jamora RDG
    250. Javidnia J
    251. Jawaid T
    252. Jenney AWJ
    253. Jeon HJ
    254. Jokar M
    255. Jomehzadeh N
    256. Joo T
    257. Joseph N
    258. Kamal Z
    259. Kanmodi KK
    260. Kantar RS
    261. Kapisi JA
    262. Karaye IM
    263. Khader YS
    264. Khajuria H
    265. Khalid N
    266. Khamesipour F
    267. Khan A
    268. Khan MJ
    269. Khan MT
    270. Khanal V
    271. Khidri FF
    272. Khubchandani J
    273. Khusuwan S
    274. Kim MS
    275. Kisa A
    276. Korshunov VA
    277. Krapp F
    278. Krumkamp R
    279. Kuddus M
    280. Kulimbet M
    281. Kumar D
    282. Kumaran EAP
    283. Kuttikkattu A
    284. Kyu HH
    285. Landires I
    286. Lawal BK
    287. Le TTT
    288. Lederer IM
    289. Lee M
    290. Lee SW
    291. Lepape A
    292. Lerango TL
    293. Ligade VS
    294. Lim C
    295. Lim SS
    296. Limenh LW
    297. Liu C
    298. Liu X
    299. Liu X
    300. Loftus MJ
    301. M Amin HI
    302. Maass KL
    303. Maharaj SB
    304. Mahmoud MA
    305. Maikanti-Charalampous P
    306. Makram OM
    307. Malhotra K
    308. Malik AA
    309. Mandilara GD
    310. Marks F
    311. Martinez-Guerra BA
    312. Martorell M
    313. Masoumi-Asl H
    314. Mathioudakis AG
    315. May J
    316. McHugh TA
    317. Meiring J
    318. Meles HN
    319. Melese A
    320. Melese EB
    321. Minervini G
    322. Mohamed NS
    323. Mohammed S
    324. Mohan S
    325. Mokdad AH
    326. Monasta L
    327. Moodi Ghalibaf A
    328. Moore CE
    329. Moradi Y
    330. Mossialos E
    331. Mougin V
    332. Mukoro GD
    333. Mulita F
    334. Muller-Pebody B
    335. Murillo-Zamora E
    336. Musa S
    337. Musicha P
    338. Musila LA
    339. Muthupandian S
    340. Nagarajan AJ
    341. Naghavi P
    342. Nainu F
    343. Nair TS
    344. Najmuldeen HHR
    345. Natto ZS
    346. Nauman J
    347. Nayak BP
    348. Nchanji GT
    349. Ndishimye P
    350. Negoi I
    351. Negoi RI
    352. Nejadghaderi SA
    353. Nguyen QP
    354. Noman EA
    355. Nwakanma DC
    356. O’Brien S
    357. Ochoa TJ
    358. Odetokun IA
    359. Ogundijo OA
    360. Ojo-Akosile TR
    361. Okeke SR
    362. Okonji OC
    363. Olagunju AT
    364. Olivas-Martinez A
    365. Olorukooba AA
    366. Olwoch P
    367. Onyedibe KI
    368. Ortiz-Brizuela E
    369. Osuolale O
    370. Ounchanum P
    371. Oyeyemi OT
    372. P A MP
    373. Paredes JL
    374. Parikh RR
    375. Patel J
    376. Patil S
    377. Pawar S
    378. Peleg AY
    379. Peprah P
    380. Perdigão J
    381. Perrone C
    382. Petcu I-R
    383. Phommasone K
    384. Piracha ZZ
    385. Poddighe D
    386. Pollard AJ
    387. Poluru R
    388. Ponce-De-Leon A
    389. Puvvula J
    390. Qamar FN
    391. Qasim NH
    392. Rafai CD
    393. Raghav P
    394. Rahbarnia L
    395. Rahim F
    396. Rahimi-Movaghar V
    397. Rahman M
    398. Rahman MA
    399. Ramadan H
    400. Ramasamy SK
    401. Ramesh PS
    402. Ramteke PW
    403. Rana RK
    404. Rani U
    405. Rashidi M-M
    406. Rathish D
    407. Rattanavong S
    408. Rawaf S
    409. Redwan EMM
    410. Reyes LF
    411. Roberts T
    412. Robotham JV
    413. Rosenthal VD
    414. Ross AG
    415. Roy N
    416. Rudd KE
    417. Sabet CJ
    418. Saddik BA
    419. Saeb MR
    420. Saeed U
    421. Saeedi Moghaddam S
    422. Saengchan W
    423. Safaei M
    424. Saghazadeh A
    425. Saheb Sharif-Askari N
    426. Sahebkar A
    427. Sahoo SS
    428. Sahu M
    429. Saki M
    430. Salam N
    431. Saleem Z
    432. Saleh MA
    433. Samodra YL
    434. Samy AM
    435. Saravanan A
    436. Satpathy M
    437. Schumacher AE
    438. Sedighi M
    439. Seekaew S
    440. Shafie M
    441. Shah PA
    442. Shahid S
    443. Shahwan MJ
    444. Shakoor S
    445. Shalev N
    446. Shamim MA
    447. Shamshirgaran MA
    448. Shamsi A
    449. Sharifan A
    450. Shastry RP
    451. Shetty M
    452. Shittu A
    453. Shrestha S
    454. Siddig EE
    455. Sideroglou T
    456. Sifuentes-Osornio J
    457. Silva LMLR
    458. Simões EAF
    459. Simpson AJH
    460. Singh A
    461. Singh S
    462. Sinto R
    463. Soliman SSM
    464. Soraneh S
    465. Stoesser N
    466. Stoeva TZ
    467. Swain CK
    468. Szarpak L
    469. T Y SS
    470. Tabatabai S
    471. Tabche C
    472. Taha ZM-A
    473. Tan K-K
    474. Tasak N
    475. Tat NY
    476. Thaiprakong A
    477. Thangaraju P
    478. Tigoi CC
    479. Tiwari K
    480. Tovani-Palone MR
    481. Tran TH
    482. Tumurkhuu M
    483. Turner P
    484. Udoakang AJ
    485. Udoh A
    486. Ullah N
    487. Ullah S
    488. Vaithinathan AG
    489. Valenti M
    490. Vos T
    491. Vu HTL
    492. Waheed Y
    493. Walker AS
    494. Walson JL
    495. Wangrangsimakul T
    496. Weerakoon KG
    497. Wertheim HFL
    498. Williams PCM
    499. Wolde AA
    500. Wozniak TM
    501. Wu F
    502. Wu Z
    503. Yadav MKK
    504. Yaghoubi S
    505. Yahaya ZS
    506. Yarahmadi A
    507. Yezli S
    508. Yismaw YE
    509. Yon DK
    510. Yuan C-W
    511. Yusuf H
    512. Zakham F
    513. Zamagni G
    514. Zhang H
    515. Zhang Z-J
    516. Zielińska M
    517. Zumla A
    518. Zyoud SHH
    519. Zyoud SH
    520. Hay SI
    521. Stergachis A
    522. Sartorius B
    523. Cooper BS
    524. Dolecek C
    525. Murray CJL
    (2024) Global burden of bacterial antimicrobial resistance 1990–2021: a systematic analysis with forecasts to 2050
    The Lancet 404:1199–1226.
    https://doi.org/10.1016/S0140-6736(24)01867-1
    1. Piracha Z
    2. Saeed U
    3. Khurshid A
    4. Chaudhary WN
    (2014)
    Isolation and partial characterization of virulent phage specific against Pseudomonas aeruginosa
    Global Journal of Medical Research 14:1–8.
  3. Software
    1. R Development Core Team
    (2010) R: a language and environment for statistical computing
    R Foundation for Statistical Computing, Vienna, Austria.
  4. Book
    1. Van Twest R
    2. Kropinski AM
    (2009) Bacteriophage enrichment from water and soil
    In: Clokie MRJ, Kropinski AM, editors. Bacteriophages. Humana Press. pp. 15–21.
    https://doi.org/10.1007/978-1-60327-164-6_2

Article and author information

Author details

  1. Camilla Do

    Department of Molecular Virology and Microbiology, Baylor College of Medicine, Houston, United States
    Contribution
    Conceptualization, Resources, Data curation, Formal analysis, Validation, Investigation, Visualization, Methodology, Writing – original draft, Writing – review and editing
    For correspondence
    camilla.do@bcm.edu
    Competing interests
    No competing interests declared
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0002-8798-908X
  2. Keiko Christine Salazar

    1. Department of Molecular Virology and Microbiology, Baylor College of Medicine, Houston, United States
    2. TAILΦR LABS, Baylor College of Medicine, Houston, United States
    Contribution
    Resources, Writing – review and editing
    Competing interests
    No competing interests declared
  3. James D Chang

    Department of Molecular Virology and Microbiology, Baylor College of Medicine, Houston, United States
    Contribution
    Resources, Formal analysis, Writing – review and editing
    Competing interests
    No competing interests declared
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0003-3211-3299
  4. Justin R Clark

    1. Department of Molecular Virology and Microbiology, Baylor College of Medicine, Houston, United States
    2. TAILΦR LABS, Baylor College of Medicine, Houston, United States
    Contribution
    Resources, Writing – review and editing
    Competing interests
    No competing interests declared
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0003-1590-6828
  5. Austen Lee Terwilliger

    1. Department of Molecular Virology and Microbiology, Baylor College of Medicine, Houston, United States
    2. TAILΦR LABS, Baylor College of Medicine, Houston, United States
    Contribution
    Resources, Writing – review and editing
    Competing interests
    No competing interests declared
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0001-8740-0290
  6. Paul Ruchhoeft

    Department of Electrical & Computer Engineering, University of Houston, Houston, United States
    Contribution
    Resources, Writing – review and editing
    Competing interests
    No competing interests declared
  7. Paul Nicholls

    Section of Infection Diseases, Department of Medicine, Baylor College of Medicine, Houston, United States
    Contribution
    Conceptualization, Resources, Data curation, Formal analysis, Validation, Investigation, Visualization, Methodology, Writing – review and editing
    Competing interests
    No competing interests declared
  8. Anthony W Maresso

    1. Department of Molecular Virology and Microbiology, Baylor College of Medicine, Houston, United States
    2. TAILΦR LABS, Baylor College of Medicine, Houston, United States
    Contribution
    Resources, Supervision, Funding acquisition, Methodology, Writing – review and editing
    Competing interests
    No competing interests declared
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0002-4452-3490

Funding

National Institute of Health Sciences (5 U19 AI157981)

  • Camilla Do
  • Keiko Christine Salazar
  • James D Chang
  • Justin R Clark
  • Austen Lee Terwilliger
  • Paul Ruchhoeft
  • Paul Nicholls
  • Anthony W Maresso

National Institute of Health Sciences (T32 AI0554413 Fellowship)

  • Paul Nicholls

National Institute of Health Sciences (K08 AI173452-01A1)

  • Paul Nicholls

Baylor College of Medicine (Intramural Funds)

  • Paul Nicholls

Robert J. Kleberg, Jr. and Helen C. Kleberg Foundation (Philanthropic)

  • Camilla Do
  • Keiko Christine Salazar
  • James D Chang
  • Justin R Clark
  • Austen Lee Terwilliger
  • Paul Ruchhoeft
  • Paul Nicholls
  • Anthony W Maresso

Levy-Longenbaugh Fund (Philanthropic)

  • Camilla Do
  • Keiko Christine Salazar
  • James D Chang
  • Justin R Clark
  • Austen Lee Terwilliger
  • Paul Ruchhoeft
  • Paul Nicholls
  • Anthony W Maresso

The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.

Acknowledgements

This work was supported by the National Institutes of Health (grant number 5 U19 AI157981) with additional funding from the Levy-Longenbaugh Fund and Kleberg Foundation. PN received additional support from T32 AI0554413 fellowship, K08 AI173452-01A1, and BCM Intramural funds. We thank Alleigh Nicholls, Tai Gip, Christian Maresso, Dylan Chirman, Ling Qiu, Ellen Vaughan, John Taylor and Rachel Lahowetz for their assistance on in-field sampling. CryoEM data was collected at the Baylor College of Medicine CryoEM ATC, which includes equipment purchased under support of CPRIT Core Facility Award RP190602.

Version history

  1. Sent for peer review:
  2. Preprint posted:
  3. Reviewed Preprint version 1:
  4. Reviewed Preprint version 2:
  5. Version of Record published:

Cite all versions

You can cite all versions using the DOI https://doi.org/10.7554/eLife.109259. This DOI represents all versions, and will always resolve to the latest one.

Copyright

© 2026, Do et al.

This article is distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use and redistribution provided that the original author and source are credited.

Metrics

  • 497
    views
  • 18
    downloads
  • 0
    citations

Views, downloads and citations are aggregated across all versions of this paper published by eLife.

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Camilla Do
  2. Keiko Christine Salazar
  3. James D Chang
  4. Justin R Clark
  5. Austen Lee Terwilliger
  6. Paul Ruchhoeft
  7. Paul Nicholls
  8. Anthony W Maresso
(2026)
Pathogen-phage geomapping to overcome resistance
eLife 15:RP109259.
https://doi.org/10.7554/eLife.109259.3

Share this article

https://doi.org/10.7554/eLife.109259