A new diagnostic modality for bovine tuberculosis: accurate and robust classification of infected cattle using transcriptomics and machine learning

  1. UCD School of Agriculture and Food Science, University College Dublin, Dublin, Ireland
  2. Animal Genomics, ETH Zurich, Universitaetstrasse 2, Zurich, Switzerland
  3. Department of Biomedical Informatics, University of Colorado Anschutz Medical Campus, Aurora, United States
  4. Irish Blood Transfusion Service, National Blood Centre, Dublin, Ireland
  5. Children’s Health Ireland, Dublin, Ireland
  6. The Roslin Institute and Royal (Dick) School of Veterinary Studies, University of Edinburgh, Edinburgh, United Kingdom
  7. Centre for Tropical Livestock Genetics and Health (CTLGH), Roslin Institute, University of Edinburgh, Edinburgh, United Kingdom
  8. European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Cambridge United Kingdom
  9. Animal Genomics, ETH Zurich, Zurich, Switzerland
  10. UCD Conway Institute of Biomolecular and Biomedical Research, University College Dublin, Dublin, Ireland
  11. UCD One Health Centre, University College Dublin, Dublin, Ireland
  12. UCD Institute of Food and Health, University College Dublin, Dublin, Ireland
  13. UCD School of Mathematics and Statistics, University College Dublin, Dublin, Ireland
  14. UCD School of Veterinary Medicine, University College Dublin, Dublin, Ireland

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Bavesh Kana
    University of the Witwatersrand, Johannesburg, South Africa
  • Senior Editor
    Bavesh Kana
    University of the Witwatersrand, Johannesburg, South Africa

Reviewer #1 (Public review):

Summary:

The control of bovine tuberculosis in managed populations such as Ireland and Great Britain is unusual in that demonstrably sick animals are rarely, if ever, seen in herds. Control is therefore focused on the identification and removal of animals that test positive to the tuberculin skin test (the legal definition of infection). Despite over a century of study, the relationship between tuberculin test status, infection and most importantly infectiousness is still poorly quantified. Different formats of the tuberculin skin test are acknowledged to have both poor sensitivity and compromised specificity, although the characteristics of these tests are likely to vary considerably between contexts due to both biological variation and discretion in measurements by testers. There is an urgent need for new, more reliable and cheaper diagnostics to address the failures of existing control programs and to enable control in emerging markets that do not currently control the disease.

Strengths:

A key strength of this study is the use of samples from both naturally infected and experimentally infected animals. This data set is used to perform a careful and exhaustive evaluation of the extent to which patterns of transcriptomic expression can be used to classify between disease free animals and those infected with bovine tuberculosis.

The experimentally infected animal samples provide evidence that expression patterns of infected animals vary with respect to the time from infection. The authors highlight that this suggests transcriptomic markers may be able to detect infection earlier than tuberculin and IGRA tests that target cell-mediated immune responses. However, these methods could potentially provide a valuable new tool for quantifying the role of individual variation and progression for a disease where the individual life-history is still frustratingly mysterious.

Weaknesses:

However, the high levels of individual variation - and in particular differences in patterns of expression between naturally and experimentally infected animals do raise questions about how diagnostic tests developed from these tools would be used in practice. In particular, while many of the models considered achieved high sensitivity - estimated specificity is consistently lower than current diagnostic tests and considerably lower than that necessary for screening tests given the frequency of testing carried out as part of statutory control programs.

Expanding the number of samples may help to address these issues, but I would have liked to see some discussion of the extent to which the level of biological variation observed in this study may limit the precision of diagnostic tests developed using these tools. Given the likely characteristics of tests based on these methods, I would be interested to hear how the authors think they could fit within current statutory programs, either as supplementary or replacement tests?

Reviewer #2 (Public review):

Summary:

This study evaluates whether peripheral blood transcriptomic profiles can be used to classify cattle infected with Mycobacterium bovis using a range of machine-learning approaches. By integrating RNA-seq datasets from naturally and experimentally infected animals, the authors develop and test predictive models capable of distinguishing infected from uninfected cattle and assess their ability to differentiate bovine tuberculosis from other infectious diseases. The study addresses an important challenge in bovine tuberculosis control and presents evidence that host transcriptional signatures may have utility as an adjunct diagnostic approach.

Strengths:

- The study combines data from multiple independent cohorts, including both naturally and experimentally infected cattle, which increases the biological relevance of the findings.

- The analytical workflow is comprehensive, scientifically sound and clearly described. Multiple machine-learning approaches are evaluated and compared rather than relying on a single modelling strategy.

- The inclusion of a held-out test set, especially because such data is limited, provides a useful assessment of model performance beyond cross-validation alone.
- Thorough evaluation against datasets from cattle infected with MAP, BoHV-1 and BRSV is a valuable addition and provides useful information regarding the specificity of the identified transcriptional signatures.

- The authors acknowledge important limitations, including batch effects and the need for additional validation.

- All underlying data and code are made publicly available

Weaknesses:

- My main concern relates to generalisability. Although a separate testing dataset was used, the training and testing datasets were generated through random partitioning of samples from the same underlying studies. As a result, classifier performance in a completely independent external cohort remains unclear. Discussion of this limitation, and whether alternative validation strategies such as leave-one-study-out analyses were considered, would strengthen the manuscript.

- The authors identify substantial study-specific batch effects following dataset integration and appropriately account for these in the modelling framework. However, given the magnitude of the reported batch structure, additional discussion regarding the potential influence of residual between-study variation on classifier performance would be helpful.

- The manuscript is framed in the context of global bovine tuberculosis control, yet the practical implementation of a transcriptomic diagnostic approach is not discussed in great detail. Since bovine tuberculosis remains a significant challenge in many low- and middle-income settings, further consideration of the feasibility, cost, infrastructure requirements, and potential translation of these signatures into more deployable diagnostic platforms would improve the broader relevance of the study.

- The datasets used for classifier development are derived primarily from Ireland, the UK and the United States. It would be useful to discuss whether differences in circulating M. bovis lineages, cattle populations, management systems, or co-infection pressures could influence host transcriptional responses and therefore the performance of the proposed classifiers in other epidemiological settings. This ties to the previous comment, since epidemiological settings in LMIC countries with a high burden of M.bovis disease would be vastly different from where the data was sourced.

Reviewer #3 (Public review):

Summary:

This is an excellent piece of work which sheds greater light on the responses of cattle to both experimental and natural infection in cattle with Mycobacterium bovis infection, using data from different experimental and field groups.

Strengths:

The work is based on robust analysis of a range of highly relevant experimental and field sample sets, using transcriptomic approaches. It provides insight into pathogenesis and disease responses, as well as some evidence regarding potential future diagnostic advances.

Weaknesses:

I have some simple, but important comments on how the work is discussed. (Consequently, most comments focus on the discussion section). In particular, I suggest that the authors have confused or conflated the great progress that they have made in improving the understanding of the responses to M. bovis infection with an improved ability to practically improve the diagnosis of the infection in the field. Minor differences in specificity and sensitivity - and predictive values - of tests or assays being used can have profound effects in different prevalence settings on the farm, and I don't feel that this understanding is adequately reflected in the discussion in particular. In this respect, the authors really should, in my view, focus not on the outstanding results that come in or from their model fitting approaches, to focus on their model testing results in relation to extrapolated meaning / external validity. The text uses words like 'robust' and 'highly accurate' which are meaningless in the context of test interpretation in the field.

I get their enthusiasm, based on really interesting findings in relation to disease progression and immune and inflammatory responses, but these indistinct claims rather devalue the quality of the rest of their work in my view.

One great challenge in work of this nature, using natural cases from farms, is that there is no gold standard for identifying the cases that current diagnostic approaches miss and which they hope their new approaches can help with. This is not discussed or mentioned in their enthusiasm for what they have achieved. Their field datasets are based on the current, insensitively detected cases. It misses the 'occult' cases that are present but undiagnosed.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation