Proceedings · Session S-133 · filed September 30, 2026

Physical Sciences ResearchSession paper

Machine learning classifies tuberculosis from breath VOCs at 93% accuracy

BIUST physicists classified active TB, MDR-TB and healthy controls from breath GC-MS data at 93% accuracy using an SVM on 1,867 raw spectral features from 87 samples.

By Amara Osei3 min read615 words

Summary

  • A linear SVM classified active TB, MDR-TB and healthy cohorts from breath GC-MS data with 93% accuracy, using 1,867 spectral features from 87 samples.
  • The study was published in Discover Artificial Intelligence by George Chimowa's group at BIUST, Botswana, with South African collaborators.
  • The model located the most diagnostic VOCs — including tridecane, decane and O-cymene — in a 10–30 minute retention-time window, enabling specialized, higher-throughput diagnostic hardware.
AI sniffs out tuberculosis in breath’s chemical signature
FigureAI sniffs out tuberculosis in breath’s chemical signature — AI-generated

A linear support vector machine has distinguished active tuberculosis, multidrug-resistant tuberculosis and healthy controls in breath samples with 93% accuracy, using 1,867 raw spectral features from 87 samples with no aggressive pre-processing. The work, published in Discover Artificial Intelligence by George Chimowa's Nano-Breath Taking group at Botswana International University of Science and Technology (BIUST), together with collaborators in South Africa, points toward breath-based TB diagnostics that could cut instrument cost and throughput requirements in settings with limited laboratory infrastructure.

The analytical problem. Exhaled breath carries thousands of volatile organic compounds (VOCs) produced by human metabolism and by pathogens such as Mycobacterium tuberculosis. Gas chromatography–mass spectrometry separates these mixtures by volatility and mass-to-charge ratio, but a single breath sample generates thousands of overlapping spectral peaks. Manual annotation of that data is slow and has long been the bottleneck between sampling and diagnosis.

What the team did. The researchers collected breath samples from three cohorts: patients with active TB, patients with multidrug-resistant (MDR) TB, and healthy volunteers. Rather than applying the smoothing and simplification steps standard in GC-MS pipelines — steps that risk erasing low-concentration molecular features with diagnostic value — they fed the high-dimensional dataset directly into four supervised algorithms: decision trees, random forest, k-nearest neighbours and support vector machines (SVM).

The linear SVM came out on top at 93% accuracy across the three-way classification task. The reason is structural. Distance-metric models degrade in high-dimensional spaces, where experimental noise gets misread as diagnostic signal. SVM instead performs geometric boundary optimization: it identifies a hyperplane maximizing the margin between classes and depends only on the sparse subset of boundary points — the support vectors — rather than the full dataset. That sparsity insulates the classifier from noise in the surrounding feature space.

A hardware-relevant result. The more commercially significant finding concerns where the diagnostic information sits. The model pinpointed a retention-time window between 10 and 30 minutes in the gas chromatogram as the region containing the highest-variance, most diagnostic VOCs. These are medium- and large-molecule compounds, including TB-associated biomarkers such as tridecane, decane and O-cymene.

For instrument developers, that concentration of signal in a 20-minute slice changes the engineering brief. A future diagnostic device would not need to acquire and process the full spectrum; it could be specialized around this elution window, shrinking data-processing loads and maximizing sample throughput in clinical settings.

Context and caveats. TB remains one of the deadliest infectious diseases worldwide, and current diagnosis leans on slow sputum cultures and backlogged laboratory queues that delay treatment. A non-invasive breath test built around an optimized classifier would be a plausible fit for resource-limited clinics lacking full laboratory infrastructure.

The researchers themselves flag the limits of the current evidence. The dataset is 87 samples across three cohorts — small by clinical validation standards — and the 93% figure is a classification result on that cohort set, not a measured clinical sensitivity or specificity from a blinded trial. Larger validation trials are required before any diagnostic claim can transfer to practice.

The group also frames the work as part of a broader methodological trend, drawing a parallel to data-modelling techniques used in engineering, such as predicting structural degradation in self-healing aerospace composites. The common thread is fusing analytical physics with machine learning to extract decision-grade signal from complex physical measurements.

If larger cohorts confirm the biomarker window, the path forward is a specialized, automated bedside breath analyzer driven by the mathematical profile of a patient's exhaled VOCs — with the BIUST results defining the retention-time range such hardware would need to cover.

via arise.aasciences.app (Original)

Filed under

  • machine-learning
  • tuberculosis
  • gc-ms
  • breath-diagnostics
  • vocs
Share this article:

More from Amara Osei

Amara Osei

Show full bio

News editor covering business strategy at Hypothesis Wire.

80 articles

References

  1. Researchers Find Hundreds of Altered Western Blots in Thermo Fisher Catalogue
  2. ICON Moves AI Agents Into Clinical-Trial Production With Anthropic and Microsoft
  3. JNC Reports 91% Virus Recovery with Large-Pore Cellulose Resin at BPI 2026
  4. Quantum Simulation Passes 12,000-Atom Mark; Lab Filters Flagged
  5. Talus Bio Releases Structure-Free AI Model for Disordered Proteome

« Previous articleNext article »