<?xml version="1.0" ?><!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.3 20210610//EN"  "JATS-archivearticle1-mathml3.dtd"><article xmlns:ali="http://www.niso.org/schemas/ali/1.0/" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.3" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="nlm-ta">elife</journal-id>
<journal-id journal-id-type="publisher-id">eLife</journal-id>
<journal-title-group>
<journal-title>eLife</journal-title>
</journal-title-group>
<issn publication-format="electronic" pub-type="epub">2050-084X</issn>
<publisher>
<publisher-name>eLife Sciences Publications, Ltd</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">89862</article-id>
<article-id pub-id-type="doi">10.7554/eLife.89862</article-id>
<article-id pub-id-type="doi" specific-use="version">10.7554/eLife.89862.1</article-id>
<article-version-alternatives>
<article-version article-version-type="publication-state">reviewed preprint</article-version>
<article-version article-version-type="preprint-version">1.2</article-version>
</article-version-alternatives>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Ecology</subject>
</subj-group>
<subj-group subj-group-type="heading">
<subject>Computational and Systems Biology</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Microbes with higher metabolic independence are enriched in human gut microbiomes under stress</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0003-2390-5286</contrib-id>
<name>
<surname>Veseli</surname>
<given-names>Iva</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
<xref ref-type="aff" rid="a2">2</xref>
<xref ref-type="corresp" rid="cor1">*</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Chen</surname>
<given-names>Yiqun T.</given-names>
</name>
<xref ref-type="aff" rid="a3">3</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0002-8435-3203</contrib-id>
<name>
<surname>Schechter</surname>
<given-names>Matthew S.</given-names>
</name>
<xref ref-type="aff" rid="a2">2</xref>
<xref ref-type="aff" rid="a4">4</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0002-1124-1147</contrib-id>
<name>
<surname>Vanni</surname>
<given-names>Chiara</given-names>
</name>
<xref ref-type="aff" rid="a5">5</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0002-8957-9922</contrib-id>
<name>
<surname>Fogarty</surname>
<given-names>Emily C.</given-names>
</name>
<xref ref-type="aff" rid="a2">2</xref>
<xref ref-type="aff" rid="a4">4</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0003-0128-6795</contrib-id>
<name>
<surname>Watson</surname>
<given-names>Andrea R.</given-names>
</name>
<xref ref-type="aff" rid="a2">2</xref>
<xref ref-type="aff" rid="a4">4</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0001-7427-4424</contrib-id>
<name>
<surname>Jabri</surname>
<given-names>Bana</given-names>
</name>
<xref ref-type="aff" rid="a2">2</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0003-3218-613X</contrib-id>
<name>
<surname>Blekhman</surname>
<given-names>Ran</given-names>
</name>
<xref ref-type="aff" rid="a2">2</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0002-2802-4317</contrib-id>
<name>
<surname>Willis</surname>
<given-names>Amy D.</given-names>
</name>
<xref ref-type="aff" rid="a6">6</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yu</surname>
<given-names>Michael K.</given-names>
</name>
<xref ref-type="aff" rid="a7">7</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0002-8679-490X</contrib-id>
<name>
<surname>Fernàndez-Guerra</surname>
<given-names>Antonio</given-names>
</name>
<xref ref-type="aff" rid="a8">8</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0002-4210-2318</contrib-id>
<name>
<surname>Füssel</surname>
<given-names>Jessika</given-names>
</name>
<xref ref-type="aff" rid="a2">2</xref>
<xref ref-type="aff" rid="a9">9</xref>
<xref ref-type="corresp" rid="cor1">*</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0001-9013-4827</contrib-id>
<name>
<surname>Eren</surname>
<given-names>A. Murat</given-names>
</name>
<xref ref-type="aff" rid="a2">2</xref>
<xref ref-type="aff" rid="a9">9</xref>
<xref ref-type="aff" rid="a10">10</xref>
<xref ref-type="aff" rid="a11">11</xref>
<xref ref-type="aff" rid="a12">12</xref>
<xref ref-type="corresp" rid="cor1">*</xref>
</contrib>
<aff id="a1"><label>1</label><institution>Biophysical Sciences Program, The University of Chicago</institution>, Chicago, IL 60637, <country>USA</country></aff>
<aff id="a2"><label>2</label><institution>Department of Medicine, The University of Chicago</institution>, Chicago, IL 60637, <country>USA</country></aff>
<aff id="a3"><label>3</label><institution>Data Science Institute and Department of Biomedical Data Science, Stanford University</institution>, Stanford, CA, 94305, <country>USA</country></aff>
<aff id="a4"><label>4</label><institution>Committee on Microbiology, The University of Chicago</institution>, Chicago, IL 60637, <country>USA</country></aff>
<aff id="a5"><label>5</label><institution>MARUM Center for Marine Environmental Sciences, University of Bremen</institution>, Bremen, <institution>Germany;</institution></aff>
<aff id="a6"><label>6</label><institution>Department of Biostatistics, University of Washington</institution>, Seattle, WA, 98195, <country>USA</country></aff>
<aff id="a7"><label>7</label><institution>Toyota Technological Institute at Chicago</institution>, Chicago, IL 60605, <country>USA</country></aff>
<aff id="a8"><label>8</label><institution>Lundbeck Foundation GeoGenetics Centre, GLOBE Institute, University of Copenhagen</institution>, Copenhagen, <country>Denmark</country></aff>
<aff id="a9"><label>9</label><institution>Institute for Chemistry and Biology of the Marine Environment, University of Oldenburg</institution>, Oldenburg, <country>Germany</country></aff>
<aff id="a10"><label>10</label><institution>Marine ‘Omics Bridging Group, Max Planck Institute for Marine Microbiology</institution>, 28359 Bremen, <country>Germany</country></aff>
<aff id="a11"><label>11</label><institution>Alfred Wegener Institute for Polar and Marine Research</institution>, Bremerhaven, <country>Germany</country></aff>
<aff id="a12"><label>12</label><institution>Helmholtz Institute for Functional Marine Biodiversity</institution>, 26129, Oldenburg, <country>Germany</country></aff>
</contrib-group>
<contrib-group content-type="section">
<contrib contrib-type="editor">
<name>
<surname>Turnbaugh</surname>
<given-names>Peter J</given-names>
</name>
<role>Reviewing Editor</role>
<aff>
<institution-wrap>
<institution>University of California, San Francisco</institution>
</institution-wrap>
<city>San Francisco</city>
<country>United States of America</country>
</aff>
</contrib>
<contrib contrib-type="senior_editor">
<name>
<surname>Garrett</surname>
<given-names>Wendy S</given-names>
</name>
<role>Senior Editor</role>
<aff>
<institution-wrap>
<institution>Harvard T.H. Chan School of Public Health</institution>
</institution-wrap>
<city>Boston</city>
<country>United States of America</country>
</aff>
</contrib>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>*</label> Corresponding authors: <email>iveseli@uchicago.edu</email>, <email>jessika.fuessel@uni-oldenburg.de</email>, <email>meren@hifmb.de</email></corresp>
</author-notes>
<pub-date date-type="original-publication" iso-8601-date="2023-09-08">
<day>08</day>
<month>09</month>
<year>2023</year>
</pub-date>
<volume>12</volume>
<elocation-id>RP89862</elocation-id>
<history>
<date date-type="sent-for-review" iso-8601-date="2023-06-14">
<day>14</day>
<month>06</month>
<year>2023</year>
</date>
</history>
<pub-history>
<event>
<event-desc>Preprint posted</event-desc>
<date date-type="preprint" iso-8601-date="2023-05-26">
<day>26</day>
<month>05</month>
<year>2023</year>
</date>
<self-uri content-type="preprint" xlink:href="https://doi.org/10.1101/2023.05.10.540289"/>
</event>
</pub-history>
<permissions>
<copyright-statement>© 2023, Veseli et al</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Veseli et al</copyright-holder>
<ali:free_to_read/>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<ali:license_ref>https://creativecommons.org/licenses/by/4.0/</ali:license_ref>
<license-p>This article is distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License</ext-link>, which permits unrestricted use and redistribution provided that the original author and source are credited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="elife-preprint-89862-v1.pdf"/>
<abstract>
<title>Abstract</title>
<p>A wide variety of human diseases are associated with loss of microbial diversity in the human gut, inspiring a great interest in the diagnostic or therapeutic potential of the microbiota. However, the ecological forces that drive diversity reduction in disease states remain unclear, rendering it difficult to ascertain the role of the microbiota in disease emergence or severity. One hypothesis to explain this phenomenon is that microbial diversity is diminished as disease states select for microbial populations that are more fit to survive environmental stress caused by inflammation or other host factors. Here, we tested this hypothesis on a large scale, by developing a software framework to quantify the enrichment of microbial metabolisms in complex metagenomes as a function of microbial diversity. We applied this framework to over 400 gut metagenomes from individuals who are healthy or diagnosed with inflammatory bowel disease (IBD). We found that high metabolic independence (HMI) is a distinguishing characteristic of microbial communities associated with individuals diagnosed with IBD. A classifier we trained using the normalized copy numbers of 33 HMI-associated metabolic modules not only distinguished states of health versus IBD, but also tracked the recovery of the gut microbiome following antibiotic treatment, suggesting that HMI is a hallmark of microbial communities in stressed gut environments.</p>
</abstract>
<kwd-group kwd-group-type="author">
<title>Keywords</title>
<kwd>microbial metabolism</kwd>
<kwd>metabolic reconstruction</kwd>
<kwd>inflammatory bowel disease</kwd>
<kwd>gut microbiome</kwd>
</kwd-group>

</article-meta>
<notes>
<notes notes-type="competing-interest-statement">
<title>Competing Interest Statement</title><p>The authors have declared no competing interest.</p></notes>
<fn-group content-type="summary-of-updates">
<title>Summary of Updates:</title>
<fn fn-type="update"><p>Updated abstract; revised section on key biosynthetic pathways and section on classifier to clarify interpretation and significance of results; added summary paragraph prior to conclusion; revised conclusion; Supplementary Information file updated.</p></fn>
</fn-group>
<fn-group content-type="external-links">
<fn fn-type="dataset"><p>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.6084/m9.figshare.22679080">https://doi.org/10.6084/m9.figshare.22679080</ext-link>
</p></fn>
<fn fn-type="dataset"><p>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.6084/m9.figshare.22776701">https://doi.org/10.6084/m9.figshare.22776701</ext-link>
</p></fn>
<fn fn-type="dataset"><p>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.7872967">https://doi.org/10.5281/zenodo.7872967</ext-link>
</p></fn>
<fn fn-type="dataset"><p>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.7883421">https://doi.org/10.5281/zenodo.7883421</ext-link>
</p></fn>
<fn fn-type="dataset"><p>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.7897987">https://doi.org/10.5281/zenodo.7897987</ext-link>
</p></fn>
</fn-group>
</notes>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>The human gut is home to a diverse assemblage of microbial cells that form complex communities (<xref ref-type="bibr" rid="c14">Coyte, Schluter, and Foster 2015</xref>). This gut microbial ecosystem is established almost immediately after birth and plays a lifelong role in human wellbeing by contributing to immune system maturation and functioning (<xref ref-type="bibr" rid="c5">Belkaid and Hand 2014</xref>; <xref ref-type="bibr" rid="c64">Maynard et al. 2012</xref>), extracting dietary nutrients (<xref ref-type="bibr" rid="c33">Hijova 2019</xref>), providing protection against pathogens (<xref ref-type="bibr" rid="c45">Khosravi and Mazmanian 2013</xref>), metabolizing drugs (M. <xref ref-type="bibr" rid="c108">Zimmermann et al. 2019</xref>), and more (<xref ref-type="bibr" rid="c46">Knight et al. 2017</xref>). There is no universal definition of a healthy gut microbiome (<xref ref-type="bibr" rid="c20">Fan and Pedersen 2021</xref>), but associations between host disease states and changes in microbial community composition have sparked great interest in the therapeutic potential of gut microbes (<xref ref-type="bibr" rid="c10">Cani 2018</xref>; <xref ref-type="bibr" rid="c96">Sorbara and Pamer 2022</xref>) and led to the emergence of hypotheses that directly link disruptions of the gut microbiome to non-communicable diseases of complex etiology (<xref ref-type="bibr" rid="c9">Byndloss and Bäumler 2018</xref>).</p>
<p>Inflammatory bowel diseases (IBDs), which describe a heterogeneous group of chronic inflammatory disorders (<xref ref-type="bibr" rid="c94">Shan, Lee, and Chang 2022</xref>), represent an increasingly common health risk around the globe (<xref ref-type="bibr" rid="c41">Kaplan 2015</xref>). Understanding the role of gut microbiota in IBD has been a major area of focus in human microbiome research. Studies focusing on individual microbial taxa that typically change in relative abundance in IBD patients have proposed a range of host-microbe interactions that may contribute to disease manifestation and progression (<xref ref-type="bibr" rid="c38">Joossens et al. 2011</xref>; <xref ref-type="bibr" rid="c90">Schirmer et al. 2019</xref>; <xref ref-type="bibr" rid="c32">Henke et al. 2019</xref>; <xref ref-type="bibr" rid="c60">Machiels et al. 2014</xref>). However, even within well-constrained cohorts, a large proportion of variability in the taxonomic composition of the microbiota is unexplained, and the proportion of variability explained by disease status is low (<xref ref-type="bibr" rid="c28">Gevers et al. 2014</xref>; <xref ref-type="bibr" rid="c89">Schirmer et al. 2018</xref>; <xref ref-type="bibr" rid="c57">Lloyd-Price et al. 2019</xref>; <xref ref-type="bibr" rid="c44">Khan et al. 2019</xref>). As neither individual taxa nor broad changes in microbial community composition yield effective predictors of disease (<xref ref-type="bibr" rid="c47">Knox et al. 2019</xref>; M. <xref ref-type="bibr" rid="c54">Lee and Chang 2021</xref>), the role of gut microbes in the etiology of IBD – or the extent to which they are bystanders to disease – remains unclear (<xref ref-type="bibr" rid="c44">Khan et al. 2019</xref>).</p>
<p>The marked decrease in microbial diversity in IBD is often associated with the loss of Firmicutes populations and an increased representation of a relatively small number of taxa, such as <italic>Bacteroides</italic>, <italic>Enterococcaceae</italic>, and others (<xref ref-type="bibr" rid="c79">Prindiville et al. 2000</xref>; <xref ref-type="bibr" rid="c87">Saitoh et al. 2002</xref>; Sartor 2006; <xref ref-type="bibr" rid="c86">Rhodes 2007</xref>; <xref ref-type="bibr" rid="c16">Devkota et al. 2012</xref>; <xref ref-type="bibr" rid="c60">Machiels et al. 2014</xref>; <xref ref-type="bibr" rid="c99">Vineis et al. 2016</xref>; <xref ref-type="bibr" rid="c57">Lloyd-Price et al. 2019</xref>). Why a handful of taxa that also typically occur in healthy individuals in lower abundances (M. <xref ref-type="bibr" rid="c54">Lee and Chang 2021</xref>; <xref ref-type="bibr" rid="c69">Nishida et al. 2018</xref>) tend to dominate the IBD microbiome is a fundamental but open question to gain insights into the ecological underpinnings of the gut microbial ecosystem under IBD. Going beyond taxonomic summaries, a recent metagenome-wide metabolic modeling study revealed a significant loss of cross-feeding partners as a hallmark of IBD, where microbial interactions were disrupted in IBD-associated microbial communities compared to those found in healthy individuals (<xref ref-type="bibr" rid="c62">Marcelino et al. 2023</xref>). This observation is in line with another recent work that proposed that the extent of ‘metabolic independence’ (characterized by the genomic presence of a set of key metabolic modules for the synthesis of essential nutrients) is a determinant of microbial survival in IBD (<xref ref-type="bibr" rid="c100">Watson et al. 2023</xref>). It is conceivable that the disrupted metabolic interactions among microbes observed in IBD (<xref ref-type="bibr" rid="c62">Marcelino et al. 2023</xref>) indicates an environment that lacks the ecosystem services provided by a complex network of microbial interactions, and selects for those organisms that harness high metabolic independence (HMI) (<xref ref-type="bibr" rid="c100">Watson et al. 2023</xref>). This interpretation offers an ecological mechanism to explain the dominance of populations with specific metabolic features in IBD. However, this proposed mechanism warrants further investigation.</p>
<p>Here we implemented a high-throughput strategy to estimate metabolic capabilities of microbial communities directly from metagenomes and investigate whether the enrichment of populations with high metabolic independence predicts IBD in the human gut. We benchmarked our findings using representative genomes associated with the human gut and their distribution in healthy individuals and those who have been diagnosed with IBD. Our results suggest that high metabolic potential (indicated by a set of 33 largely biosynthetic metabolic pathways) provides enough signal to consistently distinguish gut microbiomes under stress from those that are in homeostasis, providing deeper insights into adaptive processes initiated by stress conditions that promote the dominance of rare members of gut microbiota during disease.</p>
</sec>
<sec id="s2">
<title>Results and Discussion</title>
<p>We compiled 2,893 publicly-available stool metagenomes from 13 different studies, 5 of which explicitly studied the IBD gut microbiome (Supplementary Table 1a-c). The average sequencing depth varied across individual datasets (4.2 Mbp to 60.3 Mbp, with a median value of 21.4 Mbp, Supplementary Table 1c). To improve the sensitivity and accuracy of our downstream analyses that depend on metagenomic assembly, we excluded samples with less than 25 million reads, resulting in a set of 408 relatively deeply-sequenced metagenomes from 10 studies (26.4 Mbp to 61.9 Mbp, with a median value of 37.0 Mbp, Supplementary Table 1b, Supplementary Information, Methods), which we <italic>de novo</italic> assembled individually. The final dataset included individuals who were healthy (n=229), diagnosed with IBD (n=101), or suffered from other gastrointestinal conditions (“non-IBD”, n=78). In accordance with previous observations of reduced microbial diversity in IBD (<xref ref-type="bibr" rid="c50">Kostic, Xavier, and Gevers 2014</xref>; <xref ref-type="bibr" rid="c68">Nagalingam and Lynch 2012</xref>; <xref ref-type="bibr" rid="c47">Knox et al. 2019</xref>), the estimated number of populations based on the occurrence of bacterial single-copy core genes present in these metagenomes was higher in healthy individuals than those diagnosed with IBD (<xref rid="figs1" ref-type="fig">Supplementary Figure 1</xref>, Supplementary Table 1).</p>
<sec id="s2a">
<title>Estimating normalized copy numbers of metabolic pathways from metagenomic assemblies</title>
<p>Gaining insights into microbial metabolism requires accurate estimates of pathway presence/absence and completion. While a myriad of tools address this task for single genomes (<xref ref-type="bibr" rid="c59">Machado et al. 2018</xref>; <xref ref-type="bibr" rid="c4">Aziz et al. 2008</xref>; <xref ref-type="bibr" rid="c2">Arkin et al. 2018</xref>; <xref ref-type="bibr" rid="c72">Palù et al. 2022</xref>; <xref ref-type="bibr" rid="c92">Shaffer et al. 2020</xref>; <xref ref-type="bibr" rid="c27">Geller-McGrath et al. 2023</xref>; <xref ref-type="bibr" rid="c110">Zorrilla et al. 2021</xref>; Zhou et al., n.d.; J. <xref ref-type="bibr" rid="c107">Zimmermann, Kaleta, and Waschina 2021</xref>), working with complex environmental metagenomes poses additional challenges due to the large number of organisms that are present in metagenomic assemblies. A few tools can estimate community-level metabolic potential from metagenomes without relying on the reconstruction of individual population genomes or reference-based approaches (<xref ref-type="bibr" rid="c105">Ye and Doak 2009</xref>; <xref ref-type="bibr" rid="c42">Karp et al. 2021</xref>) (Supplementary Table 5). These high-level summaries of pathway presence and redundancy in a given environment are suitable for most surveys of metabolic capacity, particularly for microbial communities of similar richness. However, since the frequency of observed metabolic modules increases as microbial diversity increases, investigations of metabolic determinants of survival across environmental conditions with substantial differences in microbial richness requires quantitative insights into the extent of enrichment of metabolic capabilities in relation to microbial diversity. For instance, the estimated copy number of a given metabolic module may be identical between two metagenomes, but one metagenome can have a lower alpha diversity and thus have a higher selection for this module. To quantify the differential abundance of metabolic modules between metagenomes generated from healthy individuals and those from individuals diagnosed with IBD, we implemented a new software framework (<ext-link ext-link-type="uri" xlink:href="https://anvio.org/m/anvi-estimate-metabolism">https://anvio.org/m/anvi-estimate-metabolism</ext-link>) that reconstructs metabolic modules from genomes and metagenomes and then calculates the per-population copy number (PPCN) of modules in metagenomes (Methods, Supplementary Information). Briefly, the PPCN estimates the proportion of microbes in a community with a particular metabolic capacity (<xref rid="fig1" ref-type="fig">Figure 1</xref>, <xref rid="figs2" ref-type="fig">Supplementary Figure 2</xref>). We estimate the number of microbial populations using single-copy core genes (SCGs) instead of reconstructing individual genomes first, thus maximizing the <italic>de novo</italic> recovery of gene content.</p>
<fig id="fig1" position="float" orientation="portrait" fig-type="figure">
<label>Figure 1.</label>
<caption><title>Conceptual diagram of per-population copy number (PPCN) calculation.</title>
<p>Each step of the calculation is demonstrated in (<bold>A</bold>) for a sample with high diversity (6 microbial populations) and in (<bold>B</bold>) for a sample with low diversity (3 populations). Metagenome sequences are shown as black lines. The left panel shows the single-copy core genes annotated in the metagenome (indicated by letters), with a barplot showing the counts for different SCGs. The dashed black line indicates the mode of the counts, which is taken as the estimate of the number of populations. The middle panel shows the annotations of metabolic pathways (indicated by boxes and numerically labeled), with a barplot showing the copy number of each pathway (for more details on how this copy number is computed, see Supplementary Information and <xref rid="figs2" ref-type="fig">Supplementary Figure 2</xref>). The right panel shows the equation for per-population copy number (PPCN), with the barplots indicating the PPCN values for each metabolic pathway in each sample and arrows differentiating between different types of modules based on the comparison of their normalized copy numbers between samples.</p></caption>
<graphic xlink:href="540289v2_fig1.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
</sec>
<sec id="s2b">
<title>Key biosynthetic pathways are enriched in microbial populations from IBD samples</title>
<p>To gain insight into potential metabolic determinants of microbial survival in the IBD gut environment, we assessed the distribution of metabolic modules within samples from each group (IBD and healthy) with and without using PPCN normalization. A set of 33 metabolic modules were significantly enriched in metagenomes obtained from individuals diagnosed with IBD when PPCN normalization was applied (<xref rid="fig2" ref-type="fig">Figure 2d</xref>, 2e). Each metabolic module had an FDR-adjusted p &lt; 2e-10 and an effect size &gt; 0.12 from a Wilcoxon Rank Sum Test comparing IBD and healthy samples. The set included 17 modules that were previously associated with high metabolic independence (<xref ref-type="bibr" rid="c100">Watson et al. 2023</xref>) (<xref rid="fig2" ref-type="fig">Figure 2f</xref>). However, without PPCN normalization, the signal was masked by the overall higher copy numbers in healthy samples, and the same analysis did not detect higher metabolic potential in microbial populations associated with individuals diagnosed with IBD (<xref rid="fig2" ref-type="fig">Figure 2a</xref>), showing weaker differential occurrence between cohorts (<xref rid="fig2" ref-type="fig">Figure 2b</xref>, 2c, <xref rid="figs3" ref-type="fig">Supplementary Figure 3</xref>). This result suggests that the PPCN normalization is an essential step in comparative analyses of metabolisms between samples with disparate levels of diversity.</p>
<fig id="fig2" position="float" orientation="portrait" fig-type="figure">
<label>Figure 2.</label>
<caption><title>Comparison of metabolic potential across healthy and IBD cohorts.</title>
<p>Panels <bold>A – C</bold> show unnormalized copy number data and the remaining panels show normalized per-population copy number (PPCN) data. <bold>A)</bold> Scatterplot of module copy number in IBD samples (x-axis) and healthy samples (y-axis). Size of points indicates fraction of samples in which the module has a non-zero copy number, transparency of points indicates the p-value of the module in a Wilcoxon Rank Sum test for enrichment (based on PPCN data), and color indicates whether the module is enriched in the IBD samples (in this study), enriched in the good colonizers from the fecal microbiota transplant (FMT) study (<xref ref-type="bibr" rid="c100">Watson et al. 2023</xref>), or enriched in both. <bold>B)</bold> Heatmap of unnormalized copy numbers for all modules. IBD-enriched modules are highlighted by the red bar on the left. Sample group is indicated by the blue (healthy) and red (IBD) bars on the bottom. <bold>C)</bold> Boxplots of median copy number for each module enriched in the FMT colonizers from (<xref ref-type="bibr" rid="c100">Watson et al. 2023</xref>) in the healthy samples (blue) and the IBD samples (red). Solid lines connect the same module in each plot. <bold>D)</bold> Scatterplot of module PPCN values in IBD samples (x-axis) and healthy samples (y-axis). Size, transparency and color of points are defined as in panel (A). The pink dashed line indicates the effect size threshold applied to modules when determining their enrichment in IBD. <bold>E)</bold> Heatmap of PPCN values for all modules. Side bars defined as in (B). <bold>F)</bold> Boxplots of median PPCN values for modules enriched in the FMT colonizers from (<xref ref-type="bibr" rid="c100">Watson et al. 2023</xref>) in the healthy samples (blue) and the IBD samples (red). Lines defined as in (D). Modules that were also enriched in the IBD samples (in this study) are highlighted in red. <bold>G)</bold> Boxplots of PPCN values for individual modules in the healthy samples (blue) and the IBD samples (red). All example modules were enriched in both this study and in (<xref ref-type="bibr" rid="c100">Watson et al. 2023</xref>). A high-resolution version of this figure is available at <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.6084/m9.figshare.22851173">https://doi.org/10.6084/m9.figshare.22851173</ext-link>.</p></caption>
<graphic xlink:href="540289v2_fig2.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
<p>The majority of the metabolic modules that were enriched in the microbiomes of IBD patients encoded biosynthetic capabilities (23 out of 33) that resolved to amino acid metabolism (33%), carbohydrate metabolism (21%), cofactor and vitamin biosynthesis (15%), nucleotide biosynthesis (12%), lipid biosynthesis (6%) and energy metabolism (6%) (Supplementary Table 2a). In contrast to previous reports based on reference genomes (<xref ref-type="bibr" rid="c28">Gevers et al. 2014</xref>; <xref ref-type="bibr" rid="c67">Morgan et al. 2012</xref>), amino acid synthesis and carbohydrate metabolism were not reduced in the IBD gut microbiome in our dataset. Rather, our results were in accordance with a more recent finding that predicted amino acid secretion potential is increased in the microbiomes of individuals with IBD (<xref ref-type="bibr" rid="c31">Heinken, Hertel, and Thiele 2021</xref>).</p>
<p>The metagenome-level enrichment of several key biosynthesis pathways supports the hypothesis that high metabolic independence (HMI) is a determinant of survival for microbial populations in the IBD gut environment. We investigated whether biosynthetic capacity in general was enriched in IBD samples, and 62 out of 88 (70%) biosynthesis pathways described in the KEGG database had a significant enrichment in the IBD sample group at an FDR-adjusted 5% significance level (<xref rid="figs5" ref-type="fig">Supplementary Figure 5d</xref>). However, a similar proportion of non-biosynthetic pathways, 63 out of 91 (69%), were also significantly increased in the IBD samples. While biosynthetic capacity is not over-represented in the IBD sample group compared to other types of metabolism, the high proportion of enriched pathways associated with biosynthesis suggests that biosynthetic capacity is important for microbial resilience.</p>
<p>Within our set of 33 pathways that were enriched in IBD, it is notable that all the biosynthesis and central carbohydrate pathways are directly or indirectly linked via shared enzymes and metabolites. Each enriched module shared on average 25.6% of its enzymes and 40.2% of metabolites with the other enriched modules, and overall 18.2% of enzymes and 20.4% of compounds across these pathways were shared (Supplementary Table 2a). Thus, modules may be enriched not just due to the importance of their immediate end products, but also because of their role in the larger metabolic network. The few standalone modules that were enriched included the efflux pump MepA and the beta-Lactam resistance system, which are associated with drug resistance. These capacities may provide an advantage since antibiotics are a common treatment for IBDs (<xref ref-type="bibr" rid="c70">Nitzan et al. 2016</xref>), but are not related to the systematic enrichment of biosynthesis pathways that likely provide resilience to general environmental stress rather than to a specific stressor such as antibiotics.</p>
<p>While so far we divided samples into two groups, our dataset also includes individuals who do not suffer from IBD, yet are not healthy either. A recent study using flux balance analysis to model metabolite secretion potential in the dysbiotic, non-dysbiotic, and control gut communities of Crohn’s Disease patients found that several predicted microbial metabolic activities align with gradients of host health (<xref ref-type="bibr" rid="c31">Heinken, Hertel, and Thiele 2021</xref>). To test whether the HMI signal captures gradients in host health, we included the ‘non-IBD’ group of patients that suffer from gastrointestinal conditions other than IBD in our analysis. The set of 78 samples classified as ‘non-IBD’ indeed represent an intermediate group between healthy individuals and those diagnosed with IBD (<xref rid="figs5" ref-type="fig">Supplementary Figure 5b</xref>). While the HMI signal was reduced in ‘non-IBD’ patients, 75% of the pathways enriched in IBD patients were also enriched in the ‘non-IBD’ group compared to healthy individuals. Similarly, when sorting each individual cohort along a health gradient based on cohort descriptions in their respective studies (Supplementary Information), the relative proportion of metabolic pathways indicative of HMI increased as a function of increasing disease severity (<xref rid="figs6" ref-type="fig">Supplementary Figure 6a</xref>). These findings suggest that the HMI signal is sufficiently sensitive to resolve gradients in host health and could serve as a diagnostic tool to monitor changing stress levels in a single individual over time.</p>
<p>Microbiome data generated by different groups can result in systematic biases that may outweigh biological differences between otherwise similar samples (<xref ref-type="bibr" rid="c58">Lozupone et al. 2013</xref>; <xref ref-type="bibr" rid="c95">Sinha et al. 2017</xref>; <xref ref-type="bibr" rid="c13">Clausen and Willis 2022</xref>). The potential impact of such biases constitutes an important consideration for meta-analyses such as ours that analyze publicly available metagenomes from multiple sources. To account for cohort biases, we conducted an analysis of our data on a per-cohort basis, which showed robust differences between the sample groups across multiple cohorts (<xref rid="figs6" ref-type="fig">Supplementary Figure 6b</xref>, 6c). Another source of potential bias in our results is due to the representation of microbial functions in genomes in publicly available databases. For instance, we noticed that, independent of the annotation strategy, a smaller proportion of genes resolved to known functions in metagenomic assemblies of the healthy samples compared to the assemblies we generated from the IBD group (<xref rid="figs4" ref-type="fig">Supplementary Figure 4</xref>). This highlights the possibility that healthy samples merely appear to harbor less metabolic capabilities due to missing annotations. Indeed, we found that the normalized copy numbers of most metabolic modules were reduced in the healthy group, where 84% of KEGG modules (98 out of 118) have significantly lower median copy numbers (<xref rid="figs5" ref-type="fig">Supplementary Figure 5c</xref>, Supplementary Information). While the presence of a bias between the two cohorts is clear, the source of this bias and its implications are not as clear. One hypothesis that could explain this phenomenon is that the increased proportion of unknown functions in environments where populations with low metabolic independence (LMI) thrive is due to our inability to identify distant homologs of even well-studied functions in poorly studied novel genomes through public databases. If true, this would indeed impair our ability to annotate genes using state-of-the-art functional databases, and bias metabolic module completion estimates. Such a limitation would indeed warrant a careful reconsideration of common workflows and studies that rely on public resources to characterize gene function in complex environments. Another hypothesis that could explain our observation is that the general absence in culture of microbes with smaller genomes (that likely fare better in diverse gut ecosystems) had a historical impact on the characterization of novel functions that represent a relatively larger fraction of their gene repertoire. If true, this would suggest that the unknown functions are unlikely essential for well-studied metabolic capabilities. Furthermore, HMI and LMI genomes may be indistinguishable with respect to the distribution of such novel genes, but the increased number of genes in HMI genomes that resolve to well-studied metabolisms would reduce the proportion of known functions in LMI genomes, and thus in metagenomes where they thrive. While testing these hypotheses falls outside the scope of our work, we find the latter hypothesis more likely due to examples in existing literature that have successfully identified genes that belong to known metabolisms in some of the most obscure organisms via annotation strategies similar to those we have used in our work (<xref ref-type="bibr" rid="c37">Jaffe et al. 2020</xref>; <xref ref-type="bibr" rid="c21">Farag et al. 2020</xref>).</p>
<p>Taken together, these results (1) demonstrate that the PPCN normalization is an important consideration for investigations of metabolic enrichment in complex microbial communities as a function of microbial diversity, and (2) reveal that the enrichment of HMI populations in an environment offers a high-resolution marker to resolve different levels of environmental stress.</p>
</sec>
<sec id="s2c">
<title>Reference genomes with higher metabolic independence are over-represented in the gut metagenomes of individuals with IBD</title>
<p>So far, our findings demonstrate an overall, metagenome-level trend of increasing HMI within gut microbial communities as a function of IBD status without considering the individual genomes that contribute to this signal. Since the extent of metabolic independence of a microbial genome is a quantifiable trait, we considered a genome-based approach to validating our findings. Given the metagenome-level trends, we expected that the microbial genomes that encode a high number of metabolic modules associated with HMI should be more commonly detected in metagenomes from individuals diagnosed with IBD.</p>
<p>While publicly available reference genomes for microbial taxa will unlikely capture the diversity of individual gut metagenomes, we cast a broad net by surveying the ecology of 19,226 genomes in the Genome Taxonomy Database (GTDB) (<xref ref-type="bibr" rid="c75">Parks et al. 2022</xref>) that belong to three major phyla associated with the human gut environment: Bacteroidetes, Firmicutes, and Proteobacteria (<xref ref-type="bibr" rid="c103">Woting and Blaut 2016</xref>; <xref ref-type="bibr" rid="c98">Turnbaugh et al. 2009</xref>). We used Human Microbiome Project data (Human Microbiome Project Consortium 2012) to characterize the distribution of these genomes across healthy human gut metagenomes. We used their single-copy core genes to identify genomes that were representative of microbial clades that are systematically detected in the healthy human gut (<xref rid="fig3" ref-type="fig">Figure 3a</xref>) and kept those that also occurred in at least 2% of samples from our set of 330 healthy and IBD metagenomes (see Methods). Selection of genomes that are relatively well-detected in the HMP dataset effectively removed taxa that primarily occur outside of the human gut. Of the final set of 338 reference genomes that passed our filters, 258 (76.3%) resolved to Firmicutes, 60 (17.8%) to Bacteroidetes, and 20 (5.9%) to Proteobacteria. Most of these genomes resolved to families common to the colonic microbiota, such as <italic>Lachnospiraceae</italic> (30.0%), <italic>Ruminococcaceae</italic> / <italic>Oscillospiraceae</italic> (23.1%), and <italic>Bacteroidaceae</italic> (10.1%) (<xref ref-type="bibr" rid="c3">Arumugam et al. 2011</xref>), while 5.9% belonged to poorly-studied families with temporary code names (Supplementary Table 3a). Finally, we performed a more comprehensive read recruitment analysis on this smaller set of genomes using all deeply-sequenced metagenomes from cohorts that included healthy, non-IBD, and IBD samples (<xref rid="fig3" ref-type="fig">Figure 3</xref>). This provided us with a quantitative summary of the detection patterns of GTDB genome representatives common to the human gut across our dataset.</p>
<fig id="fig3" position="float" orientation="portrait" fig-type="figure">
<label>Figure 3.</label>
<caption><title>Identification of HMI genomes and their distribution across gut samples.</title>
<p><bold>A)</bold> Histogram of Ribosomal Protein S6 gene clusters (94% ANI) for which at least 50% of the representative gene sequence is covered by at least 1 read (&gt;= 50% ‘detection’) in fecal metagenomes from the Human Microbiome Project (HMP) (Human Microbiome Project Consortium 2012). The dashed line indicates our threshold for reaching at least 50% detection in at least 10% of the HMP samples; gray bars indicate the 11,145 gene clusters that do not meet this threshold while purple bars indicate the 836 clusters that do. The subplot shows data for the 836 genomes whose Ribosomal Protein S6 sequences belonged to one of the passing (purple) gene clusters. The y-axis indicates the number of healthy/IBD gut metagenomes from our set of 330 in which the full genome sequence has at least 50% detection, and the x-axis indicates the genome’s maximum detection across all 330 samples. The dashed line indicates our threshold for reaching at least 50% genome detection in at least 2% of samples; the 338 genomes that pass this threshold are tan and those that do not are purple. The phylogeny of these 338 genomes is shown in <bold>B)</bold> along with the following data, from top to bottom: taxonomic classification as assigned by GTDB; proportion of healthy samples with at least 50% detection of the genome sequence; proportion of IBD samples with at least 50% detection of the genome sequence; square-root normalized ratio of percent abundance in IBD samples to percent abundance in healthy samples; metabolic independence score (sum of completeness scores of 33 HMI-associated metabolic pathways); whether (red) or not (white) the genome is classified as having HMI with a threshold score of 26.4; heatmap of completeness scores for each of the 33 HMI-associated metabolic pathways (0% completeness is white and 100% completeness is black). Pathway name is shown on the right and colored according to its category of metabolism. <bold>C)</bold> Boxplot showing the proportion of healthy (blue) or IBD (red) samples in which genomes of each class are detected &gt;= 50%, with p-values from a Wilcoxon Rank-Sum test on the underlying data. <bold>D)</bold> Barplot showing the proportion of detected genomes (with &gt;= 50% genome sequence covered by at least 1 read) in each sample that are classified as HMI, for each group of samples. The black lines show the median for each group: 37.0% for IBD samples, 25.5% for non-IBD samples, and 18.4% for healthy samples. A high-resolution version of this figure is available at <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.6084/m9.figshare.22851173">https://doi.org/10.6084/m9.figshare.22851173</ext-link>.</p></caption>
<graphic xlink:href="540289v2_fig3.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
<p>We classified each genome as HMI if its average completeness of the 33 HMI-associated metabolic pathways was at least 80%, equivalent to a summed metabolic independence score of 26.4 (Methods). Across all genomes, the mean metabolic independence score was 24.0 (Q1: 19.9, Q3: 25.7). We identified 17.5% (59) of the reference genomes as HMI. HMI genomes were on average substantially larger (3.8 Mbp) than non-HMI genomes (2.9 Mbp) and encoded more genes (3,634 vs. 2,683 genes, respectively), which is in accordance with the reduced metabolic potential of non-HMI populations (Supplementary Table 3a). Our read recruitment analysis showed that HMI reference genomes were present in a significantly higher proportion of IBD samples compared to non-HMI genomes (<xref rid="fig3" ref-type="fig">Figure 3c</xref>, p &lt; 1e-5, Wilcoxon Rank Sum test). Similarly, the fraction of HMI populations was significantly higher within a given IBD sample compared to samples classified as ‘non-IBD’ and those from healthy individuals (<xref rid="fig3" ref-type="fig">Figure 3d</xref>, p &lt; 1e-24, Kruskal-Wallis Rank Sum test). In contrast, the detection of HMI populations and non-HMI populations was similar in healthy individuals (<xref rid="fig3" ref-type="fig">Figure 3c</xref>, p = 0.267, Wilcoxon Rank Sum test). The intestinal environment of healthy individuals likely supports both HMI and non-HMI populations, wherein ‘metabolic diversity’ is maintained by metabolic interactions such as cross feeding. Indeed, loss of cross-feeding interactions in the gut microbiome appears to be associated with a number of human diseases, including IBD (<xref ref-type="bibr" rid="c62">Marcelino et al. 2023</xref>). This interpretation is further supported by the fact that the top two HMI-associated pathways are required for the synthesis of cobalamin from glutamate. Auxotrophy for cobalamin biosynthesis is common among gut bacteria that rely on cross-feeding for this essential cofactor ((<xref ref-type="bibr" rid="c15">Degnan et al. 2014</xref>; <xref ref-type="bibr" rid="c61">Magnúsdóttir et al. 2015</xref>; <xref ref-type="bibr" rid="c43">Kelly et al. 2019</xref>)) (Supplementary information).</p>
<p>Overall, the classification of reference gut genomes as HMI and their enrichment in individuals diagnosed with IBD strongly supports the contribution of HMI to stress resilience of individual microbial populations. We note that survival in a disturbed gut environment will likely require a wide variety of additional functions that are not covered in the list of metabolic modules we consider to determine HMI status – for examples, see (<xref ref-type="bibr" rid="c15">Degnan et al. 2014</xref>; <xref ref-type="bibr" rid="c63">Martens et al. 2014</xref>; <xref ref-type="bibr" rid="c109">Zong et al. 2020</xref>; L. <xref ref-type="bibr" rid="c22">Feng et al. 2020</xref>; <xref ref-type="bibr" rid="c29">Goodman et al. 2009</xref>; <xref ref-type="bibr" rid="c78">Powell et al. 2016</xref>). Indeed, there may be many ways for a microbe to be metabolically independent, and our strategy likely failed to identify some HMI populations. Nonetheless, these data suggest that HMI serves as a reliable proxy for the identification of microbial populations that are particularly resilient.</p>
</sec>
<sec id="s2d">
<title>HMI-associated metabolic potential predicts general stress on gut microbes</title>
<p>Our analysis identified HMI as an emergent property of gut microbial communities associated with individuals diagnosed with IBD. This community-level signal translates to individual microbial populations and provides insights into the microbial ecology of stressed gut environments. HMI-associated metabolic pathways were enriched at the community level, and microbial populations encoding these modules were more prevalent in individuals with IBD than in healthy individuals. Furthermore, the copy number of these pathways and the proportion of HMI populations reflect the severity of environmental stress and translate to host health states (<xref rid="figs5" ref-type="fig">Supplementary Figure 5b</xref>, <xref rid="fig3" ref-type="fig">Figure 3d</xref>). The ecological implications of these observations suggest that HMI may serve as a predictor of general stress in the human gut environment.</p>
<p>So far, efforts to diagnose IBD using microbial markers have presented classifiers based on (1) taxonomy in pediatric IBD patients (<xref ref-type="bibr" rid="c73">Papa et al. 2012</xref>; <xref ref-type="bibr" rid="c28">Gevers et al. 2014</xref>), (2) community composition in combination with clinical data (<xref ref-type="bibr" rid="c30">Halfvarson et al. 2017</xref>), (3) untargeted metabolomics and/or species-level relative abundance from metagenomes (<xref ref-type="bibr" rid="c25">Franzosa et al. 2019</xref>) and (4) k-mer-based sequence variants in metagenomes that can be linked to microbial genomes associated with IBD (<xref ref-type="bibr" rid="c85">Reiter et al. 2022</xref>). Performance varied both between and within studies according to the target classes and data types used for training and validation of each classifier (Supplementary Table 4a). For those studies reporting accuracy, a maximum accuracy of 77% was achieved based on either metabolite profiles (for prediction of IBD-subtype) (<xref ref-type="bibr" rid="c25">Franzosa et al. 2019</xref>) or k-mer-based sequence variants (for differentiating between IBD and non-IBD samples) (<xref ref-type="bibr" rid="c85">Reiter et al. 2022</xref>). Some studies reported performance as area under the receiver operating characteristic curve (AUROCC), a typical measure of classifier utility describing both sensitivity (ability to correctly identify the disease) and specificity (ability to correctly identify absence of disease). For this metric the highest value was 0.92, achieved by (<xref ref-type="bibr" rid="c25">Franzosa et al. 2019</xref>) when using metabolite profiles, with or without species abundance data, for classifying IBD vs non-IBD. However, the majority of these classifiers were trained and tested on relatively small groups of individuals that all come from the same region, i.e. clinical studies confined to a specific hospital. Though some had high performance, they either relied on data that are inaccessible to most laboratories and clinics considering that untargeted metabolomics analyses are difficult to reproduce (<xref ref-type="bibr" rid="c48">Koek et al. 2011</xref>; <xref ref-type="bibr" rid="c56">Lin et al. 2020</xref>), or they required complex k-mer-based models without the resolution to differentiate gradients in host health (<xref ref-type="bibr" rid="c85">Reiter et al. 2022</xref>). These classifiers thus have limited translational potential across global clinical settings and do not provide an ecological framework to explain the observed shifts in community composition and activity. For practical use as a diagnostic tool, a microbiome-based classifier for IBD should rely on an ecologically meaningful, easy to measure, and high-level signal that is robust to host variables like lifestyle, geographical location, and ethnicity. High metabolic independence could potentially fill this gap as a metric related to the ecological filtering that defines microbial community changes in the IBD gut microbiome.</p>
<p>We trained a logistic regression classifier to explore the applicability of HMI as a non-invasive diagnostic tool for IBD. The classifier’s predictors were the per-population copy numbers of IBD-enriched metabolic pathways in a given metagenome. Across the 330 deeply-sequenced IBD and healthy samples included in this analysis, the classifier had high sensitivity and specificity (<xref rid="fig4" ref-type="fig">Figure 4</xref>). It correctly identified (on average) 76.8% of samples from individuals diagnosed with IBD and 89.5% of samples representing healthy individuals, for an overall accuracy of 85.6% and an average AUROCC of 0.832 (Supplementary Table 4c). Our model outperforms (<xref ref-type="bibr" rid="c28">Gevers et al. 2014</xref>; <xref ref-type="bibr" rid="c30">Halfvarson et al. 2017</xref>; <xref ref-type="bibr" rid="c85">Reiter et al. 2022</xref>) or has comparable performance to (<xref ref-type="bibr" rid="c25">Franzosa et al. 2019</xref>; <xref ref-type="bibr" rid="c73">Papa et al. 2012</xref>) the previous attempts to classify IBD from fecal samples in more restrictively-defined cohorts. It also has the advantage of being a simple model, utilizing a relatively low number of features compared to the other classifiers. Thus, HMI shows promise as an accessible diagnostic marker of IBD. Due to the lack of time-series studies that include individuals in the pre-diagnosis phase of IBD development, we cannot test the applicability of HMI to predict IBD onset (<xref ref-type="bibr" rid="c57">Lloyd-Price et al. 2019</xref>).</p>
<fig id="fig4" position="float" orientation="portrait" fig-type="figure">
<label>Figure 4.</label>
<caption><title>Performance of our metagenome classifier trained on per-population copy numbers of IBD-enriched modules.</title>
<p><bold>A)</bold> Receiver operating characteristic (ROC) curves for 25-fold cross-validation. Each fold used a random subset of 80% of the data for training and the other 20% for testing. In each fold, we calculated a set of IBD-enriched modules from the training dataset and used the PPCN of these modules to train a logistic regression model whose performance was evaluated using the test dataset. Light gray lines show the ROC curve for each fold, the dark blue line shows the mean ROC curve, the gray area delineates the confidence interval for the mean ROC, and the pink dashed line indicates the benchmark performance of a naive (random guess) classifier. <bold>B)</bold> Confusion matrix for each fold of the random cross-validation. Categories of classification, from top left to bottom right, are: true positives (correctly classified IBD samples), false positives (incorrectly classified Healthy samples), false negatives (incorrectly classified IBD samples), and true negatives (correctly classified Healthy samples). Each fold is represented by a box within each category. Opacity of the box indicates the proportion of samples in that category, and the actual proportion is written within the box with one significant digit. Underlying data for this matrix can be accessed in Supplementary Table 4d.</p></caption>
<graphic xlink:href="540289v2_fig4.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
<p>Yet, the gradient of metabolic independence reflected by per-population pathway copy number and the relative increase in the number of HMI populations detected in non-IBD samples (<xref rid="figs5" ref-type="fig">Supplementary Figure 5b</xref>, <xref rid="fig3" ref-type="fig">Figure 3d</xref>) suggests that the degree of HMI in the gut microbiome may be indicative of general gut stress, such as the stress induced by antibiotic use. Antibiotics can cause long-lasting perturbations of the gut microbiome – including reduced diversity, emergence of opportunistic pathogens, increased microbial load, and development of highly-resistant strains – with potential implications for host health (<xref ref-type="bibr" rid="c82">Ramirez et al. 2020</xref>). We applied our metabolism classifier to a metagenomic dataset that reflects the changes in the microbiome of healthy people before, during and up to 6 months following a 4-day antibiotic treatment (<xref ref-type="bibr" rid="c71">Palleja et al. 2018</xref>). The resulting pattern of sample classification corresponds to the post-treatment decline and subsequent recovery of species richness documented in the study by (<xref ref-type="bibr" rid="c71">Palleja et al. 2018</xref>).</p>
<p>All pre-treatment samples were classified as ‘healthy’ followed by a decline in the proportion of ‘healthy’ samples to a minimum 8 days post-treatment, and a gradual increase until 180 days post treatment, when over 90% of samples were classified as ‘healthy’ (<xref rid="fig5" ref-type="fig">Figure 5</xref>, Supplementary Table 4b). These observations support the role of HMI as an ecological driver of microbial resilience during gut stress caused by a variety of environmental perturbations and demonstrate its diagnostic power in reflecting gut microbiome state.</p>
<fig id="fig5" position="float" orientation="portrait" fig-type="figure">
<label>Figure 5.</label>
<caption><title>Classification results on an antibiotic time-series dataset from</title>
<p>(<xref ref-type="bibr" rid="c71">Palleja et al. 2018</xref>). Note that antibiotic treatment was taken on days 1-4. <bold>A)</bold> Samples collected per subject during the time series. <bold>B)</bold> Species richness data (figure reproduced from (<xref ref-type="bibr" rid="c71">Palleja et al. 2018</xref>)). <bold>C)</bold> Classification of each sample by the metabolism classifier profiled in <xref rid="fig4" ref-type="fig">Figure 4</xref>. Samples with insufficient sequencing depth were not classified. <bold>D)</bold> Proportion of classes assigned to samples per day in the time series.</p></caption>
<graphic xlink:href="540289v2_fig5.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
</sec>
<sec id="s2e">
<title>Conclusions</title>
<p>Overall, our observations that stem from the analysis of hundreds of reference genomes, deeply-sequenced gut metagenomes, and multiple categories of human disease states suggest that environmental stress in the human gut – whether it is associated with inflammation, cancer, or antibiotic use – promotes the survival and relative expansion of microbial populations with high metabolic independence. These results establish HMI as a high-level metric to classify gradients of human health states through the gut microbiota that is robust to ethnic, geographical or lifestyle factors. Taken together with recent evidence that models altered ecological relationships within gut microbiomes under stress due to disrupted metabolic cross-feeding (<xref ref-type="bibr" rid="c31">Heinken, Hertel, and Thiele 2021</xref>; <xref ref-type="bibr" rid="c62">Marcelino et al. 2023</xref>), our data support the hypothesis that the reduction in microbial diversity, or more generally ‘dysbiosis’, is an emergent property of microbial communities responding to disease pathogenesis or other external factors such as antibiotic use that disrupt the gut microbial ecosystem. This paradigm depicts microbes as bystanders by default, rather than perpetrators or drivers of noncommunicable human diseases, and provides an ecological framework to explain the frequently observed reduction in microbial diversity associated with IBD and other noncommunicable human diseases and disorders.</p>
</sec>
</sec>
<sec id="s3">
<title>Methods</title>
<p>A bioinformatics workflow that further details all analyses described below and gives access to reproducible data products is available at the URL <ext-link ext-link-type="uri" xlink:href="https://merenlab.org/data/ibd-gut-metabolism/">https://merenlab.org/data/ibd-gut-metabolism/</ext-link>.</p>
<sec id="s3a">
<title>A new framework for metabolism estimation</title>
<p>We developed a new program ‘anvi-estimate-metabolism’ (<ext-link ext-link-type="uri" xlink:href="https://anvio.org/m/anvi-estimate-metabolism">https://anvio.org/m/anvi-estimate-metabolism</ext-link>), which uses gene annotations to estimate ‘completeness’ and ‘copy number’ of metabolic pathways that are defined in terms of enzyme accession numbers. By default, this tool works on metabolic modules from the KEGG MODULE database (<xref ref-type="bibr" rid="c40">Kanehisa et al. 2012</xref>, <xref ref-type="bibr" rid="c39">2023</xref>) which are defined by KEGG KOfams (<xref ref-type="bibr" rid="c1">Aramaki et al. 2019</xref>), but user-defined modules based on a variety of functional annotation sources are also accepted as input. Completeness estimates describe the percentage of steps (typically, enzymatic reactions) in a given metabolic pathway that are encoded in a genome or a metagenome. Likewise, copy number summarizes the number of distinct sets of enzyme annotations that collectively encode the complete pathway. This program offers two strategies for estimating metabolic potential: a ‘stepwise’ strategy with equivalent treatment for alternative enzymes – i.e, enzymes that can catalyze the same reaction in a given metabolic pathway – and a ‘pathwise’ strategy that accounts for all possible variations of the pathway. The Supplementary Information file includes more information on these two strategies and the completeness/copy number calculations. For the analysis of metagenomes, we used stepwise copy number of KEGG modules. Briefly, the calculation of stepwise copy number is done as follows: the copy number of each step in a pathway (typically, one chemical reaction or conversion) is individually evaluated by translating the step definition into an arithmetic expression that summarizes the number of annotations for each required enzyme. In cases where multiple enzymes or an enzyme complex are needed to catalyze the reaction, we take the minimum number of annotations across these components. In cases where there are alternative enzymes that can each catalyze the reaction individually, we sum the number of annotations for each alternative. Once the copy number of each step is computed, we then calculate the copy number of the entire pathway by taking the minimum copy number across all the individual steps. The use of minimums results in a conservative estimate of pathway copy number such that only copies of the pathway with all enzymes present are counted. For the analysis of genomes, we calculated the stepwise completeness of KEGG modules. This calculation is similar to the one described above for copy number, except that the step definition is translated into a boolean expression that, once evaluated, indicates the presence or absence of each step in the pathway. Then, the completeness of the modules is computed as the proportion of present steps in the pathway.</p>
</sec>
<sec id="s3b">
<title>Metagenomic Datasets and Sample Groups</title>
<p>We acquired publicly-available gut metagenomes from 13 different studies (<xref ref-type="bibr" rid="c52">Le Chatelier et al. 2013</xref>; Q. <xref ref-type="bibr" rid="c23">Feng et al. 2015</xref>; <xref ref-type="bibr" rid="c25">Franzosa et al. 2019</xref>; <xref ref-type="bibr" rid="c57">Lloyd-Price et al. 2019</xref>; <xref ref-type="bibr" rid="c80">Qin et al. 2012</xref>; <xref ref-type="bibr" rid="c81">Quince et al. 2015</xref>; <xref ref-type="bibr" rid="c83">Rampelli et al. 2015</xref>; <xref ref-type="bibr" rid="c84">Raymond et al. 2016</xref>; <xref ref-type="bibr" rid="c89">Schirmer et al. 2018</xref>; <xref ref-type="bibr" rid="c99">Vineis et al. 2016</xref>; “BioProject” n.d.; <xref ref-type="bibr" rid="c101">Wen et al. 2017</xref>; <xref ref-type="bibr" rid="c104">Xie et al. 2016</xref>). The studies were chosen based on the following criteria: (1) they included shotgun metagenomes of fecal matter (primarily stool, but some ileal pouch luminal aspirate samples (<xref ref-type="bibr" rid="c99">Vineis et al. 2016</xref>) are also included); (2) they sampled from people living in industrialized countries (in the case where a study (<xref ref-type="bibr" rid="c83">Rampelli et al. 2015</xref>) included samples from hunter-gatherer populations, only the samples from industrialized areas were included in our analysis); (3) they included samples from people with IBD and/or they included samples from people without gastrointestinal (GI) disease or inflammation; and (4) clear metadata differentiating between case and control samples was available. A full description of the studies and samples can be found in Supplementary Table 1a-c. We grouped samples according to the health status of the sample donor. Briefly, the ‘IBD’ group of samples includes those from people diagnosed with Crohn’s disease (CD), ulcerative colitis (UC), or pouchitis. The ‘non-IBD’ group contains non-IBD controls, which includes both healthy people presenting for routine cancer screenings as well as people with benign or non-specific symptoms that are not clinically diagnosed with IBD. Colorectal cancer patients from (Q. <xref ref-type="bibr" rid="c23">Feng et al. 2015</xref>) were also put into the ‘non-IBD’ group on the basis that tumors in the GI tract may arise from local inflammation (<xref ref-type="bibr" rid="c51">Kraus and Arber 2009</xref>) and represent a source of gut stress without an accompanying diagnosis of IBD. Finally, the ‘HEALTHY’ group contains samples from people without GI-related diseases or inflammation. Note that only control or pre-treatment samples were taken from the studies covering type 2 diabetes (<xref ref-type="bibr" rid="c80">Qin et al. 2012</xref>), ankylosing spondylitis (<xref ref-type="bibr" rid="c101">Wen et al. 2017</xref>), antibiotic treatment (<xref ref-type="bibr" rid="c84">Raymond et al. 2016</xref>), and dietary intervention (“BioProject” n.d.); these controls were all assigned to the ‘HEALTHY’ group. At least one study (<xref ref-type="bibr" rid="c52">Le Chatelier et al. 2013</xref>) included samples from obese people, and these were also included in the ‘HEALTHY’ group.</p>
</sec>
<sec id="s3c">
<title>Processing of metagenomes</title>
<p>We made single assemblies of most gut metagenomes using the anvi’o metagenomics workflow implemented in the program ‘anvi-run-workflow’ (<xref ref-type="bibr" rid="c93">Shaiber et al. 2020</xref>). This workflow uses Snakemake (<xref ref-type="bibr" rid="c49">Köster and Rahmann 2012</xref>), and a tutorial is available at the URL <ext-link ext-link-type="uri" xlink:href="https://merenlab.org/2018/07/09/anvio-snakemake-workflows/">https://merenlab.org/2018/07/09/anvio-snakemake-workflows/</ext-link>. Briefly, the workflow includes quality filtering using ‘iu-filter-quality-minochè (<xref ref-type="bibr" rid="c19">Eren et al. 2013</xref>); assembly with IDBA-UD (<xref ref-type="bibr" rid="c77">Peng et al. 2012</xref>) (using a minimum contig length of 1000); gene calling with Prodigal v2.6.3 (<xref ref-type="bibr" rid="c36">Hyatt et al. 2010</xref>); tRNA identification with tRNAscan-SE v2.0.7 (<xref ref-type="bibr" rid="c12">Chan and Lowe 2019</xref>); and gene annotation of ribosomal proteins (<xref ref-type="bibr" rid="c91">Seemann n.d.</xref>), single-copy core gene sets (M. D. <xref ref-type="bibr" rid="c53">Lee 2019</xref>), KEGG KOfams (<xref ref-type="bibr" rid="c1">Aramaki et al. 2019</xref>), NCBI COGs (<xref ref-type="bibr" rid="c26">Galperin et al. 2021</xref>), and Pfam (release 33.1, (<xref ref-type="bibr" rid="c66">Mistry et al. 2021</xref>)). The aforementioned annotation was done with programs that relied on HMMER v3.3.2 (<xref ref-type="bibr" rid="c17">Eddy 2011</xref>) as well as Diamond v0.9.14.115 (<xref ref-type="bibr" rid="c8">Buchfink, Xie, and Huson 2015</xref>). As part of this workflow, all single assemblies were converted into anvi’o contigs databases. Samples from (<xref ref-type="bibr" rid="c99">Vineis et al. 2016</xref>) were processed differently because they contained merged reads rather than individual paired-end reads: no further quality filtering was run on these samples, we assembled them individually using MEGAHIT (<xref ref-type="bibr" rid="c55">Li et al. 2015</xref>), and we used the anvi’o contigs workflow to perform all subsequent steps described for the metagenomics workflow above. Note that we used a version of KEGG downloaded in December 2020 (for reproducibility, the hash of the KEGG snapshot available via ‘anvi-setup-kegg-kofams’ is 45b7cc2e4fdc). Additionally, the annotation program ‘anvi-run-kegg-kofams’ includes a heuristic for annotating hits with bitscores that are just below the KEGG-defined threshold, which is described at <ext-link ext-link-type="uri" xlink:href="https://anvio.org/m/anvi-run-kegg-kofams/">https://anvio.org/m/anvi-run-kegg-kofams/</ext-link>.</p>
</sec>
<sec id="s3d">
<title>Genomic Dataset</title>
<p>We also analyzed microbial genomes from the Genome Taxonomy Database (GTDB), release 95.0 (<xref ref-type="bibr" rid="c76">Parks et al. 2018</xref>, <xref ref-type="bibr" rid="c74">2020</xref>). We downloaded all reference genome sequences for the species cluster representatives.</p>
</sec>
<sec id="s3e">
<title>Processing of GTDB genomes</title>
<p>We converted all GTDB genomes into anvi’o contigs databases and annotated them using the anvi’o contigs workflow, which is similar to the metagenomics workflow described above and uses the same programs for gene identification and annotation.</p>
</sec>
<sec id="s3f">
<title>Estimation of the number of microbial populations per metagenome</title>
<p>We used single-copy core gene (SCG) sets belonging to each domain of microbial life (Bacteria, Archaea, Protista) (M. D. <xref ref-type="bibr" rid="c53">Lee 2019</xref>) to estimate the number of populations from each domain present in a given metagenomic sample. For each domain, we calculated the number of populations by taking the mode of the number of copies of each SCG in the set. We then summed the number of populations from each domain to get a total number of microbial populations within each sample. We accomplished this using SCG annotations provided by ‘anvi-run-hmms’ (which was run during metagenome processing) and a custom script relying on the anvi’o class ‘NumGenomesEstimator’ (see reproducible workflow).</p>
</sec>
<sec id="s3g">
<title>Removal of samples with low sequencing depth</title>
<p>We observed that, at lower sequencing depths, our estimates for the number of populations in a metagenomic sample were moderately correlated with sequencing depth (<xref rid="figs1" ref-type="fig">Supplementary Figure 1</xref>, R &gt; 0.5). These estimates rely on having accurate counts of single-copy core genes (SCGs), so we hypothesized that lower-depth samples were systematically missing SCGs, especially from populations with lower abundance. Since accurate population number estimates are critical for proper normalization of pathway copy numbers, keeping these lower-depth samples would have introduced a bias into our metabolism analyses. To address this, we removed samples with low sequencing depth from downstream analyses using a sequencing depth threshold of 25 million reads, such that the remaining samples exhibited a weaker correlation (R &lt; 0.5) between sequencing depth and number of estimated populations. We kept samples for which both the R1 file and the R2 file contained at least 25 million reads (and for the (<xref ref-type="bibr" rid="c99">Vineis et al. 2016</xref>) dataset, we kept samples containing at least 25 million merged reads). This produced our final sample set of 408 metagenomes.</p>
</sec>
<sec id="s3h">
<title>Estimation of normalized pathway copy numbers in metagenomes</title>
<p>We ran ‘anvi-estimate-metabolism’, in genome mode and with the ‘--add-copy-number’ flag, on each individual metagenome assembly to compute stepwise copy numbers for KEGG modules from the combined gene annotations of all populations present in the sample. We then divided these copy numbers by the number of estimated populations within each sample to obtain a per-population copy number (PPCN) for each pathway.</p>
</sec>
<sec id="s3i">
<title>Selection of IBD-enriched pathways</title>
<p>We used a one-sided Mann-Whitney-Wilcoxon test with a FDR-adjusted p-value threshold of p &lt;= 2e-10 on the per-sample PPCN values for each module individually to identify the pathways that were most significantly enriched in the IBD sample group compared to the healthy group. We calculated the median per-population copy number of each metabolic pathway in the IBD samples, and again in the healthy samples. After filtering for p-values &lt;= 2e-10, we also applied a minimum effect size threshold based on the median per-population copy number in each group (M<sub>IBD</sub> – M<sub>Healthy</sub> &gt;= 0.12) – this threshold was calculated by taking the mean effect size over all pathways that passed the p-value threshold. The set originally contained 34 pathways that passed both thresholds, but we removed one redundant module (M00006) which represents the first half of another module in the set (M00004).</p>
</sec>
<sec id="s3j">
<title>Test for enrichment of biosynthesis pathways</title>
<p>We used a one-sided Fisher’s exact test (also known as hypergeometric test, see e.g., (<xref ref-type="bibr" rid="c111">Boyle et al. 2004</xref>)) for testing the independence between the metabolic pathways identified to be IBD-enriched (i.e., using the methods described in “Selection of IBD-enriched pathways) and functionality (i.e., pathways annotated to be involved in biosynthesis).</p>
</sec>
<sec id="s3k">
<title>Pathway comparisons</title>
<p>Because the 33 IBD-enriched pathways were selected using PPCNs of healthy and IBD samples, statistical tests comparing PPCN distributions for these modules need to be interpreted with care, because the hypotheses were selected and tested on the same dataset (<xref ref-type="bibr" rid="c24">Fithian, Sun, and Taylor 2014</xref>). Therefore, to assess the statistical validity of the identified IBD-enriched modules, we performed the following repeated sample-split analysis: we first randomly split the IBD and healthy samples into the equal-sized training and validation sets. We select IBD-enriched modules in the training set using the Mann-Whitney-Wilcoxon test, and then compute the p-values on the validation set. We repeat this sample split analysis 1,000 times with an FDR-adjusted p-value threshold of 1e-10 on the first split; most identified modules (89.4%; 95% CI: [87.5%, 91.3%]) on the training sets remain significant at a slightly less stringent threshold (1e-8) on the validation sets. This indicates that the approach we used to identify IBD-enriched modules yields stable and statistically significant results on this dataset.</p>
</sec>
<sec id="s3l">
<title>Metagenome classification</title>
<p>We trained logistic regression models to classify samples as ‘IBD’ or ‘healthy’ using per-population copy numbers of IBD-enriched modules as features. We ran a 25-fold cross-validation pipeline on the set of 330 healthy and IBD metagenomes in our analysis, using an 80% train – 20% test random split of the data in each fold. The pipeline included selection of IBD-enriched pathways within the training samples using the same strategy as described above, followed by training and testing of a logistic regression model as implemented in the ‘sklearn’ Python package. We set the ‘penalty’ parameter of the model to “None” and the ‘max_iter’ parameter to 20,000 iterations, and we used the same random state in each fold to ensure changes in performance only come from differences in the training data rather than differences in model initialization. To summarize the overall performance of the classifier, we took the mean (over all folds) of each performance metric.</p>
<p>We trained a final classifier using the 33 IBD-enriched pathways selected earlier from the entire set of 330 healthy and IBD metagenomes. We then applied this classifier to the metagenomic samples from (<xref ref-type="bibr" rid="c71">Palleja et al. 2018</xref>), which we processed in the same way as the other samples in our analysis (including removal of samples with low sequencing depth and calculation of PPCNs of KEGG modules for use as input features to the classifier model).</p>
</sec>
<sec id="s3m">
<title>Identification of gut microbial genomes from the GTDB</title>
<p>We took 19,226 representative genomes from the GTDB species clusters belonging to the phyla Firmicutes, Bacteroidetes, and Proteobacteria, which are most common in the human gut microbiome (<xref ref-type="bibr" rid="c103">Woting and Blaut 2016</xref>). To evaluate which of these genomes might represent gut microbes in a computationally-tractable manner, we ran the anvi’o ‘EcoPhylo’ workflow (<ext-link ext-link-type="uri" xlink:href="https://anvio.org/m/ecophylo">https://anvio.org/m/ecophylo</ext-link>) to contextualize these populations within 150 healthy gut metagenomes from the Human Microbiome Project (HMP) (Human Microbiome Project Consortium 2012). Briefly, the EcoPhylo workflow (1) recovers sequences of a gene family of interest from each genome and metagenomic sample in the analysis, (2) clusters resulting sequences and picks representative sequences using mmseqs2 (<xref ref-type="bibr" rid="c97">Steinegger and Söding 2017</xref>), and (3) uses the representative sequences to rapidly summarize the distribution of each population cluster across the metagenomic samples through metagenomic read recruitment analyses. Here, we used the ribosomal protein S6 as our gene of interest, since it was the most frequently-assembled single-copy-core gene in our set of GTDB genomes. We clustered the Ribosomal Protein S6 sequences from GTDB genomes at 94% nucleotide identity.</p>
<p>To identify genomes that were likely to represent gut microbes, we selected genomes whose ribosomal protein S6 belonged to a gene cluster where at least 50% of the representative sequence was covered (i.e. detection &gt;= 0.5x) in more than 10% of samples (i.e. n &gt; 15). There are 100 distinct individuals represented in the 150 HMP gut metagenomes – 56 of which were sampled just once and 46 of which were sampled at 2 or 3 time points – so this threshold is equivalent to detecting the genome in 5% – 15% of individuals. From this selection we obtained a set of 836 genomes; however, these were not exclusively gut microbes, as some non-gut populations have similar ribosomal protein S6 sequences to gut microbes and can therefore pass this selection step. To eliminate these, we mapped our set of 330 healthy and IBD metagenomes to the 836 genomes using the anvi’o metagenomics workflow, and extracted genomes whose entire sequence was at least 50% covered (i.e. detection &gt;= 0.5x) in over 2% (n &gt; 6) of these samples. Our final set of 338 genomes was used in downstream analysis.</p>
</sec>
<sec id="s3n">
<title>Genome phylogeny</title>
<p>To create the phylogeny, we identified the following ribosomal proteins that were annotated in at least 90% (n = 304) of the genomes: Ribosomal_S6, Ribosomal_S16, Ribosomal_L19, Ribosomal_L27, Ribosomal_S15, Ribosomal_S20p, Ribosomal_L13, Ribosomal_L21p, Ribosomal_L20, and Ribosomal_L9_C. We used ‘anvi-get-sequences-for-hmm-hits’ to extract the amino acid sequences for these genes, align the sequences using MUSCLE v3.8.1551 (<xref ref-type="bibr" rid="c18">Edgar 2004</xref>), and concatenate the alignments. We used trimAl v1.4.rev15 (<xref ref-type="bibr" rid="c11">Capella-Gutiérrez, Silla-Martínez, and Gabaldón 2009</xref>) to remove any positions containing more than 50% of gap characters from the final alignment. Finally, we built the tree with IQtree v2.2.0.3 (<xref ref-type="bibr" rid="c65">Minh et al. 2020</xref>), using the WAG model and running 1,000 bootstraps.</p>
</sec>
<sec id="s3o">
<title>Determination of HMI status for genomes</title>
<p>We estimated metabolic potential for each genome with ‘anvi-estimate-metabolism’ (in genome mode) to get stepwise completeness scores for each KEGG module, and then we used the script ‘anvi-script-estimate-metabolic-independencè to give each genome a metabolic independence score based on completeness of the 33 IBD-enriched pathways. Briefly, the latter script calculates the score by summing the completeness scores of each pathway of interest. Genomes were classified as having high metabolic independence (HMI) if their score was greater than or equal to 26.4. We calculated this threshold by requiring these 33 pathways to be, on average, at least 80% complete in a given genome.</p>
</sec>
<sec id="s3p">
<title>Genome distribution across sample groups</title>
<p>We mapped the gut metagenomes from the healthy, non-IBD, and IBD groups to each genome using the anvi’o metagenomics workflow in reference mode. We used ‘anvi-summarizè to obtain a matrix of genome detection across all samples. We summarized this data as follows: for each genome, we computed the proportion of samples in each group in which at least 50% of the genome sequence was covered by at least 1 read (&gt;= 50% detection). For each sample, we calculated the proportion of detected genomes that were classified as HMI. We also computed the percent abundance of each genome in each sample by dividing the number of reads mapping to that genome by the total number of reads in the sample.</p>
</sec>
<sec id="s3q">
<title>Visualizations</title>
<p>We used ggplot2 (<xref ref-type="bibr" rid="c102">Wickham 2016</xref>) to generate most of the initial data visualizations. The phylogeny and heatmap in <xref rid="fig3" ref-type="fig">Figure 3</xref> were generated by the anvi’o interactive interface and the ROC curves in <xref rid="fig4" ref-type="fig">Figure 4</xref> were generated using the pyplot package of matplotlib (<xref ref-type="bibr" rid="c35">Hunter 2007</xref>). These visualizations were refined for publication using Inkscape, an open-source graphical editing software that is available at <ext-link ext-link-type="uri" xlink:href="https://inkscape.org/">https://inkscape.org/</ext-link>.</p>
</sec>
</sec>
</body>
<back>
<sec id="s4">
<title>Data Availability</title>
<p>Accession numbers for publicly available data are listed in our Supplementary Tables at doi: 10.6084/m9.figshare.22679080. Our Supplementary Information file is also available at doi: 10.6084/m9.figshare.22679080. Contigs databases of our assemblies for the 408 deeply-sequenced metagenomes can be accessed at doi: 10.5281/zenodo.7872967, and databases for our assemblies of the (<xref ref-type="bibr" rid="c71">Palleja et al. 2018</xref>) metagenomes can be accessed at doi: 10.5281/zenodo.7897987. Contigs databases of the 338 GTDB gut reference genomes are available at doi: 10.5281/zenodo.7883421.</p>
</sec>
<ack>
<title>Acknowledgements</title>
<p>We thank Chris Quince for advice on statistical significance testing. IV acknowledges support from the National Science Foundation Graduate Research Fellowship under Grant No. 1746045; ADW acknowledges support from the National Institutes of General Medical Sciences under R35 GM133420. YTC acknowledges support from the Stanford Data Science Postdoctoral Fellowship. RB acknowledges support from the National Institutes of Health under R35 GM128716. ECF acknowledges support from the University of Chicago International Student Fellowship. Additional support for ECF and AME came from an NIH NIDDK grant (RC2 DK122394) to AME.</p>
</ack>
<sec id="s5">
<title>Author Contributions</title>
<p>IV, JF and AME conceived the study. IV developed methodology. IV, MSS and AME developed computational analysis tools. IV, JF, and AME performed formal analyses. YTC supported statistical analyses. RB, BJ, ADW and MKY conducted investigations. ECF, ARW, CV and AFG provided resources. ECF, ARW and IV curated data. IV and AME prepared figures. IV, JF and AME wrote the paper with critical input from all authors. JF and AME supervised the project.</p>
</sec>
<sec id="s6">
<title>Ethics declarations</title>
<sec id="s6a">
<title>Competing interests</title>
<p>The authors declare that they have no competing interests.</p>
</sec>
</sec>
<ref-list>
<title>References</title>
<ref id="c1"><mixed-citation publication-type="journal"><string-name><surname>Aramaki</surname>, <given-names>Takuya</given-names></string-name>, <string-name><given-names>Romain</given-names> <surname>Blanc-Mathieu</surname></string-name>, <string-name><given-names>Hisashi</given-names> <surname>Endo</surname></string-name>, <string-name><given-names>Koichi</given-names> <surname>Ohkubo</surname></string-name>, <string-name><given-names>Minoru</given-names> <surname>Kanehisa</surname></string-name>, <string-name><given-names>Susumu</given-names> <surname>Goto</surname></string-name>, and <string-name><given-names>Hiroyuki</given-names> <surname>Ogata</surname></string-name>. <year>2019</year>. “<article-title>KofamKOALA: KEGG Ortholog Assignment Based on Profile HMM and Adaptive Score Threshold</article-title>.” <source>Bioinformatics</source> <volume>36</volume> (<issue>7</issue>): <fpage>2251</fpage>–<lpage>52</lpage>.</mixed-citation></ref>
<ref id="c2"><mixed-citation publication-type="journal"><string-name><surname>Arkin</surname>, <given-names>Adam P.</given-names></string-name>, <string-name><given-names>Robert W.</given-names> <surname>Cottingham</surname></string-name>, <string-name><given-names>Christopher S.</given-names> <surname>Henry</surname></string-name>, <string-name><given-names>Nomi L.</given-names> <surname>Harris</surname></string-name>, <string-name><given-names>Rick L.</given-names> <surname>Stevens</surname></string-name>, <string-name><given-names>Sergei</given-names> <surname>Maslov</surname></string-name>, <string-name><given-names>Paramvir</given-names> <surname>Dehal</surname></string-name>, <etal>et al.</etal> <year>2018</year>. “<article-title>KBase: The United States Department of Energy Systems Biology Knowledgebase</article-title>.” <source>Nature Biotechnology</source> <volume>36</volume> (<issue>7</issue>): <fpage>566</fpage>–<lpage>69</lpage>.</mixed-citation></ref>
<ref id="c3"><mixed-citation publication-type="journal"><string-name><surname>Arumugam</surname>, <given-names>Manimozhiyan</given-names></string-name>, <string-name><given-names>Jeroen</given-names> <surname>Raes</surname></string-name>, <string-name><given-names>Eric</given-names> <surname>Pelletier</surname></string-name>, <string-name><given-names>Denis</given-names> <surname>Le Paslier</surname></string-name>, <string-name><given-names>Takuji</given-names> <surname>Yamada</surname></string-name>, <string-name><given-names>Daniel R.</given-names> <surname>Mende</surname></string-name>, <string-name><given-names>Gabriel R.</given-names> <surname>Fernandes</surname></string-name>, <etal>et al.</etal> <year>2011</year>. “<article-title>Enterotypes of the Human Gut Microbiome</article-title>.” <source>Nature</source> <volume>473</volume> (<issue>7346</issue>): <fpage>174</fpage>–<lpage>80</lpage>.</mixed-citation></ref>
<ref id="c4"><mixed-citation publication-type="journal"><string-name><surname>Aziz</surname>, <given-names>Ramy K.</given-names></string-name>, <string-name><given-names>Daniela</given-names> <surname>Bartels</surname></string-name>, <string-name><given-names>Aaron A.</given-names> <surname>Best</surname></string-name>, <string-name><given-names>Matthew</given-names> <surname>DeJongh</surname></string-name>, <string-name><given-names>Terrence</given-names> <surname>Disz</surname></string-name>, <string-name><given-names>Robert A.</given-names> <surname>Edwards</surname></string-name>, <string-name><given-names>Kevin</given-names> <surname>Formsma</surname></string-name>, <etal>et al.</etal> <year>2008</year>. “<article-title>The RAST Server: Rapid Annotations Using Subsystems Technology</article-title>.” <source>BMC Genomics</source> <volume>9</volume> (<issue>February</issue>): <fpage>75</fpage>.</mixed-citation></ref>
<ref id="c5"><mixed-citation publication-type="journal"><string-name><surname>Belkaid</surname>, <given-names>Yasmine</given-names></string-name>, and <string-name><given-names>Timothy W.</given-names> <surname>Hand</surname></string-name>. <year>2014</year>. “<article-title>Role of the Microbiota in Immunity and Inflammation</article-title>.” <source>Cell</source> <volume>157</volume> (<issue>1</issue>): <fpage>121</fpage>–<lpage>41</lpage>.</mixed-citation></ref>
<ref id="c6"><mixed-citation publication-type="web"><collab>“BioProject.”</collab> n.d. Accessed <date-in-citation content-type="access-date">September 23, 2022</date-in-citation>. <ext-link ext-link-type="uri" xlink:href="https://www.ncbi.nlm.nih.gov/bioproject/PRJEB6092/">https://www.ncbi.nlm.nih.gov/bioproject/PRJEB6092/</ext-link>.</mixed-citation></ref>
<ref id="c7"><mixed-citation publication-type="journal"><string-name><surname>Boyle</surname>, <given-names>Elizabeth I.</given-names></string-name>, <string-name><given-names>Shuai</given-names> <surname>Weng</surname></string-name>, <string-name><given-names>Jeremy</given-names> <surname>Gollub</surname></string-name>, <string-name><given-names>Heng</given-names> <surname>Jin</surname></string-name>, <string-name><given-names>David</given-names> <surname>Botstein</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Michael Cherry</surname></string-name>, and <string-name><given-names>Gavin</given-names> <surname>Sherlock</surname></string-name>. <year>2004</year>. “<article-title>GO::TermFinder--Open Source Software for Accessing Gene Ontology Information and Finding Significantly Enriched Gene Ontology Terms Associated with a List of Genes</article-title>.” <source>Bioinformatics</source> <volume>20</volume> (<issue>18</issue>): <fpage>3710</fpage>–<lpage>15</lpage>.</mixed-citation></ref>
<ref id="c8"><mixed-citation publication-type="journal"><string-name><surname>Buchfink</surname>, <given-names>Benjamin</given-names></string-name>, <string-name><given-names>Chao</given-names> <surname>Xie</surname></string-name>, and <string-name><given-names>Daniel H.</given-names> <surname>Huson</surname></string-name>. <year>2015</year>. “<article-title>Fast and Sensitive Protein Alignment Using DIAMOND</article-title>.” <source>Nature Methods</source> <volume>12</volume> (<issue>1</issue>): <fpage>59</fpage>–<lpage>60</lpage>.</mixed-citation></ref>
<ref id="c9"><mixed-citation publication-type="journal"><string-name><surname>Byndloss</surname>, <given-names>Mariana X.</given-names></string-name>, and <string-name><given-names>Andreas J.</given-names> <surname>Bäumler</surname></string-name>. <year>2018</year>. “<article-title>The Germ-Organ Theory of Non-Communicable Diseases</article-title>.” <source>Nature Reviews. Microbiology</source> <volume>16</volume> (<issue>2</issue>): <fpage>103</fpage>–<lpage>10</lpage>.</mixed-citation></ref>
<ref id="c10"><mixed-citation publication-type="journal"><string-name><surname>Cani</surname>, <given-names>Patrice D</given-names></string-name>. <year>2018</year>. “<article-title>Human Gut Microbiome: Hopes, Threats and Promises</article-title>.” <source>Gut</source> <volume>67</volume> (<issue>9</issue>): <fpage>1716</fpage>–<lpage>25</lpage>.</mixed-citation></ref>
<ref id="c11"><mixed-citation publication-type="journal"><string-name><surname>Capella-Gutiérrez</surname>, <given-names>Salvador</given-names></string-name>, <string-name><given-names>José M.</given-names> <surname>Silla-Martínez</surname></string-name>, and <string-name><given-names>Toni</given-names> <surname>Gabaldón</surname></string-name>. <year>2009</year>. “<article-title>trimAl: A Tool for Automated Alignment Trimming in Large-Scale Phylogenetic Analyses</article-title>.” <source>Bioinformatics</source> <volume>25</volume> (<issue>15</issue>): <fpage>1972</fpage>–<lpage>73</lpage>.</mixed-citation></ref>
<ref id="c12"><mixed-citation publication-type="journal"><string-name><surname>Chan</surname>, <given-names>Patricia P.</given-names></string-name>, and <string-name><given-names>Todd M.</given-names> <surname>Lowe</surname></string-name>. <year>2019</year>. “<article-title>tRNAscan-SE: Searching for tRNA Genes in Genomic Sequences</article-title>.” <source>Methods in Molecular Biology</source> <volume>1962</volume>: <fpage>1</fpage>–<lpage>14</lpage>.</mixed-citation></ref>
<ref id="c13"><mixed-citation publication-type="journal"><string-name><surname>Clausen</surname>, <given-names>David S.</given-names></string-name>, and <string-name><given-names>Amy D.</given-names> <surname>Willis</surname></string-name>. <year>2022</year>. “<article-title>Evaluating Replicability in Microbiome Data</article-title>.” <source>Biostatistics</source> <volume>23</volume> (<issue>4</issue>): <fpage>1099</fpage>–<lpage>1114</lpage>.</mixed-citation></ref>
<ref id="c14"><mixed-citation publication-type="journal"><string-name><surname>Coyte</surname>, <given-names>Katharine Z.</given-names></string-name>, <string-name><given-names>Jonas</given-names> <surname>Schluter</surname></string-name>, and <string-name><given-names>Kevin R.</given-names> <surname>Foster</surname></string-name>. <year>2015</year>. “<article-title>The Ecology of the Microbiome: Networks, Competition, and Stability</article-title>.” <source>Science</source> <volume>350</volume> (<issue>6261</issue>): <fpage>663</fpage>–<lpage>66</lpage>.</mixed-citation></ref>
<ref id="c15"><mixed-citation publication-type="journal"><string-name><surname>Degnan</surname>, <given-names>Patrick H.</given-names></string-name>, <string-name><given-names>Natasha A.</given-names> <surname>Barry</surname></string-name>, <string-name><given-names>Kenny C.</given-names> <surname>Mok</surname></string-name>, <string-name><given-names>Michiko E.</given-names> <surname>Taga</surname></string-name>, and <string-name><given-names>Andrew L.</given-names> <surname>Goodman</surname></string-name>. <year>2014</year>. “<article-title>Human Gut Microbes Use Multiple Transporters to Distinguish Vitamin B₁₂ Analogs and Compete in the Gut</article-title>.” <source>Cell Host &amp; Microbe</source> <volume>15</volume> (<issue>1</issue>): <fpage>47</fpage>–<lpage>57</lpage>.</mixed-citation></ref>
<ref id="c16"><mixed-citation publication-type="journal"><string-name><surname>Devkota</surname>, <given-names>Suzanne</given-names></string-name>, <string-name><given-names>Yunwei</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Mark W.</given-names> <surname>Musch</surname></string-name>, <string-name><given-names>Vanessa</given-names> <surname>Leone</surname></string-name>, <string-name><given-names>Hannah</given-names> <surname>Fehlner-Peach</surname></string-name>, <string-name><given-names>Anuradha</given-names> <surname>Nadimpalli</surname></string-name>, <string-name><given-names>Dionysios A.</given-names> <surname>Antonopoulos</surname></string-name>, <string-name><given-names>Bana</given-names> <surname>Jabri</surname></string-name>, and <string-name><given-names>Eugene B.</given-names> <surname>Chang</surname></string-name>. <year>2012</year>. “<article-title>Dietary-Fat-Induced Taurocholic Acid Promotes Pathobiont Expansion and Colitis in Il10-/-Mice</article-title>.” <source>Nature</source> <volume>487</volume> (<issue>7405</issue>): <fpage>104</fpage>–<lpage>8</lpage>.</mixed-citation></ref>
<ref id="c17"><mixed-citation publication-type="journal"><string-name><surname>Eddy</surname>, <given-names>Sean R</given-names></string-name>. <year>2011</year>. “<article-title>Accelerated Profile HMM Searches</article-title>.” <source>PLoS Computational Biology</source> <volume>7</volume> (<issue>10</issue>): <fpage>e1002195</fpage>.</mixed-citation></ref>
<ref id="c18"><mixed-citation publication-type="journal"><string-name><surname>Edgar</surname>, <given-names>Robert C</given-names></string-name>. <year>2004</year>. “<article-title>MUSCLE: Multiple Sequence Alignment with High Accuracy and High Throughput</article-title>.” <source>Nucleic Acids Research</source> <volume>32</volume> (<issue>5</issue>): <fpage>1792</fpage>–<lpage>97</lpage>.</mixed-citation></ref>
<ref id="c19"><mixed-citation publication-type="journal"><string-name><surname>Eren</surname>, <given-names>A. Murat</given-names></string-name>, <string-name><given-names>Joseph H.</given-names> <surname>Vineis</surname></string-name>, <string-name><given-names>Hilary G.</given-names> <surname>Morrison</surname></string-name>, and <string-name><given-names>Mitchell L.</given-names> <surname>Sogin</surname></string-name>. <year>2013</year>. “<article-title>A Filtering Method to Generate High Quality Short Reads Using Illumina Paired-End Technology</article-title>.” <source>PloS One</source> <volume>8</volume> (<issue>6</issue>): <fpage>e66643</fpage>.</mixed-citation></ref>
<ref id="c20"><mixed-citation publication-type="journal"><string-name><surname>Fan</surname>, <given-names>Yong</given-names></string-name>, and <string-name><given-names>Oluf</given-names> <surname>Pedersen</surname></string-name>. <year>2021</year>. “<article-title>Gut Microbiota in Human Metabolic Health and Disease</article-title>.” <source>Nature Reviews. Microbiology</source> <volume>19</volume> (<issue>1</issue>): <fpage>55</fpage>–<lpage>71</lpage>.</mixed-citation></ref>
<ref id="c21"><mixed-citation publication-type="journal"><string-name><surname>Farag</surname>, <given-names>Ibrahim F.</given-names></string-name>, <string-name><given-names>Jennifer F.</given-names> <surname>Biddle</surname></string-name>, <string-name><given-names>Rui</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>Amanda J.</given-names> <surname>Martino</surname></string-name>, <string-name><given-names>Christopher H.</given-names> <surname>House</surname></string-name>, and <string-name><given-names>Rosa I.</given-names> <surname>León-Zayas</surname></string-name>. <year>2020</year>. “<article-title>Metabolic Potentials of Archaeal Lineages Resolved from Metagenomes of Deep Costa Rica Sediments</article-title>.” <source>The ISME Journal</source> <volume>14</volume> (<issue>6</issue>): <fpage>1345</fpage>–<lpage>58</lpage>.</mixed-citation></ref>
<ref id="c22"><mixed-citation publication-type="journal"><string-name><surname>Feng</surname>, <given-names>Lihui</given-names></string-name>, <string-name><given-names>Arjun S.</given-names> <surname>Raman</surname></string-name>, <string-name><given-names>Matthew C.</given-names> <surname>Hibberd</surname></string-name>, <string-name><given-names>Jiye</given-names> <surname>Cheng</surname></string-name>, <string-name><given-names>Nicholas W.</given-names> <surname>Griffin</surname></string-name>, <string-name><given-names>Yangqing</given-names> <surname>Peng</surname></string-name>, <string-name><given-names>Semen A.</given-names> <surname>Leyn</surname></string-name>, <string-name><given-names>Dmitry A.</given-names> <surname>Rodionov</surname></string-name>, <string-name><given-names>Andrei L.</given-names> <surname>Osterman</surname></string-name>, and <string-name><given-names>Jeffrey I.</given-names> <surname>Gordon</surname></string-name>. <year>2020</year>. “<article-title>Identifying Determinants of Bacterial Fitness in a Model of Human Gut Microbial Succession</article-title>.” <source>Proceedings of the National Academy of Sciences of the United States of America</source> <volume>117</volume> (<issue>5</issue>): <fpage>2622</fpage>–<lpage>33</lpage>.</mixed-citation></ref>
<ref id="c23"><mixed-citation publication-type="journal"><string-name><surname>Feng</surname>, <given-names>Qiang</given-names></string-name>, <string-name><given-names>Suisha</given-names> <surname>Liang</surname></string-name>, <string-name><given-names>Huijue</given-names> <surname>Jia</surname></string-name>, <string-name><given-names>Andreas</given-names> <surname>Stadlmayr</surname></string-name>, <string-name><given-names>Longqing</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>Zhou</given-names> <surname>Lan</surname></string-name>, <string-name><given-names>Dongya</given-names> <surname>Zhang</surname></string-name>, <etal>et al.</etal> <year>2015</year>. “<article-title>Gut Microbiome Development along the Colorectal Adenoma-Carcinoma Sequence</article-title>.” <source>Nature Communications</source> <volume>6</volume> (<issue>March</issue>): <fpage>6528</fpage>.</mixed-citation></ref>
<ref id="c24"><mixed-citation publication-type="web"><string-name><surname>Fithian</surname>, <given-names>William</given-names></string-name>, <string-name><given-names>Dennis</given-names> <surname>Sun</surname></string-name>, and <string-name><given-names>Jonathan</given-names> <surname>Taylor</surname></string-name>. <year>2014</year>. “<article-title>Optimal Inference After Model Selection</article-title>.” <italic>arXiv [math.ST]</italic>. arXiv. <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1410.2597">http://arxiv.org/abs/1410.2597</ext-link>.</mixed-citation></ref>
<ref id="c25"><mixed-citation publication-type="journal"><string-name><surname>Franzosa</surname>, <given-names>Eric A.</given-names></string-name>, <string-name><given-names>Alexandra</given-names> <surname>Sirota-Madi</surname></string-name>, <string-name><given-names>Julian</given-names> <surname>Avila-Pacheco</surname></string-name>, <string-name><given-names>Nadine</given-names> <surname>Fornelos</surname></string-name>, <string-name><given-names>Henry J.</given-names> <surname>Haiser</surname></string-name>, <string-name><given-names>Stefan</given-names> <surname>Reinker</surname></string-name>, <string-name><given-names>Tommi</given-names> <surname>Vatanen</surname></string-name>, <etal>et al.</etal> <year>2019</year>. “<article-title>Gut Microbiome Structure and Metabolic Activity in Inflammatory Bowel Disease</article-title>.” <source>Nature Microbiology</source> <volume>4</volume> (<issue>2</issue>): <fpage>293</fpage>–<lpage>305</lpage>.</mixed-citation></ref>
<ref id="c26"><mixed-citation publication-type="journal"><string-name><surname>Galperin</surname>, <given-names>Michael Y.</given-names></string-name>, <string-name><given-names>Yuri I.</given-names> <surname>Wolf</surname></string-name>, <string-name><given-names>Kira S.</given-names> <surname>Makarova</surname></string-name>, <string-name><given-names>Roberto Vera</given-names> <surname>Alvarez</surname></string-name>, <string-name><given-names>David</given-names> <surname>Landsman</surname></string-name>, and <string-name><given-names>Eugene V.</given-names> <surname>Koonin</surname></string-name>. <year>2021</year>. “<article-title>COG Database Update: Focus on Microbial Diversity, Model Organisms, and Widespread Pathogens</article-title>.” <source>Nucleic Acids Research</source> <volume>49</volume> (<issue>D1</issue>): <fpage>D274</fpage>–<lpage>81</lpage>.</mixed-citation></ref>
<ref id="c27"><mixed-citation publication-type="journal"><string-name><surname>Geller-McGrath</surname>, <given-names>D.</given-names></string-name>, <string-name><given-names>Kishori M.</given-names> <surname>Konwar</surname></string-name>, <string-name><given-names>V. P.</given-names> <surname>Edgcomb</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Pachiadaki</surname></string-name>, <string-name><given-names>J. W.</given-names> <surname>Roddy</surname></string-name>, <string-name><given-names>T. J.</given-names> <surname>Wheeler</surname></string-name>, and <string-name><given-names>J. E.</given-names> <surname>McDermott</surname></string-name>. <year>2023</year>. “<article-title>MetaPathPredict: A Machine Learning-Based Tool for Predicting Metabolic Modules in Incomplete Bacterial Genomes</article-title>.” <source>bioRxiv</source>. <pub-id pub-id-type="doi">10.1101/2022.12.21.521254</pub-id>.</mixed-citation></ref>
<ref id="c28"><mixed-citation publication-type="journal"><string-name><surname>Gevers</surname>, <given-names>Dirk</given-names></string-name>, <string-name><given-names>Subra</given-names> <surname>Kugathasan</surname></string-name>, <string-name><given-names>Lee A.</given-names> <surname>Denson</surname></string-name>, <string-name><given-names>Yoshiki</given-names> <surname>Vázquez-Baeza</surname></string-name>, <string-name><given-names>Will</given-names> <surname>Van Treuren</surname></string-name>, <string-name><given-names>Boyu</given-names> <surname>Ren</surname></string-name>, <string-name><given-names>Emma</given-names> <surname>Schwager</surname></string-name>, <etal>et al.</etal> <year>2014</year>. “<article-title>The Treatment-Naive Microbiome in New-Onset Crohn’s Disease</article-title>.” <source>Cell Host &amp; Microbe</source> <volume>15</volume> (<issue>3</issue>): <fpage>382</fpage>–<lpage>92</lpage>.</mixed-citation></ref>
<ref id="c29"><mixed-citation publication-type="journal"><string-name><surname>Goodman</surname>, <given-names>Andrew L.</given-names></string-name>, <string-name><given-names>Nathan P.</given-names> <surname>McNulty</surname></string-name>, <string-name><given-names>Yue</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>Douglas</given-names> <surname>Leip</surname></string-name>, <string-name><given-names>Robi D.</given-names> <surname>Mitra</surname></string-name>, <string-name><given-names>Catherine A.</given-names> <surname>Lozupone</surname></string-name>, <string-name><given-names>Rob</given-names> <surname>Knight</surname></string-name>, and <string-name><given-names>Jeffrey I.</given-names> <surname>Gordon</surname></string-name>. <year>2009</year>. “<article-title>Identifying Genetic Determinants Needed to Establish a Human Gut Symbiont in Its Habitat</article-title>.” <source>Cell Host &amp; Microbe</source> <volume>6</volume> (<issue>3</issue>): <fpage>279</fpage>– <lpage>89</lpage>.</mixed-citation></ref>
<ref id="c30"><mixed-citation publication-type="journal"><string-name><surname>Halfvarson</surname>, <given-names>Jonas</given-names></string-name>, <string-name><given-names>Colin J.</given-names> <surname>Brislawn</surname></string-name>, <string-name><given-names>Regina</given-names> <surname>Lamendella</surname></string-name>, <string-name><given-names>Yoshiki</given-names> <surname>Vázquez-Baeza</surname></string-name>, <string-name><given-names>William A.</given-names> <surname>Walters</surname></string-name>, <string-name><given-names>Lisa M.</given-names> <surname>Bramer</surname></string-name>, <string-name><given-names>Mauro</given-names> <surname>D’Amato</surname></string-name>, <etal>et al.</etal> <year>2017</year>. “<article-title>Dynamics of the Human Gut Microbiome in Inflammatory Bowel Disease</article-title>.” <source>Nature Microbiology</source> <volume>2</volume> (<issue>February</issue>): <fpage>17004</fpage>.</mixed-citation></ref>
<ref id="c31"><mixed-citation publication-type="journal"><string-name><surname>Heinken</surname>, <given-names>Almut</given-names></string-name>, <string-name><given-names>Johannes</given-names> <surname>Hertel</surname></string-name>, and <string-name><given-names>Ines</given-names> <surname>Thiele</surname></string-name>. <year>2021</year>. “<article-title>Metabolic Modelling Reveals Broad Changes in Gut Microbial Metabolism in Inflammatory Bowel Disease Patients with Dysbiosis</article-title>.” <source>NPJ Systems Biology and Applications</source> <volume>7</volume> (<issue>1</issue>): <fpage>19</fpage>.</mixed-citation></ref>
<ref id="c32"><mixed-citation publication-type="journal"><string-name><surname>Henke</surname>, <given-names>Matthew T.</given-names></string-name>, <string-name><given-names>Douglas J.</given-names> <surname>Kenny</surname></string-name>, <string-name><given-names>Chelsi D.</given-names> <surname>Cassilly</surname></string-name>, <string-name><given-names>Hera</given-names> <surname>Vlamakis</surname></string-name>, <string-name><given-names>Ramnik J.</given-names> <surname>Xavier</surname></string-name>, and <string-name><given-names>Jon</given-names> <surname>Clardy</surname></string-name>. <year>2019</year>. “<article-title>Ruminococcus Gnavus, a Member of the Human Gut Microbiome Associated with Crohn’s Disease, Produces an Inflammatory Polysaccharide</article-title>.” <source>Proceedings of the National Academy of Sciences of the United States of America</source> <volume>116</volume> (<issue>26</issue>): <fpage>12672</fpage>–<lpage>77</lpage>.</mixed-citation></ref>
<ref id="c33"><mixed-citation publication-type="journal"><string-name><surname>Hijova</surname>, <given-names>E</given-names></string-name>. <year>2019</year>. “<article-title>Gut Bacterial Metabolites of Indigestible Polysaccharides in Intestinal Fermentation as Mediators of Public Health</article-title>.” <source>Bratislavske Lekarske Listy</source> <volume>120</volume> (<issue>11</issue>): <fpage>807</fpage>–<lpage>12</lpage>.</mixed-citation></ref>
<ref id="c34"><mixed-citation publication-type="journal"><collab>Human Microbiome Project Consortium.</collab> <year>2012</year>. “<article-title>A Framework for Human Microbiome Research</article-title>.” <source>Nature</source> <volume>486</volume> (<issue>7402</issue>): <fpage>215</fpage>–<lpage>21</lpage>.</mixed-citation></ref>
<ref id="c35"><mixed-citation publication-type="other"><collab>Hunter</collab>. <year>2007</year>. “<article-title>Matplotlib: A 2D Graphics Environment</article-title> ” 9 (May): <fpage>90</fpage>–<lpage>95</lpage>.</mixed-citation></ref>
<ref id="c36"><mixed-citation publication-type="journal"><string-name><surname>Hyatt</surname>, <given-names>Doug</given-names></string-name>, <string-name><given-names>Gwo-Liang</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Philip F.</given-names> <surname>Locascio</surname></string-name>, <string-name><given-names>Miriam L.</given-names> <surname>Land</surname></string-name>, <string-name><given-names>Frank W.</given-names> <surname>Larimer</surname></string-name>, and <string-name><given-names>Loren J.</given-names> <surname>Hauser</surname></string-name>. <year>2010</year>. “<article-title>Prodigal: Prokaryotic Gene Recognition and Translation Initiation Site Identification</article-title>.” <source>BMC Bioinformatics</source> <volume>11</volume> (<issue>March</issue>): <fpage>119</fpage>.</mixed-citation></ref>
<ref id="c37"><mixed-citation publication-type="journal"><string-name><surname>Jaffe</surname>, <given-names>Alexander L.</given-names></string-name>, <string-name><given-names>Cindy J.</given-names> <surname>Castelle</surname></string-name>, <string-name><given-names>Paula B. Matheus</given-names> <surname>Carnevali</surname></string-name>, <string-name><given-names>Simonetta</given-names> <surname>Gribaldo</surname></string-name>, and <string-name><given-names>Jillian F.</given-names> <surname>Banfield</surname></string-name>. <year>2020</year>. “<article-title>The Rise of Diversity in Metabolic Platforms across the Candidate Phyla Radiation</article-title>.” <source>BMC Biology</source> <volume>18</volume> (<issue>1</issue>): <fpage>69</fpage>.</mixed-citation></ref>
<ref id="c38"><mixed-citation publication-type="journal"><string-name><surname>Joossens</surname>, <given-names>Marie</given-names></string-name>, <string-name><given-names>Geert</given-names> <surname>Huys</surname></string-name>, <string-name><given-names>Margo</given-names> <surname>Cnockaert</surname></string-name>, <string-name><given-names>Vicky</given-names> <surname>De Preter</surname></string-name>, <string-name><given-names>Kristin</given-names> <surname>Verbeke</surname></string-name>, <string-name><given-names>Paul</given-names> <surname>Rutgeerts</surname></string-name>, <string-name><given-names>Peter</given-names> <surname>Vandamme</surname></string-name>, and <string-name><given-names>Severine</given-names> <surname>Vermeire</surname></string-name>. <year>2011</year>. “<article-title>Dysbiosis of the Faecal Microbiota in Patients with Crohn’s Disease and Their Unaffected Relatives</article-title>.” <source>Gut</source> <volume>60</volume> (<issue>5</issue>): <fpage>631</fpage>–<lpage>37</lpage>.</mixed-citation></ref>
<ref id="c39"><mixed-citation publication-type="journal"><string-name><surname>Kanehisa</surname>, <given-names>Minoru</given-names></string-name>, <string-name><given-names>Miho</given-names> <surname>Furumichi</surname></string-name>, <string-name><given-names>Yoko</given-names> <surname>Sato</surname></string-name>, <string-name><given-names>Masayuki</given-names> <surname>Kawashima</surname></string-name>, and <string-name><given-names>Mari Ishiguro-</given-names> <surname>Watanabe</surname></string-name>. <year>2023</year>. “<article-title>KEGG for Taxonomy-Based Analysis of Pathways and Genomes</article-title>.” <source>Nucleic Acids Research</source> <volume>51</volume> (<issue>D1</issue>): <fpage>D587</fpage>–<lpage>92</lpage>.</mixed-citation></ref>
<ref id="c40"><mixed-citation publication-type="journal"><string-name><surname>Kanehisa</surname>, <given-names>Minoru</given-names></string-name>, <string-name><given-names>Susumu</given-names> <surname>Goto</surname></string-name>, <string-name><given-names>Yoko</given-names> <surname>Sato</surname></string-name>, <string-name><given-names>Miho</given-names> <surname>Furumichi</surname></string-name>, and <string-name><given-names>Mao</given-names> <surname>Tanabe</surname></string-name>. <year>2012</year>. “<article-title>KEGG for Integration and Interpretation of Large-Scale Molecular Data Sets</article-title>.” <source>Nucleic Acids Research 40</source> (Database issue): <fpage>D109</fpage>–<lpage>14</lpage>.</mixed-citation></ref>
<ref id="c41"><mixed-citation publication-type="journal"><string-name><surname>Kaplan</surname>, <given-names>Gilaad G</given-names></string-name>. <year>2015</year>. “<article-title>The Global Burden of IBD: From 2015 to 2025</article-title>.” <source>Nature Reviews. Gastroenterology &amp; Hepatology</source> <volume>12</volume> (<issue>12</issue>): <fpage>720</fpage>–<lpage>27</lpage>.</mixed-citation></ref>
<ref id="c42"><mixed-citation publication-type="journal"><string-name><surname>Karp</surname>, <given-names>Peter D.</given-names></string-name>, <string-name><given-names>Peter E.</given-names> <surname>Midford</surname></string-name>, <string-name><given-names>Richard</given-names> <surname>Billington</surname></string-name>, <string-name><given-names>Anamika</given-names> <surname>Kothari</surname></string-name>, <string-name><given-names>Markus</given-names> <surname>Krummenacker</surname></string-name>, <string-name><given-names>Mario</given-names> <surname>Latendresse</surname></string-name>, <string-name><given-names>Wai Kit</given-names> <surname>Ong</surname></string-name>, <etal>et al.</etal> <year>2021</year>. “<article-title>Pathway Tools Version 23.0 Update: Software for Pathway/genome Informatics and Systems Biology</article-title>.” <source>Briefings in Bioinformatics</source> <volume>22</volume> (<issue>1</issue>): <fpage>109</fpage>–<lpage>26</lpage>.</mixed-citation></ref>
<ref id="c43"><mixed-citation publication-type="journal"><string-name><surname>Kelly</surname>, <given-names>Caleb J.</given-names></string-name>, <string-name><given-names>Erica E.</given-names> <surname>Alexeev</surname></string-name>, <string-name><given-names>Linda</given-names> <surname>Farb</surname></string-name>, <string-name><given-names>Thad W.</given-names> <surname>Vickery</surname></string-name>, <string-name><given-names>Leon</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>Campbell</given-names> <surname>Eric L</surname></string-name>, <string-name><given-names>David A.</given-names> <surname>Kitzenberg</surname></string-name>, <etal>et al.</etal> <year>2019</year>. “<article-title>Oral Vitamin B12 Supplement Is Delivered to the Distal Gut, Altering the Corrinoid Profile and Selectively Depleting Bacteroides in C57BL/6 Mice</article-title>.” <source>Gut Microbes</source> <volume>10</volume> (<issue>6</issue>): <fpage>654</fpage>–<lpage>62</lpage>.</mixed-citation></ref>
<ref id="c44"><mixed-citation publication-type="journal"><string-name><surname>Khan</surname>, <given-names>Israr</given-names></string-name>, <string-name><given-names>Naeem</given-names> <surname>Ullah</surname></string-name>, <string-name><given-names>Lajia</given-names> <surname>Zha</surname></string-name>, <string-name><given-names>Yanrui</given-names> <surname>Bai</surname></string-name>, <string-name><given-names>Ashiq</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>Tang</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>Tuanjie</given-names> <surname>Che</surname></string-name>, and <string-name><given-names>Chunjiang</given-names> <surname>Zhang</surname></string-name>. <year>2019</year>. “<article-title>Alteration of Gut Microbiota in Inflammatory Bowel Disease (IBD): Cause or Consequence? IBD Treatment Targeting the Gut Microbiome</article-title>.” <source>Pathogens</source> <volume>8</volume> (<fpage>3</fpage>). <pub-id pub-id-type="doi">10.3390/pathogens8030126</pub-id>.</mixed-citation></ref>
<ref id="c45"><mixed-citation publication-type="journal"><string-name><surname>Khosravi</surname>, <given-names>Arya</given-names></string-name>, and <string-name><given-names>Sarkis K.</given-names> <surname>Mazmanian</surname></string-name>. <year>2013</year>. “<article-title>Disruption of the Gut Microbiome as a Risk Factor for Microbial Infections</article-title>.” <source>Current Opinion in Microbiology</source> <volume>16</volume> (<issue>2</issue>): <fpage>221</fpage>–<lpage>27</lpage>.</mixed-citation></ref>
<ref id="c46"><mixed-citation publication-type="journal"><string-name><surname>Knight</surname>, <given-names>Rob</given-names></string-name>, <string-name><given-names>Chris</given-names> <surname>Callewaert</surname></string-name>, <string-name><given-names>Clarisse</given-names> <surname>Marotz</surname></string-name>, <string-name><given-names>Embriette R.</given-names> <surname>Hyde</surname></string-name>, <string-name><given-names>Justine W.</given-names> <surname>Debelius</surname></string-name>, <string-name><given-names>Daniel</given-names> <surname>McDonald</surname></string-name>, and <string-name><given-names>Mitchell L.</given-names> <surname>Sogin</surname></string-name>. <year>2017</year>. “<article-title>The Microbiome and Human Biology</article-title>.” <source>Annual Review of Genomics and Human Genetics</source> <volume>18</volume> (<issue>August</issue>): <fpage>65</fpage>–<lpage>86</lpage>.</mixed-citation></ref>
<ref id="c47"><mixed-citation publication-type="journal"><string-name><surname>Knox</surname>, <given-names>Natalie C.</given-names></string-name>, <string-name><given-names>Jessica D.</given-names> <surname>Forbes</surname></string-name>, <string-name><given-names>Gary</given-names> <surname>Van Domselaar</surname></string-name>, and <string-name><given-names>Charles N.</given-names> <surname>Bernstein</surname></string-name>. <year>2019</year>. “<article-title>The Gut Microbiome as a Target for IBD Treatment: Are We There Yet?</article-title>” <source>Current Treatment Options in Gastroenterology</source> <volume>17</volume> (<issue>1</issue>): <fpage>115</fpage>–<lpage>26</lpage>.</mixed-citation></ref>
<ref id="c48"><mixed-citation publication-type="journal"><string-name><surname>Koek</surname>, <given-names>Maud M.</given-names></string-name>, <string-name><given-names>Renger H.</given-names> <surname>Jellema</surname></string-name>, <string-name><given-names>Jan</given-names> <surname>van der Greef</surname></string-name>, <string-name><given-names>Albert C.</given-names> <surname>Tas</surname></string-name>, and <string-name><given-names>Thomas</given-names> <surname>Hankemeier</surname></string-name>. <year>2011</year>. “<article-title>Quantitative Metabolomics Based on Gas Chromatography Mass Spectrometry: Status and Perspectives</article-title>.” <source>Metabolomics: Official Journal of the Metabolomic Society</source> <volume>7</volume> (<issue>3</issue>): <fpage>307</fpage>–<lpage>28</lpage>.</mixed-citation></ref>
<ref id="c49"><mixed-citation publication-type="journal"><string-name><surname>Köster</surname>, <given-names>Johannes</given-names></string-name>, and <string-name><given-names>Sven</given-names> <surname>Rahmann</surname></string-name>. <year>2012</year>. “<article-title>Snakemake--a Scalable Bioinformatics Workflow Engine</article-title>.” <source>Bioinformatics</source> <volume>28</volume> (<issue>19</issue>): <fpage>2520</fpage>–<lpage>22</lpage>.</mixed-citation></ref>
<ref id="c50"><mixed-citation publication-type="journal"><string-name><surname>Kostic</surname>, <given-names>Aleksandar D.</given-names></string-name>, <string-name><given-names>Ramnik J.</given-names> <surname>Xavier</surname></string-name>, and <string-name><given-names>Dirk</given-names> <surname>Gevers</surname></string-name>. <year>2014</year>. “<article-title>The Microbiome in Inflammatory Bowel Disease: Current Status and the Future Ahead</article-title>.” <source>Gastroenterology</source> <volume>146</volume> (<issue>6</issue>): <fpage>1489</fpage>–<lpage>99</lpage>.</mixed-citation></ref>
<ref id="c51"><mixed-citation publication-type="journal"><string-name><surname>Kraus</surname>, <given-names>Sarah</given-names></string-name>, and <string-name><given-names>Nadir</given-names> <surname>Arber</surname></string-name>. <year>2009</year>. “<article-title>Inflammation and Colorectal Cancer</article-title>.” <source>Current Opinion in Pharmacology</source> <volume>9</volume> (<issue>4</issue>): <fpage>405</fpage>–<lpage>10</lpage>.</mixed-citation></ref>
<ref id="c52"><mixed-citation publication-type="journal"><string-name><surname>Le Chatelier</surname>, <given-names>Emmanuelle</given-names></string-name>, <string-name><given-names>Trine</given-names> <surname>Nielsen</surname></string-name>, <string-name><given-names>Junjie</given-names> <surname>Qin</surname></string-name>, <string-name><given-names>Edi</given-names> <surname>Prifti</surname></string-name>, <string-name><given-names>Falk</given-names> <surname>Hildebrand</surname></string-name>, <string-name><given-names>Gwen</given-names> <surname>Falony</surname></string-name>, <string-name><given-names>Mathieu</given-names> <surname>Almeida</surname></string-name>, <etal>et al.</etal> <year>2013</year>. “<article-title>Richness of Human Gut Microbiome Correlates with Metabolic Markers</article-title>.” <source>Nature</source> <volume>500</volume> (<issue>7464</issue>): <fpage>541</fpage>–<lpage>46</lpage>.</mixed-citation></ref>
<ref id="c53"><mixed-citation publication-type="journal"><string-name><surname>Lee</surname>, <given-names>Michael D</given-names></string-name>. <year>2019</year>. “<article-title>GToTree: A User-Friendly Workflow for Phylogenomics</article-title>.” <source>Bioinformatics</source> <volume>35</volume> (<issue>20</issue>): <fpage>4162</fpage>–<lpage>64</lpage>.</mixed-citation></ref>
<ref id="c54"><mixed-citation publication-type="journal"><string-name><surname>Lee</surname>, <given-names>Mirae</given-names></string-name>, and <string-name><given-names>Eugene B.</given-names> <surname>Chang</surname></string-name>. <year>2021</year>. “<article-title>Inflammatory Bowel Diseases (IBD) and the Microbiome—Searching the Crime Scene for Clues</article-title>.” <source>Gastroenterology</source> <volume>160</volume> (<issue>2</issue>): <fpage>524</fpage>–<lpage>37</lpage>.</mixed-citation></ref>
<ref id="c55"><mixed-citation publication-type="journal"><string-name><surname>Li</surname>, <given-names>Dinghua</given-names></string-name>, <string-name><given-names>Chi-Man</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Ruibang</given-names> <surname>Luo</surname></string-name>, <string-name><given-names>Kunihiko</given-names> <surname>Sadakane</surname></string-name>, and <string-name><given-names>Tak-Wah</given-names> <surname>Lam</surname></string-name>. <year>2015</year>. “<article-title>MEGAHIT: An Ultra-Fast Single-Node Solution for Large and Complex Metagenomics Assembly via Succinct de Bruijn Graph</article-title>.” <source>Bioinformatics</source> <volume>31</volume> (<issue>10</issue>): <fpage>1674</fpage>–<lpage>76</lpage>.</mixed-citation></ref>
<ref id="c56"><mixed-citation publication-type="journal"><string-name><surname>Lin</surname>, <given-names>Yanping</given-names></string-name>, <string-name><given-names>Gary W.</given-names> <surname>Caldwell</surname></string-name>, <string-name><given-names>Ying</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Wensheng</given-names> <surname>Lang</surname></string-name>, and <string-name><given-names>John</given-names> <surname>Masucci</surname></string-name>. <year>2020</year>. “<article-title>Inter-Laboratory Reproducibility of an Untargeted Metabolomics GC-MS Assay for Analysis of Human Plasma</article-title>.” <source>Scientific Reports</source> <volume>10</volume> (<issue>1</issue>): <fpage>10918</fpage>.</mixed-citation></ref>
<ref id="c57"><mixed-citation publication-type="journal"><string-name><surname>Lloyd-Price</surname>, <given-names>Jason</given-names></string-name>, <string-name><given-names>Cesar</given-names> <surname>Arze</surname></string-name>, <string-name><given-names>Ashwin N.</given-names> <surname>Ananthakrishnan</surname></string-name>, <string-name><given-names>Melanie</given-names> <surname>Schirmer</surname></string-name>, <string-name><given-names>Julian Avila-</given-names> <surname>Pacheco</surname></string-name>, <string-name><given-names>Tiffany W.</given-names> <surname>Poon</surname></string-name>, <string-name><given-names>Elizabeth</given-names> <surname>Andrews</surname></string-name>, <etal>et al.</etal> <year>2019</year>. “<article-title>Multi-Omics of the Gut Microbial Ecosystem in Inflammatory Bowel Diseases</article-title>.” <source>Nature</source> <volume>569</volume> (<issue>7758</issue>): <fpage>655</fpage>–<lpage>62</lpage>.</mixed-citation></ref>
<ref id="c58"><mixed-citation publication-type="journal"><string-name><surname>Lozupone</surname>, <given-names>Catherine A.</given-names></string-name>, <string-name><given-names>Jesse</given-names> <surname>Stombaugh</surname></string-name>, <string-name><given-names>Antonio</given-names> <surname>Gonzalez</surname></string-name>, <string-name><given-names>Gail</given-names> <surname>Ackermann</surname></string-name>, <string-name><given-names>Doug</given-names> <surname>Wendel</surname></string-name>, <string-name><given-names>Yoshiki</given-names> <surname>Vázquez-Baeza</surname></string-name>, <string-name><given-names>Janet K.</given-names> <surname>Jansson</surname></string-name>, <string-name><given-names>Jeffrey I.</given-names> <surname>Gordon</surname></string-name>, and <string-name><given-names>Rob</given-names> <surname>Knight</surname></string-name>. <year>2013</year>. “<article-title>Meta-Analyses of Studies of the Human Microbiota</article-title>.” <source>Genome Research</source> <volume>23</volume> (<issue>10</issue>): <fpage>1704</fpage>–<lpage>14</lpage>.</mixed-citation></ref>
<ref id="c59"><mixed-citation publication-type="journal"><string-name><surname>Machado</surname>, <given-names>Daniel</given-names></string-name>, <string-name><given-names>Sergej</given-names> <surname>Andrejev</surname></string-name>, <string-name><given-names>Melanie</given-names> <surname>Tramontano</surname></string-name>, and <string-name><given-names>Kiran Raosaheb</given-names> <surname>Patil</surname></string-name>. <year>2018</year>. “<article-title>Fast Automated Reconstruction of Genome-Scale Metabolic Models for Microbial Species and Communities</article-title>.” <source>Nucleic Acids Research</source> <volume>46</volume> (<issue>15</issue>): <fpage>7542</fpage>–<lpage>53</lpage>.</mixed-citation></ref>
<ref id="c60"><mixed-citation publication-type="journal"><string-name><surname>Machiels</surname>, <given-names>Kathleen</given-names></string-name>, <string-name><given-names>Marie</given-names> <surname>Joossens</surname></string-name>, <string-name><given-names>João</given-names> <surname>Sabino</surname></string-name>, <string-name><given-names>Vicky</given-names> <surname>De Preter</surname></string-name>, <string-name><given-names>Ingrid</given-names> <surname>Arijs</surname></string-name>, <string-name><given-names>Venessa</given-names> <surname>Eeckhaut</surname></string-name>, <string-name><given-names>Vera</given-names> <surname>Ballet</surname></string-name>, <etal>et al.</etal> <year>2014</year>. “<article-title>A Decrease of the Butyrate-Producing Species Roseburia Hominis and Faecalibacterium Prausnitzii Defines Dysbiosis in Patients with Ulcerative Colitis</article-title>.” <source>Gut</source> <volume>63</volume> (<issue>8</issue>): <fpage>1275</fpage>–<lpage>83</lpage>.</mixed-citation></ref>
<ref id="c61"><mixed-citation publication-type="journal"><string-name><surname>Magnúsdóttir</surname>, <given-names>Stefanía</given-names></string-name>, <string-name><given-names>Dmitry</given-names> <surname>Ravcheev</surname></string-name>, <string-name><given-names>Valérie</given-names> <surname>de Crécy-Lagard</surname></string-name>, and <string-name><given-names>Ines</given-names> <surname>Thiele</surname></string-name>. <year>2015</year>. “<article-title>Systematic Genome Assessment of B-Vitamin Biosynthesis Suggests Co-Operation among Gut Microbes</article-title>.” <source>Frontiers in Genetics</source> <volume>6</volume> (<issue>April</issue>): <fpage>148</fpage>.</mixed-citation></ref>
<ref id="c62"><mixed-citation publication-type="journal"><string-name><surname>Marcelino</surname>, <given-names>Vanessa R.</given-names></string-name>, <string-name><given-names>Caitlin</given-names> <surname>Welsh</surname></string-name>, <string-name><given-names>Christian</given-names> <surname>Diener</surname></string-name>, <string-name><given-names>Emily L.</given-names> <surname>Gulliver</surname></string-name>, <string-name><given-names>Emily L.</given-names> <surname>Rutten</surname></string-name>, <string-name><given-names>Remy B.</given-names> <surname>Young</surname></string-name>, <string-name><given-names>Edward M.</given-names> <surname>Giles</surname></string-name>, <string-name><given-names>Sean M.</given-names> <surname>Gibbons</surname></string-name>, <string-name><given-names>Chris</given-names> <surname>Greening</surname></string-name>, and <string-name><given-names>Samuel C.</given-names> <surname>Forster</surname></string-name>. <year>2023</year>. “<article-title>Disease-Specific Loss of Microbial Cross-Feeding Interactions in the Human Gut</article-title>.” <source>bioRxiv</source>. <pub-id pub-id-type="doi">10.1101/2023.02.17.528570</pub-id>.</mixed-citation></ref>
<ref id="c63"><mixed-citation publication-type="journal"><string-name><surname>Martens</surname>, <given-names>Eric C.</given-names></string-name>, <string-name><given-names>Amelia G.</given-names> <surname>Kelly</surname></string-name>, <string-name><given-names>Alexandra S.</given-names> <surname>Tauzin</surname></string-name>, and <string-name><given-names>Harry</given-names> <surname>Brumer</surname></string-name>. <year>2014</year>. “<article-title>The Devil Lies in the Details: How Variations in Polysaccharide Fine-Structure Impact the Physiology and Evolution of Gut Microbes</article-title>.” <source>Journal of Molecular Biology</source> <volume>426</volume> (<issue>23</issue>): <fpage>3851</fpage>–<lpage>65</lpage>.</mixed-citation></ref>
<ref id="c64"><mixed-citation publication-type="journal"><string-name><surname>Maynard</surname>, <given-names>Craig L.</given-names></string-name>, <string-name><given-names>Charles O.</given-names> <surname>Elson</surname></string-name>, <string-name><given-names>Robin D.</given-names> <surname>Hatton</surname></string-name>, and <string-name><given-names>Casey T.</given-names> <surname>Weaver</surname></string-name>. <year>2012</year>. “<article-title>Reciprocal Interactions of the Intestinal Microbiota and Immune System</article-title>.” <source>Nature</source> <volume>489</volume> (<issue>7415</issue>): <fpage>231</fpage>–<lpage>41</lpage>.</mixed-citation></ref>
<ref id="c65"><mixed-citation publication-type="journal"><string-name><surname>Minh</surname>, <given-names>Bui Quang</given-names></string-name>, <string-name><given-names>Heiko A.</given-names> <surname>Schmidt</surname></string-name>, <string-name><given-names>Olga</given-names> <surname>Chernomor</surname></string-name>, <string-name><given-names>Dominik</given-names> <surname>Schrempf</surname></string-name>, <string-name><given-names>Michael D.</given-names> <surname>Woodhams</surname></string-name>, <string-name><given-names>Arndt</given-names> <surname>von Haeseler</surname></string-name>, and <string-name><given-names>Robert</given-names> <surname>Lanfear</surname></string-name>. <year>2020</year>. “<article-title>IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era</article-title>.” <source>Molecular Biology and Evolution</source> <volume>37</volume> (<issue>5</issue>): <fpage>1530</fpage>–<lpage>34</lpage>.</mixed-citation></ref>
<ref id="c66"><mixed-citation publication-type="journal"><string-name><surname>Mistry</surname>, <given-names>Jaina</given-names></string-name>, <string-name><given-names>Sara</given-names> <surname>Chuguransky</surname></string-name>, <string-name><given-names>Lowri</given-names> <surname>Williams</surname></string-name>, <string-name><given-names>Matloob</given-names> <surname>Qureshi</surname></string-name>, <string-name><given-names>Gustavo A.</given-names> <surname>Salazar</surname></string-name>, <string-name><given-names>Erik L. L.</given-names> <surname>Sonnhammer</surname></string-name>, <string-name><given-names>Silvio C. E.</given-names> <surname>Tosatto</surname></string-name>, <etal>et al.</etal> <year>2021</year>. “<article-title>Pfam: The Protein Families Database in 2021</article-title>.” <source>Nucleic Acids Research</source> <volume>49</volume> (<issue>D1</issue>): <fpage>D412</fpage>–<lpage>19</lpage>.</mixed-citation></ref>
<ref id="c67"><mixed-citation publication-type="journal"><string-name><surname>Morgan</surname>, <given-names>Xochitl C.</given-names></string-name>, <string-name><given-names>Timothy L.</given-names> <surname>Tickle</surname></string-name>, <string-name><given-names>Harry</given-names> <surname>Sokol</surname></string-name>, <string-name><given-names>Dirk</given-names> <surname>Gevers</surname></string-name>, <string-name><given-names>Kathryn L.</given-names> <surname>Devaney</surname></string-name>, <string-name><given-names>Doyle V.</given-names> <surname>Ward</surname></string-name>, <string-name><given-names>Joshua A.</given-names> <surname>Reyes</surname></string-name>, <etal>et al.</etal> <year>2012</year>. “<article-title>Dysfunction of the Intestinal Microbiome in Inflammatory Bowel Disease and Treatment</article-title>.” <source>Genome Biology</source> <volume>13</volume> (<issue>9</issue>): <fpage>R79</fpage>.</mixed-citation></ref>
<ref id="c68"><mixed-citation publication-type="journal"><string-name><surname>Nagalingam</surname>, <given-names>Nabeetha A.</given-names></string-name>, and <string-name><given-names>Susan V.</given-names> <surname>Lynch</surname></string-name>. <year>2012</year>. “<article-title>Role of the Microbiota in Inflammatory Bowel Diseases</article-title>.” <source>Inflammatory Bowel Diseases</source> <volume>18</volume> (<issue>5</issue>): <fpage>968</fpage>–<lpage>84</lpage>.</mixed-citation></ref>
<ref id="c69"><mixed-citation publication-type="journal"><string-name><surname>Nishida</surname>, <given-names>Atsushi</given-names></string-name>, <string-name><given-names>Ryo</given-names> <surname>Inoue</surname></string-name>, <string-name><given-names>Osamu</given-names> <surname>Inatomi</surname></string-name>, <string-name><given-names>Shigeki</given-names> <surname>Bamba</surname></string-name>, <string-name><given-names>Yuji</given-names> <surname>Naito</surname></string-name>, and <string-name><given-names>Akira</given-names> <surname>Andoh</surname></string-name>. <year>2018</year>. “<article-title>Gut Microbiota in the Pathogenesis of Inflammatory Bowel Disease</article-title>.” <source>Clinical Journal of Gastroenterology</source> <volume>11</volume> (<issue>1</issue>): <fpage>1</fpage>–<lpage>10</lpage>.</mixed-citation></ref>
<ref id="c70"><mixed-citation publication-type="journal"><string-name><surname>Nitzan</surname>, <given-names>Orna</given-names></string-name>, <string-name><given-names>Mazen</given-names> <surname>Elias</surname></string-name>, <string-name><given-names>Avi</given-names> <surname>Peretz</surname></string-name>, and <string-name><given-names>Walid</given-names> <surname>Saliba</surname></string-name>. <year>2016</year>. “<article-title>Role of Antibiotics for Treatment of Inflammatory Bowel Disease</article-title>.” <source>World Journal of Gastroenterology: WJG</source> <volume>22</volume> (<issue>3</issue>): <fpage>1078</fpage>–<lpage>87</lpage>.</mixed-citation></ref>
<ref id="c71"><mixed-citation publication-type="journal"><string-name><surname>Palleja</surname>, <given-names>Albert</given-names></string-name>, <string-name><given-names>Kristian H.</given-names> <surname>Mikkelsen</surname></string-name>, <string-name><given-names>Sofia K.</given-names> <surname>Forslund</surname></string-name>, <string-name><given-names>Alireza</given-names> <surname>Kashani</surname></string-name>, <string-name><given-names>Kristine H.</given-names> <surname>Allin</surname></string-name>, <string-name><given-names>Trine</given-names> <surname>Nielsen</surname></string-name>, <string-name><given-names>Tue H.</given-names> <surname>Hansen</surname></string-name>, <etal>et al.</etal> <year>2018</year>. “<article-title>Recovery of Gut Microbiota of Healthy Adults Following Antibiotic Exposure</article-title>.” <source>Nature Microbiology</source> <volume>3</volume> (<issue>11</issue>): <fpage>1255</fpage>–<lpage>65</lpage>.</mixed-citation></ref>
<ref id="c72"><mixed-citation publication-type="journal"><string-name><surname>Palù</surname>, <given-names>Matteo</given-names></string-name>, <string-name><given-names>Arianna</given-names> <surname>Basile</surname></string-name>, <string-name><given-names>Guido</given-names> <surname>Zampieri</surname></string-name>, <string-name><given-names>Laura</given-names> <surname>Treu</surname></string-name>, <string-name><given-names>Alessandro</given-names> <surname>Rossi</surname></string-name>, <string-name><given-names>Maria Silvia</given-names> <surname>Morlino</surname></string-name>, and <string-name><given-names>Stefano</given-names> <surname>Campanaro</surname></string-name>. <year>2022</year>. “<article-title>KEMET – A Python Tool for KEGG Module Evaluation and Microbial Genome Annotation Expansion</article-title>.” <source>Computational and Structural Biotechnology Journal</source>. <pub-id pub-id-type="doi">10.1016/j.csbj.2022.03.015</pub-id>.</mixed-citation></ref>
<ref id="c73"><mixed-citation publication-type="journal"><string-name><surname>Papa</surname>, <given-names>Eliseo</given-names></string-name>, <string-name><given-names>Michael</given-names> <surname>Docktor</surname></string-name>, <string-name><given-names>Christopher</given-names> <surname>Smillie</surname></string-name>, <string-name><given-names>Sarah</given-names> <surname>Weber</surname></string-name>, <string-name><given-names>Sarah P.</given-names> <surname>Preheim</surname></string-name>, <string-name><given-names>Dirk</given-names> <surname>Gevers</surname></string-name>, <string-name><given-names>Georgia</given-names> <surname>Giannoukos</surname></string-name>, <etal>et al.</etal> <year>2012</year>. “<article-title>Non-Invasive Mapping of the Gastrointestinal Microbiota Identifies Children with Inflammatory Bowel Disease</article-title>.” <source>PloS One</source> <volume>7</volume> (<issue>6</issue>): <fpage>e39242</fpage>.</mixed-citation></ref>
<ref id="c74"><mixed-citation publication-type="journal"><string-name><surname>Parks</surname>, <given-names>Donovan H.</given-names></string-name>, <string-name><given-names>Maria</given-names> <surname>Chuvochina</surname></string-name>, <string-name><given-names>Pierre-Alain</given-names> <surname>Chaumeil</surname></string-name>, <string-name><given-names>Christian</given-names> <surname>Rinke</surname></string-name>, <string-name><given-names>Aaron J.</given-names> <surname>Mussig</surname></string-name>, and <string-name><given-names>Philip</given-names> <surname>Hugenholtz</surname></string-name>. <year>2020</year>. “<article-title>A Complete Domain-to-Species Taxonomy for Bacteria and Archaea</article-title>.” <source>Nature Biotechnology</source> <volume>38</volume> (<issue>9</issue>): <fpage>1079</fpage>–<lpage>86</lpage>.</mixed-citation></ref>
<ref id="c75"><mixed-citation publication-type="journal"><string-name><surname>Parks</surname>, <given-names>Donovan H.</given-names></string-name>, <string-name><given-names>Maria</given-names> <surname>Chuvochina</surname></string-name>, <string-name><given-names>Christian</given-names> <surname>Rinke</surname></string-name>, <string-name><given-names>Aaron J.</given-names> <surname>Mussig</surname></string-name>, <string-name><given-names>Pierre-Alain</given-names> <surname>Chaumeil</surname></string-name>, and <string-name><given-names>Philip</given-names> <surname>Hugenholtz</surname></string-name>. <year>2022</year>. “<article-title>GTDB: An Ongoing Census of Bacterial and Archaeal Diversity through a Phylogenetically Consistent, Rank Normalized and Complete Genome-Based Taxonomy</article-title>.” <source>Nucleic Acids Research</source> <volume>50</volume> (<issue>D1</issue>): <fpage>D785</fpage>–<lpage>94</lpage>.</mixed-citation></ref>
<ref id="c76"><mixed-citation publication-type="journal"><string-name><surname>Parks</surname>, <given-names>Donovan H.</given-names></string-name>, <string-name><given-names>Maria</given-names> <surname>Chuvochina</surname></string-name>, <string-name><given-names>David W.</given-names> <surname>Waite</surname></string-name>, <string-name><given-names>Christian</given-names> <surname>Rinke</surname></string-name>, <string-name><given-names>Adam</given-names> <surname>Skarshewski</surname></string-name>, <string-name><given-names>Pierre-Alain</given-names> <surname>Chaumeil</surname></string-name>, and <string-name><given-names>Philip</given-names> <surname>Hugenholtz</surname></string-name>. <year>2018</year>. “<article-title>A Standardized Bacterial Taxonomy Based on Genome Phylogeny Substantially Revises the Tree of Life</article-title>.” <source>Nature Biotechnology</source> <volume>36</volume> (<issue>10</issue>): <fpage>996</fpage>–<lpage>1004</lpage>.</mixed-citation></ref>
<ref id="c77"><mixed-citation publication-type="journal"><string-name><surname>Peng</surname>, <given-names>Yu</given-names></string-name>, <string-name><given-names>Henry C. M.</given-names> <surname>Leung</surname></string-name>, <string-name><given-names>S. M.</given-names> <surname>Yiu</surname></string-name>, and <string-name><given-names>Francis Y. L.</given-names> <surname>Chin</surname></string-name>. <year>2012</year>. “<article-title>IDBA-UD: A de Novo Assembler for Single-Cell and Metagenomic Sequencing Data with Highly Uneven Depth</article-title>.” <source>Bioinformatics</source> <volume>28</volume> (<issue>11</issue>): <fpage>1420</fpage>–<lpage>28</lpage>.</mixed-citation></ref>
<ref id="c78"><mixed-citation publication-type="journal"><string-name><surname>Powell</surname>, <given-names>J. Elijah</given-names></string-name>, <string-name><given-names>Sean P.</given-names> <surname>Leonard</surname></string-name>, <string-name><given-names>Waldan K.</given-names> <surname>Kwong</surname></string-name>, <string-name><given-names>Philipp</given-names> <surname>Engel</surname></string-name>, and <string-name><given-names>Nancy A.</given-names> <surname>Moran</surname></string-name>. <year>2016</year>. “<article-title>Genome-Wide Screen Identifies Host Colonization Determinants in a Bacterial Gut Symbiont</article-title>.” <source>Proceedings of the National Academy of Sciences of the United States of America</source> <volume>113</volume> (<issue>48</issue>): <fpage>13887</fpage>–<lpage>92</lpage>.</mixed-citation></ref>
<ref id="c79"><mixed-citation publication-type="journal"><string-name><surname>Prindiville</surname>, <given-names>T. P.</given-names></string-name>, <string-name><given-names>R. A.</given-names> <surname>Sheikh</surname></string-name>, <string-name><given-names>S. H.</given-names> <surname>Cohen</surname></string-name>, <string-name><given-names>Y. J.</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>M. C.</given-names> <surname>Cantrell</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Silva Jr</surname></string-name>. <year>2000</year>. “<article-title>Bacteroides Fragilis Enterotoxin Gene Sequences in Patients with Inflammatory Bowel Disease</article-title>.” <source>Emerging Infectious Diseases</source> <volume>6</volume> (<issue>2</issue>): <fpage>171</fpage>–<lpage>74</lpage>.</mixed-citation></ref>
<ref id="c80"><mixed-citation publication-type="journal"><string-name><surname>Qin</surname>, <given-names>Junjie</given-names></string-name>, <string-name><given-names>Yingrui</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Zhiming</given-names> <surname>Cai</surname></string-name>, <string-name><given-names>Shenghui</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Jianfeng</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>Fan</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Suisha</given-names> <surname>Liang</surname></string-name>, <etal>et al.</etal> <year>2012</year>. “<article-title>A Metagenome-Wide Association Study of Gut Microbiota in Type 2 Diabetes</article-title>.” <source>Nature</source> <volume>490</volume> (<issue>7418</issue>): <fpage>55</fpage>–<lpage>60</lpage>.</mixed-citation></ref>
<ref id="c81"><mixed-citation publication-type="journal"><string-name><surname>Quince</surname>, <given-names>Christopher</given-names></string-name>, <string-name><given-names>Umer Zeeshan</given-names> <surname>Ijaz</surname></string-name>, <string-name><given-names>Nick</given-names> <surname>Loman</surname></string-name>, <string-name><given-names>Murat A.</given-names> <surname>Eren</surname></string-name>, <string-name><given-names>Delphine</given-names> <surname>Saulnier</surname></string-name>, <string-name><given-names>Julie</given-names> <surname>Russell</surname></string-name>, <string-name><given-names>Sarah J.</given-names> <surname>Haig</surname></string-name>, <etal>et al.</etal> <year>2015</year>. “<article-title>Extensive Modulation of the Fecal Metagenome in Children With Crohn’s Disease During Exclusive Enteral Nutrition</article-title>.” <source>American Journal of Gastroenterology</source>. <pub-id pub-id-type="doi">10.1038/ajg.2015.357</pub-id>.</mixed-citation></ref>
<ref id="c82"><mixed-citation publication-type="journal"><string-name><surname>Ramirez</surname>, <given-names>Jaime</given-names></string-name>, <string-name><given-names>Francisco</given-names> <surname>Guarner</surname></string-name>, <string-name><given-names>Luis Bustos</given-names> <surname>Fernandez</surname></string-name>, <string-name><given-names>Aldo</given-names> <surname>Maruy</surname></string-name>, <string-name><given-names>Vera Lucia</given-names> <surname>Sdepanian</surname></string-name>, and <string-name><given-names>Henry</given-names> <surname>Cohen</surname></string-name>. <year>2020</year>. “<article-title>Antibiotics as Major Disruptors of Gut Microbiota</article-title>.” <source>Frontiers in Cellular and Infection Microbiology</source> <volume>10</volume> (<issue>November</issue>): <fpage>572912</fpage>.</mixed-citation></ref>
<ref id="c83"><mixed-citation publication-type="journal"><string-name><surname>Rampelli</surname>, <given-names>Simone</given-names></string-name>, <string-name><given-names>Stephanie L.</given-names> <surname>Schnorr</surname></string-name>, <string-name><given-names>Clarissa</given-names> <surname>Consolandi</surname></string-name>, <string-name><given-names>Silvia</given-names> <surname>Turroni</surname></string-name>, <string-name><given-names>Marco</given-names> <surname>Severgnini</surname></string-name>, <string-name><given-names>Clelia</given-names> <surname>Peano</surname></string-name>, <string-name><given-names>Patrizia</given-names> <surname>Brigidi</surname></string-name>, <string-name><given-names>Alyssa N.</given-names> <surname>Crittenden</surname></string-name>, <string-name><given-names>Amanda G.</given-names> <surname>Henry</surname></string-name>, and <string-name><given-names>Marco</given-names> <surname>Candela</surname></string-name>. <year>2015</year>. “<article-title>Metagenome Sequencing of the Hadza Hunter-Gatherer Gut Microbiota</article-title>.” <source>Current Biology: CB</source> <volume>25</volume> (<issue>13</issue>): <fpage>1682</fpage>–<lpage>93</lpage>.</mixed-citation></ref>
<ref id="c84"><mixed-citation publication-type="journal"><string-name><surname>Raymond</surname>, <given-names>Frédéric</given-names></string-name>, <string-name><given-names>Amin A.</given-names> <surname>Ouameur</surname></string-name>, <string-name><given-names>Maxime</given-names> <surname>Déraspe</surname></string-name>, <string-name><given-names>Naeem</given-names> <surname>Iqbal</surname></string-name>, <string-name><given-names>Hélène</given-names> <surname>Gingras</surname></string-name>, <string-name><given-names>Bédis</given-names> <surname>Dridi</surname></string-name>, <string-name><given-names>Philippe</given-names> <surname>Leprohon</surname></string-name>, <etal>et al.</etal> <year>2016</year>. “<article-title>The Initial State of the Human Gut Microbiome Determines Its Reshaping by Antibiotics</article-title>.” <source>The ISME Journal</source> <volume>10</volume> (<issue>3</issue>): <fpage>707</fpage>–<lpage>20</lpage>.</mixed-citation></ref>
<ref id="c85"><mixed-citation publication-type="journal"><string-name><surname>Reiter</surname>, <given-names>Taylor E.</given-names></string-name>, <string-name><given-names>Luiz</given-names> <surname>Irber</surname></string-name>, <string-name><given-names>Alicia A.</given-names> <surname>Gingrich</surname></string-name>, <string-name><given-names>Dylan</given-names> <surname>Haynes</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Tessa Pierce-Ward</surname></string-name>, <string-name><given-names>Phillip T.</given-names> <surname>Brooks</surname></string-name>, <string-name><given-names>Yosuke</given-names> <surname>Mizutani</surname></string-name>, <etal>et al.</etal> <year>2022</year>. “<article-title>Meta-Analysis of Metagenomes via Machine Learning and Assembly Graphs Reveals Strain Switches in Crohn’s Disease</article-title>.” <source>bioRxiv</source>. <pub-id pub-id-type="doi">10.1101/2022.06.30.498290</pub-id>.</mixed-citation></ref>
<ref id="c86"><mixed-citation publication-type="journal"><string-name><surname>Rhodes</surname>, <given-names>Jonathan M</given-names></string-name>. <year>2007</year>. “<article-title>The Role of Escherichia Coli in Inflammatory Bowel Disease</article-title>.” <source>Gut</source>.</mixed-citation></ref>
<ref id="c87"><mixed-citation publication-type="journal"><string-name><surname>Saitoh</surname>, <given-names>Shin</given-names></string-name>, <string-name><given-names>Satoshi</given-names> <surname>Noda</surname></string-name>, <string-name><given-names>Yuji</given-names> <surname>Aiba</surname></string-name>, <string-name><given-names>Atsushi</given-names> <surname>Takagi</surname></string-name>, <string-name><given-names>Mitsuo</given-names> <surname>Sakamoto</surname></string-name>, <string-name><given-names>Yoshimi</given-names> <surname>Benno</surname></string-name>, and <string-name><given-names>Yasuhiro</given-names> <surname>Koga</surname></string-name>. <year>2002</year>. “<article-title>Bacteroides Ovatus as the Predominant Commensal Intestinal Microbe Causing a Systemic Antibody Response in Inflammatory Bowel Disease</article-title>.” <source>Clinical and Diagnostic Laboratory Immunology</source> <volume>9</volume> (<issue>1</issue>): <fpage>54</fpage>–<lpage>59</lpage>.</mixed-citation></ref>
<ref id="c88"><mixed-citation publication-type="journal"><string-name><given-names>Sartor R.</given-names> <surname>Balfour</surname></string-name>. <year>2006</year>. “<article-title>Mechanisms of Disease: Pathogenesis of Crohn’s Disease and Ulcerative Colitis</article-title>.” <source>Nature Clinical Practice. Gastroenterology &amp; Hepatology</source> <volume>3</volume> (<issue>7</issue>): <fpage>390</fpage>–<lpage>407</lpage>.</mixed-citation></ref>
<ref id="c89"><mixed-citation publication-type="journal"><string-name><surname>Schirmer</surname>, <given-names>Melanie</given-names></string-name>, <string-name><given-names>Eric A.</given-names> <surname>Franzosa</surname></string-name>, <string-name><given-names>Jason</given-names> <surname>Lloyd-Price</surname></string-name>, <string-name><given-names>Lauren J.</given-names> <surname>McIver</surname></string-name>, <string-name><given-names>Randall</given-names> <surname>Schwager</surname></string-name>, <string-name><given-names>Tiffany W.</given-names> <surname>Poon</surname></string-name>, <string-name><given-names>Ashwin N.</given-names> <surname>Ananthakrishnan</surname></string-name>, <etal>et al.</etal> <year>2018</year>. “<article-title>Dynamics of Metatranscription in the Inflammatory Bowel Disease Gut Microbiome</article-title>.” <source>Nature Microbiology</source> <volume>3</volume> (<issue>3</issue>): <fpage>337</fpage>–<lpage>46</lpage>.</mixed-citation></ref>
<ref id="c90"><mixed-citation publication-type="journal"><string-name><surname>Schirmer</surname>, <given-names>Melanie</given-names></string-name>, <string-name><given-names>Ashley</given-names> <surname>Garner</surname></string-name>, <string-name><given-names>Hera</given-names> <surname>Vlamakis</surname></string-name>, and <string-name><given-names>Ramnik J.</given-names> <surname>Xavier</surname></string-name>. <year>2019</year>. “<article-title>Microbial Genes and Pathways in Inflammatory Bowel Disease</article-title>.” <source>Nature Reviews. Microbiology</source> <volume>17</volume> (<issue>8</issue>): <fpage>497</fpage>–<lpage>511</lpage>.</mixed-citation></ref>
<ref id="c91"><mixed-citation publication-type="web"><string-name><surname>Seemann</surname>, <given-names>Torsten</given-names></string-name>. n.d. <italic>Barrnap: Bacterial Ribosomal RNA Predictor</italic>. Github. Accessed <date-in-citation content-type="access-date">September 30, 2022</date-in-citation>. <ext-link ext-link-type="uri" xlink:href="https://github.com/tseemann/barrnap">https://github.com/tseemann/barrnap</ext-link>.</mixed-citation></ref>
<ref id="c92"><mixed-citation publication-type="journal"><string-name><surname>Shaffer</surname>, <given-names>Michael</given-names></string-name>, <string-name><given-names>Mikayla A.</given-names> <surname>Borton</surname></string-name>, <string-name><given-names>Bridget B.</given-names> <surname>McGivern</surname></string-name>, <string-name><given-names>Ahmed A.</given-names> <surname>Zayed</surname></string-name>, <string-name><given-names>Sabina Leanti</given-names> <surname>La Rosa</surname></string-name>, <string-name><given-names>Lindsey M.</given-names> <surname>Solden</surname></string-name>, <string-name><given-names>Pengfei</given-names> <surname>Liu</surname></string-name>, <etal>et al.</etal> <year>2020</year>. “<article-title>DRAM for Distilling Microbial Metabolism to Automate the Curation of Microbiome Function</article-title>.” <source>Nucleic Acids Research</source> <volume>48</volume> (<issue>16</issue>): <fpage>8883</fpage>–<lpage>8900</lpage>.</mixed-citation></ref>
<ref id="c93"><mixed-citation publication-type="journal"><string-name><surname>Shaiber</surname>, <given-names>Alon</given-names></string-name>, <string-name><given-names>Amy D.</given-names> <surname>Willis</surname></string-name>, <string-name><given-names>Tom O.</given-names> <surname>Delmont</surname></string-name>, <string-name><given-names>Simon</given-names> <surname>Roux</surname></string-name>, <string-name><given-names>Lin-Xing</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Abigail C.</given-names> <surname>Schmid</surname></string-name>, <string-name><given-names>Mahmoud</given-names> <surname>Yousef</surname></string-name>, <etal>et al.</etal> <year>2020</year>. “<article-title>Functional and Genetic Markers of Niche Partitioning among Enigmatic Members of the Human Oral Microbiome</article-title>.” <source>Genome Biology</source> <volume>21</volume> (<issue>1</issue>): <fpage>292</fpage>.</mixed-citation></ref>
<ref id="c94"><mixed-citation publication-type="journal"><string-name><surname>Shan</surname>, <given-names>Yue</given-names></string-name>, <string-name><given-names>Mirae</given-names> <surname>Lee</surname></string-name>, and <string-name><given-names>Eugene B.</given-names> <surname>Chang</surname></string-name>. <year>2022</year>. “<article-title>The Gut Microbiome and Inflammatory Bowel Diseases</article-title>.” <source>Annual Review of Medicine</source> <volume>73</volume> (<issue>January</issue>): <fpage>455</fpage>–<lpage>68</lpage>.</mixed-citation></ref>
<ref id="c95"><mixed-citation publication-type="journal"><string-name><surname>Sinha</surname>, <given-names>Rashmi</given-names></string-name>, <string-name><given-names>Galeb</given-names> <surname>Abu-Ali</surname></string-name>, <string-name><given-names>Emily</given-names> <surname>Vogtmann</surname></string-name>, <string-name><given-names>Anthony A.</given-names> <surname>Fodor</surname></string-name>, <string-name><given-names>Boyu</given-names> <surname>Ren</surname></string-name>, <string-name><given-names>Amnon</given-names> <surname>Amir</surname></string-name>, <string-name><given-names>Emma</given-names> <surname>Schwager</surname></string-name>, <etal>et al.</etal> <year>2017</year>. “<article-title>Assessment of Variation in Microbial Community Amplicon Sequencing by the Microbiome Quality Control (MBQC) Project Consortium</article-title>.” <source>Nature Biotechnology</source> <volume>35</volume> (<issue>11</issue>): <fpage>1077</fpage>–<lpage>86</lpage>.</mixed-citation></ref>
<ref id="c96"><mixed-citation publication-type="journal"><string-name><surname>Sorbara</surname>, <given-names>Matthew T.</given-names></string-name>, and <string-name><given-names>Eric G.</given-names> <surname>Pamer</surname></string-name>. <year>2022</year>. “<article-title>Microbiome-Based Therapeutics</article-title>.” <source>Nature Reviews. Microbiology</source> <volume>20</volume> (<issue>6</issue>): <fpage>365</fpage>–<lpage>80</lpage>.</mixed-citation></ref>
<ref id="c97"><mixed-citation publication-type="journal"><string-name><surname>Steinegger</surname>, <given-names>Martin</given-names></string-name>, and <string-name><given-names>Johannes</given-names> <surname>Söding</surname></string-name>. <year>2017</year>. “<article-title>MMseqs2 Enables Sensitive Protein Sequence Searching for the Analysis of Massive Data Sets</article-title>.” <source>Nature Biotechnology</source> <volume>35</volume> (<issue>11</issue>): <fpage>1026</fpage>–<lpage>28</lpage>.</mixed-citation></ref>
<ref id="c98"><mixed-citation publication-type="journal"><string-name><surname>Turnbaugh</surname>, <given-names>Peter J.</given-names></string-name>, <string-name><given-names>Micah</given-names> <surname>Hamady</surname></string-name>, <string-name><given-names>Tanya</given-names> <surname>Yatsunenko</surname></string-name>, <string-name><given-names>Brandi L.</given-names> <surname>Cantarel</surname></string-name>, <string-name><given-names>Alexis</given-names> <surname>Duncan</surname></string-name>, <string-name><given-names>Ruth E.</given-names> <surname>Ley</surname></string-name>, <string-name><given-names>Mitchell L.</given-names> <surname>Sogin</surname></string-name>, <etal>et al.</etal> <year>2009</year>. “<article-title>A Core Gut Microbiome in Obese and Lean Twins</article-title>.” <source>Nature</source> <volume>457</volume> (<issue>7228</issue>): <fpage>480</fpage>–<lpage>84</lpage>.</mixed-citation></ref>
<ref id="c99"><mixed-citation publication-type="other"><string-name><surname>Vineis</surname>, <given-names>J. H.</given-names></string-name>, <string-name><given-names>D. L.</given-names> <surname>Ringus</surname></string-name>, <string-name><given-names>H. G.</given-names> <surname>Morrison</surname></string-name>, <string-name><given-names>T. O.</given-names> <surname>Delmont</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Dalal</surname></string-name>, <string-name><given-names>L. H.</given-names> <surname>Raffals</surname></string-name>, <string-name><given-names>D. A.</given-names> <surname>Antonopoulos</surname></string-name>, <etal>et al.</etal> <year>2016</year>. “<article-title>Patient-Specific Bacteroides Genome Variants in Pouchitis</article-title>. <source>mBio 7</source>.”</mixed-citation></ref>
<ref id="c100"><mixed-citation publication-type="journal"><string-name><surname>Watson</surname>, <given-names>Andrea R.</given-names></string-name>, <string-name><given-names>Jessika</given-names> <surname>Füssel</surname></string-name>, <string-name><given-names>Iva</given-names> <surname>Veseli</surname></string-name>, <string-name><given-names>Johanna Zaal</given-names> <surname>DeLongchamp</surname></string-name>, <string-name><given-names>Marisela</given-names> <surname>Silva</surname></string-name>, <string-name><given-names>Florian</given-names> <surname>Trigodet</surname></string-name>, <string-name><given-names>Karen</given-names> <surname>Lolans</surname></string-name>, <etal>et al.</etal> <year>2023</year>. “<article-title>Metabolic Independence Drives Gut Microbial Colonization and Resilience in Health and Disease</article-title>.” <source>Genome Biology</source> <volume>24</volume> (<issue>1</issue>): <fpage>78</fpage>.</mixed-citation></ref>
<ref id="c101"><mixed-citation publication-type="journal"><string-name><surname>Wen</surname>, <given-names>Chengping</given-names></string-name>, <string-name><given-names>Zhijun</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>Tiejuan</given-names> <surname>Shao</surname></string-name>, <string-name><given-names>Lin</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Zhijun</given-names> <surname>Xie</surname></string-name>, <string-name><given-names>Emmanuelle</given-names> <surname>Le Chatelier</surname></string-name>, <string-name><given-names>Zhixing</given-names> <surname>He</surname></string-name>, <etal>et al.</etal> <year>2017</year>. “<article-title>Quantitative Metagenomics Reveals Unique Gut Microbiome Biomarkers in Ankylosing Spondylitis</article-title>.” <source>Genome Biology</source> <volume>18</volume> (<issue>1</issue>): <fpage>142</fpage>.</mixed-citation></ref>
<ref id="c102"><mixed-citation publication-type="book"><string-name><surname>Wickham</surname>, <given-names>Hadley</given-names></string-name>. <year>2016</year>. <source>ggplot2: Elegant Graphics for Data Analysis</source>. <publisher-name>Springer-Verlag</publisher-name> <publisher-name>New York</publisher-name>.</mixed-citation></ref>
<ref id="c103"><mixed-citation publication-type="journal"><string-name><surname>Woting</surname>, <given-names>Anni</given-names></string-name>, and <string-name><given-names>Michael</given-names> <surname>Blaut</surname></string-name>. <year>2016</year>. “<article-title>The Intestinal Microbiota in Metabolic Disease</article-title>.” <source>Nutrients</source> <volume>8</volume> (<issue>4</issue>): <fpage>202</fpage>.</mixed-citation></ref>
<ref id="c104"><mixed-citation publication-type="journal"><string-name><surname>Xie</surname>, <given-names>Hailiang</given-names></string-name>, <string-name><given-names>Ruijin</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>Huanzi</given-names> <surname>Zhong</surname></string-name>, <string-name><given-names>Qiang</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>Zhou</given-names> <surname>Lan</surname></string-name>, <string-name><given-names>Bingcai</given-names> <surname>Qin</surname></string-name>, <string-name><given-names>Kirsten J.</given-names> <surname>Ward</surname></string-name>, <etal>et al.</etal> <year>2016</year>. “<article-title>Shotgun Metagenomics of 250 Adult Twins Reveals Genetic and Environmental Impacts on the Gut Microbiome</article-title>.” <source>Cell Systems</source> <volume>3</volume> (<issue>6</issue>): <fpage>572</fpage>–<lpage>84</lpage>.e3.</mixed-citation></ref>
<ref id="c105"><mixed-citation publication-type="journal"><string-name><surname>Ye</surname>, <given-names>Yuzhen</given-names></string-name>, and <string-name><given-names>Thomas G.</given-names> <surname>Doak</surname></string-name>. <year>2009</year>. “<article-title>A Parsimony Approach to Biological Pathway Reconstruction/inference for Genomes and Metagenomes</article-title>.” <source>PLoS Computational Biology</source> <volume>5</volume> (<issue>8</issue>): <fpage>e1000465</fpage>.</mixed-citation></ref>
<ref id="c106"><mixed-citation publication-type="other"><string-name><surname>Zhou</surname>, <given-names>Zhichao</given-names></string-name>, <string-name><given-names>Patricia Q.</given-names> <surname>Tran</surname></string-name>, <string-name><given-names>Adam M.</given-names> <surname>Breister</surname></string-name>, <string-name><given-names>Yang</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Kristopher</given-names> <surname>Kieft</surname></string-name>, <string-name><given-names>Elise S.</given-names> <surname>Cowley</surname></string-name>, <string-name><given-names>Ulas</given-names> <surname>Karaoz</surname></string-name>, and <string-name><given-names>Karthik</given-names> <surname>Anantharaman</surname></string-name>. n.d. “<article-title>METABOLIC: High-Throughput Profiling of Microbial Genomes for Functional Traits, Metabolism, Biogeochemistry, and Community-Scale Functional Networks</article-title>.” <pub-id pub-id-type="doi">10.21203/rs.3.rs-965097/v1</pub-id>.</mixed-citation></ref>
<ref id="c107"><mixed-citation publication-type="journal"><string-name><surname>Zimmermann</surname>, <given-names>Johannes</given-names></string-name>, <string-name><given-names>Christoph</given-names> <surname>Kaleta</surname></string-name>, and <string-name><given-names>Silvio</given-names> <surname>Waschina</surname></string-name>. <year>2021</year>. “<article-title>Gapseq: Informed Prediction of Bacterial Metabolic Pathways and Reconstruction of Accurate Metabolic Models</article-title>.” <source>Genome Biology</source> <volume>22</volume> (<issue>1</issue>): <fpage>81</fpage>.</mixed-citation></ref>
<ref id="c108"><mixed-citation publication-type="journal"><string-name><surname>Zimmermann</surname>, <given-names>Michael</given-names></string-name>, <string-name><given-names>Maria</given-names> <surname>Zimmermann-Kogadeeva</surname></string-name>, <string-name><given-names>Rebekka</given-names> <surname>Wegmann</surname></string-name>, and <string-name><given-names>Andrew L.</given-names> <surname>Goodman</surname></string-name>. <year>2019</year>. “<article-title>Mapping Human Microbiome Drug Metabolism by Gut Bacteria and Their Genes</article-title>.” <source>Nature</source> <volume>570</volume> (<issue>7762</issue>): <fpage>462</fpage>–<lpage>67</lpage>.</mixed-citation></ref>
<ref id="c109"><mixed-citation publication-type="journal"><string-name><surname>Zong</surname>, <given-names>Xin</given-names></string-name>, <string-name><given-names>Jie</given-names> <surname>Fu</surname></string-name>, <string-name><given-names>Bocheng</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>Yizhen</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>Mingliang</given-names> <surname>Jin</surname></string-name>. <year>2020</year>. “<article-title>Interplay between Gut Microbiota and Antimicrobial Peptides</article-title>.” <source>Animal Nutrition (Zhongguo Xu Mu Shou Yi Xue Hui</source><italic>)</italic> <volume>6</volume> (<issue>4</issue>): <fpage>389</fpage>–<lpage>96</lpage>.</mixed-citation></ref>
<ref id="c110"><mixed-citation publication-type="journal"><string-name><surname>Zorrilla</surname>, <given-names>Francisco</given-names></string-name>, <string-name><given-names>Filip</given-names> <surname>Buric</surname></string-name>, <string-name><given-names>Kiran R.</given-names> <surname>Patil</surname></string-name>, and <string-name><given-names>Aleksej</given-names> <surname>Zelezniak</surname></string-name>. <year>2021</year>. “<article-title>metaGEM: Reconstruction of Genome Scale Metabolic Models Directly from Metagenomes</article-title>.” <source>Nucleic Acids Research</source> <volume>49</volume> (<issue>21</issue>): <fpage>e126</fpage>.</mixed-citation></ref>
<ref id="c111"><mixed-citation publication-type="journal"><string-name><surname>Boyle</surname>, <given-names>Elizabeth I</given-names></string-name>, <string-name><given-names>Shuai</given-names> <surname>Weng</surname></string-name>, <string-name><given-names>Jeremy</given-names> <surname>Gollub</surname></string-name>, <string-name><given-names>Heng</given-names> <surname>Jin</surname></string-name>, <string-name><given-names>David</given-names> <surname>Botstein</surname></string-name>, <string-name><given-names>J</given-names> <surname>Michael Cherry</surname></string-name>, and <string-name><given-names>Gavin</given-names> <surname>Sherlock</surname></string-name>. <year>2004</year>. <article-title>“GO::TermFinder–open Source Software for Accessing Gene Ontology Information and Finding Significantly Enriched Gene Ontology Terms Associated with a List of Genes.” Bioinformatics (Oxford</article-title>, <source>England</source>) <volume>20</volume> (<issue>18</issue>): <fpage>3710</fpage>–<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bth456</pub-id>.</mixed-citation></ref>
</ref-list>
<sec id="d1e6260">
<title>Supplemental Figures</title>
<fig id="figs1" position="float" orientation="portrait" fig-type="figure">
<label>Supplementary Figure 1.</label>
<caption><title>Scatterplot of sequencing depth vs estimated number of microbial populations in each of 2,893 stool metagenomes.</title>
<p>Sequencing depth is represented by the number of R1 reads, except for (<xref ref-type="bibr" rid="c99">Vineis et al. 2016</xref>) samples, in which case it is the number of merged paired-end reads. The vertical line indicates our sequencing depth threshold of 25 million reads. Per-group Spearman’s correlation coefficients and p-values are shown for the subset of samples with depth &lt; 25 million reads (top left) and for the subset with depth &gt;= 25 million reads (top right). Regression lines are shown for each group in each subset, with standard error indicated by the colored background.</p></caption>
<graphic xlink:href="540289v2_figs1.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
<fig id="figs2" position="float" orientation="portrait" fig-type="figure">
<label>Supplementary Figure 2.</label>
<caption><title>Technical details of the metabolism reconstruction software framework in anvi’o.</title>
<p><bold>A)</bold> Workflow of metabolism reconstruction programs and their inputs/outputs. Dark arrows indicate the primary analysis path utilized in this study. Blue background indicates optional features in the framework. A demonstration of completeness score and copy number calculations for metabolic pathways (performed by the program ‘anvi-estimate-metabolism’ is shown using example enzyme annotation data in panels B – E (for a theoretical pathway) and F – I (for a real pathway). <bold>B)</bold> Theoretical metabolic pathway, where hexagons represent metabolites, arrows represent chemical reactions, letters represent enzymes (subscripts indicate enzyme components), and the example number of gene annotation hits for each enzyme is written in gray. <bold>C)</bold> The definition of the theoretical pathway from panel B, written in terms of the required enzymes. <bold>D)</bold> Table showing the major steps in the pathway and example calculations for step presence and copy number. Step presence is calculated by evaluating a boolean expression created from the step definition in which enzymes with &gt; 0 hits are replaced with True (T) and the others with False (F). Step copy number is calculated by evaluating the corresponding arithmetic expression in which the enzymes are replaced with their annotation counts. <bold>E)</bold> Final calculations of completeness score (fraction of present steps) and copy number for the theoretical metabolic pathway. <bold>F – I)</bold> Same as panels B – E, but for KEGG module M00043. A high-resolution version of this figure is available at <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.6084/m9.figshare.22851173">https://doi.org/10.6084/m9.figshare.22851173</ext-link>.</p></caption>
<graphic xlink:href="540289v2_figs2.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
<fig id="figs3" position="float" orientation="portrait" fig-type="figure">
<label>Supplementary Figure 3.</label>
<caption><title>Comparison of unnormalized copy number data and normalized (per-population copy number, or PPCN) data for the IBD-enriched modules.</title>
<p><bold>A)</bold> Boxplot of median copy numbers for each module in the healthy samples (blue) and IBD samples (red). <bold>B)</bold> Boxplots of median PPCN for each module in the healthy samples (blue) and IBD samples (red). Lines connect data points for the same module in each plot. The gray dashed line in each plot indicates the overall median value.</p></caption>
<graphic xlink:href="540289v2_figs3.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
<fig id="figs4" position="float" orientation="portrait" fig-type="figure">
<label>Supplementary Figure 4.</label>
<caption><title>Histograms of annotations per gene call</title>
<p>from <bold>A,B)</bold> NCBI COGs; <bold>C,D)</bold> KEGG KOfams; and <bold>E,F)</bold> Pfams. Panels A, C, and E show data for metagenomes in the subset of 330 deeply-sequenced samples from healthy people and people with IBD, and panels B, D, and F show data for all 2,893 samples including those from non-IBD controls.</p></caption>
<graphic xlink:href="540289v2_figs4.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
<fig id="figs5" position="float" orientation="portrait" fig-type="figure">
<label>Supplementary Figure 5.</label>
<caption><title>Additional boxplots of median per-population copy number for various subsets of metabolic pathways and metagenome samples.</title>
<p><bold>A)</bold> 33 modules enriched in HMI populations from (<xref ref-type="bibr" rid="c100">Watson et al. 2023</xref>) compared to the 33 IBD-enriched modules from this study, with medians computed in the set of deeply-sequenced healthy (n = 229) and IBD (n = 101) samples. <bold>B)</bold> The 33 IBD-enriched modules from this study, with medians computed in the set of deeply-sequenced healthy (n = 229), non-IBD (n = 78), and IBD (n = 101) samples. <bold>C)</bold> All KEGG modules (n = 117) with non-zero copy number in at least one sample, with medians computed in the set of deeply-sequenced healthy (n = 229) and IBD (n = 101) samples. <bold>D)</bold> All biosynthesis modules (n = 88) from the KEGG MODULE database, with medians computed in the set of deeply-sequenced healthy (n = 229) and IBD (n = 101) samples. Where applicable, dashed lines indicate the overall median for all modules, and solid lines connect the points for the same module in each sample group. The IBD sample group is highlighted in red, the NONIBD group in pink, and the HEALTHY group in blue.</p></caption>
<graphic xlink:href="540289v2_figs5.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
<fig id="figs6" position="float" orientation="portrait" fig-type="figure">
<label>Supplementary Figure 6.</label>
<caption><title>Boxplots of median per-population copy number of 33 IBD-enriched modules for samples from each individual cohort,</title>
<p><bold>A)</bold> with medians computed within each sample (ie, one point per sample) and <bold>B)</bold> with medians computed for each IBD-enriched module (ie, one point per module). The x-axis indicates study of origin. <bold>C)</bold> Boxplots of median per-population copy number of 33 IBD-enriched modules for the 115 samples in the deeply-sequenced set that are not from (<xref ref-type="bibr" rid="c52">Le Chatelier et al. 2013</xref>) or (<xref ref-type="bibr" rid="c99">Vineis et al. 2016</xref>). The dashed line indicates the overall median for all 33 modules, and solid lines connect the points for the same module in each sample group.</p></caption>
<graphic xlink:href="540289v2_figs6.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
<fig id="figs7" position="float" orientation="portrait" fig-type="figure">
<label>Supplementary Figure 7.</label>
<caption><title>Assessing batch effect of the IBD-enrichment study.</title>
<p><bold>A)</bold> Scatter plot comparing the module ranks of Wilcoxon-Mann-Whitney p-values comparing IBD and healthy subjects on (<xref ref-type="bibr" rid="c52">Le Chatelier et al. 2013</xref>) and (<xref ref-type="bibr" rid="c99">Vineis et al. 2016</xref>) (y axis) and the rest of our dataset (x axis). <bold>B)</bold> Venn diagram displaying the overlap of IBD-enriched modules identified by the 33 smallest p-values in (<xref ref-type="bibr" rid="c52">Le Chatelier et al. 2013</xref>) and (<xref ref-type="bibr" rid="c99">Vineis et al. 2016</xref>) and the rest of our dataset. There is good agreement (20 out of 33) between the two sets of modules, indicating generalizability of the signals across studies used in our sample set.</p></caption>
<graphic xlink:href="540289v2_figs7.tif" mimetype="image" mime-subtype="tiff"/>
</fig>
</sec>
<sec id="d1e6420">
<title>Supplemental Tables</title>
<p>Supplementary Tables and our Supplementary Information can be accessed at doi:10.6084/m9.figshare.22679080.</p>
<p><bold>Table 1: Samples and cohorts used in this study.</bold> a) Description of studies/cohorts providing publicly-available gut metagenomes from healthy people, non-IBD controls, and people with IBD. For each study, we note the sample groups it contributes metagenomes to; whether or not those samples were sufficiently deeply sequenced to be included in the main analyses; the country of origin of the samples; the sample type (fecal metagenome or ileal pouch luminal aspirate); the number of samples it contributes to each group before and after applying the sequencing depth threshold; and cohort details/exclusions as described within the study. b) Description of 408 samples included in the primary analyses of this manuscript (ie, those with sufficient sequencing depth of &gt;= 25 million reads), including their associated diagnosis (ulcerative colitis (UC), Crohn’s disease (CD), non-IBD, healthy, colorectal cancer with adenoma (CRC_ADENOMA), or colorectal cancer with carcinoma (CRC_CARCINOMA)); study of origin; sample group; sequencing depth; and number of microbial populations estimated to be represented within the metagenome. c) Description of all samples initially considered and their SRA accession numbers. d) The number of gene calls and the number/proportion of annotations per gene call for KOfams, COGs, and Pfams in each sample. e) Description of the 57 antibiotic time-series gut metagenomes from (<xref ref-type="bibr" rid="c71">Palleja et al. 2018</xref>) used for classifier testing, including SRA accession number; sampling day in the time series; sequencing depth; and estimated numbers of microbial populations represented in the sample.</p>
<p><bold>Table 2: Metabolism data in metagenomes.</bold> a) Description of the 33 KEGG modules enriched in IBD samples, including: module name, KEGG categorization, and definition; their median per-population copy numbers (PPCN) in the healthy sample group and IBD sample group; the p-value, FDR-adjusted p-value, and W statistic from the per-module Wilcoxon Rank Sum test used to determine enrichment in IBD; the difference between its median PPCN in IBD samples and median PPCN in healthy samples (‘effect size’); the fraction of samples in which the module occurs with non-zero copy number; whether the module is also enriched in the HMI populations analyzed in (<xref ref-type="bibr" rid="c100">Watson et al. 2023</xref>); the number of total enzymes in the module; the number of total compounds in the modules; and the numbers and proportions of shared enzymes or compounds between this module and the other IBD-enriched modules. b) Description of all 179 KEGG modules with non-zero copy number in at least one metagenome. Most of the columns match the corresponding column in sheet (b) with the exception of the ‘enrichment status’ column, which indicates whether the module was found to be enriched in the IBD samples in this study (‘IBD_ENRICHED’), in the high-metabolic independence genomes in (<xref ref-type="bibr" rid="c100">Watson et al. 2023</xref>) (‘HMI_ENRICHED’), in both (‘HMI_AND_IBD’), or in neither (‘OTHER’). c) Matrix of stepwise copy number of each module in each deeply-sequenced gut metagenome. d) Per-population copy number of each module in each deeply-sequenced gut metagenome in the IBD, non-IBD and healthy sample groups. e) Per-population copy number of each module in each antibiotic time-series sample from (<xref ref-type="bibr" rid="c71">Palleja et al. 2018</xref>).</p>
<p><bold>Table 3: GTDB genome data.</bold> a) List of 338 GTDB representative genomes identified as gut microbes, their taxonomy, metabolic independence score, classification as high metabolic independence (‘HMI’) or not (‘non-HMI’), genome length in base pairs, and number of gene calls. b) Matrix of stepwise completeness of each module in each genome. c) Matrix of genome detection in each deeply-sequenced gut metagenome in the IBD, non-IBD, and healthy sample groups. d) Percent abundance of each genome in each deeply-sequenced gut metagenome. e) Per-genome proportion of samples from each sample group that the genome is detected in using a threshold of 50% (ie, at least half of the genome sequence is covered by at least one sequencing read in a given sample). f) Per-sample proportion of detected genomes that are classified as HMI. g) Average completion of each IBD-enriched module within the HMI genome group and the non-HMI genome group, as well as the difference between these values.</p>
<p><bold>Table 4: Metagenome classifier information.</bold> a) Details and performance of previously-published classifiers for IBD and IBD subtypes. For each classifier, we summarize the cohort details as described by the study; the size of training datasets and validation datasets (if any); the type(s) of samples, data, and extracted features used for classification; the target classes (that is, what the samples were being classified as); the classifier type and training/validation strategy; and the performance metrics as reported by the study. b) Classification of each (<xref ref-type="bibr" rid="c71">Palleja et al. 2018</xref>) metagenome by our logistic regression model trained for distinguishing IBD vs healthy samples on the basis of PPCN data for IBD-enriched modules. This table describes whether the sample was classified as healthy (‘HEALTHY’) or stressed (‘IBD’, which we consider to be equivalent to an identification of gut stress), and also whether the sample had low sequencing depth (&lt; 25 million reads) or not. c) Summary of the performance of our metagenome classifier across different training/validation strategies using the IBD and healthy metagenome samples. It also includes the details of our final classifier trained on all 330 samples, though performance data is not available for this model since there were no IBD/healthy samples left for validation – however, see manuscript for its performance on the (<xref ref-type="bibr" rid="c71">Palleja et al. 2018</xref>) antibiotic time-series dataset. The subsequent sheets include per-fold data and performance information for each train-test strategy: d) random split cross-validation (25-fold) on PPCN data; e) leave-two-studies-out cross-validation (24-fold); and f) (10-fold) cross-validation leaving out samples from the two dominating studies in our dataset, (<xref ref-type="bibr" rid="c52">Le Chatelier et al 2013</xref>) and (<xref ref-type="bibr" rid="c99">Vineis et al 2016</xref>).</p>
<p><bold>Table 5: Details of available software for metabolism estimation.</bold> For each tool (including the one published in this study), we summarize: the software category (based upon the tool’s architecture and mode of use); its metabolism reconstruction strategy (whether it is a pathway prediction tool or a modeling tool or both); the data source(s) it uses for enzyme and metabolic pathway information; how it calculates pathway completeness or generates models (depending on reconstruction strategy); what input and output types it accepts/generates; any additional capabilities as advertised by the tool’s publication; whether or not the tool is open-source; the program type; and what language(s) it is developed in (if known). The reference publication and code repository or webpage for each tool is also included.</p>
</sec>
</back>
<sub-article id="sa0" article-type="editor-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.89862.1.sa3</article-id>
<title-group>
<article-title>eLife Assessment</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Turnbaugh</surname>
<given-names>Peter J</given-names>
</name>
<role specific-use="editor">Reviewing Editor</role>
<aff>
<institution-wrap>
<institution>University of California, San Francisco</institution>
</institution-wrap>
<city>San Francisco</city>
<country>United States of America</country>
</aff>
</contrib>
</contrib-group>
<kwd-group kwd-group-type="evidence-strength">
<kwd>Compelling</kwd>
<kwd>Incomplete</kwd>
</kwd-group>
<kwd-group kwd-group-type="claim-importance">
<kwd>Important</kwd>
</kwd-group>
</front-stub>
<body>
<p>This study describes an <bold>important</bold> bioinformatics tool for normalizing gene copy number from metagenomic assemblies. The tool is used in a meta-analysis of data from inflammatory bowel disease (IBD) patients and healthy controls. While some of the evidence for the power of the method is <bold>compelling</bold>, other evidence seems <bold>incomplete</bold>. The inclusion of additional computational and/or experimental validation would markedly strengthen the study. This paper will likely be of broad interest to researchers studying the role of complex microbial communities in host health and disease.</p>
</body>
</sub-article>
<sub-article id="sa1" article-type="referee-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.89862.1.sa2</article-id>
<title-group>
<article-title>Reviewer #1 (Public Review):</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<anonymous/>
<role specific-use="referee">Reviewer</role>
</contrib>
</contrib-group>
</front-stub>
<body>
<p>In this work, Veseli et al. present a computational framework to infer the functional diversity of microbiomes in relation to microbial diversity directly from metagenomic data. The framework reconstructs metabolic modules from metagenomes and calculates the per-population copy number of each module, resulting in the proportion of microbes in the sample carrying certain genes. They applied this framework to a dataset of gut microbiomes from 109 inflammatory bowel disease (IBD) patients, 78 patients with other gastrointestinal conditions, and 229 healthy controls. They found that the microbiomes of IBD patients were enriched in a high fraction of metabolic pathways, including biosynthesis pathways such as those for amino acids, vitamins, nucleotides, and lipids. Hence, they had higher metabolic independence compared with healthy controls. To an extent, the authors also found a pathway enrichment suggesting higher metabolic independence in patients with gastrointestinal conditions other than IBD indicating this could be a signal for a general loss in host health. Finally, a machine learning classifier using high metabolic independence in microbiomes could predict IBD with good accuracy. Overall, this is an interesting and well-written article and presents a novel workflow that enables a comprehensive characterization of microbiome cohorts.</p>
</body>
</sub-article>
<sub-article id="sa2" article-type="referee-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.89862.1.sa1</article-id>
<title-group>
<article-title>Reviewer #2 (Public Review):</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<anonymous/>
<role specific-use="referee">Reviewer</role>
</contrib>
</contrib-group>
</front-stub>
<body>
<p>This study builds upon the team's recent discovery that antibiotic treatment and other disturbances favour the persistence of bacteria with genomes that encode complete modules for the synthesis of essential metabolites (Watson et al. 2023). Veseli and collaborators now provide an in-depth analysis of metabolic pathway completeness within microbiomes, finding strong evidence for an enrichment of bacteria with high metabolic independence in the microbiomes associated with IBD and other gastrointestinal disorders. Importantly, this study provides new open-source software to facilitate the reconstruction of metabolic pathways, estimate their completeness and normalize their results according to species diversity. Finally, this study also shows that the metabolic independence of microbial communities can be used as a marker of dysbiosis. The function-based health index proposed here is more robust to individuals' lifestyles and geographic origin than previously proposed methods based on bacterial taxonomy.</p>
<p>The implications of this study have the potential to spur a paradigm shift in the field. It shows that certain bacterial taxa that have been consistently associated with disease might not be harmful to their host as previously thought. These bacteria seem to be the only species that are able to survive in a stressed gut environment. They might even be important to rebuild a healthy microbiome (although the authors are careful not to make this speculation).</p>
<p>This paper provides an in-depth discussion of the results, and limitations are clearly addressed throughout the manuscript. Some of the potential limitations relate to the use of large publicly available datasets, where sample processing and the definition of healthy status varies between studies. The authors have recognised these issues and their results were robust to analyses performed on a per-cohort basis. These potential limitations, therefore, are unlikely to have affected the conclusions of this study.</p>
<p>Overall, this manuscript is a magnificent contribution to the field, likely to inspire many other studies to come.</p>
</body>
</sub-article>
<sub-article id="sa3" article-type="referee-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.89862.1.sa0</article-id>
<title-group>
<article-title>Reviewer #3 (Public Review):</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<anonymous/>
<role specific-use="referee">Reviewer</role>
</contrib>
</contrib-group>
</front-stub>
<body>
<p>The major strength of this manuscript is the &quot;anvi-estimate-metabolism' tool, which is already accessible online, extensively documented, and potentially broadly useful to microbial ecologists. However, the context for this tool and its validation is lacking in the current version of the manuscript. It is unclear whether similar tools exist; if so, it would help to benchmark this new tool against prior methods. Simulated datasets could be used to validate the approach and test its robustness to different levels of bacterial richness, genome sizes, and annotation level.</p>
<p>The concept of metabolic independence was intriguing, although it also raises some concerns about the overinterpretation of metagenomic data. As mentioned by the authors, IBD is associated with taxonomic shifts that could confound the copy number estimates that are the primary focus of this analysis. It is unclear if the current results can be explained by IBD-associated shifts in taxonomic composition and/or average genome size. The level of prior knowledge varies a lot between taxa; especially for the IBD-associated gamma-Proteobacteria. It can be difficult to distinguish genes for biosynthesis and catabolism just from the KEGG module names and the new normalization tool proposed herein markedly affects the results relative to more traditional analyses. As such, it seems safer to view the current analysis as hypothesis-generating, requiring additional data to assess the degree to which metabolic dependencies are linked to IBD.</p>
</body>
</sub-article>
</article>