<?xml version="1.0" ?><!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.3 20210610//EN"  "JATS-archivearticle1-mathml3.dtd"><article xmlns:ali="http://www.niso.org/schemas/ali/1.0/" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.3" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="nlm-ta">elife</journal-id>
<journal-id journal-id-type="publisher-id">eLife</journal-id>
<journal-title-group>
<journal-title>eLife</journal-title>
</journal-title-group>
<issn publication-format="electronic" pub-type="epub">2050-084X</issn>
<publisher>
<publisher-name>eLife Sciences Publications, Ltd</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">100485</article-id>
<article-id pub-id-type="doi">10.7554/eLife.100485</article-id>
<article-id pub-id-type="doi" specific-use="version">10.7554/eLife.100485.1</article-id>
<article-version-alternatives>
<article-version article-version-type="publication-state">reviewed preprint</article-version>
<article-version article-version-type="preprint-version">1.1</article-version>
</article-version-alternatives>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Computational and Systems Biology</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Establishing comprehensive quaternary structural proteomes from genome sequence</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Catoiu</surname>
<given-names>Edward Alexander</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Mih</surname>
<given-names>Nathan</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Lu</surname>
<given-names>Maxwell</given-names>
</name>
<xref ref-type="aff" rid="a2">2</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0003-2357-6785</contrib-id>
<name>
<surname>Palsson</surname>
<given-names>Bernhard</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
<xref ref-type="aff" rid="a3">3</xref>
<email>bpalsson@ucsd.edu</email>
</contrib>
<aff id="a1"><label>1</label><institution>Department of Bioengineering, University of California</institution>, San Diego, La Jolla, CA 92101</aff>
<aff id="a2"><label>2</label><institution>Omnicorp Inc. (Pilot AI)</institution>, San Francisco, CA 94129</aff>
<aff id="a3"><label>3</label><institution>The Novo Nordisk Foundation (NNF) Center for Biosustainability, The Technical University of Denmark</institution>, Kongens Lyngby 2800, <country>Denmark</country></aff>
</contrib-group>
<contrib-group content-type="section">
<contrib contrib-type="editor">
<name>
<surname>Graña</surname>
<given-names>Martin</given-names>
</name>
<role>Reviewing Editor</role>
<aff>
<institution-wrap>
<institution>Institut Pasteur de Montevideo</institution>
</institution-wrap>
<city>Montevideo</city>
<country>Uruguay</country>
</aff>
</contrib>
<contrib contrib-type="senior_editor">
<name>
<surname>Cui</surname>
<given-names>Qiang</given-names>
</name>
<role>Senior Editor</role>
<aff>
<institution-wrap>
<institution>Boston University</institution>
</institution-wrap>
<city>Boston</city>
<country>United States of America</country>
</aff>
</contrib>
</contrib-group>
<pub-date date-type="original-publication" iso-8601-date="2024-09-25">
<day>25</day>
<month>09</month>
<year>2024</year>
</pub-date>
<volume>13</volume>
<elocation-id>RP100485</elocation-id>
<history>
<date date-type="sent-for-review" iso-8601-date="2024-06-13">
<day>13</day>
<month>06</month>
<year>2024</year>
</date>
</history>
<pub-history>
<event>
<event-desc>Preprint posted</event-desc>
<date date-type="preprint" iso-8601-date="2024-04-28">
<day>28</day>
<month>04</month>
<year>2024</year>
</date>
<self-uri content-type="preprint" xlink:href="https://doi.org/10.1101/2024.04.24.590993"/>
</event>
</pub-history>
<permissions>
<copyright-statement>© 2024, Catoiu et al</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Catoiu et al</copyright-holder>
<ali:free_to_read/>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<ali:license_ref>https://creativecommons.org/licenses/by/4.0/</ali:license_ref>
<license-p>This article is distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License</ext-link>, which permits unrestricted use and redistribution provided that the original author and source are credited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="elife-preprint-100485-v1.pdf"/>
<abstract>
<title>Abstract</title><p>A critical body of knowledge has developed through advances in protein microscopy, protein-fold modeling, structural biology software, availability of sequenced bacterial genomes, large-scale mutation databases, and genome-scale models. Based on these recent advances, we develop a computational framework that; i) identifies the oligomeric structural proteome encoded by an organism’s genome from available structural resources; ii) maps multi-strain alleleomic variation, resulting in the structural proteome for a species; and iii) calculates the 3D orientation of proteins across subcellular compartments with residue-level precision. Using the platform, we; iv) compute the quaternary <italic>E. coli</italic> K-12 MG1655 structural proteome; v) use a dataset of 12,000 mutations to build Random Forest classifiers that can predict the severity of mutations; and, in combination with a genome-scale model that computes proteome allocation, vi) obtain the spatial allocation of the <italic>E. coli</italic> proteome. Thus, in conjunction with relevant datasets and increasingly accurate computational models, we can now annotate quaternary structural proteomes, at genome-scale, to obtain a molecular-level understanding of whole-cell functions.</p>
</abstract>
<abstract>
<title>Significance</title>
<p>Advancements in experimental and computational methods have revealed the shapes of multi-subunit proteins. The absence of a unified platform that maps actionable datatypes onto these increasingly accurate structures creates a barrier to structural analyses, especially at the genome-scale. Here, we describe QSPACE, a computational annotation platform that evaluates existing resources to identify the best-available structure for each protein in a user’s query, maps the 3D location of actionable datatypes (<italic>e.g.</italic>, active sites, published mutations) onto the selected structures, and uses third-party APIs to determine the subcellular compartment of all amino acids of a protein. As proof-of-concept, we deployed QSPACE to generate the quaternary structural proteome of <italic>E. coli</italic> MG1655 and demonstrate two use-cases involving large-scale mutant analysis and genome-scale modelling.</p>
</abstract>
<custom-meta-group>
<custom-meta specific-use="meta-only">
<meta-name>publishing-route</meta-name>
<meta-value>prc</meta-value>
</custom-meta>
</custom-meta-group>
</article-meta>
<notes>
<notes notes-type="competing-interest-statement">
<title>Competing Interest Statement</title><p>The authors have declared no competing interest.</p></notes>
<fn-group content-type="external-links">
<fn fn-type="dataset"><p>
<ext-link ext-link-type="uri" xlink:href="https://www.github.com/EdwardCatoiu/QSPACE">https://www.github.com/EdwardCatoiu/QSPACE</ext-link>
</p></fn>
</fn-group>
</notes>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>The proteome of the cell is responsible for metabolite uptake and secretion, genetic information processing and replication, energy production, and all other processes required for maintaining cellular homeostasis. Before becoming a functional unit of this multi-scale system, a protein must fold properly into its native three-dimensional shape. This folding process is also multi-scale. A peptide sequence (primary sequence) associates locally to form small recognizable patterns (e.g., alpha-helices and beta-sheets). Often stabilized by disulfide bridges and physiochemical attraction, these secondary structures fold onto each other to form larger recognizable domains, resulting in the three-dimensional structure of the protein monomer (tertiary structure). These protein monomers often oligomerize and form multi-subunit enzymes (quaternary structures) that carry out the functions in the cell.</p>
<p>Structural biology—the study of protein shape and function—has advanced rapidly in recent years. For proteins that form large multi-subunit complexes and for those spanning the cell membrane, the three-dimensional shape was particularly difficult to study with classical crystallographic techniques. The development of cryogenic electron microscopy – a method that images thin slices of a protein frozen in its native state (much like a biopsy)—has drastically increased the speed and ease by which these previously unknowable protein structures can be resolved<sup><xref ref-type="bibr" rid="c1">1</xref>–<xref ref-type="bibr" rid="c5">5</xref></sup>. Concurrently, computational methods have also experienced increasing success in accurately predicting protein structures<sup><xref ref-type="bibr" rid="c6">6</xref>–<xref ref-type="bibr" rid="c10">10</xref></sup>. Most recently, deep learning algorithms<sup><xref ref-type="bibr" rid="c11">11</xref>–<xref ref-type="bibr" rid="c12">12</xref></sup> (e.g., AlphaFold) have utilized multiple sequence alignments and incorporated biophysical knowledge about protein structure to predict the shape of proteins without homologous structures in the Protein Data Bank<sup><xref ref-type="bibr" rid="c13">13</xref></sup>. Even more promising, these algorithms can be “hacked” to predict the structures of oligomeric assemblies of protein complexes<sup><xref ref-type="bibr" rid="c14">14</xref></sup>. Recent benchmarking efforts have confirmed the accuracy of homodimer AF-multimer models<sup><xref ref-type="bibr" rid="c15">15</xref></sup> and have subsequently been used for homo-oligomeric predictions<sup><xref ref-type="bibr" rid="c16">16</xref></sup>.</p>
<p>Although structural biology can offer molecular insights into a protein’s shape and function, mutations in key domains can change the enzymatic properties and modulate a protein’s function. Changes in protein function can be either beneficial (e.g., an increase in stability of the active form) or detrimental (e.g., a loss of substrate binding efficiency). Protein engineering employs a variety of techniques to find mutations that produce a desired phenotype. One such technique, multiplex automated genomic engineering (MAGE)<sup><xref ref-type="bibr" rid="c17">17</xref></sup>, can introduce many mutations with unknown effects at specific sites in the genome. Mutations that result in the desired phenotype can then be selected for. Laboratory evolution, an experimental approach involving the serial passage of a cell population in an increasingly stringent selection pressure, can speed up the evolutionary process and beneficial mutations can be identified by sequencing the endpoint strain<sup><xref ref-type="bibr" rid="c18">18</xref>–<xref ref-type="bibr" rid="c20">20</xref></sup>. The dramatic decrease in sequencing costs in the last ten years has allowed for many mutations identified by these experimental techniques to be collected in databases.</p>
<p>Concurrent with advancements in structural biology and the formation of large-scale mutation databases, systems biology—the study of systems-level cellular behavior— was driven by the development of genome-scale models (GEMs) of cellular metabolism that predict gene essentiality, growth phenotypes, and proteome allocation in a few organisms<sup><xref ref-type="bibr" rid="c21">21</xref>–<xref ref-type="bibr" rid="c26">26</xref></sup>. Software (e.g., ssbio<sup><xref ref-type="bibr" rid="c27">27</xref></sup>) was developed to map available structural information to these modeled proteomes. Structural systems biology—the study of structural biology at the systems-level—has incorporated protein structures into genome-scale models (GEM-PROs) to study protein-fold evolution and investigate structural differences between organisms<sup><xref ref-type="bibr" rid="c28">28</xref>–<xref ref-type="bibr" rid="c32">32</xref></sup>. Notwithstanding the incorporation of protein information at the monomer-level, the use of GEM-PROs is the most recent step towards building genome-scale models that reflect the physical nature of the cellular proteome.</p>
<p>Given the availability of high-quality protein structures and structural models that capture the shape of multi-subunit complexes, the deposition of mutation-phenotype information into large-scale databases, and the development of genome-scale models of cellular proteome allocation, the creation of a <italic>genome annotation platform</italic> with interoperability between structural, functional, mutational, and systems-level information is now possible.</p>
<p>In this study, we present the Quaternary Structural Proteome Atlas of a CEll (QSPACE) — a computational annotation platform that 1) utilizes state-of-the-art modeling software (e.g. Alphafold<sup><xref ref-type="bibr" rid="c11">11</xref></sup> &amp; Alphafold Multimer<sup><xref ref-type="bibr" rid="c14">14</xref></sup>) and the latest crystallographic depositions to identify a three-dimensional structural representation that accounts for the multi-subunit assembly of the cellular proteome; 2) calculates structural properties of the proteome; 3) provides a three-dimensional context to map functional information including enzymatic domains, binding sites, and protein interfaces; 4) draws mutational information from large-scale databases of laboratory-acquired mutations<sup><xref ref-type="bibr" rid="c20">20</xref>,<xref ref-type="bibr" rid="c33">33</xref>–<xref ref-type="bibr" rid="c35">35</xref></sup> and of the wild-type natural sequence diversity (alleleome) of <italic>E. coli</italic><sup><xref ref-type="bibr" rid="c36">36</xref></sup>; and 5) calculates the subcellular compartmentalization of the proteome with residue-level resolution.</p>
<p>The QSPACE platform allows users to rapidly interact with protein structural data for biological inquiries ranging from the single-protein to the genome-scale (GS). Using <italic>E. coli</italic> as an example, we present two separate genome-scale applications of the QSPACE platform to demonstrate its broad applicability. First, we exploit QSPACE’s superimposition of mutant datasets and annotated functional domains on the protein structure to calculate 100 residue-level features for over 12,000 published <italic>E. coli</italic> mutations in UniProt<sup><xref ref-type="bibr" rid="c37">37</xref></sup>, allowing us to build RF-classifiers capable of predicting the severity of amino acid substitutions. Second, we showcase how QSPACE’s subcellular compartmentalization of the protein structures advances genome-scale modelling efforts. By calculating the size (volume, and, when applicable, the cross-sectional area of membrane proteins) of <italic>E. coli</italic> protein structures and incorporating them into iJL1678b<sup><xref ref-type="bibr" rid="c26">26</xref></sup>—a genome-scale model that predicts the macromolecular expression (80%, by mass) of <italic>E. coli</italic> MG1655—we are able to predict the physical space (across multiple subcellular compartments) required by the computed proteome of <italic>Escherichia coli</italic> K-12 MG1655 at optimal growth rate. To our knowledge, this QSPACE/GEM-PRO is the most comprehensive whole-cell approach that captures the 3D nature of the <italic>E. coli</italic> structural proteome. As structural, mutational, and functional knowledge is discovered, and GEMs are developed with increasing specificity, QSPACE can provide a method to rapidly integrate all information related to the structural proteome for an increasing number of organisms. QSPACE can be deployed for any organism following the tutorial python notebook available at <ext-link ext-link-type="uri" xlink:href="https://github.com/EdwardCatoiu/QSPACE/">https://github.com/EdwardCatoiu/QSPACE/</ext-link>.</p>
</sec>
<sec id="s2">
<title>Results</title>
<sec id="s2a">
<title>Overview of the QSPACE platform</title>
<p>The Quaternary Structural Proteome Atlas of a Cell (QSPACE) is an annotation platform that compiles available structural data from the latest structural biology efforts to obtain a 3D representation of all codon positions in a genome – complete with residue-level biophysical, chemical, and mutational data (see Table S1 for details). The QSPACE of <italic>E. coli</italic> is presented as a CSV file in Dataset S1. The two user-defined inputs to the QSPACE platform (<xref rid="fig1" ref-type="fig">Fig. 1a</xref>) are i) a list of gene IDs and ii) a dictionary of protein complexes and the associated stoichiometric ratio of the genes that make up each complex (Dataset S2A). QSPACE automatically downloads all protein structures (and homology models) from RCSB-PDB<sup><xref ref-type="bibr" rid="c13">13</xref></sup>, ITASSER<sup><xref ref-type="bibr" rid="c8">8</xref></sup>, SWISS-MODEL<sup><xref ref-type="bibr" rid="c10">10</xref></sup> &amp; AlphaFold<sup><xref ref-type="bibr" rid="c12">12</xref></sup> that correspond to any of the genes in the user-defined inputs. QSPACE then finds the 3D coordinate file (i.e. “structure”) that best reflects the user-defined (input #2) multi-subunit protein assembly (<xref rid="fig1" ref-type="fig">Fig. 1b</xref>, details in <xref rid="fig2" ref-type="fig">Fig. 2</xref>). When no available structures can accurately reflect the gene-stoichiometry of a protein complex, QSPACE will attempt to generate models for the protein structure using an external GoogleColab notebook running AlphaFold Multimer<sup><xref ref-type="bibr" rid="c14">14</xref></sup> (v2.0 via ColabFold<sup><xref ref-type="bibr" rid="c38">38</xref></sup>).</p>
<fig id="fig1" position="float" orientation="portrait" fig-type="figure">
<label>Figure 1:</label>
<caption><title>The Quaternary Structural Proteome Atlas of a CEll is a genome-scale annotation platform that was applied to the <italic>E. coli</italic> proteome.</title>
<p><bold>(a)</bold> The QSPACE platform requires two user-defined inputs: a list of gene(s) and dictionary of proteins and their associated gene-stoichiometric ratios. User-defined proteins can be oligomeric or monomeric. QSPACE accommodates residue-level sequence variation (the alleleome<sup><xref ref-type="bibr" rid="c36">36</xref></sup> was used in this study). <bold>(b)</bold> QPACE identifies (or generates w/AF-multimer) the protein structure/model that best reflects the gene-stoichiometry of each user-defined protein complex (details in <xref rid="fig2" ref-type="fig">Figure 2a</xref>). The resulting structures <bold>(c)</bold> are analyzed using various software packages to calculate physicochemical properties and to identify evolutionary variable regions and functional domains (details in <xref rid="fig3" ref-type="fig">Figure 3a</xref>). <bold>(d)</bold> The structures are localized to their subcellular compartments and membrane-embedded structures are oriented across the membrane, resulting in a three-dimensional representation of <bold>(e)</bold> the structural proteome of an organism. <bold>(f)</bold> The amino acids of protein complexes are mapped to protein structures with varying levels of coverage (mean = 0.94) in <italic>E. coli</italic>. <bold>(g)</bold> Genome-scale counts of the unique codon positions belonging to various computed categories are shown.</p></caption>
<graphic xlink:href="590993v1_fig1.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="fig2" position="float" orientation="portrait" fig-type="figure">
<label>Figure 2:</label>
<caption><title>QSPACE yields 3D structures that reflect the oligomeric nature of multi-subunit proteins.</title>
<p>(<bold>a</bold>) There are three ways that QSPACE’s protein-structure module finds the 3D structural representation for the user-defined gene-stoichiometry of a protein (target). (i) For each protein target (<italic>left</italic>), all structures that share genes with the protein are identified (<italic>center left</italic>) and combined (if applicable) to recreate the user-defined gene-stoichiometry (<italic>center right</italic>). If multiple matches are found, the structure that most accurately reflects the complete protein is selected (<italic>right</italic>). (ii) Sometimes, only higher-order structures of the protein-target are identified (<italic>center left</italic>). After QCQA of the relevant quality metrics associated with each structure (see Fig. S9, Methods 3.1), the highest confidence oligomeric structure is used to redefine the gene stoichiometry for the protein (‘new complex’, <italic>right</italic>). (iii) If identified structures are unable to recreate the protein in its entirety (‘missing subunits’, <italic>center left)</italic>, Alphafold Multimer is used to predict the structure (&lt;2000 AAs) and the quality of the resulting oligomeric models are assessed (see Fig. S11, Dataset S9). (<bold>b</bold>) To achieve a multi-subunit representation of the <italic>E. coli</italic> proteome, structures or models from various sources are used. Unlike previous GEM-PROs, QSPACE accommodates monomeric and k-meric Alphafold models. (<bold>c</bold>) The protein-to-structure module yields 3D structures that represent the oligomeric nature of the <italic>E. coli</italic> proteome. This representation is further improved using structural data to correct existing protein-gene stoichiometry and with Alphafold Multimer to calculate novel oligomeric structures, (<bold>d</bold>) allowing for a truer accounting of higher-order oligomeric proteins. (<bold>e</bold>) The QSPACE platform generated 50 novel structures for ABC-transporters in <italic>E. coli</italic>. The pLDDT scores are used to color the AF-Multimer models. The associated AlphaFold Model Score (see Fig. S11, Dataset S3) is displayed in the bottom right. Asterisks are used to denote incomplete models resulting from incorrectly defined gene-stoichiometries or from protein size limitations of ColabFold. Existing PDB structures are shaded.</p></caption>
<graphic xlink:href="590993v1_fig2.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<p>The thresholds used by QSPACE to assess the accuracy of selected protein structures are described in the text accompanying <xref rid="fig2" ref-type="fig">Figure 2</xref> and in the methods section. All quality metrics related to protein structures for the <italic>E. coli</italic> QSPACE are provided in Dataset S3. By exploiting previously published repositories of protein structures, QSPACE reduces the threshold of interacting with genome-scale structural data to the order of days. Depending on user-preferences, QSPACE can function with all or some of the structural repositories used in this manuscript and can easily accommodate structural data from new sources as they become available.</p>
<p>Once the structure file representing the quaternary assembly of each protein is determined, multiple software packages and databases (see Table S1) are used to map physio-chemical, evolutionary, and functional information to the protein structures (<xref rid="fig1" ref-type="fig">Fig. 1c</xref>). The 3D overlay of multiple data types (details in <xref rid="fig3" ref-type="fig">Fig. 3</xref>) creates potential for many analysis tools (e.g., <xref rid="fig4" ref-type="fig">Fig. 4</xref>). The amino acids in each protein are then assigned to one of twelve subcellular compartments; and those representing the membrane fraction of the proteome are oriented across one of the <italic>E. coli</italic> membranes (<xref rid="fig1" ref-type="fig">Fig. 1d</xref>, details in <xref rid="fig5" ref-type="fig">Fig. 5</xref>). These structures can be integrated with genome-scale systems models to add a 3D understanding of the biophysical/spatial allocation of the proteome in a functioning cell (see <xref rid="fig6" ref-type="fig">Fig. 6</xref>). Users of QSPACE can bypass the mapping of any of these datasets if they are not relevant to their research, or not available for their organism.</p>
<fig id="fig3" position="float" orientation="portrait" fig-type="figure">
<label>Figure 3:</label>
<caption><title>Multi-dimensional QSPACE features to predict mutant phenotypes.</title>
<p><bold>(a)</bold> QSPACE calculates 100 properties (residue-level, sequence-level, and structure-level) for all amino acids. (right) The mapping of the three-dimensional location of functional domains on the protein structure enables the calculation of a mutation’s proximity to important protein regions. <bold>(b)</bold> (left) UniProt mutations and their annotated phenotypes, are mapped to QSPACE and (right) can be classified into general categories that reflect mutant severity. <bold>(c)</bold> Annotated phenotypes and multi-dimensional properties for 12,000 UniProt mutations mapped to the <italic>E. coli</italic> QSPACE (Dataset S5) can be used to train RF-classifiers that can predict the effect of a mutation (see <xref rid="fig4" ref-type="fig">Fig. 4</xref>). We suggest the application of these classifiers to novel mutant datasets, (e.g. mutations in adaptive laboratory evolution experiments (ALE), the long-term evolution experiment (LTEE) and the alleleome are already mapped to the <italic>E. coli</italic> QSPACE in Dataset S1).</p></caption>
<graphic xlink:href="590993v1_fig3.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="fig4" position="float" orientation="portrait" fig-type="figure">
<label>Figure 4:</label>
<caption><title>Random-forest prediction of mutant phenotypes in UniProtKB.</title>
<p><bold>(a)</bold> QSPACE finds 4,299 mutants with known phenotypes in proteins containing <italic>active sites</italic>. The mutational properties (‘features’) of the amino acid residue (grey), the 5 amino acid long sequence centered at the mutation (yellow), and the local 3D protein structure (green) are calculated (see <xref rid="fig3" ref-type="fig">Fig. 3</xref>). The numbers inside the boxes describe the number of features used. Random Forest Classifiers are trained on (top, “0D”) residue-level parameters; (middle, “1D”) residue and sequence-level parameters; and (bottom, “3D”) residue, sequence, and structure-level parameters. For RF-classifiers initially trained on more than 30 parameters, the least predictive parameter is removed until the 30 most-predictive parameters are identified. <bold>(b)</bold> “One vs Rest” receiver operating characteristic curves and precision-recall curves are calculated from the averages of 100 RF-classifiers trained on the 3 different sets of parameters. The shaded region represents 1 standard deviation. <bold>(c)</bold> The importance of individual features for “3D” RF-classifiers is determined. MUT- and MUT+ reflect pre- and post-mutation sequence properties, respectively. The radius of the 3D environment is described where applicable. The interoperability of multiple datatypes in QPSACE allows for the calculation of a mutation’s proximity to the nearest active site—the third most important feature in determining mutant severity. <bold>(d)</bold> QSPACE can be used to analyze mutations found in proteins containing various functional domains. <bold>(e)</bold> The weighted (by phenotype class) area under the curve (AUC) is calculated from the ROC and P-R curves in Panel B for RF-classifiers (3 parameter sets x 100 RF-classifiers) for each functional domain. <bold>(f)</bold> The cumulative importance of residue, sequence, and structure features in “3D” RF-classifiers of mutations in proteins containing each domain.</p></caption>
<graphic xlink:href="590993v1_fig4.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="fig5" position="float" orientation="portrait" fig-type="figure">
<label>Figure 5:</label>
<caption><title>The membrane module orients proteins across the membrane to identify residue-level subcellular compartments of the <italic>E. coli</italic> proteome.</title>
<p><bold>(a)</bold> In the metadata provided by EcoCyc (cellular compartment), Gene Ontology Terms (pathway, function, compartment), iML1515 (metabolic subsystem), and UniProt (topological &amp; transmembrane domains) databases, there are 1,777 protein structures mapped to at least one gene that is associated with the <italic>E. coli</italic> membrane. <bold>(b)</bold> Membrane-crossing residues are identified by the amino acid sequence information provided by UniProt, predicted by DeepTMHMM, and calculated by OPM. From these residues, a plane of best fit is calculated. <bold>(c)</bold> Structures with two calculated membrane planes pass the QCQA analysis if i) the angle between the planes is less than 35°, ii) the thickness of the membrane embedded region is between 12 and 45 Angstroms, and iii) the cross-sectional area of the membrane embedded region is less than 10,000 Å<sup><xref ref-type="bibr" rid="c2">2</xref></sup>. <bold>(d)</bold> Membrane proteins are oriented using the topological information provided by UniProt (if available) or manually using common protein motifs (see Dataset S6-S7) such that <bold>(e)</bold> the subcellular compartment of every amino acid of the <italic>E. coli</italic> proteome can be determined.</p></caption>
<graphic xlink:href="590993v1_fig5.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="fig6" position="float" orientation="portrait" fig-type="figure">
<label>Figure 6:</label>
<caption><title>QSPACE integrates with genome-scale models to predict the physical space required by the <italic>E. coli</italic> proteome at optimal growth rate.</title>
<p><bold>(a)</bold> The compartmentalization of each amino acid of the <italic>E. coli</italic> proteome allows for the calculation of geometric properties of all proteins. <bold>(b)</bold> The volume of ATP-synthase and <bold>(c)</bold> its cross-sectional membrane area are shown. <bold>(d)</bold> The integration of QSPACE with genome-scale models (iJL1678b-ME, in this case) of metabolism (M-matrix) and macromolecular expression (E-matrix) (ME-models) can be used to calculate <bold>(e)</bold> the proteome allocation, <bold>(f)</bold> the volumetric allocation, and <bold>(g)</bold> the membrane composition of <italic>E. coli</italic> at optimal growth rate. The expression, volume, and membrane area allocated to ATP-synthase is shown (Panels E-G, cyan). The calculated spatial allocation (Panels F-G) of the macromolecular expression predicted by existing ME-models (in Panel E) is a fundamental advancement towards building genome-scale biophysical whole-cell models.</p></caption>
<graphic xlink:href="590993v1_fig6.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<p>As an example, we apply QSPACE to the genome of <italic>E. coli</italic> K-12 MG1655 (<xref rid="fig1" ref-type="fig">Fig. 1e</xref>, Dataset S1) and identify the quaternary structural representation of its oligomeric proteome—as defined by the multi-decade bibliomic curation available in EcoCyc<sup><xref ref-type="bibr" rid="c39">39</xref></sup> and in the <italic>E. coli</italic> genome-scale model iJL1678b<sup><xref ref-type="bibr" rid="c26">26</xref></sup>. These gene-stoichiometric inputs to the <italic>E. coli</italic> QSPACE are provided in Dataset S2A. Selecting from both experimentally resolved structures deposited in the Protein DataBank (RCSB-PDB)<sup><xref ref-type="bibr" rid="c13">13</xref></sup> and from structural models calculated using protein modeling methods (ITASSER<sup><xref ref-type="bibr" rid="c8">8</xref></sup>, SWISS<sup><xref ref-type="bibr" rid="c10">10</xref></sup>, AlphaFold<sup><xref ref-type="bibr" rid="c12">12</xref></sup> &amp; AlphaFold Multimer<sup><xref ref-type="bibr" rid="c14">14</xref></sup>) (details in <xref rid="fig2" ref-type="fig">Fig. 2</xref>), QSPACE can map the 3D position of 94% (on average) of the amino acids belonging to 3,985 annotated <italic>E. coli</italic> proteins (<xref rid="fig1" ref-type="fig">Fig. 1f</xref>, Dataset S3A). The set of structures that QSPACE maps to the <italic>E. coli</italic> structural proteome can be used as 3D scaffolds to map multiple structural, functional, mutational, and spatial data types (<xref rid="fig1" ref-type="fig">Fig. 1g</xref>).</p>
</sec>
<sec id="s2b">
<title>Structural representation of multi-subunit proteins</title>
<p>Proteins often require oligomerization to function properly. The fundamental advancement of the <italic>E. coli</italic> QSPACE over existing genome-scale models with protein structures (e.g. iML1515-GP<sup><xref ref-type="bibr" rid="c40">40</xref></sup> see Fig. S1) is that it can be used to identify structures that represent the quaternary shapes of multi-subunit proteins.</p>
<p>To ensure that the user-defined multi-subunit proteins are accurately reflected in the structural data, we designed a pipeline to identify the best available protein structure for a target oligomeric protein, to suggest changes to the user-defined gene stoichiometry when the existing structural data suggests oligomerization, and to generate <italic>de novo</italic> structural models for oligomeric enzymes whose subunits cannot be fully represented by the structures in the PDB. A simplified representation is shown in <xref rid="fig2" ref-type="fig">Figure 2a</xref>.</p>
<p>The input to the QSPACE pipeline is a user-defined dictionary of protein complexes and their associated gene-stoichiometries. For <italic>E. coli</italic>, this information is the result of multi-decade bibliomic evidence that has been annotated in the EcoCyc database<sup><xref ref-type="bibr" rid="c39">39</xref></sup> and in the genome-scale model iJL1678b-ME<sup><xref ref-type="bibr" rid="c26">26</xref></sup> (Dataset S2A). Across these resources of annotated protein-complexes, 31% (1,334/4,309) of <italic>E. coli</italic> genes participate in 1,047 oligomeric complexes, 667 genes are annotated as monomers, and 2,308 genes are not included (i.e. assumed to be monomers) (Fig. S9A-B). In the set of annotated or assumed monomers, QSPACE identified structures (in the PDB or SWISS-MODEL repository) containing one or more oligomeric conformations for 983 of these genes (<xref rid="fig2" ref-type="fig">Fig. 2a</xref>.ii &amp; Fig. S9C). QPACE uses a semi-automated pipeline that relies on various structure-derived quality metrics to assess the accuracy of PDB and SWISS-MODELs before redefining the existing monomeric annotation for these genes (see Methods 3.1).</p>
<p>The accuracy of quaternary structures (experimental and modelled) has been the focus of many community-wide structural biology efforts. Previous studies have estimated that the accuracy of the quaternary structures in the PDB (‘biological assemblies’) is in the range of 80-90% <sup><xref ref-type="bibr" rid="c41">41</xref>–<xref ref-type="bibr" rid="c44">44</xref></sup>, and the accuracy of PISA-generated homo-oligomers to be 85%<sup><xref ref-type="bibr" rid="c45">45</xref></sup>. QSPACE uses PDB biological assemblies that are author-defined, software-defined (by PISA), or both (see Dataset S8). In cases where the PDB structures suggested oligomerization (contrary to the existing monomeric annotation), we reviewed the publication(s) associated with each PDB structure to confirm the oligomeric structure is believed (by the authors) to be biologically relevant (case IV-V in Fig. S9C-D).</p>
<p>Since 2017, QSQE-scores have been used to assess the quality of oligomeric SWISS-MODELs<sup><xref ref-type="bibr" rid="c46">46</xref></sup>. Recently, the SWISS-MODEL QSQE-score was shown to distinguish between biologically relevant and non-relevant homodimer structures at a rate of 0.79<sup><xref ref-type="bibr" rid="c15">15</xref></sup>. Although other modelling platforms perform slightly better<sup><xref ref-type="bibr" rid="c15">15</xref></sup>, SWISS-MODELs are precomputed and readily available, making them a convenient choice for rapid integration into the QSPACE annotation platform. Thus, in cases where SWISS-MODELs provided structural evidence of oligomerization (cases I-III in Fig. S9C-D), QSPACE relies on the established metrics and thresholds (QSQE<sup><xref ref-type="bibr" rid="c46">46</xref></sup> &gt; 0.5, GMQE<sup><xref ref-type="bibr" rid="c47">47</xref></sup> &gt; 0.5, and QMN4<sup><xref ref-type="bibr" rid="c47">47</xref></sup> &gt; -4) to assess the accuracy of each oligomeric SWISS-MODEL. SWISS-MODELs with scores exceeding these thresholds are used to redefine the oligomerization state of the user-defined monomers.</p>
<p>All oligomeric structures that were considered for changing the annotated <italic>E. coli</italic> monomers are provided in Dataset S2B. The relevant quality metrics associated with each structure ultimately selected in the <italic>E. coli</italic> QSPACE are provided in Dataset S3B.</p>
<p>When structures (in the PDB and/or SWISS-MODEL) are unable to fully reflect the gene-stoichiometry of a user-defined oligomer, the QSPACE platform relies on Alphafold Multimer<sup><xref ref-type="bibr" rid="c14">14</xref></sup> (v2.0, via ColabFold<sup><xref ref-type="bibr" rid="c38">38</xref></sup>) to generate <italic>de novo</italic> structures for desired protein oligomers (<xref rid="fig2" ref-type="fig">Fig. 2</xref>.a.iii). Alphafold Multimer (v2.0) was shown to outperform existing methods in modelling physiological homodimers<sup><xref ref-type="bibr" rid="c15">15</xref></sup> and has been reported to generate high-confidence homo-oligomeric structures for various organisms, including <italic>E. coli</italic><sup><xref ref-type="bibr" rid="c16">16</xref></sup>. QSPACE assigns confidence to AF-Multimer models using an established scoring metric<sup><xref ref-type="bibr" rid="c14">14</xref></sup> (0.8*iPTM + 0.2*PTM ≥ 0.8) (Fig. S11). It is important to note that the iPTM thresholds were shown to correlate with biologically relevant homo-dimer models<sup><xref ref-type="bibr" rid="c15">15</xref></sup>.</p>
<p>Furthermore, we confirmed the physiological relevance of 86% (841/973) of the homo-oligomeric structures that QSPACE ultimately selects to represent <italic>E. coli</italic> proteome (Fig. S10) using QSalignWeb<sup><xref ref-type="bibr" rid="c45">45</xref></sup> — a webserver that uses superposition of structures to infer the physiological relevance of a quaternary structure. We provide all relevant quality metrics associated with each structure, and the QSalign inferred relevance (when applicable) for all proteins in the <italic>E. coli</italic> QSPACE in Dataset S3B.</p>
<p>The final structural representation of the <italic>E. coli</italic> proteome is a collection of experimental structures (deposited in the PDB) and models (generated by SWISS-MODEL, I-TASSER, AlphaFold, and AlphaFold Multimer) (<xref rid="fig2" ref-type="fig">Fig. 2b</xref>). The collection of structures identified by QSPACE captures the multi-subunit assembly of 1,473 oligomeric proteins (<xref rid="fig2" ref-type="fig">Fig. 2c</xref>). Proteins that are not known to oligomerize and that have no structural evidence of oligomerization are mapped to their respective monomeric structures as in previous GEM-PRO formulations<sup><xref ref-type="bibr" rid="c27">27</xref>–<xref ref-type="bibr" rid="c31">31</xref></sup>. We show that QSPACE identifies the structures of higher-order oligomeric enzymes (<xref rid="fig2" ref-type="fig">Fig. 2d</xref>). Among these oligomers, the QSPACE platform identifies high-confidence structures for 51/54 ATP-binding cassette (ABC) transporters in <italic>E. coli</italic> (defined in EcoCyc<sup><xref ref-type="bibr" rid="c39">39</xref></sup> and/or iJL1678b<sup><xref ref-type="bibr" rid="c26">26</xref></sup>). Only 4 of these transporters have experimentally resolved structures in the PDB (2QI9, 3RLF, 7CGE &amp; GMHU). We present high-confidence novel structures QSPACE generated with AF-multimer for the 47/50 remaining ABC-transporters in <xref rid="fig2" ref-type="fig">Figure 2e</xref>. Incomplete AF-multimer models (3/50, <italic>asterisks</italic> in <xref rid="fig2" ref-type="fig">Fig. 2e</xref>) provide obvious suggestions for the correct gene-stoichiometry of ABC-transporters that were incorrectly annotated at the time of publication (e.g. putative ABC-55 transporter is missing an ATP-binding subunit).</p>
<p>When compared to the latest <italic>E. coli</italic> genome-scale model with protein structures (iML1515-GP<sup><xref ref-type="bibr" rid="c40">40</xref></sup>), QSPACE improves the oligomeric structural annotation for 70% of genes in iML1515, while offering a 2.86-fold increase in gene coverage and higher quality structures (Fig. S1). To our knowledge this result is the most advanced genome-scale structural representation of the <italic>E. coli</italic> proteome and <italic>de facto</italic> represents a major advancement in genome annotation.</p>
</sec>
<sec id="s2c">
<title>Interoperable data types form the basis for predicting mutant phenotypes</title>
<p>An accurate 3D structural representation of the proteome can serve as a scaffold for mapping multiple data types, thus providing a structured approach to data integration. The interoperability of multiple datatypes can accelerate our understanding of structure-function relationships and mechanisms. To this end, QSPACE uses third-party software (Table S1) to map residue-level, sequence-level, and protein-level properties (columns in Dataset S1) to all amino acids of the <italic>E. coli</italic> proteome. To illustrate the extensive functional content contained in the <italic>E. coli</italic> QSPACE, we provide a global accounting of all functionally important regions of the <italic>E. coli</italic> proteome (Fig. S4 and Dataset S10).</p>
<p>Non-synonymous mutation—the swapping of one amino acid residue for another—provides an opportunity for QSPACE to be used for mutant analysis. Residue-level properties (e.g., the Grantham score<sup><xref ref-type="bibr" rid="c48">48</xref></sup>) of each mutation are calculated. The physio-chemical properties (e.g., the hydrophobicity) of the local sequence (i.e., 5 amino acids centered at the mutation) of each mutation are also determined. Using the protein structure, QSPACE can calculate the properties of the local 3D environment (all amino acids within a fixed radius) of a mutation. The interoperable mapping of multiple datatypes onto the protein structure also allows for the calculation of unique properties (e.g., the distance between a mutation and the nearest protein active site). A graphical summary of mutant-specific properties can be found in <xref rid="fig3" ref-type="fig">Figure 3a</xref>.</p>
<p>The UniProt knowledgebase<sup><xref ref-type="bibr" rid="c37">37</xref></sup> contains annotated phenotypes for over 12,000 non-synonymous <italic>E. coli</italic> mutations. We use keyword phrases (Dataset S4) to assign each mutations annotated phenotype to one of eight phenotype classes of varying severity (<xref rid="fig3" ref-type="fig">Fig. 3b</xref>). Combined with the 100 mutant-specific properties calculated by QSPACE (Dataset S5), the mutant-phenotype UniProt dataset can be used to train Random Forest (RF) classifiers that can predict the severity of mutations in novel mutational databases (e.g., from adaptive laboratory evolutions, the long-term evolution experiment, or the natural sequence variants) (<xref rid="fig3" ref-type="fig">Fig. 3c</xref>).</p>
</sec>
<sec id="s2d">
<title>Random-forest classification of mutant phenotypes</title>
<p>We investigated the accuracy with which Random Forest (RF) classifiers predicted mutant phenotypes and the relative importance of higher-dimensional features. To this end, we selected all UniProt mutations found in proteins containing annotated <italic>active sites</italic> (e.g., <xref rid="fig4" ref-type="fig">Fig. 4a</xref>) and calculated 100 mutant-specific properties for each mutation (see <xref rid="fig3" ref-type="fig">Fig. 3</xref>, Dataset S5). To quantify the importance of higher-dimensional properties, we trained three sets of RF-classifiers on varying combinations of residue (“0D”, gray), sequence (“1D”, yellow), and structure (“3D”, green) features, and iteratively removed the least-predictive feature after 100 train/test cycles until the 30 most-predictive features were identified (<xref rid="fig4" ref-type="fig">Fig. 4a</xref>). For each set of RF-classifiers, we quantified model performance (accuracy, precision, and recall) using “One vs Rest” validation for each phenotype class (<xref rid="fig4" ref-type="fig">Fig. 4b</xref>). The importance of individual features used to train the “3D” RF-classifiers is shown (<xref rid="fig4" ref-type="fig">Fig. 4c</xref>).</p>
<p>Mutations in the UniProt dataset are not limited to proteins containing active sites (<xref rid="fig4" ref-type="fig">Fig. 4d</xref>). Thus, we followed the procedure described in <xref rid="fig4" ref-type="fig">Figure 4a-c</xref> to obtain a global assessment of our ability to predict mutant phenotypes found in proteins containing various functionally important annotations (UniProt “Feature”). The Area Under the Curve (AUC) for the Receiver Operating Characteristic (ROC) and Precision-Recall (P-R) curves were weighted by the relative occurrence of each phenotype class (i.e., horizontal line in <xref rid="fig4" ref-type="fig">Figure 4b</xref>, right) and plotted for RF-classifiers trained for each functional class (<xref rid="fig4" ref-type="fig">Fig. 4e</xref>). For each functional annotation, the cumulative importance of residue-level, sequence-level, and structural features in “3D” RF-classifiers supports the use of protein structures as a context to study interoperable datatypes and mutations (<xref rid="fig4" ref-type="fig">Fig. 4f</xref>).</p>
</sec>
<sec id="s2e">
<title>The membrane module yields angstrom-level subcellular compartmentalization of the <italic>E. coli</italic> proteome</title>
<p>While mapping data types to individual protein complex structures can prove useful, understanding the location and space that these protein complexes occupy in the cell is important for building a genome-scale representation that reflects the physical embodiment of a proteome. To date, genomic databases of <italic>E. coli</italic> (e.g., EcoCyc) assign the entire gene to a subcellular compartment. UniProt sometimes offers sequence annotation of transmembrane and topological (‘bulb’) domains, however, these annotations may be inaccurate (see <xref rid="fig5" ref-type="fig">Fig. 5b</xref>) or missing entirely. Since the structure is not used to determine a protein’s subcellular compartment, assigned cellular compartments can often be incomplete (e.g. there is no distinction between membrane proteins that contain and those that do not contain membrane-spanning regions) and the residue-level orientation of a protein across the cell membrane cannot be achieved. Likewise, sequence-based prediction software (e.g., DeepTMHMM<sup><xref ref-type="bibr" rid="c49">49</xref></sup>) and structure-based prediction software (e.g., OPM<sup><xref ref-type="bibr" rid="c50">50</xref></sup>) are agnostic to membrane orientation and can also generate erroneous results.</p>
<p>To achieve a residue-level representation of the <italic>E. coli</italic> proteome, we use a structure-guided approach that combines and assesses all available annotations and predictions (from UniProt, DeepTMHMM, and OPM) to better identify the integration and orientation of the membrane-embedded proteome.</p>
<p>QSPACE queries the available gene-level subcellular compartment information provided by Ecocyc<sup><xref ref-type="bibr" rid="c39">39</xref></sup>, UniProt<sup><xref ref-type="bibr" rid="c37">37</xref></sup>, Gene Ontology<sup><xref ref-type="bibr" rid="c51">51</xref></sup>, and genome-scale model iML1515<sup><xref ref-type="bibr" rid="c40">40</xref></sup> to identify all potential membrane-embedded protein structures (<xref rid="fig5" ref-type="fig">Fig. 5a</xref>). For each identified structure, QSPACE determines the membrane-spanning residues for each subunit using the sequence annotations provided in UniProt, the sequence-based predictions generated by DeepTMHMM<sup><xref ref-type="bibr" rid="c49">49</xref></sup>, and the structure-based calculation of the membrane planes predicted by OPM<sup><xref ref-type="bibr" rid="c50">50</xref></sup>. For each of the three sources of residue information (when available), QSPACE calculates the normal vectors of the corresponding membrane planes (<xref rid="fig5" ref-type="fig">Fig. 5b</xref>). For each pair of membrane planes, the angle between the planes, and the thickness and area of the membrane-embedded region are used to determine whether the calculated membranes are viable (<xref rid="fig5" ref-type="fig">Fig. 5c</xref>).</p>
<p>QSPACE segregates each viable membrane protein into three sections: a membrane-embedded region and two ‘bulbous’ regions. Each bulb is automatically assigned (Dataset S6) to either the cytoplasmic, periplasmic, or extracellular side of either the inner or outer membrane, using the annotated topological domains in UniProt or manually assigned (Dataset S7) using common 3D motifs in the protein structures (<xref rid="fig5" ref-type="fig">Fig. 5d</xref>). Proteins annotated to the cell membrane (<xref rid="fig5" ref-type="fig">Fig. 5a</xref>) that do not contain a membrane-embedded region are considered ‘membrane-associated’ and tagged to their respective membrane while those tagged to the cytoplasm or periplasm are left unchanged. The gene ontology (GO) terms of genes mapped to non-membrane proteins were used to assign proteins to the cytoplasm, periplasm, or extracellular space.</p>
<p>In <italic>E. coli</italic>, QSPACE was able to assign 86% of proteins (89% of AAs) to one of twelve subcellular compartments (<xref rid="fig5" ref-type="fig">Fig. 5e</xref>), resulting in a residue-level annotation of cellular compartmentalization of the <italic>E. coli</italic> proteome across both cellular membranes. The membrane integration for an additional 5% of proteins (2% of AAs) is known (<xref rid="fig5" ref-type="fig">Fig. 5e</xref>, compartments #13-17), however there is insufficient information to properly orient these proteins across the membrane (e.g. short, single-pass transmembrane helix proteins, see Dataset S7). Incorporated into genome-scale models that compute protein expression (or proteomic datasets), the residue-level compartmentalization of each protein structure provides a first-principles approach to compute the location and size taken up by a cell’s proteome.</p>
</sec>
<sec id="s2f">
<title>Computing the physical space required by the <italic>E. coli</italic> proteome</title>
<p>The multi-subunit protein complexes carry out metabolic reactions, transport nutrients across the cell membrane, maintain cellular homeostasis, replicate the cellular genome, and even synthesize other proteins. Considering all these functions simultaneously calls for the use of computational models. As genome-scale models (GEMs) have increased in scope and mechanistic detail<sup><xref ref-type="bibr" rid="c22">22</xref>–<xref ref-type="bibr" rid="c26">26</xref>,<xref ref-type="bibr" rid="c52">52</xref></sup>, they require the biosynthesis and proper assembly of multi-subunit complexes to drive the reactions in their reconstructed metabolic networks.</p>
<p>While genome-scale models using protein structures (GEM-PROs) have been used for a variety of applications<sup><xref ref-type="bibr" rid="c53">53</xref></sup> (e.g., contextualization of disease-associated human mutations<sup><xref ref-type="bibr" rid="c32">32</xref></sup>, identification of protein-fold conservation in similar metabolic reactions<sup><xref ref-type="bibr" rid="c31">31</xref></sup>, prediction of thermosensitivity in a metabolic network<sup><xref ref-type="bibr" rid="c30">30</xref></sup>, comparative structural analyses of multiple organisms<sup><xref ref-type="bibr" rid="c28">28</xref></sup>), the promise of a complete physical representation of a functioning cellular proteome has yet to be delivered. QSPACE moves us close to this goal by calculating the subcellular compartment of every amino acid across the proteome.</p>
<p>The successful annotation and 3D orientation of proteins across the subcellular compartments is crucial for building genome-scale models that can predict the physical distribution of the cellular proteome. A geometric analysis of the compartmentalized proteome (<xref rid="fig6" ref-type="fig">Fig. 6a</xref>) allows us to calculate the volume occupied (<xref rid="fig6" ref-type="fig">Fig. 6b</xref>) as well as the membrane area required (if applicable) (<xref rid="fig6" ref-type="fig">Fig. 6c</xref>) by each protein. Genome-scale models of metabolism and macromolecular expression (ME-models) (<xref rid="fig6" ref-type="fig">Fig. 6d</xref>) predict the proteome allocation required to sustain growth in optimally growing bacterial cells (<xref rid="fig6" ref-type="fig">Fig. 6e</xref>). In calculating the physical space required by each protein, the spatial requirements of model-predicted proteomes can also be determined (<xref rid="fig6" ref-type="fig">Fig. 6f-g</xref>). Thus, it is now possible to compute the composition and location of the structural proteome. A more detailed supra-protein-complex-level 3D arrangement requires additional considerations<sup><xref ref-type="bibr" rid="c54">54</xref>–<xref ref-type="bibr" rid="c55">55</xref></sup>.</p>
</sec>
</sec>
<sec id="s3">
<title>Discussion</title>
<p>The 3D visualization and modeling of the structural proteome of a functioning cell has been an implicit goal of genome-scale annotations and computational biology methods. QSPACE, introduced here, rapidly identifies and annotates multi-subunit protein structures (including <italic>de novo</italic> annotations of protein-complex assemblies and <italic>de novo</italic> structural models) at the genome-scale for computational modeling and structural analyses. In conjunction with mutational databases, functional annotations, and other data types, the oligomeric structures identified through QSPACE can be used to obtain a deeper understanding of whole-cell functions.</p>
<p>To achieve a physical representation of the cellular proteome, the structure of each individual protein complex in its native state is needed. To this end, QSPACE allows for multi-gene mapping to oligomeric crystallographic depositions (e.g., PDB bioassemblies), existing homo-oligomeric structural models (e.g., high-quality<sup><xref ref-type="bibr" rid="c47">47</xref></sup> SWISS-PROT models<sup><xref ref-type="bibr" rid="c10">10</xref></sup> with high QS-scores<sup><xref ref-type="bibr" rid="c46">46</xref></sup>), and <italic>de novo</italic> high quality oligomeric models (from Alphafold Multimer<sup><xref ref-type="bibr" rid="c14">14</xref></sup>/ColabFold<sup><xref ref-type="bibr" rid="c38">38</xref></sup>). Unlike a purely annotative workflow, QSPACE uses a structure-guided assessment to identify previously unannotated oligomeric assemblies and generate <italic>de novo</italic> structural models when the existing structural data for a protein complex is incomplete. As an example, we present the novel structures of 50/54 ABC-transporters in <italic>E.coli</italic>, and show that even incomplete models 3/50 can provide clues to the correct oligomerization of a protein (<xref rid="fig2" ref-type="fig">Fig. 2e</xref>). In the <italic>E. coli</italic> QSPACE, we confirmed the physiological relevance for 86% of the homo-oligomeric structures with QSalignWeb<sup><xref ref-type="bibr" rid="c45">45</xref></sup> (Fig. S10). Thus, QSPACE achieves the structural representation of multi-subunit protein complexes, a significant advancement over existing genome-scale models using structural biology software (ssbio<sup><xref ref-type="bibr" rid="c27">27</xref></sup>, see Fig. S1).</p>
<p>The protein structures identified by QSPACE are a well-suited 3D scaffold on which to calculate protein properties, identify enzymatic domains, and analyze impactful mutations. QSPACE’s interoperability of various data types (columns in Dataset S1), can drive biological discovery. In this study, we showed how the residue-level, sequence-level, and protein-level properties calculated by QSPACE (<xref rid="fig3" ref-type="fig">Fig. 3a</xref>) for the mutations annotated in the UniProt knowledgebase (<xref rid="fig3" ref-type="fig">Fig. 3b</xref>) can be used to accurately predict mutant phenotypes (<xref rid="fig4" ref-type="fig">Fig. 4b</xref> &amp; 4e). Interestingly, when we iteratively removed the least predictive properties from the RF-classifiers during the training phase (<xref rid="fig4" ref-type="fig">Fig. 4a</xref>), we found that the predictive power of RF models was overwhelmingly the result of structure-level features (<xref rid="fig4" ref-type="fig">Fig. 4c</xref> &amp; 4f). Thus, QSPACE provides users a rapid way to interact with relevant structures and interoperable datatypes to elucidate structure-function relationships across multiple scales.</p>
<p>In addition to the annotated mutations in UniProt, QSPACE can also be used to analyze novel mutational data sets from adaptive laboratory evolutions (ALEdb<sup><xref ref-type="bibr" rid="c20">20</xref></sup>, Fig. S5), the long-term evolution experiment (LTEE<sup><xref ref-type="bibr" rid="c33">33</xref></sup>, Fig. S6), and the natural sequence variation<sup><xref ref-type="bibr" rid="c36">36</xref></sup> of <italic>E. coli</italic> in three dimensions (<xref rid="fig3" ref-type="fig">Fig. 3c</xref>). To our knowledge, the <italic>E. coli</italic> QSPACE provides the first 3D representation of the natural sequence variation of an organism at the genome-scale, and it moves the description and scale of the structural proteome to the species level.</p>
<p>QSPACE advances whole cell modeling efforts<sup><xref ref-type="bibr" rid="c54">54</xref>,<xref ref-type="bibr" rid="c56">56</xref>–<xref ref-type="bibr" rid="c57">57</xref></sup> by establishing structural annotations relevant for molecular processes. Advancements with computational genome-scale models (GEMs) over the past decade have allowed for the prediction of proteome allocation for cells at optimal growth rate<sup><xref ref-type="bibr" rid="c22">22</xref>–<xref ref-type="bibr" rid="c26">26</xref>,<xref ref-type="bibr" rid="c52">52</xref></sup>. Increasingly detailed, GEMs include reactions for protein assembly and translocation across subcellular compartments (e.g., membranes), however, previous GEM formulations with monomeric protein structures (GEM-PROs) have yet to reflect the biophysical embodiment of these <italic>in silico</italic> processes.</p>
<p>Using a structures-based approach that combines and assesses all available annotations and predictions for membrane-spanning proteins, QSPACE determines the membrane integration and orientation of proteins across both the inner and outer membrane of <italic>E. coli</italic>. In fact, QSPACE even calculates membrane integration for proteins that span both membranes (e.g., AcrAB-TolC efflux pump, PDB:5v5s, see <xref rid="fig5" ref-type="fig">Fig. 5e</xref>). In this study, QSPACE determined the subcellular compartment for 89% of amino acids in <italic>E. coli</italic>. As a proof-of-concept, we combine the protein-level information in QSPACE with a genome-scale model of macromolecular expression (iJL1678b-ME<sup><xref ref-type="bibr" rid="c26">26</xref></sup>) to calculate the physical size occupied by the predicted proteome of <italic>E. coli</italic> at optimal growth rate. To our knowledge, this first-principles approach resulted in the first GEM-PRO that embodies the spatial allocation of the <italic>E. coli</italic> proteome.</p>
<p>Taken together, the QSPACE genome annotation platform proves users a rapid method to interact with the best available quaternary structures for any list of proteins (e.g., a strain), can accommodate natural sequence variations described by the alleleome<sup><xref ref-type="bibr" rid="c36">36</xref></sup> to generate species-level structural proteomes, and enables a physical embodiment of the structural proteome against the 3D morphology of the bacterial cell. The analysis of mutant phenotypes and the size calculation for the <italic>E. coli</italic> proteome demonstrate that QSPACE is amenable to diverse applications. As structures are resolved for large protein complexes, as the scope of genome-scale models expands to include an increasing number of niche cellular mechanisms (e.g., stress responses), and as new mutations of functional importance are annotated in publicly available databases, the QSPACE platform will provide an interoperable pipeline for the structural proteomes for a growing list of organisms.</p>
</sec>
<sec id="s4">
<title>Limitations</title>
<p>We emphasize that QSPACE is a large-scale annotation platform that interfaces with numerous third-party software to quickly map multiple interoperable datasets onto relevant protein structures. QSPACE applications range from the single-protein to the genome-scale. As such, the structures identified by QSPACE reflect the gene-stoichiometry of the protein complexes defined by the user. For extensively studied organisms (e.g., <italic>E. coli</italic>), these protein complexes have been defined over decades of published work. For less-studied organisms, a structural proteome assigned by the workflow presented in this study may be incomplete. In such cases, QSPACE can still provide insights into the structural proteome. For instance, QSPACE can generate the entirety of all homo-oligomerization states of an organism’s genome. By modifying the user-defined protein complexes to reflect the gene-stoichiometry of any theoretical homo-oligomerization state of a gene(s), QSPACE will identify (or use AF-multimer to generate) structures for these oligomers, provide the user with a quality assessment of each oligomeric structure, and would thus reveal homo-oligomerization states backed by structural evidence.</p>
<p>As QSPACE relies on third-party software and repositories for the generation of novel structures, the mapping of datasets, and the calculation of structural properties, it is limited by the maintenance, capabilities and accuracy of such resources. For example, the use of AF-multimer via ColabFold allows for modelling protein complexes up to 2000 amino acids, and the quality assessment of such models is currently based on their iPTM and PTM scores, rendering QSPACE incapable of generating higher-order structures that have not been published in repositories. As new modelling platforms, better scoring methods, and larger repositories of pre-computed structures are disseminated by the structural biology community, we see potential for their incorporation into the QSPACE workflow to identify increasingly accurate structures for user-defined proteins. Thus, the maintenance of the QSPACE codebase is vital to ensure that QSPACE can provide users with the most-accurate protein structures for future applications.</p>
</sec>
</body>
<back>
<ack>
<title>Acknowledgements</title>
<p>We would like to thank Marc Abrams for assistance with manuscript editing. This work was funded by Novo Nordisk Foundation (Grant Number NNF20CC0035580) (E.A.C. and B.O.P.) and NIH (Grant R01 GM057089) (B.O.P.).</p>
</ack>
<sec id="s8">
<title>Author Contributions</title>
<p>E.A.C. and B.O.P. designed and performed the research; E.A.C., N.M. and M.L. contributed analytical tools and code; E.A.C. and B.O.P. analyzed data; E.A.C. and B.O.P. wrote the manuscript; E.A.C., N.M., M.L. and B.O.P. edited the manuscript; and B.O.P. supervised the research.</p>
</sec>
<sec id="s9">
<title>Competing Interests</title>
<p>The authors declare no competing interest.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="c1"><label>1.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Bai</surname></string-name>, <string-name><given-names>G.</given-names> <surname>McMullan</surname></string-name>, <string-name><given-names>S.H.W.</given-names> <surname>Scheres</surname></string-name></person-group>, <article-title>How cryo-EM is revolutionizing structural biology</article-title>. <source>Trends in Biochem. Sci</source>. <volume>40</volume>(<issue>1</issue>), <fpage>49</fpage>–<lpage>57</lpage> (<year>2015</year>). DOI: <pub-id pub-id-type="doi">10.1016/j.tibs.2014.10.005</pub-id>.</mixed-citation></ref>
<ref id="c2"><label>2.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Cheng</surname></string-name></person-group>, <article-title>Single-particle cryo-EM—How did it get here and where will it go</article-title>. <source>Science</source>, <volume>361</volume>(<issue>6405</issue>), <fpage>876</fpage>–<lpage>880</lpage> (<year>2018</year>). DOI: <pub-id pub-id-type="doi">10.1126/science.aat4346</pub-id>.</mixed-citation></ref>
<ref id="c3"><label>3.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.P.</given-names> <surname>Renaud</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Chari</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>H.W.</given-names> <surname>Remigy</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Start</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Wiesmann</surname></string-name></person-group>, <article-title>Cryo-EM in drug discovery: achievements, limitations and prospects</article-title>. <source>Nat. Rev. Drug Discov</source>. <volume>17</volume>, <fpage>471</fpage>–<lpage>492</lpage> (<year>2018</year>). <pub-id pub-id-type="doi">10.1038/nrd.2018.77</pub-id>.</mixed-citation></ref>
<ref id="c4"><label>4.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Nakane</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Kotecha</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Sente</surname></string-name>, <string-name><given-names>G.</given-names> <surname>McMullan</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Masiulis</surname></string-name>, <string-name><given-names>P.M.G.E.</given-names> <surname>Brown</surname></string-name>, <string-name><given-names>I.T.</given-names> <surname>Grigoras</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Malinauskaite</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Malinauskas</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Miehling</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Uchański</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Karia</surname></string-name>, <string-name><given-names>E.V.</given-names> <surname>Pechnikova</surname></string-name>, <string-name><given-names>E.</given-names> <surname>de Jong</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Keizer</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Bischoff</surname></string-name>, <string-name><given-names>J.</given-names> <surname>McCormack</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Tiemeijer</surname></string-name>, <string-name><given-names>S.W.</given-names> <surname>Hardwick</surname></string-name>, <string-name><given-names>D.Y.</given-names> <surname>Chirgadze</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Murshudov</surname></string-name>, <string-name><given-names>A.R.</given-names> <surname>Aricescu</surname></string-name>, <string-name><given-names>S.H.W.</given-names> <surname>Scheres</surname></string-name></person-group>, <article-title>Single-particle cryo-EM at atomic resolution</article-title>. <source>Nature</source>. <volume>587</volume>, <fpage>152</fpage>–<lpage>156</lpage> (<year>2020</year>). <pub-id pub-id-type="doi">10.1038/s41586-020-2829-0</pub-id></mixed-citation></ref>
<ref id="c5"><label>5.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Cheng</surname></string-name></person-group>, <article-title>Membrane protein structural biology in the era of single particle cryo-EM</article-title>. <source>Curr. Opin. Struct. Biol</source>. <volume>52</volume>, <fpage>58</fpage>–<lpage>63</lpage> (<year>2018</year>). <pub-id pub-id-type="doi">10.1016/j.sbi.2018.08.008</pub-id></mixed-citation></ref>
<ref id="c6"><label>6.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Pearce</surname></string-name>, <string-name><given-names>E.W.</given-names> <surname>Bell</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name></person-group>, <article-title>Folding non-homology proteins by coupling deep-learning contact maps with I-TASSER assembly simulations</article-title>. <source>Cell Rep</source>. <volume>1</volume>, (<year>2021</year>). <pub-id pub-id-type="doi">10.1016/j.crmeth.2021.100014</pub-id></mixed-citation></ref>
<ref id="c7"><label>7.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Yan</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Roy</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Poisson</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name></person-group>. <article-title>The I-TASSER Suite: Protein structure and function prediction</article-title>. <source>Nat. Methods</source>, <volume>12</volume>, <fpage>7</fpage>–<lpage>8</lpage> (<year>2015</year>). <pub-id pub-id-type="doi">10.1038/nmeth.3213</pub-id></mixed-citation></ref>
<ref id="c8"><label>8.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name></person-group>. <article-title>I-TASSER server: new development for protein structure and function predictions</article-title>. <source>Nucleic Acids Res</source>., <volume>43</volume>, <fpage>W174</fpage>–<lpage>W181</lpage> (<year>2015</year>). <pub-id pub-id-type="doi">10.1093/nar/gkv342</pub-id></mixed-citation></ref>
<ref id="c9"><label>9.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Waterhouse</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Bertoni</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Bienert</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Studer</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Tauriello</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Gumienny</surname></string-name>, <string-name><given-names>F.T.</given-names> <surname>Heer</surname></string-name>, <string-name><given-names>T.A.P</given-names> <surname>de Beer</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Rempfer</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Bordoli</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Lepore</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Schwede</surname></string-name></person-group>, <article-title>SWISS-MODEL: homology modeling of protein structures and complexes</article-title>. <source>Nucleic Acids Res</source>. <volume>46</volume>, <fpage>W296</fpage>–<lpage>W303</lpage> (<year>2018</year>). <pub-id pub-id-type="doi">10.1093/nar/gky427</pub-id></mixed-citation></ref>
<ref id="c10"><label>10.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Bienert</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Waterhouse</surname></string-name>, <string-name><given-names>T.A.P.</given-names> <surname>de Beer</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Tauriello</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Studer</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Bordoli</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Schwede</surname></string-name></person-group>, <article-title>The SWISS-MODEL Repository - new features and functionality</article-title>. <source>Nucleic Acids Res</source>. <volume>45</volume>, <fpage>D313</fpage>–<lpage>D319</lpage> (<year>2017</year>). <pub-id pub-id-type="doi">10.1093/nar/gkw1132</pub-id></mixed-citation></ref>
<ref id="c11"><label>11.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Jumper</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Evans</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Pritzel</surname></string-name>, <etal>et al.</etal></person-group> <article-title>Highly accurate protein structure prediction with AlphaFold</article-title>. <source>Nature</source> (<year>2021</year>). <pub-id pub-id-type="doi">10.1038/s41586-021-03819-2</pub-id></mixed-citation></ref>
<ref id="c12"><label>12.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Varadi</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Anyango</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Deshpande</surname></string-name>, <etal>et al.</etal></person-group> <article-title>AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models</article-title>. <source>Nucleic Acids Res</source>. <volume>50</volume>(<issue>D1</issue>), <fpage>D439</fpage>–<lpage>D444</lpage> (<year>2022</year>). <pub-id pub-id-type="doi">10.1093/nar/gkab1061</pub-id></mixed-citation></ref>
<ref id="c13"><label>13.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.M.</given-names> <surname>Berman</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Westbrook</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Gilliland</surname></string-name>, <string-name><given-names>T.N.</given-names> <surname>Bhat</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Weissig</surname></string-name>, <string-name><given-names>I.N.</given-names> <surname>Shindyalov</surname></string-name>, <string-name><given-names>P.E.</given-names> <surname>Bourne</surname></string-name></person-group>. <article-title>The Protein Data Bank</article-title>. <source>Nucleic Acids Res</source>. <volume>28</volume>, <fpage>235</fpage>–<lpage>242</lpage> (<year>2000</year>).<pub-id pub-id-type="doi">10.1093/nar/28.1.235</pub-id></mixed-citation></ref>
<ref id="c14"><label>14.</label><mixed-citation publication-type="preprint"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Evans</surname></string-name>, <string-name><given-names>M.</given-names> <surname>O’Neill</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Pritzel</surname></string-name>, <etal>et al.</etal></person-group>, <article-title>Protein complex prediction with AlphaFold-Multimer</article-title>. <source>bioRxiv</source>. <pub-id pub-id-type="doi">10.1101/2021.10.04.463034</pub-id> <year>2021</year></mixed-citation></ref>
<ref id="c15"><label>15.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Schweke</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Tauriello</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Pantolini</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Schwede</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Cazals</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Lhéritier</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Fernandez-Recio</surname></string-name>, <string-name><given-names>L.A.</given-names> <surname>Rodríguez-Lumbreras</surname></string-name>, <string-name><given-names>O.</given-names> <surname>Schueler-Furman</surname></string-name>, <string-name><given-names>J.K.</given-names> <surname>Varga</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Jiménez-García</surname></string-name>, <string-name><given-names>M.F.</given-names> <surname>Réau</surname></string-name>, <string-name><given-names>A.M.J.J.</given-names> <surname>Bonvin</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Savojardo</surname></string-name>, <string-name><given-names>P.L.</given-names> <surname>Martelli</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Casadio</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Tubiana</surname></string-name>, <string-name><given-names>H.J.</given-names> <surname>Wolfson</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Oliva</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Barradas-Bautista</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Ricciardelli</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Cavallo</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Venclovas</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Olechnovič</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Guerois</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Andreani</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Martin</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Terashi</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Sarkar</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Christoffer</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Aderinwale</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Verburgt</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Kihara</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Marchand</surname></string-name>, <string-name><given-names>B.E.</given-names> <surname>Correia</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Duan</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Qiu</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zou</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Dey</surname></string-name>, <string-name><given-names>R.L.</given-names> <surname>Dunbrack</surname></string-name>, <string-name><given-names>E.D.</given-names> <surname>Levy</surname></string-name>, <string-name><given-names>S.J.</given-names> <surname>Wodak</surname></string-name></person-group>, <article-title>Discriminating physiological from non-physiological interfaces in structures of protein complexes: A community-wide study</article-title>. <source>Proteomics</source>. <volume>23</volume>(<issue>17</issue>) (<year>2023</year>). <pub-id pub-id-type="doi">10.1002/pmic.202200323</pub-id>.</mixed-citation></ref>
<ref id="c16"><label>16.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Schweke</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Pacesa</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Levin</surname></string-name>, <string-name><given-names>C.A.</given-names> <surname>Goverde</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Duhoo</surname></string-name>, <string-name><given-names>L.J</given-names> <surname>Dornfeld</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Dubreuil</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Georgeon</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Ovchinnikov</surname></string-name>, <string-name><given-names>D. N</given-names> <surname>Woolfson</surname></string-name>, <string-name><given-names>B.E.</given-names> <surname>Correia</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Dey</surname></string-name>, <string-name><given-names>E.D.</given-names> <surname>Levy</surname></string-name></person-group>. <article-title>An atlas of protein homo-oligomerization across domains of life</article-title>. <source>Cell</source> (<year>2024</year>).</mixed-citation></ref>
<ref id="c17"><label>17.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.H.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>F. J.</given-names> <surname>Isaacs</surname></string-name>, <string-name><given-names>P. A.</given-names> <surname>Carr</surname></string-name>, <string-name><given-names>Z.Z.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>C.R.</given-names> <surname>Forest</surname></string-name>, <string-name><given-names>G.M.</given-names> <surname>Church</surname></string-name></person-group>, <article-title>Programming cells by multiplex genome engineering and accelerated evolution</article-title>. <source>Nature</source>, <fpage>460</fpage>(<lpage>7257</lpage>), 894–898 (<year>2009</year>). <pub-id pub-id-type="doi">10.1038/nature08187</pub-id></mixed-citation></ref>
<ref id="c18"><label>18.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.E.</given-names> <surname>Sandberg</surname></string-name>, <string-name><given-names>M.J.</given-names> <surname>Salazar</surname></string-name>, <string-name><given-names>L.L.</given-names> <surname>Weng</surname></string-name>, <string-name><given-names>B.O.</given-names> <surname>Palsson</surname></string-name>, <string-name><given-names>A.M.</given-names> <surname>Feist</surname></string-name></person-group>, <article-title>The emergence of adaptive laboratory evolution as an efficient tool for biological discovery and industrial biotechnology</article-title>. <source>Metab. Eng</source>. <volume>56</volume>, <fpage>1</fpage>–<lpage>16</lpage> (<year>2019</year>). <pub-id pub-id-type="doi">10.1016/j.ymben.2019.08.004</pub-id>.</mixed-citation></ref>
<ref id="c19"><label>19.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Kang</surname></string-name>, <string-name><given-names>SH</given-names> <surname>Cho</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Yoo</surname></string-name>, <string-name><given-names>U.</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Cho</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Palsson</surname></string-name>, <string-name><given-names>BK</given-names> <surname>Cho</surname></string-name></person-group>, <article-title>Minireview: Engineering evolution to reconfigure phenotypic traits in microbes for biotechnological applications</article-title>, <source>Comput. Struct. Biotechnol. J</source>. <volume>21</volume>, <fpage>563</fpage>–<lpage>573</lpage> (<year>2023</year>). <pub-id pub-id-type="doi">10.1016/j.csbj.2022.12.042</pub-id>.</mixed-citation></ref>
<ref id="c20"><label>20.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.V.</given-names> <surname>Phaneuf</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Gosting</surname></string-name>, <string-name><given-names>B.O.</given-names> <surname>Palsson</surname></string-name>, <string-name><given-names>A.M.</given-names> <surname>Feist</surname></string-name></person-group>, <article-title>ALEdb 1.0: a database of mutations from adaptive laboratory evolution experimentation</article-title>. <source>Nucleic Acids Res</source>. <volume>47</volume>(<issue>D1</issue>) <fpage>D1164</fpage>–<lpage>D1171</lpage> (<year>2019</year>). <pub-id pub-id-type="doi">10.1093/nar/gky983</pub-id></mixed-citation></ref>
<ref id="c21"><label>21.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.D.</given-names> <surname>Tibocha-Bonilla</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Zuñiga</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Lekbua</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Lloyd</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Rychel</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Short</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Zengler</surname></string-name></person-group>, <article-title>Predicting stress response and improved protein overproduction in Bacillus subtilis</article-title>. <source>NPJ Syst. Biol. Appl</source>. <volume>8</volume>, (<year>2022</year>). <pub-id pub-id-type="doi">10.1038/s41540-022-00259-0</pub-id></mixed-citation></ref>
<ref id="c22"><label>22.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.J.</given-names> <surname>O’Brien</surname></string-name>, <string-name><given-names>J.A.</given-names> <surname>Lerman</surname></string-name>, <string-name><given-names>R.L.</given-names> <surname>Chang</surname></string-name>, <string-name><given-names>D.R.</given-names> <surname>Hyduke</surname></string-name>, <string-name><given-names>B.Ø.</given-names> <surname>Palsson</surname></string-name></person-group>, <article-title>Genome-scale models of metabolism and gene expression extend and refine growth phenotype prediction</article-title>. <source>Mol. Syst. Biol</source>. <volume>9</volume>, (<year>2013</year>). <ext-link ext-link-type="uri" xlink:href="https://www.embopress.org/doi/pdf/10.1038/msb.2013.52#sec-18">https://www.embopress.org/doi/pdf/10.1038/msb.2013.52#sec-18</ext-link></mixed-citation></ref>
<ref id="c23"><label>23.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Du</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>C.J.</given-names> <surname>Lloyd</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Fang</surname></string-name>, <string-name><given-names>B.O.</given-names> <surname>Palsson</surname></string-name></person-group>, <article-title>Genome-scale model of metabolism and gene expression provides a multi-scale description of acid stress responses in Escherichia coli</article-title>. <source>PLOS Comp. Bio</source>. <volume>15</volume>(<issue>12</issue>), (<year>2019</year>). <pub-id pub-id-type="doi">10.1371/journal.pcbi.1007525</pub-id></mixed-citation></ref>
<ref id="c24"><label>24.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Mih</surname></string-name>, <string-name><given-names>BO</given-names> <surname>Palsson</surname></string-name></person-group>, <article-title>Thermosensitiviy of growth is determined by chaperone-mediated proteome reallocation</article-title>. <source>PNAS</source>, <volume>114</volume>(<issue>43</issue>), <fpage>11548</fpage>–<lpage>11553</lpage> (<year>2017</year>). <pub-id pub-id-type="doi">10.1073/pnas.1705524114</pub-id></mixed-citation></ref>
<ref id="c25"><label>25.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Nih</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Anand</surname></string-name>, <string-name><given-names>J.H.</given-names> <surname>Park</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Tan</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yurkovich</surname></string-name>, <string-name><given-names>J.M.</given-names> <surname>Monk</surname></string-name>, <string-name><given-names>C.J.</given-names> <surname>Lloyd</surname></string-name>, <string-name><given-names>T.E.</given-names> <surname>Sandberg</surname></string-name>, <string-name><given-names>S.W.</given-names> <surname>Seo</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>A.V.</given-names> <surname>Sastry</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Phaneuf</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>J.R.</given-names> <surname>Broddrick</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Heckmann</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Szubin</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Hefner</surname></string-name>, <string-name><given-names>A.M.</given-names> <surname>Feist</surname></string-name>, <string-name><given-names>B.O.</given-names> <surname>Palsson</surname></string-name></person-group>, <article-title>Cellular responses to reactive oxygen species are predicted from molecular mechanisms</article-title>. <source>PNAS</source>. <volume>116</volume>(<issue>28</issue>), <fpage>14368</fpage>–<lpage>14373</lpage> (<year>2019</year>). <pub-id pub-id-type="doi">10.1073/pnas.1905039116</pub-id></mixed-citation></ref>
<ref id="c26"><label>26.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.J.</given-names> <surname>Lloyd</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Ebrahim</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Z.A.</given-names> <surname>King</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Catoiu</surname></string-name>, <string-name><given-names>E.J.</given-names> <surname>Obrien</surname></string-name>, <string-name><given-names>J.K.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>B.O.</given-names> <surname>Palsson</surname></string-name></person-group>, <article-title>COBRAme: A computational framework for genome-scale models of metabolism and gene expression</article-title>. <source>PLOS Comp. Bio</source>, <volume>14</volume>(<issue>7</issue>), (<year>2018</year>). <pub-id pub-id-type="doi">10.1371/journal.pcbi.1006302</pub-id></mixed-citation></ref>
    <ref id="c27"><label>27.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Mih</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Brunk</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Catoiu</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Sastry</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Kavvas</surname></string-name>, <string-name><given-names>J.M.</given-names> <surname>Monk</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>B.O.</given-names>,<surname>Palsson</surname></string-name></person-group>, <article-title>, ssbio: a Python framework for structural systems biology</article-title>. <source>Bioinformatics</source>. <volume>34</volume>(<issue>12</issue>), <fpage>2155</fpage>–<lpage>2157</lpage> (<year>2018</year>). <pub-id pub-id-type="doi">10.1093/bioinformatics/bty077</pub-id></mixed-citation></ref>
<ref id="c28"><label>28.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Brunk</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Mih</surname></string-name>, <string-name><given-names>J.M.</given-names> <surname>Monk</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>E.J.</given-names> <surname>Obrien</surname></string-name>, <string-name><given-names>S.E.</given-names> <surname>Bliven</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>R.L.</given-names> <surname>Chang</surname></string-name>, <string-name><given-names>P.E.</given-names> <surname>Bourne</surname></string-name>, <string-name><given-names>B.O.</given-names> <surname>Palsson</surname></string-name></person-group>, <article-title>Systems biology of the structural proteome</article-title>. <source>BMC Syst. Biol</source>. <volume>10</volume>(<issue>26</issue>), (<year>2013</year>). <pub-id pub-id-type="doi">10.1186/s12918-016-0271-6</pub-id></mixed-citation></ref>
<ref id="c29"><label>29.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.L.</given-names> <surname>Chang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Xie</surname></string-name>, <string-name><given-names>P.E.</given-names> <surname>Bourne</surname></string-name>, <string-name><given-names>B.O.</given-names> <surname>Palsson</surname></string-name></person-group>, <article-title>Drug off-target effects predicted using structural analysis in the context of a metabolic network model</article-title>. <source>PLOS Comp. Biol</source>. <volume>6</volume>(<issue>9</issue>), (<year>2010</year>). <pub-id pub-id-type="doi">10.1371/journal.pcbi.1000938</pub-id></mixed-citation></ref>
<ref id="c30"><label>30.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.L</given-names> <surname>Chang</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Andrews</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Godzik</surname></string-name>, <string-name><given-names>B.O.</given-names> <surname>Palsson</surname></string-name></person-group>, <article-title>Structural systems biology evaluation of metabolic thermotolerance in Escherichia coli</article-title>. <source>Science</source>. <fpage>34</fpage>(<lpage>6137</lpage>), 1220-1223 (<year>2013</year>). <pub-id pub-id-type="doi">10.1126/science.1234012</pub-id></mixed-citation></ref>
<ref id="c31"><label>31.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Theile</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Weekes</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Jaroszewski</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Ginalski</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Deacon</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wooley</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Lesley</surname></string-name>, <string-name><given-names>I.A.</given-names> <surname>Wilson</surname></string-name>, <string-name><given-names>B.O.</given-names> <surname>Palsson</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Osterman</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Godzik</surname></string-name></person-group>, <article-title>Three-dimensional structural view of the central metabolic network of Thermotoga maritima</article-title>. <source>Science</source>, <volume>325</volume>(<issue>5947</issue>), <fpage>1544</fpage>-<lpage>1549</lpage> (<year>2009</year>). <pub-id pub-id-type="doi">10.1126/science.1174671</pub-id></mixed-citation></ref>
<ref id="c32"><label>32.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Brunk</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Sahoo</surname></string-name>, <string-name><given-names>D.C.</given-names> <surname>Zielinski</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Altunkaya</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Dräger</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Mih</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Gatto</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Nilsson</surname></string-name>, <string-name><given-names>G.A.</given-names> <surname>Preciat-Gonzalez</surname></string-name>, <string-name><given-names>M.K.</given-names> <surname>Aurich</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Prlić</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Sastry</surname></string-name>, <string-name><given-names>A.D.</given-names> <surname>Danielsdottir</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Heinken</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Noronha</surname></string-name>, <string-name><given-names>P.W.</given-names> <surname>Rose</surname></string-name>, <string-name><given-names>S.K.</given-names> <surname>Burley</surname></string-name>, <string-name><given-names>R.M.T.</given-names> <surname>Fleming</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Nielsen</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Thiele</surname></string-name>, <string-name><given-names>B.O.</given-names> <surname>Palsson</surname></string-name></person-group>, <article-title>Recon3D enables a three-dimensional view of gene variation in human metabolism</article-title>. <source>Nat. Biotechnol</source>. <volume>36</volume>, <fpage>272</fpage>–<lpage>281</lpage> (<year>2018</year>). <pub-id pub-id-type="doi">10.1038/nbt.4072</pub-id></mixed-citation></ref>
<ref id="c33"><label>33.</label><mixed-citation publication-type="data"><person-group person-group-type="author"><collab>Barrick Lab</collab></person-group>, <data-title>LTEE-Ecoli</data-title>. <source>Barrick Lab</source> [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://barricklab.org/shiny/LTEE-Ecoli/">https://barricklab.org/shiny/LTEE-Ecoli/</ext-link> [Accessed 1 April (2022]. <year>2022</year></mixed-citation></ref>
<ref id="c34"><label>34.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.E.</given-names> <surname>Lenski</surname></string-name>, <string-name><given-names>M.R.</given-names> <surname>Rose</surname></string-name>, <string-name><given-names>S.C.</given-names> <surname>Simpson</surname></string-name>, <string-name><given-names>S.C.</given-names> <surname>Tadler</surname></string-name></person-group>, <article-title>Long-term experimental evolution in Escherichia coli adaptation and divergence during 2,000 generations</article-title>, <source>Am. Nat</source>. <volume>138</volume>(<issue>6</issue>) <fpage>1315</fpage>–<lpage>1341</lpage> (<year>1991</year>).</mixed-citation></ref>
<ref id="c35"><label>35.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>O.</given-names> <surname>Tenaillon</surname></string-name>, <string-name><given-names>J. E.</given-names> <surname>Barrick</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Ribeck</surname></string-name>, <string-name><given-names>D. E.</given-names> <surname>Deatherage</surname></string-name>, <string-name><given-names>J. L.</given-names> <surname>Blanchard</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Dasgupta</surname></string-name>, <string-name><given-names>G.C.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Wielgoss</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Cruveiller</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Médigue</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Schneider</surname></string-name>, <string-name><given-names>R. E.</given-names> <surname>Lenski</surname></string-name></person-group>, <article-title>Tempo and mode of genome evolution in a 50,000-generation experiment</article-title>. <source>Nature</source>. <volume>536</volume>, <fpage>165</fpage>–<lpage>170</lpage> (<year>2017</year>).</mixed-citation></ref>
<ref id="c36"><label>36.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.A.</given-names> <surname>Catoiu</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Phaneuf</surname></string-name>, <string-name><given-names>J.M.</given-names> <surname>Monk</surname></string-name>, <string-name><given-names>B.O.</given-names> <surname>Palsson</surname></string-name></person-group>, <article-title>Whole genome sequences from wild-type and laboratory evolved strains define the alleleome and establish its hallmarks</article-title>. <source>PNAS</source>. <volume>120</volume>(<issue>15</issue>), <year>2023</year>. <pub-id pub-id-type="doi">10.1073/pnas.221883512</pub-id></mixed-citation></ref>
<ref id="c37"><label>37.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Apweiler</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Bairoch</surname></string-name>, <string-name><given-names>C.H.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>W.C</given-names> <surname>Barker</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Boeckmann</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Ferro</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Gasteiger</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Lopez</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Magrane</surname></string-name>, <string-name><given-names>M.J</given-names> <surname>Martin</surname></string-name>, <string-name><given-names>D.A</given-names> <surname>Natale</surname></string-name>, <string-name><given-names>C.</given-names> <surname>O’Donovan</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Redaschi</surname></string-name>, <string-name><given-names>L.S.</given-names> <surname>Yeh</surname></string-name></person-group>. <article-title>UniProt: the Universal Protein knowledgebase</article-title>. <source>Nucleic Acids Res</source>. <volume>32</volume>, <fpage>D115</fpage>–<lpage>D119</lpage> (<year>2004</year>). <pub-id pub-id-type="doi">10.1093/nar/gkh131</pub-id></mixed-citation></ref>
<ref id="c38"><label>38.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Mirdita</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Schütze</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Moriwaki</surname></string-name>, <etal>et al.</etal></person-group> <article-title>ColabFold: making protein folding accessible to all</article-title>. <source>Nat Methods</source>. <volume>19</volume>, <fpage>679</fpage>–<lpage>682</lpage> (<year>2022</year>). <pub-id pub-id-type="doi">10.1038/s41592-022-01488-1</pub-id></mixed-citation></ref>
<ref id="c39"><label>39.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.M.</given-names> <surname>Keseler</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Collado-Vides</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Santos-Zavaleta</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Peralta-Gil</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Gama-Castro</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Muñiz-Rascado</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Bonavides-Martinez</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Paley</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Krummenacker</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Altman</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Kaipa</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Spaulding</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Pacheco</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Latendresse</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Fulcher</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Sarker</surname></string-name>, <string-name><given-names>A.G.</given-names> <surname>Shearer</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Mackie</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Paulsen</surname></string-name>, <string-name><given-names>R.P.</given-names> <surname>Gunsalus</surname></string-name>, <string-name><given-names>P.D.</given-names> <surname>Karp</surname></string-name></person-group>, <article-title>Ecocyc: a comprehensive database of Escherichia coli biology</article-title>. <source>Nucleic Acids Res</source>. <volume>39</volume>, <fpage>D583</fpage>–<lpage>590</lpage> (<year>2011</year>). <pub-id pub-id-type="doi">10.1093/nar/gkq1143</pub-id></mixed-citation></ref>
<ref id="c40"><label>40.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Monk</surname></string-name>, <string-name><given-names>C.J.</given-names> <surname>Lloyd</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Brunk</surname></string-name>, <etal>et al.</etal></person-group> <article-title>iML1515, a knowledgebase that computes Escherichia coli traits</article-title>. <source>Nat Biotechnol</source>. <volume>35</volume>, <fpage>904</fpage>–<lpage>908</lpage> (<year>2017</year>). <pub-id pub-id-type="doi">10.1038/nbt.3956</pub-id></mixed-citation></ref>
<ref id="c41"><label>41.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Krissinel</surname></string-name> &amp; <string-name><given-names>K.</given-names> <surname>Henrick</surname></string-name></person-group>, <article-title>Inference of macromolecular assemblies from crystalline state</article-title>. <source>J. Mol. Biol</source>. <volume>372</volume>, <fpage>774</fpage>–<lpage>797</lpage> (<year>2007</year>). <pub-id pub-id-type="doi">10.1016/j.jmb.2007.05.022</pub-id></mixed-citation></ref>
<ref id="c42"><label>42.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>A.A.</given-names> <surname>Canutescu</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Shapovalov</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Obradovic</surname></string-name> &amp; <string-name><given-names>R.L.</given-names> <surname>Dunbrack</surname></string-name></person-group>, <article-title>Statistical analysis of interface similarity in crystals of homologous proteins</article-title>. <source>J. Mol. Biol</source>. <volume>381</volume>, <fpage>487</fpage>–<lpage>507</lpage> (<year>2008</year>). <pub-id pub-id-type="doi">10.1016/j.jmb.2008.06.002</pub-id></mixed-citation></ref>
<ref id="c43"><label>43.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Baskaran</surname></string-name>, <string-name><given-names>J.M.</given-names> <surname>Duarte</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Biyani</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Bliven</surname></string-name> &amp; <string-name><given-names>G.</given-names> <surname>Capitani</surname></string-name></person-group>, <article-title>A PDB-wide, evolution-based assessment of protein–protein interfaces</article-title>. <source>BMC Struct. Biol</source>. <volume>14</volume>(<issue>22</issue>) (<year>2014</year>). <pub-id pub-id-type="doi">10.1186/s12900-014-0022-0</pub-id></mixed-citation></ref>
<ref id="c44"><label>44.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.D.</given-names> <surname>Levy</surname></string-name></person-group>, <article-title>PiQSi: protein quaternary structure investigation</article-title>. <source>Structure</source> <volume>15</volume>(<issue>11</issue>), <fpage>1364</fpage>– <lpage>1367</lpage> (<year>2007</year>). <pub-id pub-id-type="doi">10.1016/j.str.2007.09.019</pub-id>.</mixed-citation></ref>
<ref id="c45"><label>45.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Dey</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Priluskiy</surname></string-name> &amp; <string-name><given-names>E.D.</given-names> <surname>Levy</surname></string-name></person-group>, <article-title>QSalignWeb: A server to predict and analyze protein quaternary structure</article-title>. <source>Front Mol Biosci</source> (<year>2022</year>). <pub-id pub-id-type="doi">10.3389/fmolb.2021.787510</pub-id></mixed-citation></ref>
<ref id="c46"><label>46.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Bertoni</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Kiefer</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Biasini</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Bordoli</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Schwede</surname></string-name></person-group>, <article-title>Modeling protein quaternary structure of homo- and hetero-oligomers beyond binary interactions by homology</article-title>. <source>Sci. Rep</source>. <volume>7</volume> (<year>2017</year>). <pub-id pub-id-type="doi">10.1038/s41598-017-09654-8</pub-id></mixed-citation></ref>
<ref id="c47"><label>47.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Benkert</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Biasini</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Schwede</surname></string-name></person-group>, <article-title>Toward the estimation of the absolute quality of individual protein structure models</article-title>. <source>Bioinformatics</source> <volume>27</volume>, <fpage>343</fpage>–<lpage>350</lpage> (<year>2011</year>). <pub-id pub-id-type="doi">10.1093/bioinformatics/btq662</pub-id></mixed-citation></ref>
<ref id="c48"><label>48.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Grantham</surname></string-name></person-group>, <article-title>Amino acid difference formula to explain protein evolution</article-title>, <source>Science</source>. <volume>185</volume>, <fpage>862</fpage>–<lpage>864</lpage> (<year>1974</year>). doi:<pub-id pub-id-type="doi">10.1126/science.185.4154.862</pub-id></mixed-citation></ref>
<ref id="c49"><label>49.</label><mixed-citation publication-type="preprint"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Hallgren</surname></string-name>, <string-name><given-names>K.D.</given-names> <surname>Tsirigos</surname></string-name>, <string-name><given-names>M.D.</given-names> <surname>Pedersen</surname></string-name>, <string-name><given-names>J.J.A.</given-names> <surname>Armenteros</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Marcatili</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Nielsen</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Krogh</surname></string-name>, <string-name><given-names>O.</given-names> <surname>Winther</surname></string-name></person-group>. <article-title>DeepTMHMM predicts alpha and beta transmembrane proteins using deep neural networks</article-title>. <source>bioRxiv</source> (<year>2022</year>). doi: <pub-id pub-id-type="doi">10.1101/2022.04.08.487609</pub-id></mixed-citation></ref>
<ref id="c50"><label>50.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.A</given-names> <surname>Lomize</surname></string-name>, <string-name><given-names>I.D.</given-names> <surname>Pogozheva</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Joo</surname></string-name>, <string-name><given-names>H.I.</given-names> <surname>Mosberg</surname></string-name>, <string-name><given-names>A.L</given-names> <surname>Lomize</surname></string-name></person-group>, <article-title>OPM database and PPM web server: resources for positioning of proteins in membranes</article-title>. <source>Nucleic Acids Res</source>. <volume>40</volume>, <fpage>D370</fpage>–<lpage>D376</lpage> (<year>2012</year>). <pub-id pub-id-type="doi">10.1093/nar/gkr703</pub-id></mixed-citation></ref>
<ref id="c51"><label>51.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Ashburner</surname></string-name>, <string-name><given-names>C.A.</given-names> <surname>Ball</surname></string-name>, <string-name><given-names>J. A.</given-names> <surname>Blake</surname></string-name>, <etal>et al.</etal></person-group> <article-title>Gene ontology: tool for the unification of biology. The Gene Ontology Consortium</article-title>. <source>Nat. Genet</source>. <volume>25</volume>(<issue>1</issue>), <fpage>25</fpage>–<lpage>29</lpage> (<year>2000</year>). <pub-id pub-id-type="doi">10.1038/75556</pub-id></mixed-citation></ref>
<ref id="c52"><label>52.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.K.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>E.J.</given-names> <surname>O’Brien</surname></string-name>, <string-name><given-names>J.A.</given-names> <surname>Lerman</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Zengler</surname></string-name>, <string-name><given-names>B.O.</given-names> <surname>Palsson</surname></string-name> &amp; <string-name><given-names>A.M.</given-names> <surname>Feist</surname></string-name></person-group>, <article-title>Reconstruction and modeling protein translocation and compartmentalization in Escherichia coli at the genome-scale</article-title>. <source>BMC Syst. Biol</source>. <volume>8</volume>(<issue>110</issue>), <year>2014</year>. <pub-id pub-id-type="doi">10.1186/s12918-014-0110-6</pub-id></mixed-citation></ref>
<ref id="c53"><label>53.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Mih</surname></string-name>, <string-name><given-names>B.O.</given-names> <surname>Palsson</surname></string-name></person-group>, <article-title>Expanding the uses of genome-scale models with protein structures</article-title>. <source>Mol. Cyst. Biol</source>. <volume>15</volume>(<issue>11</issue>), (<year>2019</year>). <pub-id pub-id-type="doi">10.15252/msb.20188601</pub-id></mixed-citation></ref>
<ref id="c54"><label>54.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.R.</given-names> <surname>Thornburg</surname></string-name>, <string-name><given-names>D.M.</given-names> <surname>Bianchi</surname></string-name>, <string-name><given-names>T.A.</given-names> <surname>Brier</surname></string-name>, <string-name><given-names>B.R.</given-names> <surname>Gilbert</surname></string-name>, <string-name><given-names>T.M.</given-names> <surname>Earnest</surname></string-name>, <string-name><given-names>M.C.R.</given-names> <surname>Melo</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Safronova</surname></string-name>, <string-name><given-names>J.P.</given-names> <surname>Sáenz</surname></string-name>, <string-name><given-names>A.T.</given-names> <surname>Cook</surname></string-name>, <string-name><given-names>K.S.</given-names> <surname>Wise</surname></string-name>, <string-name><given-names>C.A.</given-names> <surname>Hutchison</surname></string-name>, <string-name><given-names>H.O.</given-names> <surname>Smith</surname></string-name>, <string-name><given-names>J.I.</given-names> <surname>Glass</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Luthey-Schulten</surname></string-name></person-group>, <article-title>Fundamental behaviors emerge from simulations of a living minimal cell</article-title>. <source>Cell</source>. <volume>185</volume>(<issue>2</issue>), <fpage>345</fpage>–<lpage>360</lpage>, <year>2022</year> <pub-id pub-id-type="doi">10.1016/j.cell.2021.12.025</pub-id>.</mixed-citation></ref>
<ref id="c55"><label>55.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Maritan</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Autin</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Karr</surname></string-name>, <string-name><given-names>M.W.</given-names> <surname>Covert</surname></string-name>, <string-name><given-names>A.J.</given-names> <surname>Olson</surname></string-name>, <string-name><given-names>D.S.</given-names> <surname>Goodsell</surname></string-name></person-group>, <article-title>Building Structural Models of a Whole Mycoplasma Cell</article-title>. <source>J. Mol. Biol</source>. <volume>434</volume>(<issue>2</issue>) <year>2022</year>. <pub-id pub-id-type="doi">10.1016/j.jmb.2021.167351</pub-id>.</mixed-citation></ref>
<ref id="c56"><label>56.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.J.</given-names> <surname>O’Brien</surname></string-name>, <string-name><given-names>J.M.</given-names> <surname>Monk</surname></string-name>, <string-name><given-names>B.O</given-names> <surname>Palsson</surname></string-name></person-group>. <article-title>Using Genome-scale Models to Predict Biological Capabilities</article-title>. <source>Cell</source>. <volume>161</volume>(<issue>5</issue>), <fpage>971</fpage>–<lpage>987</lpage> (<year>2015</year>). doi: <pub-id pub-id-type="doi">10.1016/j.cell.2015.05.019</pub-id>.</mixed-citation></ref>
<ref id="c57"><label>57.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.R</given-names> <surname>Karr</surname></string-name>, <string-name><given-names>J.C.</given-names> <surname>Sanghvi</surname></string-name>, <string-name><given-names>D.N.</given-names> <surname>Macklin</surname></string-name>, <string-name><given-names>M.V.</given-names> <surname>Gutschow</surname></string-name>, <string-name><given-names>J.M.</given-names> <surname>Jacobs</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Bolival Jr</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Assad-Garcia</surname></string-name>, <string-name><given-names>J.I.</given-names> <surname>Glass</surname></string-name>, <string-name><given-names>M.W.</given-names> <surname>Covert</surname></string-name></person-group>. <article-title>A whole-cell computational model predicts phenotype from genotype</article-title>. <source>Cell</source>. <volume>150</volume>(<issue>2</issue>), <fpage>389</fpage>–<lpage>401</lpage> (<year>2012</year>). doi: <pub-id pub-id-type="doi">10.1016/j.cell.2012.05.044</pub-id>.</mixed-citation></ref>
<ref id="c58"><label>58.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Rose</surname></string-name>, <string-name><given-names>J.M.</given-names> <surname>Duarte</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Lowe</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Segura</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Bi</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Bhikadiya</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>A.S.</given-names> <surname>Rose</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Bittrich</surname></string-name>, <string-name><given-names>S.K.</given-names> <surname>Burley</surname></string-name>, <string-name><given-names>J.D.</given-names> <surname>Westbrook</surname></string-name></person-group>. <article-title>RCSB Protein Data Bank: Architectural Advances Towards Integrated Searching and Efficient Access to Macromolecular Structure Data from the PDB Archive</article-title>, <source>J. Mol. Biol</source>. <volume>433</volume>(<issue>11</issue>), (<year>2021</year>). doi: <pub-id pub-id-type="doi">10.1016/j.jmb.2020.11.003</pub-id></mixed-citation></ref>
<ref id="c59"><label>59.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Linding</surname></string-name>, <string-name><given-names>L.J.</given-names> <surname>Jensen</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Diella</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Bork</surname></string-name>, <string-name><given-names>T.J.</given-names> <surname>Gibson</surname></string-name>, <string-name><given-names>R.B.</given-names> <surname>Russell</surname></string-name></person-group>. <article-title>Protein disorder prediction: implications for structural proteomics</article-title>. <source>Structure</source>. <volume>11</volume>(<issue>11</issue>), <fpage>1453</fpage>–<lpage>1459</lpage> (<year>2003</year>). doi: <pub-id pub-id-type="doi">10.1016/j.str.2003.10.002</pub-id>.</mixed-citation></ref>
<ref id="c60"><label>60.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Tubiana</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Schneidman-Duhovny</surname></string-name>, <string-name><given-names>H.J.</given-names> <surname>Wolfson</surname></string-name></person-group>, <article-title>ScanNet: an interpretable geometric deep learning model for structure-based protein binding site prediction</article-title>. <source>Nat Methods</source>. <volume>19</volume>, <fpage>730</fpage>–<lpage>739</lpage> (<year>2022</year>). <pub-id pub-id-type="doi">10.1038/s41592-022-01490-7</pub-id></mixed-citation></ref>
<ref id="c61"><label>61.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Cheng</surname></string-name>, <string-name><given-names>A.Z.</given-names> <surname>Randall</surname></string-name>, <string-name><given-names>M.J.</given-names> <surname>Sweredoski</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Baldi</surname></string-name></person-group>, <article-title>SCRATCH: a protein structure and structural feature prediction server</article-title>. <source>Nucleic Acids Res</source>. <volume>33</volume>(<issue>W</issue>), <fpage>72</fpage>–<lpage>76</lpage> (<year>2005</year>). doi: <pub-id pub-id-type="doi">10.1093/nar/gki396</pub-id>.</mixed-citation></ref>
<ref id="c62"><label>62.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Kabsch</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Sander</surname></string-name></person-group>, <article-title>Dictionary of protein secondary structure: Pattern recognition of hydrogen-bonded and geometrical features</article-title>. <source>Biopolymers</source>, <volume>22</volume>, <fpage>2577</fpage>–<lpage>2637</lpage> (<year>1983</year>). <pub-id pub-id-type="doi">10.1002/bip.360221211</pub-id></mixed-citation></ref>
<ref id="c63"><label>63.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.F.</given-names> <surname>Sanner</surname></string-name>, <string-name><given-names>A.J.</given-names> <surname>Olson</surname></string-name>, <string-name><given-names>J.C.</given-names> <surname>Spehner</surname></string-name></person-group>, <article-title>Reduced Surface: An Efficient Way to Compute Molecular Surfaces</article-title>. <source>Biopolymers</source>. <volume>38</volume>, <fpage>305</fpage>–<lpage>320</lpage> (<year>1996</year>).</mixed-citation></ref>
<ref id="c64"><label>64.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Liebermeister</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Noor</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Flamholz</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Davidi</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Bernhardt</surname></string-name>, and <string-name><given-names>R.</given-names> <surname>Milo</surname></string-name></person-group>, <article-title>Visual account of protein investment in cellular functions</article-title>. <source>PNAS</source>. <volume>111</volume>(<issue>23</issue>), <fpage>8488</fpage>–<lpage>8493</lpage> (<year>2014</year>). <pub-id pub-id-type="doi">10.1073/pnas.131481011</pub-id></mixed-citation></ref>
</ref-list>
<sec id="s5">
<title>Materials and Methods</title>
<p>Detailed methods are provided in the Supplementary Appendix. Additional information can be found at github.com/EdwardCatoiu/QSPACE.</p>
<sec id="s5a">
<title>Overview of the QSPACE workflow</title>
<p>We encourage the reader to familiarize themselves with the ‘demo_QSPACE.ipynb’ tutorial notebook available at (<ext-link ext-link-type="uri" xlink:href="https://github.com/EdwardCatoiu/QSPACE">https://github.com/EdwardCatoiu/QSPACE</ext-link>). The QSPACE platform offers the user flexibility in the use of some or all structural repositories (for identifying/generating structures), third-party software (for calculating structural properties), and mutant datasets. This section will describe the generation of the <italic>E. coli</italic> QSPACE (Dataset S1) and the two applications presented in this study.</p>
<p>The overall QSPACE workflow can be summarized: 1) A list of 4,309 genes identified across 2,661 <italic>E. coli</italic> strains<sup><xref ref-type="bibr" rid="c36">36</xref></sup> and the gene-stoichiometries of <italic>E. coli</italic> proteins are annotated in EcoCyc<sup><xref ref-type="bibr" rid="c39">39</xref></sup> and iJL1678b<sup><xref ref-type="bibr" rid="c26">26</xref></sup> and serve as the user-defined input to the <italic>E. coli</italic> QSPACE; 2) UniProt IDs were identified and corresponding .fasta and .txt files were downloaded; 3) Homology models corresponding to the UniProt IDs are downloaded from various repositories; 4) PDB structures (and associated PDB bioassemblies) corresponding to the UniProt IDs and sequences were downloaded using PDB APIs; 5) For proteins that are annotated to be monomers and for those not included in EcoCyc and/or iJL1678b (assumed monomers), a semi-automated module is used to assess if the oligomeric structures (from PDB/SWISS-MODEL) provide overwhelming evidence of oligomerization; 6) For annotated oligomers (user-defined in #1) whose gene-stoichiometry is not reflected in PDB or SWISS-MODEL, AF-multimer v2.0 (via ColabFold) is used to generate oligomeric models; 7) iPTM and PTM scores are used to assess the quality of AF-multimer models; 8) The highest sequence identity structure (after structural QCQA relevant to its source) is selected for each unique gene stoichiometry. When multiple structures provide the same sequence identity for the same gene stoichiometry, preference is given to PDB, AF-Multimer, AlphaFoldDB, SWISS, and ITASSER, respectively); 9) A structure (or combination of structures) is selected as the representation of each user-defined (or re-defined in #5) protein complex; 10) A CSV file (the backbone of Dataset S1, ‘the QSPACE’) is generated where each row provides a mapping between each amino acid across 4,309 <italic>E. coli</italic> genes (user-defined in #1) and its 3D position (chain and residue number) on the protein structure that reflects the protein complex (user-defined in #1 and/or redefined in #5); 11) When possible, third-party software is used to identify membrane-embedded amino acids in proteins that are believed to be in (or associated with) the membrane; 12) Proteins with membrane-embedded regions are oriented across the membrane using available topological information (automatically) and/or by manual inspection of common motifs with known orientation; 13) Multiple third-party software is used to calculate protein properties; 14) External mutation databases and functional annotations (in UniProt) are mapped to the QSPACE CSV file (Dataset 1).</p>
<p>Notably, in steps 1-8, the highest quality structure for each unique structure-gene stoichiometry is selected from each resource of homology and experimental structures. In step 9, QSPACE selects the specific structure that best represents each user-defined gene-stoichiometries. All relevant quality metrics associated with all structures in the structure pool are defined in Dataset S9. The structures selected by QSPACE (in step 9) from this pool (and all associated quality metrics), are used to generate the residue-level CSV file described in step 10. The associated quality metrics for all selected structures in the final CSV file are provided in Dataset S8.</p>
<p>The QSPACE CSV (Dataset S1, AA-to-structure mapping in step 10, additional data mapped in steps 11-14) contains information that can be utilized for various applications, for example: 15) The severity of UniProt mutants mapped to the QSPACE was determined using a key-word search of their annotated phenotypes; 16) the AA-level mapping in QSPACE was used to calculate the aggregate properties of the local environment (i.e. the amino acids directly adjacent in sequence and in 3D space) of each mutation; 17) annotated mutant phenotypes and calculated properties were used to train RF-classifiers to predict mutant severity; 18) (Separately) the area taken up by each protein in the membrane was calculated from the membrane-embedded regions and inferred membrane planes generated in step 11; 19) The volume of each protein was calculated; 20) The geometry of all proteins were incorporated into a genome-scale model of macromolecular expression, iJL1678b-ME, to compute the physical space taken up by the model-predicted proteome of <italic>E. coli</italic>.</p>
</sec>
</sec>
<sec id="s6">
<title>Data Availability</title>
<p>All data is freely available from public sources.</p>
<p>Structures selected from the PDB were last downloaded March 5<sup>th</sup>, 2023. We show experimental structures from the PDB with accession numbers 2GRX, 5V5S, 7NYU, 1NEK, 6OQS, 6C53, 1PFK, 6V0C. Structures selected from the SWISS-MODEL <italic>E. coli</italic> repository were last downloaded December 20<sup>th</sup>, 2022. We show SWISS-MODELs with UniProt IDs <underline>P33232</underline> and <underline>P39099</underline>. Structures selected from the ITASSER <italic>E. coli</italic> database repository were last downloaded November 3<sup>rd</sup>, 2022. Structures selected from the Alphafold database were last downloaded January 15<sup>th</sup>, 2023. We show Alphafold models with UniProt IDs <underline>P33924</underline> and <underline>P30143</underline>. Structures were last modelled using ColabFold on April 6<sup>th</sup>, 2023. We show Alphafold Multimer/ColabFold models for protein complexes with EcoCyc IDs <underline>CYT-D-UBIOX-CPLX</underline> and <underline>ABC-13-CPLX</underline>.</p>
<p>Protein complex gene stoichiometry data for <italic>E. coli</italic> is provided by the Public SmartTable in EcoCyc at <ext-link ext-link-type="uri" xlink:href="https://ecocyc.org/group?id=Biocyc12-4862-3584200844">https://ecocyc.org/group?id=Biocyc12-4862-3584200844</ext-link> and by genome-scale model iJL1674b-ME at <ext-link ext-link-type="uri" xlink:href="https://github.com/SBRG/ecolime">https://github.com/SBRG/ecolime</ext-link>.</p>
<p>ALE mutation data is available at <ext-link ext-link-type="uri" xlink:href="https://aledb.org/">https://aledb.org/</ext-link>. LTEE mutation data is available at <ext-link ext-link-type="uri" xlink:href="https://barricklab.org/shiny/LTEE-Ecoli/">https://barricklab.org/shiny/LTEE-Ecoli/</ext-link>. Both mutation datasets were mapped to the <italic>E. coli</italic> genome by Catoiu <italic>et. al</italic>. 2023<sup><xref ref-type="bibr" rid="c36">36</xref></sup>.</p>
<p>Data generated in this study is provided in the Supplementary Material and/or at <ext-link ext-link-type="uri" xlink:href="https://github.com/EdwardCatoiu/QSPACE/">https://github.com/EdwardCatoiu/QSPACE/</ext-link>.</p>
<p>Select data generated in this study that exceeds the size limits of GitHub is available at <ext-link ext-link-type="uri" xlink:href="https://drive.google.com/drive/folders/1OkXnPK2YP3WAk62Mmu1p00z2dKiS1HQN?usp=sharing">https://drive.google.com/drive/folders/1OkXnPK2YP3WAk62Mmu1p00z2dKiS1HQN?usp=sharing</ext-link>.</p>
</sec>
<sec id="s7">
<title>Code Availability</title>
<p>All source code for QSPACE is provided at <ext-link ext-link-type="uri" xlink:href="https://github.com/EdwardCatoiu/QSPACE/">https://github.com/EdwardCatoiu/QSPACE/</ext-link>. The best way to build a QSPACE is to follow the detailed instructions in the iPython tutorial notebook (“demo_QSPACE.ipynb”).</p>
<p>QSPACE could not be possible without the following:</p>
<p>Python v.3.7.9 (<ext-link ext-link-type="uri" xlink:href="https://www.python.org/">https://www.python.org/</ext-link>); Ssbio v.0.9.9.8 (<ext-link ext-link-type="uri" xlink:href="https://github.com/SBRG/ssbio">https://github.com/SBRG/ssbio</ext-link>); Biopython v.1.81 (<ext-link ext-link-type="uri" xlink:href="https://github.com/biopython/biopython">https://github.com/biopython/biopython</ext-link>); ScanNet (<ext-link ext-link-type="uri" xlink:href="https://github.com/jertubiana/ScanNet">https://github.com/jertubiana/ScanNet</ext-link>); Nglview v.0.11.9 (<ext-link ext-link-type="uri" xlink:href="https://github.com/nglviewer/nglview">https://github.com/nglviewer/nglview</ext-link>); Pandas v.1.1.5 (<ext-link ext-link-type="uri" xlink:href="https://github.com/pandas-dev/pandas">https://github.com/pandas-dev/pandas</ext-link>); SciPy v.1.5.4 (<ext-link ext-link-type="uri" xlink:href="https://github.com/scipy/scipy">https://github.com/scipy/scipy</ext-link>); NumPy v.1.19.5 (<ext-link ext-link-type="uri" xlink:href="https://github.com/numpy/numpy">https://github.com/numpy/numpy</ext-link>); Matplotlib v.3.3.3 (<ext-link ext-link-type="uri" xlink:href="https://github.com/matplotlib/matplotlib">https://github.com/matplotlib/matplotlib</ext-link>); Matplotlib_venn v.0.11.9 (<ext-link ext-link-type="uri" xlink:href="https://github.com/konstantint/matplotlib-venn">https://github.com/konstantint/matplotlib-venn</ext-link>); PyVenn (<ext-link ext-link-type="uri" xlink:href="https://github.com/tctianchi/pyvenn">https://github.com/tctianchi/pyvenn</ext-link>); Seaborn v.0.10.1 (<ext-link ext-link-type="uri" xlink:href="https://github.com/mwaskom/seaborn">https://github.com/mwaskom/seaborn</ext-link>); scikit-learn v.1.0.2 (<ext-link ext-link-type="uri" xlink:href="https://scikit-learn.org/stable/">https://scikit-learn.org/stable/</ext-link>)</p>
</sec>
</back>
<sub-article id="sa0" article-type="editor-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.100485.1.sa2</article-id>
<title-group>
<article-title>eLife Assessment</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Graña</surname>
<given-names>Martin</given-names>
</name>
<role specific-use="editor">Reviewing Editor</role>
<aff>
<institution-wrap>
<institution>Institut Pasteur de Montevideo</institution>
</institution-wrap>
<city>Montevideo</city>
<country>Uruguay</country>
</aff>
</contrib>
</contrib-group>
<kwd-group kwd-group-type="claim-importance">
<kwd>Important</kwd>
</kwd-group>
<kwd-group kwd-group-type="evidence-strength">
<kwd>Incomplete</kwd>
</kwd-group>
</front-stub>
<body>
<p>This study presents an <bold>important</bold> platform for mapping mutation effects onto higher-level protein structural information, addressing a significant gap in current research. While the work is ambitious and incorporates often-overlooked aspects of higher-order structure, the strength of the evidence supporting some results seems <bold>incomplete</bold>. The quaternary structure modeling appears to underestimate oligomeric proteins compared to previous studies, and the mutation analysis lacks crucial baseline information. Despite these limitations, the method has potential for broader applications and generalization to additional organisms, warranting further development and refinement.</p>
</body>
</sub-article>
<sub-article id="sa1" article-type="referee-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.100485.1.sa1</article-id>
<title-group>
<article-title>Reviewer #1 (Public review):</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<anonymous/>
<role specific-use="referee">Reviewer</role>
</contrib>
</contrib-group>
</front-stub>
<body>
<p>Summary:</p>
<p>This work presents a computational platform that integrates currently available experimental or precomputed datasets and/or state-of-the-art modeling methods to assemble a proteome structure from a given list of genes (representing a whole proteome of an organism, or some specific subset of interest). The main advancement is that the proteome structure contains not only the tertiary structure information (such as is provided by precomputed AlphaFold predicted proteomes) but also information about the quaternary structure. Adding quaternary structure information on the whole proteomes is a challenging problem (and the manuscript would benefit from a more comprehensive introduction section presenting these challenges). Importantly, this addition of quaternary structure information is likely to significantly improve any downstream modelling or prediction. This is because most proteins form either stable or transient complexes, and a significant proportion of proteins interacts with cellular structures such as the different biological membranes. These interactions provide important context for interpreting residue-level information, such as for example the fitness/functional effects of point mutations.</p>
<p>Strengths:</p>
<p>The main strength of this work is that it approaches the question of protein quaternary structure in a comprehensive way. Namely, in addition to oligomeric state, it also includes membrane and cellular localization. It also demonstrates how to use and combine the available experimental and precomputed modelling to achieve the same for any set of genes.</p>
<p>Weaknesses:</p>
<p>The feasibility of obtaining a similar dataset (of useful/informative size) for a more complex organism is not clear.</p>
</body>
</sub-article>
<sub-article id="sa2" article-type="referee-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.100485.1.sa0</article-id>
<title-group>
<article-title>Reviewer #2 (Public review):</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<anonymous/>
<role specific-use="referee">Reviewer</role>
</contrib>
</contrib-group>
</front-stub>
<body>
<p>In this study, a methodology called QSPACE is developed and presented. It integrates structural information for a specific organism, here E. coli. The process entails the gathering of individual structures, including oligomeric information/stoichiometry, the incorporation of data on transmembrane regions, and the utilization of the resulting dataset for the analysis of mutation effects and the allocation of proteomes.</p>
<p>This work aims high, setting an ambitious goal of modeling the quaternary structure of a proteome. The method could be applied to other organisms in the future and has value in that respect. At the same time, the work tries to cover (too?) much ground and some of the results/analyses don't measure up. There are indeed a number of shortcomings and/or inconsistencies in the results presented. The comments below will help improve the work and its usefulness.</p>
<p>(1) It is described that &quot;QSPACE then finds the 3D coordinate file (i.e. &quot;structure&quot;) that best reflects the user-defined (input #2) multi-subunit protein assembly&quot;. What is meant by &quot;best reflects&quot;? What if two different structures with the same stoichiometry are available? Which one is picked?</p>
<p>(2) There appears to be a significant under-estimation of oligomer formation: it is reported that &quot;31% (1,334/4,309) of E. coli genes participate in 1,047 oligomeric complexes, 667 genes are annotated as monomers, and 2,308 genes are not included&quot;. However, it is generally observed that ~50% of E coli genes form homo-oligomers (see PMID 10940245 or more recently 38325366), and adding hetero-oligomers on top of that should increase the fraction of oligomers further. In that respect, the estimate forming the basis of this work (31% of genes participating in oligomeric complexes) seems incorrect. It is unclear why the authors did not identify more proteins as adopting a quaternary structure. It is generally hard to grasp details of the dataset, for example, the simple statistic of how many genes participate in homo- versus hetero- oligomer. Such information is partially presented in panels 2c &amp; 2d, but it is very small and hard to see (I would suggest removing the structures of the ABC transporters to make space to present this with more detail).</p>
<p>(3) There are a number of misleading statements/overstatements that I encourage the authors to revise. For example (not exhaustive):</p>
<p>
&quot;to our knowledge this result is the most advanced genome-scale structural representation of the E. coli proteome and de facto represents a major advancement in genome annotation.&quot;</p>
<p>
&quot;angstrom-level subcellular compartmentalization&quot; - Can we really talk about sub-atomic precision when even side chains can move by several angstroms?</p>
<p>
&quot;we provide a global accounting of all functionally important regions&quot; - &quot;all&quot; is not justified</p>
<p>
&quot;Incorporated into genome-scale models that compute protein expression&quot; - what does that mean? There are gene expression &amp; protein abundance datasets, why is the &quot;compute&quot; necessary?</p>
<p>
&quot;Likewise, sequence-based prediction software (e.g., DeepTMHMM49) and structure-based prediction software (e.g., OPM50) are agnostic to membrane orientation and can also generate erroneous results&quot; - what does &quot;erroenous results&quot; mean in this context? Those tools are not supposed to predict orientation.</p>
<p>(4) What was the benchmark used to estimate the accuracy of orientation assignments?</p>
<p>(5) It is not clear why structural information is required to calculate the volume taken up by different proteins across the proteome. For each protein, the expression level (copy number) is expected to have a significant effect, but I'm unsure of why oligomerization is considered key here. It will modulate the volume exclusion associated with interface contact areas, but isn't this negligible compared to other factors, in particular expression?</p>
<p>(6) Models aiming at predicting deleterious effects of mutations typically use sequence conservation, but I do not see such information used in Figure 4. Assessing the added value of structural information should include such evolutionary information (residue-level sequence conservation) in the baseline.</p>
<p>(7) The &quot;proteome allocation&quot; analysis is presented as an important result, but I did not find details of equations used to conduct this analysis. I assume that &quot;proteome allocation&quot; is based solely on expression, and that &quot;cell volume&quot; uses structural information on top of it. There is a significant difference between &quot;proteome allocation&quot; and &quot;cell volume&quot; as reflected in the proteomaps shown in panels 4e &amp; 4f, but there is no explanation for it. Are the proteins' identities the same in these two panels? Were only proteins counted or was RNA considered as well? Clarifications are needed for RNA, for example, how were volumes calculated in structures containing RNAs? Datasets used to derive these maps should also be provided to enable reproducing them.</p>
<p>(8) I did not see that the structures generated are available - they should be deposited on a permanent repository with a DOI.</p>
</body>
</sub-article>
</article>