<?xml version="1.0" ?><!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.3 20210610//EN"  "JATS-archivearticle1-mathml3.dtd"><article xmlns:ali="http://www.niso.org/schemas/ali/1.0/" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.3" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="nlm-ta">elife</journal-id>
<journal-id journal-id-type="publisher-id">eLife</journal-id>
<journal-title-group>
<journal-title>eLife</journal-title>
</journal-title-group>
<issn publication-format="electronic" pub-type="epub">2050-084X</issn>
<publisher>
<publisher-name>eLife Sciences Publications, Ltd</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">96738</article-id>
<article-id pub-id-type="doi">10.7554/eLife.96738</article-id>
<article-id pub-id-type="doi" specific-use="version">10.7554/eLife.96738.2</article-id>
<article-version-alternatives>
<article-version article-version-type="publication-state">reviewed preprint</article-version>
<article-version article-version-type="preprint-version">1.2</article-version>
</article-version-alternatives>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Chromosomes and Gene Expression</subject>
</subj-group>
<subj-group subj-group-type="heading">
<subject>Genetics and Genomics</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Regulatory genome annotation of 33 insect species</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Asma</surname>
<given-names>Hasiba</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Tieke</surname>
<given-names>Ellen</given-names>
</name>
<xref ref-type="aff" rid="a2">2</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Deem</surname>
<given-names>Kevin D</given-names>
</name>
<xref ref-type="aff" rid="a2">2</xref>
<xref ref-type="author-notes" rid="n1">§</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Rahmat</surname>
<given-names>Jabale</given-names>
</name>
<xref ref-type="aff" rid="a2">2</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Dong</surname>
<given-names>Tiffany</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Huang</surname>
<given-names>Xinbo</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0001-9824-3454</contrib-id>
<name>
<surname>Tomoyasu</surname>
<given-names>Yoshinori</given-names>
</name>
<xref ref-type="aff" rid="a2">2</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0002-4149-2705</contrib-id>
<name>
<surname>Halfon</surname>
<given-names>Marc S</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
<xref ref-type="aff" rid="a3">3</xref>
<xref ref-type="aff" rid="a4">4</xref>
<xref ref-type="aff" rid="a5">5</xref>
<xref ref-type="corresp" rid="cor1">*</xref>
</contrib>
<aff id="a1"><label>1</label><institution>Program in Genetics, Genomics, and Bioinformatics, University at Buffalo-State University of New York</institution>, Buffalo NY 14203</aff>
<aff id="a2"><label>2</label><institution>Department of Biology, Miami University</institution>, Oxford OH 45056</aff>
<aff id="a3"><label>3</label><institution>Department of Biochemistry, University at Buffalo-State University of New York</institution>, Buffalo NY 14203</aff>
<aff id="a4"><label>4</label><institution>Department of Biomedical Informatics, University at Buffalo-State University of New York</institution>, Buffalo NY 14203</aff>
<aff id="a5"><label>5</label><institution>Department of Biological Sciences, University at Buffalo-State University of New York</institution>, Buffalo NY 14260</aff>
</contrib-group>
<contrib-group content-type="section">
<contrib contrib-type="editor">
<name>
<surname>Dalal</surname>
<given-names>Yamini</given-names>
</name>
<role>Reviewing Editor</role>
<aff>
<institution-wrap>
<institution>National Cancer Institute</institution>
</institution-wrap>
<city>Bethesda</city>
<country>United States of America</country>
</aff>
</contrib>
<contrib contrib-type="senior_editor">
<name>
<surname>Dalal</surname>
<given-names>Yamini</given-names>
</name>
<role>Senior Editor</role>
<aff>
<institution-wrap>
<institution>National Cancer Institute</institution>
</institution-wrap>
<city>Bethesda</city>
<country>United States of America</country>
</aff>
</contrib>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>*</label>Author for Correspondence at: Jacobs School of Medicine and Biomedical Sciences University at Buffalo, 955 Main St., Room 5128, Buffalo, NY 14203, 716-829-3126, <email>mshalfon@buffalo.edu</email></corresp>
<fn id="n1" fn-type="present-address"><p><sup>§</sup>Present address: Department of Biology University of Rochester Rochester, NY 14627</p></fn>
</author-notes>
<pub-date date-type="original-publication" iso-8601-date="2024-06-07">
<day>07</day>
<month>06</month>
<year>2024</year>
</pub-date>
<pub-date date-type="update" iso-8601-date="2024-09-18">
<day>18</day>
<month>09</month>
<year>2024</year>
</pub-date>
<volume>13</volume>
<elocation-id>RP96738</elocation-id>
<history>
<date date-type="sent-for-review" iso-8601-date="2024-03-15">
<day>15</day>
<month>03</month>
<year>2024</year>
</date>
</history>
<pub-history>
<event>
<event-desc>Preprint posted</event-desc>
<date date-type="preprint" iso-8601-date="2024-01-25">
<day>25</day>
<month>01</month>
<year>2024</year>
</date>
<self-uri content-type="preprint" xlink:href="https://doi.org/10.1101/2024.01.23.576926"/>
</event>
<event>
<event-desc>Reviewed preprint v1</event-desc>
<date date-type="reviewed-preprint" iso-8601-date="2024-06-07">
<day>07</day>
<month>06</month>
<year>2024</year>
</date>
<self-uri content-type="reviewed-preprint" xlink:href="https://doi.org/10.7554/eLife.96738.1"/>
<self-uri content-type="editor-report" xlink:href="https://doi.org/10.7554/eLife.96738.1.sa3">eLife assessment</self-uri>
<self-uri content-type="referee-report" xlink:href="https://doi.org/10.7554/eLife.96738.1.sa2">Reviewer #1 (Public Review):</self-uri>
<self-uri content-type="referee-report" xlink:href="https://doi.org/10.7554/eLife.96738.1.sa1">Reviewer #2 (Public Review):</self-uri>
<self-uri content-type="referee-report" xlink:href="https://doi.org/10.7554/eLife.96738.1.sa0">Reviewer #3 (Public Review):</self-uri>
<self-uri content-type="author-comment" xlink:href="https://doi.org/10.7554/eLife.96738.1.sa4">Author response:</self-uri>
</event>
</pub-history>
<permissions>
<copyright-statement>© 2024, Asma et al</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Asma et al</copyright-holder>
<ali:free_to_read/>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<ali:license_ref>https://creativecommons.org/licenses/by/4.0/</ali:license_ref>
<license-p>This article is distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License</ext-link>, which permits unrestricted use and redistribution provided that the original author and source are credited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="elife-preprint-96738-v2.pdf"/>
<abstract>
<title>Abstract</title>
<p>Annotation of newly-sequenced genomes frequently includes genes, but rarely covers important non-coding genomic features such as the <italic>cis</italic>-regulatory modules—e.g., enhancers and silencers—that regulate gene expression. Here, we begin to remedy this situation by developing a workflow for rapid initial annotation of insect regulatory sequences, and provide a searchable database resource with enhancer predictions for 33 genomes. Using our previously-developed SCRMshaw computational enhancer prediction method, we predict over 2.8 million regulatory sequences along with the tissues where they are expected to be active, in a set of insect species ranging over 360 million years of evolution. Extensive analysis and validation of the data provides several lines of evidence suggesting that we achieve a high true-positive rate for enhancer prediction. One, we show that our predictions target specific loci, rather than random genomic locations. Two, we predict enhancers in orthologous loci across a diverged set of species to a significantly higher degree than random expectation would allow. Three, we demonstrate that our predictions are highly enriched for regions of accessible chromatin. Four, we achieve a validation rate in excess of 70% using in vivo reporter gene assays. As we continue to annotate both new tissues and new species, our regulatory annotation resource will provide a rich source of data for the research community and will have utility for both small-scale (single gene, single species) and large-scale (many genes, many species) studies of gene regulation. In particular, the ability to search for functionally-related regulatory elements in orthologous loci should greatly facilitate studies of enhancer evolution even among distantly related species.</p>
</abstract>
<kwd-group kwd-group-type="author">
<title>Keywords</title>
<kwd>Regulatory genomics</kwd>
<kwd>arthropods</kwd>
<kwd>enhancers</kwd>
<kwd>cis-regulation</kwd>
<kwd>genome annotation</kwd>
<kwd>functional genomics</kwd>
<kwd>enhancer prediction</kwd>
</kwd-group>
<custom-meta-group>
<custom-meta specific-use="meta-only">
<meta-name>publishing-route</meta-name>
<meta-value>prc</meta-value>
</custom-meta>
</custom-meta-group>
</article-meta>
<notes>
<notes notes-type="competing-interest-statement">
<title>Competing Interest Statement</title><p>The authors have declared no competing interest.</p></notes>
<fn-group content-type="summary-of-updates">
<title>Summary of Updates:</title>
<fn fn-type="update"><p>Additional data availability and more detailed explanations of methods; various clarifications; new data in Supplemental Tables S1 and S2</p></fn>
</fn-group>
<fn-group content-type="external-links">
<fn fn-type="dataset"><p>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5061/dryad.3j9kd51t0">https://doi.org/10.5061/dryad.3j9kd51t0</ext-link>
</p></fn>
</fn-group>
</notes>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>The past two decades have witnessed an explosive rise in sequenced metazoan genomes, from a mere handful in the first few years of the century to over 8,000 today (ref. 1; accessed 16 January 2024). This impressive statistic, however, masks the reality that these genomes exist in various stages of completion. Fewer than 30% of these genomes are assembled at the chromosome level, and only 28% of those have a comprehensive annotation (<xref ref-type="bibr" rid="c1">1</xref>). Moreover, almost none of the genome annotations include regulatory sequences (also referred to as <italic>cis</italic>-regulatory modules, or CRMs) such as enhancers and silencers. This is unfortunate, as CRMs comprise a significant percentage of the genome, and knowledge of these sequences is expected to have value comparable to that of knowing the protein-coding genes. Characterizing CRMs is critical for understanding mechanisms of gene regulation and the organization of gene regulatory networks. Moreover, the role of regulatory mutations is increasingly recognized as a driver of both evolution and disease (<xref ref-type="bibr" rid="c2">2</xref>–<xref ref-type="bibr" rid="c6">6</xref>).</p>
<p>One reason for the overall dearth of regulatory annotations is that, historically, large-scale CRM discovery has been difficult and both resource and labor intensive (<xref ref-type="bibr" rid="c7">7</xref>). For much of the last four decades, CRMs could only be identified through painstaking, low-throughput experimental assays. Although in recent years high-throughput empirical and computational CRM discovery methods have been developed, the various different methods frequently show limited agreement (<xref ref-type="bibr" rid="c8">8</xref>–<xref ref-type="bibr" rid="c10">10</xref>), with the result that comprehensive CRM annotation across all cell types and life-cycle stages has remained a challenge for all but the most exhaustively studied model organisms. The problem is particularly acute for the insects. Insects represent a species-rich class—they constitute somewhere between 65%-90% of all animal species (<xref ref-type="bibr" rid="c11">11</xref>, <xref ref-type="bibr" rid="c12">12</xref>)—and have major impacts on human health and agriculture. The early radiation of the insects, coupled with typically short generation times, means that most of the relevant biomedically and agriculturally important species share little non-coding sequence conservation with each other or with the principal insect model species, <italic>Drosophila melanogaster</italic>. Thus, common sequence-homology based CRM discovery approaches are of little use in providing regulatory insights into these species, and knowledge transfers poorly from one species to another. Furthermore, many insects have a complex and varied life cycle, making it particularly important—yet onerous—to assay for CRM function at multiple stages.</p>
<p>We previously developed a powerful computational method, SCRMshaw (“<underline>S</underline>upervised <italic><underline>C</underline>is</italic>-<underline>R</underline>egulatory <underline>M</underline>odule prediction”), for accurate prediction of CRMs, particularly enhancers (<xref ref-type="bibr" rid="c13">13</xref>–<xref ref-type="bibr" rid="c15">15</xref>). (Although we expect SCRMshaw to be adept at finding multiple CRM types, our validation efforts to date have focused solely on enhancers.) SCRMshaw requires only a sequenced genome and a “training set” of some 15-30 known enhancers that regulate a common pattern of gene expression (e.g., midgut expression) and relies on the idea that enhancers with similar function will have similar sequence characteristics, not possible to detect by eye or by traditional alignment methods, but identifiable using machine learning. Although not universally true, this assumption is robust enough to allow effective enhancer discovery without requiring knowledge of transcription factor binding sites or of the expression patterns of the genes being regulated. Importantly, we have shown that we can leverage the wealth of existing <italic>D. melanogaster</italic> enhancer data (<xref ref-type="bibr" rid="c16">16</xref>) to train models for <italic>cross-species</italic> supervised enhancer discovery in diverged (160-345 million years (MY)) insect species—including flies, mosquitoes, beetles, bees, and wasps—despite a virtually complete lack of observable alignment at the non-coding sequence level (<xref ref-type="bibr" rid="c17">17</xref>–<xref ref-type="bibr" rid="c20">20</xref>).</p>
<p>Here, we use SCRMshaw to undertake an initial regulatory annotation of 33 individual insect genomes, using a collection of 48 training sets composed of experimentally validated <italic>D. melanogaster</italic> enhancers. These species are spread across five orders spanning over 360 MY of evolution, and represent roughly 10% of annotated insect species with scaffold-level or better assembly. Annotated predicted enhancers are provided in a searchable database that allows querying by species, tissue/cell type, or potential target gene. A series of simulations as well as <italic>in silico</italic> and <italic>in vivo</italic> validation experiments demonstrate the effectiveness of our approach and place an upper bound on false-positive prediction rates. Our results represent the first release of a rich insect regulatory genome annotation resource, which will continue to grow and annotate insect regulatory genomes in parallel with the sequencing of new insect genomes.</p>
</sec>
<sec id="s2">
<title>Results</title>
<p>We previously demonstrated that SCRMshaw is remarkably effective at predicting enhancers across the entire ∼345 MY range of the holometabolous insects, using training data derived solely from <italic>D. melanogaster</italic> (<xref ref-type="bibr" rid="c17">17</xref>, <xref ref-type="bibr" rid="c19">19</xref>). SCRMshaw (<xref rid="fig1" ref-type="fig">Fig. 1A</xref>) uses training sets composed of known enhancers defined by a common functional characterization (e.g. “nervous system,” “wing disc”) to build a statistical model that captures their short DNA subsequence (<italic>kmer</italic>) count distribution. This <italic>kmer</italic> distribution is then compared to that of a set of non-enhancer “background” sequences in a machine-learning framework. The <italic>kmers</italic> likely serve as proxies for the unknown transcription factor binding sites, but these sites themselves, even when known, are not explicitly used by the algorithm. The trained model is then used to score overlapping sequence windows in the genome, and the highest-scoring windows are output as predicted enhancers (<xref ref-type="bibr" rid="c13">13</xref>, <xref ref-type="bibr" rid="c14">14</xref>, <xref ref-type="bibr" rid="c21">21</xref>). When searching the genomes of the mosquitoes <italic>Anopheles gambiae</italic> and <italic>Aedes aegypti,</italic> the red flour beetle <italic>Tribolium castaneum,</italic> the honey bee <italic>Apis mellifera,</italic> and the wasp <italic>Nasonia vitripennis,</italic> SCRMshaw successfully predicted enhancers in a cross-species fashion with an approximately 75% prediction success rate, based on reporter gene assays in xenotransgenic <italic>D. melanogaster</italic> (chosen as a pragmatic transgene host species) and comparison to already-identified enhancers in the other species (<xref ref-type="bibr" rid="c17">17</xref>–<xref ref-type="bibr" rid="c20">20</xref>). These results suggest that there are significant, albeit hidden, homologies governing the sequence characteristics of insect enhancers, at least for those involved in a substantial number of gene regulatory networks, and motivated us to apply SCRMshaw to a large and diverse set of sequenced insect genomes.</p>
<fig id="fig1" position="float" orientation="portrait" fig-type="figure">
<label>Figure 1:</label>
<caption><title>The SCRMshaw method and analysis pipeline.</title>
<p>(A) Supervised motif-blind CRM discovery (SCRMshaw). (a) SCRMshaw uses a training set of known <italic>D. melanogaster</italic> enhancers (“training sequences”), drawn from REDfly, that are defined by common functional characterization, and a 10 fold larger background set of similarly sized common functional characterization, non enhancer sequences (“background sequences”). (b) The short DNA subsequence (<italic>kmer</italic>) count distributions of these sequences are then used to train a statistical model. Note that although the pictured example shows 5-mers, <italic>kmers</italic> of different sizes are used for some of the underlying statistical models (see Methods). The trained model (c) is used to score overlapping windows in the “target genome.” (d) High-scoring regions are predicted to be functional regulatory sequences (asterisks). Figure adapted from (<xref ref-type="bibr" rid="c81">81</xref>). (B) The workflow used for the regulatory genome annotation described in this paper. The left side shows pre-processing steps, the right side, post-processing. Input to SCRMshaw consists of the genome sequence and gene annotation. A protein sequence annotation is supplied later for the orthology mapping step. Final results are made available as part of the REDfly regulatory annotation knowledgebase.</p></caption>
<graphic xlink:href="576926v2_fig1.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<sec id="s2a">
<title>A cross-species SCRMshaw pipeline</title>
<p>To facilitate application of SCRMshaw to large numbers of newly-sequenced genomes, we developed a detailed workflow to ensure proper formatting of input genomes, rapid prediction of tissue-specific enhancer sequences, evaluation of results, and annotation of loci with information drawn from the respective orthologous regions in the richly-investigated <italic>D. melanogaster</italic> genome (<xref rid="fig1" ref-type="fig">Fig. 1B</xref>; see Methods). The workflow consists of four major steps: <italic>Preflight</italic>: SCRMshaw requires two input files for each genome: a FASTA-formatted file of the genome sequence itself, and a GFFv3-formatted file of the genome annotation. The genome file is masked for tandem repeats using Tandem Repeat Finder (<xref ref-type="bibr" rid="c22">22</xref>). <italic>Preflight</italic> validates the formats of these files and produces a comprehensive log file that highlights any issues along with basic information such as the number of chromosomes/scaffolds and their sizes, data types present in the annotation (e.g., ‘gene’, ‘exon’, ‘ncRNA’, etc.), and average intergenic distances. <italic>Preflight</italic> also provides a sample output of the SCRMshaw-generated ‘gene’ and ‘exon’ files. This feature allows users to identify any discrepancies or errors stemming from the input files and to reformat these files as needed before running SCRMshaw. Any minor scaffolds that are annotated as not containing genes are discarded.</p>
<sec id="s2a1">
<title>SCRMshaw</title>
<p>SCRMshaw is run as previously described (<xref ref-type="bibr" rid="c15">15</xref>), using the “HD” variant (<xref ref-type="bibr" rid="c21">21</xref>). SCRMshaw_HD scans a genome with a 500 bp sliding window offset in 10 bp increments, using a set of three statistical models that compare the <italic>kmer</italic> composition of a set of training enhancers from <italic>Drosophila melanogaster</italic> to randomly selected <italic>D. melanogaster</italic> non-coding sequences (see Methods).</p>
</sec>
<sec id="s2a2">
<title>Post-processing</title>
<p>The raw SCRMshaw output is post-processed to determine the final set of predicted enhancers. The original post-processing procedure described in (<xref ref-type="bibr" rid="c21">21</xref>) had a tendency to predict enhancers skewed toward long lengths (median ∼ 1100 bp). We have revised that method here (see Methods) to yield predictions of more compact size (median 750 bp), which is more in keeping with empirically characterized enhancer lengths (<xref ref-type="bibr" rid="c23">23</xref>).</p>
</sec>
<sec id="s2a3">
<title>Orthology mapping</title>
<p>Each SCRMshaw prediction is assigned its closest flanking genes as putative enhancer target genes, to aid in interpretation of the SCRMshaw output. There are clear shortcomings to this approach, as many enhancer targets are not the closest genes, leading to mis-assignments (e.g. 24, 25-27). However, potentially more accurate methods for target-gene assignment rely on gene expression data, epigenetic data, and/or chromatin conformation data that are frequently not available for the species we are studying and thus difficult to incorporate into our prediction pipeline (<xref ref-type="bibr" rid="c27">27</xref>–<xref ref-type="bibr" rid="c30">30</xref>). Putative target genes are then mapped to their <italic>D. melanogaster</italic> orthologs (if an ortholog exists) using the <italic>Orthologer</italic> software from the Zdobnov lab (<xref ref-type="bibr" rid="c31">31</xref>) (see Methods). We use <italic>D. melanogaster</italic> as it has by far the most comprehensive gene annotation of the insects and thus provides the most detailed functional information for each gene. Mapping the genes from each species to a common ortholog helps us to assess whether we have obtained predicted enhancers in orthologous loci within the various species on which we have run SCRMshaw. The orthology mapping step requires that a set of predicted proteins is present as part of the existing genome annotation. Note that neither this step nor the earlier target-gene assignment step affects the enhancer predictions themselves, which are generated prior to target-gene assignment and orthology mapping and are considered equally valid regardless of putative target genes and whether or not orthologs can be matched to them.</p>
</sec>
</sec>
<sec id="s2b">
<title>Annotation of 33 insect genomes</title>
<p>We ran our annotation pipeline on an initial set of 33 genomes (additional genome annotation is ongoing). These initial genomes were chosen based on availability and to sample broadly among the holometabolous insect orders and the Hemiptera (<xref rid="fig2" ref-type="fig">Fig 2A</xref>; <xref rid="tbl1" ref-type="table">Table 1</xref>). For each genome, we ran SCRMshaw using a collection of 48 training sets (Supplemental Table S1) and all three SCRMshaw scoring methods (“IMM,” “hexMCD”, “PAC-rc”; (<xref ref-type="bibr" rid="c13">13</xref>, <xref ref-type="bibr" rid="c14">14</xref>)). For fifteen species where a protein annotation was available, we assigned <italic>D. melanogaster</italic> orthologs to each predicted locus, with an average of 54% (range 38-82%) of genes in a given species having a <italic>D. melanogaster</italic> ortholog (<xref rid="fig2" ref-type="fig">Fig. 2B</xref>). The complete data are available in multiple formats (see Data Availability).</p>
<fig id="fig2" position="float" orientation="portrait" fig-type="figure">
<label>Figure 2:</label>
<caption><title>Annotation of 33 insect genomes.</title>
<p>(A) Genomes from five insect orders were annotated in this study (more are ongoing). (B) Percentage of genes with <italic>Drosophila</italic> orthologs as mapped via our orthology pipeline (see Methods), for the 15 mapped species. “No Mapped Fly Orthologs” indicates that our orthology mapping pipeline did not identify clear <italic>D. melanogaster</italic> orthologs. For any given gene, this could reflect either a true lack of a respective ortholog, or failure of our procedure to accurately identify an existing ortholog. For complete species names, see <xref rid="tbl1" ref-type="table">Table 1</xref>. (C) Total SCRMshaw predictions for each species. For each species, the lefthand column shows cumulative results for each SCRMshaw sub-method summed over each of the 48 training sets. The righthand column shows the number of unique predictions after merging overlapping predictions from both sub-methods and training sets. Species are displayed alphabetically by taxonomic order. (See also Supplemental Table S2.) (D) Size distribution of SCRMshaw predictions, prior to merging overlapping predictions but after removing outlier predictions &gt; 2 kb in length. Species are ordered identically to panel C.</p></caption>
<graphic xlink:href="576926v2_fig2.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<table-wrap id="tbl1" orientation="portrait" position="float">
<label>Table 1:</label>
<caption><title>Species used in this study</title></caption>
<graphic xlink:href="576926v2_tbl1.tif" mime-subtype="tiff" mimetype="image"/>
<graphic xlink:href="576926v2_tbl1a.tif" mime-subtype="tiff" mimetype="image"/>
</table-wrap>
<p>Collectively, we predicted a total of 2,873,192 enhancers in these 33 species, with each species having on average approximately 87,000 predictions (Supplemental Table S2) (<xref rid="fig2" ref-type="fig">Fig. 2C</xref>). (As some enhancers are predicted by multiple scoring methods, or from more than one training set, the number of unique sequences is lower at 1,164,354; see gray bars in <xref rid="fig2" ref-type="fig">Fig. 2C</xref> and Supplemental Table S2.) The median length of the predicted enhancers across all species was 750 bp (mean 695, range 490-32500 bp). However, we noted the presence of a small number— 2642, &lt; 0.1% of the total—of unusually large regions (&gt; 2000 bp), the bulk of which were confined to just a few genomes (Supplemental Fig. S1). We therefore discarded any predictions with length greater than 1.5 times the interquartile range of the complete prediction set and re-evaluated the size distribution. This resulted in a median size of 740 bp with a mean of 676 bp and a range of 490-1120 bp, indicating that the overall impact of these outlier sequences is minimal (<xref rid="fig2" ref-type="fig">Fig. 2D</xref>). Inspection of the excessively large elements revealed that they result from SCRMshaw predictions that lie immediately adjacent to each other (without gaps) and have similar SCRMshaw scores, which thus become merged into one broad predicted element. Preliminary analysis suggests that these regions result from insufficient masking of tandem repeat regions and/or genome assembly errors, although other causes, such as extremely enhancer-dense “superenhancer” regions (reviewed by 32), cannot entirely be ruled out.</p>
</sec>
<sec id="s2c">
<title>Many loci contain multiple enhancer predictions</title>
<p>We have noted in the past that SCRMshaw often predicts multiple enhancers in a single locus (e.g. 33). This is consistent with the concept of “shadow enhancers,” sets of multiple redundant or semi-redundant enhancers regulating the same gene (reviewed by 34). On the other hand, if SCRMshaw is predicting enhancers with low specificity (i.e., largely at random), it is instead possible that larger loci may just accumulate a high number of predictions simply due to their greater length.</p>
<p>To distinguish between these possibilities, we conducted simulations on three representative genomes (<italic>D. melanogaster, C. pipiens,</italic> and <italic>A. aegypti</italic>), each with a different average intergenic region size chosen to represent small, medium and large genomes respectively (average sizes 610 bp, 5,274.5 bp, and 14,641 bp; see Methods). We randomized the location of SCRMshaw predictions across the non-coding component of each genome and compared the randomized results to our SCRMshaw output. We found that, for a subset of loci, the number of real SCRMshaw predictions per locus was consistently higher than the number obtained at random (Supplemental Table S3). For example, using the <italic>mapping1.visceral_mesoderm</italic> training set, 4.6% (15/328) of <italic>D. melanogaster</italic> loci containing one or more SCRMshaw predictions had a significantly (<italic>P</italic>&lt;0.001) larger number of predictions/locus than expected, 1.5% (10/658) of <italic>C. pipiens</italic> loci had more predictions than expected, and 1% (3/299) of <italic>A. aegypti</italic> loci had excess predicted enhancers (Supplemental Table S3). Similarly, for the <italic>mapping2.ectoderm</italic> training set, 7.6%, 1.5%, and 11% of loci in the three species, respectively, had a significantly (<italic>P</italic>&lt;0.001) greater than expected number of predictions per locus (Supplemental Table S3). Averaged over all training sets, 3-4% of loci with predictions had significantly more than expected by chance (mean values: 3.0% for <italic>D. melanogaster</italic>, 3.0% for <italic>C. pipiens</italic>, and 3.7% for <italic>A. aegypti</italic>; maximum values: 10.2% for <italic>D. melanogaster,</italic> 9.3% for <italic>C. pipiens</italic>, and 11.6% for <italic>A. aegypti</italic>). Only very few training sets did not have any loci with significantly more enhancers than expected (1 set in <italic>A. aegypti,</italic> 1 in <italic>C. pipiens</italic> and 8 in <italic>D. melanogaster</italic>; Supplemental Table S3).</p>
<p>To further ensure that these results were not influenced by locus size, we binned the loci by length and assessed the numbers of real versus simulated predictions/locus for each bin. The binned results were similar to the results using all loci, i.e., the number of SCRMshaw predictions/locus (<xref rid="fig3" ref-type="fig">Fig. 3</xref>, blue boxplots) was consistently higher than the number obtained via randomization (<xref rid="fig3" ref-type="fig">Fig. 3</xref>, pink boxplots). Results were also similar when comparing predictions only at intergenic versus only at intronic positions (Supplemental Table S3). These results suggest that SCRMshaw is not predicting enhancers randomly throughout the genome, but rather is identifying multiple related “shadow” enhancers in a subset of loci. Consistent with this interpretation, functional tests of the full set of predicted enhancers in several <italic>D. melanogaster</italic> loci has confirmed that many of the predictions act as functionally similar enhancers (T. Williams, personal communication).</p>
<fig id="fig3" position="float" orientation="portrait" fig-type="figure">
<label>Figure 3:</label>
<caption><title>SCRMshaw makes multiple predictions per locus.</title>
<p>The number of SCRMshaw predictions per locus (y-axis) are shown as boxplots for loci falling within the given size ranges (x-axis). Black boxes cover the 25<sup>th</sup>-75<sup>th</sup> percentiles, bars indicate median values and dots indicate values exceeding 1.5 times the interquartile range (boxes are not visible for all bins due to very low degrees of variation). Values in pink represent expected values drawn from randomization, while values in blue represent observed values from SCRMshaw. All values are from results with the training set “mapping1.visceral_mesoderm”; results from other training sets were similar (see Supplemental Table S3). Shown are results from the genomes of (A) <italic>D. melanogaster</italic>, (B) <italic>C. pipiens</italic>, and (C) <italic>A. aegypti</italic> representing small, medium, and large genomes, respectively.</p></caption>
<graphic xlink:href="576926v2_fig3.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
</sec>
<sec id="s2d">
<title>SCRMshaw predicts enhancers in orthologous loci across species</title>
<p>A major premise underlying cross-species applications of SCRMshaw is that conserved regulatory strategies should allow for a model trained on enhancers in one species to predict for similar enhancers in another species. We reasoned that at least some of the time, this should lead to identification of enhancers in orthologous loci, as orthologous genes are frequently involved in similar biological pathways and developmental regulatory networks (<xref ref-type="bibr" rid="c3">3</xref>). Indeed, we previously showed that SCRMshaw was able to predict enhancers in several orthologous loci, for instance those for the <italic>single-minded</italic> locus in <italic>D. melanogaster, A. gambiae,</italic> and <italic>A. mellifera,</italic> and the <italic>wingless</italic> locus in <italic>D. melanogaster, A. mellifera,</italic> and <italic>T. castaneum</italic> (<xref ref-type="bibr" rid="c17">17</xref>). To test whether this is generally true, we used the fifteen species with mapped <italic>D. melanogaster</italic> orthologs (plus <italic>D. melanogaster</italic> itself), and all 48 of our training sets, to compare the number of SCRMshaw-based versus random predictions obtained in common (orthologous) loci.</p>
<p>We observed a substantial reduction in the number of orthologous loci with predictions in common as we moved from considering a set of ten to the full set of all sixteen species (average number of common loci over all 48 training sets: 70.7 (10 species) &gt; 39.4. &gt; 19.8 &gt; 9.14 &gt; 4.12 &gt; 1.6 &gt; 0.6 (16 species) out of an average number of 640 predictions with mapped orthologs) (<xref rid="fig4" ref-type="fig">Fig. 4A</xref>, Supplemental Table S4a). This rapid decline is likely due to a variety of factors, including differences in taxonomic order (e.g., Diptera vs. Hymenoptera), the quality and degree of ortholog data for each species, and the quality of annotation for each species, all of which likely lead to an underestimation of the true numbers of common loci in our data (see Discussion). We simulated SCRMshaw predictions for all sixteen species by randomizing the SCRMshaw results and compared the number of common loci between the simulated and real data for combinations of ten to sixteen species. The number of common loci among the real SCRMshaw predictions was consistently significantly higher than the number of common loci observed in the simulated data (<italic>P</italic>&lt; 0.05; <xref rid="fig4" ref-type="fig">Fig. 4B</xref> and Supplemental Table S4). In particular, when evaluating between 10 and 12 species (data for 14-16 species are not reliable due to the very small number of observed loci in common) only a single training set, <italic>adult_PNS</italic>, did not have a significant overrepresentation of common loci (Supplemental Table S4c). We also examined the fold enrichment, i.e., the extent to which the true number was in excess of the simulated number, and found that on average there were greater than 2.4x more predictions in common loci than expected, when considering groups of 10-12 species; almost all training sets had a fold enrichment &gt;1.5x (<xref rid="fig4" ref-type="fig">Fig. 4C</xref>, Supplemental Table S4d). These results strongly suggest that SCRMshaw is successfully finding sets of conserved (or convergent) enhancers, as it is predicting sequences in orthologous loci in a training-set specific manner across multiple species significantly more often than can be accounted for by chance.</p>
<fig id="fig4" position="float" orientation="portrait" fig-type="figure">
<label>Figure 4:</label>
<caption><title>SCRMshaw predicts CRMs in orthologous loci across species.</title>
<p>(A) The number of loci in common that contain at least one SCRMshaw prediction, for 10 or more species. (B) <italic>z</italic>-scores demonstrating that the number of loci in common with one or more SCRMshaw predictions is significantly higher than expectation, based on 360 randomizations. The small number of common predictions for 14-16 species make these statistics unreliable. Dotted lines indicate <italic>z</italic>-score values representing significance at the (unadjusted) <italic>P</italic>&lt;0.005 and <italic>P</italic>&lt;0.05 levels. (C) Fold enrichment values illustrating the excess of loci in common with one or more SCRMshaw predictions compared to expectation. Dotted line shows 1.5x enrichment.</p></caption>
<graphic xlink:href="576926v2_fig4.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
</sec>
<sec id="s2e">
<title>SCRMshaw predictions correlate with regions of accessible chromatin</title>
<p>Active enhancers are frequently found in regions of accessible chromatin, as assayed by methods such as DNAse-seq, FAIRE-seq (Formaldehyde-Assisted Isolation of Regulatory Elements with sequencing), and ATAC-seq (Assay for Transposase-Accessible Chromatin by sequencing)(<xref ref-type="bibr" rid="c35">35</xref>–<xref ref-type="bibr" rid="c37">37</xref>). We confirmed previously that FAIRE-predicted and SCRMshaw-predicted enhancers in <italic>T. castaneum</italic> have a high (&gt;79%) degree of overlap (<xref ref-type="bibr" rid="c18">18</xref>). To determine whether a similar correlation exists for other species, we compared our SCRMshaw predictions with available FAIRE-seq and ATAC-seq data for seven species: <italic>D. melanogaster, T. castaneum, A. gambiae, Danaus plexippus, Vanessa cardui, Junonia coenia,</italic> and <italic>Heliconius himera</italic>. These comparisons are imperfect, as the tissues used to obtain the chromatin data do not precisely correspond to the training sequences used for SCRMshaw, and the data were obtained using a variety of methods. Nevertheless, in the majority of cases where we were able to establish a rough match between the tissues, we observed significant overlap between the two methods of enhancer detection. For example, in <italic>D. melanogaster</italic>, 40% of SCRMshaw predictions using the <italic>blastoderm.mapping1</italic> training set overlapped ATAC-seq peaks obtained from blastoderm embryos (<italic>P</italic>&lt;4.15e-112, fold enrichment 2.98) (<xref ref-type="bibr" rid="c38">38</xref>)(<xref rid="tbl2" ref-type="table">Table 2</xref> row 2). Similarly, 35% and 41% of SCRMshaw predictions from the <italic>mappng2.wing</italic> and <italic>haltere_disc</italic> sets overlapped FAIRE data that included wing and haltere cells (<xref ref-type="bibr" rid="c39">39</xref>)(<italic>P</italic>&lt;4.6e-117 and 1.5e-102, fold enrichment of 3.29 and 2.98)(<xref rid="tbl2" ref-type="table">Table 2</xref>, rows 1, 3). In the mosquito <italic>A. gambiae</italic>, we compared SCRMshaw predictions from the <italic>embryonic_midgut</italic> and <italic>mapping1.salivary</italic> training sets to ATAC-seq data from adult midgut and salivary tissues, observing overlaps of 40% and 37% respectively (<italic>P</italic>&lt;1.5e-102 and 2.15e-93, fold enrichment of 3.45 and 3.22;)(<xref rid="tbl2" ref-type="table">Table 2</xref>, rows 9, 10). When comparing our SCRMshaw predictions from the <italic>mapping2.wing</italic> set to ATAC-seq data for larval wing tissue in the butterflies <italic>D. plexippus</italic> and <italic>H. himera</italic>, we observed overlaps of 60% and 68% (<italic>P</italic>&lt;5.59e-19 and <italic>P</italic>&lt;9.92e-25, fold enrichment of 1.99, 1.66;)(<xref rid="tbl2" ref-type="table">Table 2</xref>, rows 11, 12)(<xref ref-type="bibr" rid="c40">40</xref>). The butterflies <italic>J. coenia</italic> and <italic>V. cardui</italic> are exceptions; intriguingly, they show a depletion in SCRMshaw predictions compared to expectation (<xref rid="tbl2" ref-type="table">Table 2</xref>, rows 13, 14). Whether this is due to a mismatch in the data used for comparison, the state of the genome assemblies for these two draft genomes, or some other failure of SCRMshaw to perform strongly on these species remains to be determined. Overall, the highly significant overlap we observe between SCRMshaw predictions and open chromatin regions in most species and tissues provides further strong evidence that SCRMshaw is effectively predicting enhancers across a broad range of genomes.</p>
<table-wrap id="tbl2" orientation="portrait" position="float">
<label>Table 2:</label>
<caption><title>Overlap of SCRMshaw predictions with FAIRE-seq and ATAC-seq peaks</title></caption>
<graphic xlink:href="576926v2_tbl2.tif" mime-subtype="tiff" mimetype="image"/>
</table-wrap>
</sec>
<sec id="s2f">
<title>Reporter gene analysis demonstrates that SCRMshaw predictions are functional enhancers</title>
<p>As a concrete test of our ability to use SCRMshaw to predict functional enhancers, we assayed a subset of SCRMshaw predictions by reporter gene analysis in transgenic <italic>D. melanogaster</italic>. Our previous studies have shown a high success rate for such assays, ranging from 65% through &gt;80%, depending on the species and training sets used (<xref ref-type="bibr" rid="c13">13</xref>, <xref ref-type="bibr" rid="c14">14</xref>, <xref ref-type="bibr" rid="c17">17</xref>, <xref ref-type="bibr" rid="c19">19</xref>, <xref ref-type="bibr" rid="c20">20</xref>).</p>
<p>We focused our <italic>in vivo</italic> validation experiments on SCRMshaw predictions made using the <italic>mapping2.wing, haltere_disc,</italic> and <italic>disc.mapping2</italic> training sets and a set of four species we had previously shown to be amenable to SCRMshaw prediction<italic>: D. melanogaster, A. aegypti, T. castaneum,</italic> and <italic>A. mellifera</italic>. The imaginal discs are well-established tissues for investigating gene regulation in <italic>D. melanogaster</italic> (although see Discussion for a caveat on using imaginal discs for evaluating enhancer activities in a cross-species setting), and the chosen training sets all gave significant results in the open-chromatin comparisons discussed above. We selected six sets of putative enhancers for testing (<xref rid="tbl3" ref-type="table">Table 3</xref>, <xref rid="tbl4" ref-type="table">Table 4</xref>). For the first three sets, we selected sequences where we had predictions in each of the orthologous loci for <italic>D. melanogaster, T. castaneum,</italic> and <italic>A. mellifera</italic> (<italic>A. aegypti</italic> was not considered for these sets), and where the <italic>D. melanogaster</italic> prediction mapped near a gene expressed in the wing imaginal discs (<xref rid="fig5" ref-type="fig">Fig. 5</xref>). These predictions were conducted using SCRMshaw’s IMM scoring method only, with post-processing performed using the original method described in (<xref ref-type="bibr" rid="c21">21</xref>), and the Amel_4.5 version of the <italic>A. mellifera</italic> genome. For the second three sets, we chose sequences where we had predictions in orthologous loci for at least three of the four species and where the <italic>D. melanogaster</italic> sequence had previously been tested in a reporter gene assay and was known to be active in the relevant imaginal discs (<xref rid="fig5" ref-type="fig">Fig. 5</xref>). Imposing this latter criterion allowed us to leverage existing knowledge and reduce the necessary amount of <italic>in vivo</italic> testing for each set of predictions, enabling us to test a larger set of sequences overall. For this second set of three, we used all three SCRMshaw scoring methods with a revised post-processing algorithm (see Methods), and the Amel_Hav3.1 <italic>A. mellifera</italic> genome. For all six sets, predictions were chosen for testing based on high SCRMshaw scores and overlap with open-chromatin data (where available). We also considered the position of each prediction within the locus (i.e., first intron, downstream intergenic region, etc.), favoring sequences where position was maintained among the orthologs. Selected sequences were cloned into a cross-species compatible reporter vector (<xref ref-type="bibr" rid="c18">18</xref>, <xref ref-type="bibr" rid="c41">41</xref>), and reporter gene activity was visualized either directly or by using the lineage tracing system G-TRACE (<xref ref-type="bibr" rid="c42">42</xref>). In the latter, enhancer activity is visualized through two reporters; the first reporter visualizes the direct enhancer activity while expression of the second reporter is induced and maintained in all cells that descend from a cell in which the enhancer is initially active, even if activity subsequently shuts off.</p>
<fig id="fig5" position="float" orientation="portrait" fig-type="figure">
<label>Figure 5:</label>
<caption><title>Previously described gene expression and enhancer activity for select <italic>D. melanogaster</italic> sequences predicted by SCRMshaw.</title>
<p>The lefthand column shows native <italic>D. melanogaster</italic> gene expression in imaginal discs (green), while the righthand column shows described enhancer activity (magenta). Gray shading indicates that expression has not been described. Moving clockwise from the left side of each panel are the wing, haltere, leg, and eye-antennal discs. The enhancers whose activities are described in the table are: (B) <italic>ex_BCDE</italic> (<xref ref-type="bibr" rid="c45">45</xref>), (F) <italic>hth_GMR46D04</italic> (<xref ref-type="bibr" rid="c50">50</xref>), (H) <italic>Ubx_GMR39A02</italic> (<xref ref-type="bibr" rid="c50">50</xref>), (J) <italic>psq_GMR41E12</italic> (<xref ref-type="bibr" rid="c50">50</xref>).</p></caption>
<graphic xlink:href="576926v2_fig5.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<table-wrap id="tbl3" orientation="portrait" position="float">
<label>Table 3:</label>
<caption><title>Gene loci chosen for <italic>in vivo</italic> validation</title></caption>
<graphic xlink:href="576926v2_tbl3.tif" mime-subtype="tiff" mimetype="image"/>
</table-wrap>
<table-wrap id="tbl4" orientation="portrait" position="float">
<label>Table 4:</label>
<caption><title>SCRMshaw predictions chosen for <italic>in vivo</italic> validation</title></caption>
<graphic xlink:href="576926v2_tbl4.tif" mime-subtype="tiff" mimetype="image"/>
</table-wrap>
<sec id="s2f1">
<title>Ex</title>
<p>The <italic>expanded (ex</italic>) gene plays a crucial role in tissue growth control, including wings (<xref ref-type="bibr" rid="c43">43</xref>–<xref ref-type="bibr" rid="c45">45</xref>). There was a single prediction in the <italic>D. melanogaster ex</italic> locus, <italic>Dm_ex_20p0</italic>, which falls within an open chromatin region of the third intron (Supplemental Fig. S2), and which comprises an untested subsequence of a longer sequence that acts as an imaginal disc enhancer (<xref rid="fig5" ref-type="fig">Fig. 5B</xref>; (<xref ref-type="bibr" rid="c45">45</xref>)). This sequence drove reporter activity in the wing, leg, and antennal discs (<xref rid="fig6" ref-type="fig">Fig. 6A</xref>). The <italic>T. castaneum</italic> genome also had only one prediction for the <italic>ex</italic> locus, <italic>Tc_ex_9p0</italic>, which overlaps well with a FAIRE-seq peak (Supplemental Fig. S2) within the second intron. Although no imaginal disc activity was observed (<xref rid="fig6" ref-type="fig">Fig. 6B</xref>), <italic>Tc_ex_9p0-</italic>driven reporter gene expression was observed during late pupal stages in the legs (Supplemental Fig. S4G). The <italic>A. mellifera</italic> genome (v4.5) had one prediction, <italic>Am_ex_20p3</italic>, within the fourth intron of the <italic>ex</italic> locus (Supplemental Fig. S2; however, note that subsequent prediction using the updated Amel_Hav3.1 genome and revised SCRMshaw post-processing yielded additional predictions in this locus). <italic>Am_ex_20p3</italic> drove active but variable expression in both the pouch and notum regions of the wing imaginal disc (<xref rid="fig6" ref-type="fig">Fig. 6C</xref>). Although often significantly limited to a small number of cells, the pouch expression of <italic>Am_ex_20p3</italic> was similar in pattern to that seen with <italic>Dm_ex_20p0</italic> (cf. <xref rid="fig6" ref-type="fig">Fig. 6A</xref>). Weak and inconsistent expression in the pouch region of the haltere discs was also observed. In addition, <italic>Am_ex_20p3</italic> drove expression in the leg discs (<xref rid="fig6" ref-type="fig">Fig. 6C</xref>).</p>
<fig id="fig6" position="float" orientation="portrait" fig-type="figure">
<label>Figure 6:</label>
<caption><title>Reporter gene expression for tested <italic>ex</italic>, <italic>klu</italic>, and <italic>ush</italic> predicted enhancer sequences.</title>
<p>Each row shows expression for the indicated construct in (i) wing discs, (ii) haltere discs, (iii) T1 (prothoracic) leg discs, (iv) T2 (mesothoracic) leg discs, (v) T3 (metathoracic) leg discs, and (vi) eye-antennal discs (with eye portion to the left). Positive results were obtained by the enhancers associated with the <italic>ex</italic> locus of <italic>D. melanogaster</italic> (wing, legs, antenna, A) and <italic>A. <underline>mellifera</underline></italic> (wing, haltere, legs, C); the <italic>klu</italic> locus of <italic>D. melanogaster</italic> (wing, haltere, legs, D) and <italic>A. <underline>mellifera</underline></italic> (legs, arrows in F); the <italic>ush</italic> locus of <italic>T. castaneum</italic> (wing, arrow, H(i); eye, H(vi)) and <italic>A. <underline>mellifera</underline></italic> (legs, arrows I(iii, iv, v); eye, I(vi)). Enhancer activities were visualized by UAS-tdTomato that was included in the reporter construct. Scale bar is 50 µm for each column.</p></caption>
<graphic xlink:href="576926v2_fig6.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
</sec>
<sec id="s2f2">
<title>Klu</title>
<p><italic>Klumpfuss (klu</italic>) encodes a zinc finger protein important for proper tissue specification and differentiation (<xref ref-type="bibr" rid="c46">46</xref>). The <italic>D. melanogaster</italic> genome had three predictions for the <italic>klu</italic> locus (Supplemental Fig. S2). We selected <italic>Dm_klu_16p1</italic>, which overlaps a region of open chromatin within the second intron present in most imaginal discs (most distinct in the T3 leg disc; Supplemental Fig. S2). <italic>Dm_klu_16p1</italic> drove active expression in the notum regions of the wing and haltere discs and in the leg discs (<xref rid="fig6" ref-type="fig">Fig. 6D</xref>). <italic>Tc_klu_8p6</italic> was the only prediction for the <italic>T. castaneum klu</italic> locus. It falls within the second intron, similar to <italic>D. melanogaster’s klu</italic> prediction <italic>Dm_klu_161p1</italic> (Supplemental Fig. S2), but lacked enhancer activity (<xref rid="fig6" ref-type="fig">Fig. 6E</xref>). <italic>A. mellifera</italic> had only one prediction for the <italic>klu</italic> locus (<italic>Am_klu_20p2</italic>) (Supplemental Fig. S2). We included this prediction for validation even though its location in the third intron does not match with that in the other two species. <italic>Am_klu_20p2</italic> drove weak expression in the most proximal part of the leg discs (<xref rid="fig6" ref-type="fig">Fig. 6F</xref>, arrows), but no activity was observed in the other imaginal discs. (Note that additional predictions were subsequently produced with our revised post-processing algorithm and the updated Amel_Hav3.1 genome, for both the <italic>T. castaneum</italic> and <italic>A. mellifera klu</italic> loci).</p>
</sec>
<sec id="s2f3">
<title>Ush</title>
<p>Our final choice from our first set for <italic>in vivo</italic> validation was <italic>u-shaped</italic> (<italic>ush</italic>), which encodes a transcription factor with described imaginal disc expression (<xref ref-type="bibr" rid="c47">47</xref>–<xref ref-type="bibr" rid="c49">49</xref>). The <italic>D. melanogaster</italic> genome had two predictions for the <italic>ush</italic> locus (three when using the updated post-processing algorithm) (Supplemental Fig. S2). We tested <italic>Dm_ush_16p4</italic>, but did not observe any larval disc activity (<xref rid="fig6" ref-type="fig">Fig. 6G</xref>). <italic>Tc_ush_6p8</italic> and <italic>Am_ush_20p8</italic> were the only predictions for <italic>T. castaneum</italic> and <italic>A. mellifera</italic>, respectively, both located in the second intron (again, additional predictions are found using updated methods and genome versions). <italic>Tc_ush_6p8</italic> had expression in the wing disc as well as weak activity in the peripodial membrane of the eye-antennal disc (<xref rid="fig6" ref-type="fig">Fig. 6H</xref>, Supplemental Fig. S4K). <italic>Am_ush_20p8</italic> displayed active expression in the peripodial membrane surrounding the eye disc, as well as in a single cell, or small subset of cells, located at the center of each leg disc (<xref rid="fig6" ref-type="fig">Fig. 6I</xref>, arrows).</p>
</sec>
<sec id="s2f4">
<title>Hth</title>
<p>SCRMshaw predicted ten enhancers in the locus of the <italic>D. melanogaster</italic> Hox cofactor <italic>homothorax</italic> (<italic>hth</italic>). One of the two top-scoring predictions, <italic>Dm_hth_30p1</italic>, located in the fifth intron, overlapped the known <italic>hth_GMR46D04</italic> enhancer, which has activity in all of the larval imaginal discs (<xref rid="fig5" ref-type="fig">Fig. 5F</xref>, Supplemental Fig. S3)(<xref ref-type="bibr" rid="c50">50</xref>). <italic>A. aegypti</italic> had a single prediction in the orthologous <italic>hth</italic> locus, <italic>Aa_hth_35p9</italic>, located in the first intron (Supplemental Fig. S3). This sequence failed to display activity in imaginal discs (<xref rid="fig7" ref-type="fig">Fig. 7A</xref>).</p>
<fig id="fig7" position="float" orientation="portrait" fig-type="figure">
<label>Figure 7:</label>
<caption><title>Reporter gene expression for tested <italic>hth</italic>, <italic>Ubx</italic>, and <italic>psq</italic> predicted enhancer sequences.</title>
<p>Each row shows expression for the indicated construct in (i) wing discs, (ii) haltere discs, (iii) T1 (prothoracic) leg discs, (iv) T2 (mesothoracic) leg discs, (v) T3 (metathoracic) leg discs, and (vi) eye-antennal discs (with eye portion to the left). Positive results were obtained by the enhancers associated with the <italic>hth</italic> locus of <italic>T. castaneum</italic> (the notum portion of the wing disc, legs, B); the <italic>Ubx</italic> locus of <italic>A. aegypti</italic> (legs, C), <italic>T. castaneum</italic> (legs, D) (myoblast cells in the wing disc, arrows, E9i); legs, eye, E(iii-vi)), and <italic>A. mellifera</italic> (wing, haltere, legs, F) (wing, haltere, G); the <italic>psq</italic> locus of <italic>T. castaneum</italic> (legs, eye, I). Enhancer activities were detected by the G-TRACE system; magenta represents direct enhancer activity detected by dsRed expression, while green indicates lineage-based GFP expression. Scale bar is 50 µm for each column.</p></caption>
<graphic xlink:href="576926v2_fig7.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<p><italic>T. castaneum</italic> had 23 <italic>hth</italic>-locus predictions. We selected a high-scoring (albeit not the highest-scoring) sequence, <italic>Tc_hth_15p5</italic>, due to the similarity of its position to that of the <italic>D. melanogaster</italic> enhancer—in the fifth intron—and overlapping FAIRE peak (Supplemental Fig. S3)(<xref ref-type="bibr" rid="c18">18</xref>). The <italic>Tc_hth_15p5</italic> reporter showed activity in the presumptive notum region of the wing disc and in the proximal region of the leg discs, resembling both the endogenous <italic>D. melanogaster hth</italic> expression pattern and that of the <italic>GMR46D04</italic> enhancer (<xref rid="fig7" ref-type="fig">Fig. 7B</xref>, cf. <xref rid="fig5" ref-type="fig">Fig. 5F</xref>). No enhancer activity was observed in the haltere or eye-antennal discs (<xref rid="fig7" ref-type="fig">Fig. 7B</xref>). In the pupal stage, <italic>Tc_hth_15p5</italic> exhibited active expression along the thorax, corresponding to adult <italic>hth</italic> expression in <italic>D. melanogaster</italic> (Supplemental Fig. S4I, arrows) (<xref ref-type="bibr" rid="c51">51</xref>). No <italic>hth</italic> predictions were obtained for <italic>A. mellifera</italic>.</p>
</sec>
<sec id="s2f5">
<title>Ubx</title>
<p>The classic homeotic gene <italic>Ultrabithorax</italic> (<italic>Ubx</italic>) regulates tissue identity in the thoracic and abdominal segments (<xref ref-type="bibr" rid="c52">52</xref>). Of six SCRMshaw predictions in the <italic>D. melanogaster Ubx</italic> locus, we noted that one, <italic>Dm_Ubx_36p1</italic>, overlaps a cluster of known enhancers in the third intron centered on <italic>Ubx_GMR39A02</italic> and <italic>Ubx_abx6.8</italic> (Supplemental Fig. S3). <italic>Ubx_GMR39A02</italic> mirrors the haltere activity of <italic>Ubx</italic>, but also displays ectopic activity in the pouch and notum regions of the wing disc (<xref rid="fig5" ref-type="fig">Fig. 5H</xref>)(<xref ref-type="bibr" rid="c50">50</xref>). Similarly, <italic>Ubx_abx6.8</italic> also drives both native and ectopic expression in the imaginal discs (<xref ref-type="bibr" rid="c53">53</xref>).</p>
<p>There were eight predictions in the <italic>Ubx</italic> locus of the <italic>A. aegypti</italic> genome (Supplemental Fig. S3). We selected sequence <italic>Aa_Ubx_26p0</italic>, as it was both the highest scoring prediction and was in the third intron, similar to its putative <italic>D. melanogaster</italic> counterparts. Although direct reporter expression from <italic>Aa_Ubx_26p0</italic> was too weak to observe directly, use of the lineage-tracing G-TRACE system confirmed widespread activity in all leg discs (<xref rid="fig7" ref-type="fig">Fig. 7C</xref>).</p>
<p>We chose two <italic>T. castaneum</italic> sequences for testing, out of ten predictions: <italic>Tc_Ubx_17p4</italic>, which overlapped well with an accessible chromatin region in the third thoracic epidermal tissue and the central nervous system (<xref ref-type="bibr" rid="c18">18</xref>), and <italic>Tc_Ubx_19p9</italic>, which had the highest local prediction score but only overlapped with a chromatin region accessible predominantly during early embryogenesis (Supplemental Fig. S3). <italic>Tc_Ubx_17p4</italic> displayed activity in all three leg discs (<xref rid="fig7" ref-type="fig">Fig. 7D</xref>). The expression in T2 and T3 leg discs corresponds to endogenous <italic>Ubx</italic> expression, whereas T1 leg disc expression appears to be ectopic. <italic>Tc_Ubx_19p9,</italic> predicted using the “haltere” training set, did not drive clear expression in the haltere disc but did drive G-TRACE expression in the wing disc in the adepithelial adult myoblast cells, as well as direct reporter expression in the peripodial membrane of the eye disc. Very limited expression was also observed in the leg discs when using G-TRACE (<xref rid="fig7" ref-type="fig">Fig. 7E</xref>, Supplemental Fig. S4J).</p>
<p>The <italic>A. mellifera</italic> genome had eleven predictions at the <italic>Ubx</italic> locus (Supplemental Fig. S3). <italic>Am_Ubx_37p2</italic> was chosen based on a high SCRMshaw score, although it falls within the fourth, rather than the third, intron. <italic>Am_Ubx_0p39</italic>, on the other hand, was in the corresponding third intron location. <italic>Am_Ubx_0p39</italic> had activity in the pouch region of the wing and haltere discs, as well as in the proximal region of leg discs (<xref rid="fig7" ref-type="fig">Fig. 7F</xref>). <italic>Am_Ubx_37p2</italic> drove expression in specific portions of the wing and haltere discs (<xref rid="fig7" ref-type="fig">Fig. 7G</xref>). Although <italic>Ubx</italic> is not expressed in the <italic>D. melanogaster</italic> wing disc (i.e., the forewing of the fly), wing activity has been observed previously with tested <italic>Ubx</italic> enhancer fragments (e.g. 50). Moreover, <italic>Ubx</italic> is expressed in both the forewing and hindwing discs in honeybees, making the reporter expression we observed driven by the two predicted <italic>A. mellifera</italic> enhancers consistent with their potential native activities (<xref ref-type="bibr" rid="c54">54</xref>).</p>
</sec>
<sec id="s2f6">
<title>Psq</title>
<p>We chose <italic>pipsqueak</italic> (<italic>psq</italic>), a transcription factor involved in Polycomb group gene silencing (<xref ref-type="bibr" rid="c55">55</xref>), and its orthologous loci as the final targets for enhancer validation. There were four predicted <italic>psq</italic> enhancers in <italic>D. melanogaster</italic>, one of which, <italic>Dm_psq_25p6,</italic> overlaps known enhancer <italic>psq_GMR41E12</italic> (Supplemental Fig. S3). This enhancer is located within the second <italic>psq</italic> intron and drives expression in all imaginal discs (<xref rid="fig5" ref-type="fig">Fig. 5J</xref>)(<xref ref-type="bibr" rid="c50">50</xref>). The <italic>A. aegypti</italic> genome had four predictions for the <italic>psq</italic> locus (Supplemental Fig. S3); we chose <italic>Aa_psq_21p5</italic>, located in the second intron and with the highest local SCRMshaw score, for validation (Supplemental Fig. S3). However, <italic>Aa_psq_21p5</italic> did not have observable imaginal disc expression (<xref rid="fig7" ref-type="fig">Fig. 7H</xref>).</p>
<p><italic>T. castaneum</italic> had only two predictions for the <italic>psq</italic> locus, both within the third intron (Supplemental Fig. S3). <italic>Tc_psq_19p7</italic>, which was chosen for validation due to its higher score, drove expression in the eye discs and in a very limited number of cells in the leg discs (<xref rid="fig7" ref-type="fig">Fig. 7I</xref>).</p>
<p>From the <italic>A. mellifera</italic> genome, we selected the only prediction for the <italic>psq</italic> locus, <italic>Am_psq_29p2</italic> (Supplemental Fig. S3). <italic>Am_psq_29p2</italic> reporter activity was negative in all tissues assayed (<xref rid="fig7" ref-type="fig">Fig. 7J</xref>).</p>
</sec>
<sec id="s2f7">
<title>Embryonic activity</title>
<p>Although our SCRMshaw predictions were targeted toward imaginal disc activity, activity was also observed in embryos for many of the reporter lines (Supplemental Fig. S5). Analysis of this activity was complicated by the fact that our reporter lines, while not having any basal activity in imaginal discs (Supplemental Fig. S4, A-E), displayed reporter gene expression in several tissues including hemocytes, caudal visceral mesoderm, and the proventriculus, even in the absence of a putative enhancer sequence (Supplemental Fig. S5A-D). Similar expression is seen in control embryos of the G-TRACE line alone, i.e., even in the absence of a Gal4 driver (data not shown), suggesting that the observed basal activity might be coming from the UAS construct. In those lines that had reporter gene expression in other tissues (Supplemental Fig. S5I-S, X, Y), most of that expression was not clearly associated with the expected endogenous expression of the predicted target gene, although the complex embryonic expression patterns of these genes makes a definitive assessment difficult. Moreover, for most of the species, we do not currently know the expression patterns of either the gene in its native species (as opposed to the expression of its <italic>Drosophila</italic> ortholog), or the expression patterns of other nearby potential target genes. Further analysis, including additional control experiments, use of different reporter vectors, and assessment of gene expression patterns in each of the relevant species will be necessary before drawing final conclusions as to embryonic enhancer activity.</p>
</sec>
</sec>
<sec id="s2g">
<title>An insect regulatory annotation resource</title>
<p>Taken together, the results from our simulations, our comparisons to open chromatin regions, and our <italic>in vivo</italic> validation experiments demonstrate that SCRMshaw is remarkably effective at predicting regulatory sequences across a wide range of insect species. To facilitate access to our SCRMshaw-based regulatory annotations, we created a database with the results from our predictions using all training sets and all completed species. This database, which is freely accessible as part of the REDfly insect regulatory annotation site (16, <ext-link ext-link-type="uri" xlink:href="http://redfly.ccr.buffalo.edu">http://redfly.ccr.buffalo.edu</ext-link>), contains processed final prediction data and can be searched and filtered by gene, <italic>D. melanogaster</italic> ortholog, training set, enhancer location, and various other criteria. We will continue to add to this database as additional species and training sets are run through our annotation pipeline.</p>
</sec>
</sec>
<sec id="s3">
<title>Discussion</title>
<p>The resource we introduce here is part of an ongoing effort to provide an initial regulatory annotation for all sequenced insects. These annotations are not complete, as we currently lack well-curated training sets for many tissues, including key embryonic tissues such as the central nervous system and many non-embryonic tissues. We aim to generate the necessary additional training sets over time, and add these new annotations to those presented here, along with comprehensive annotation for additional species. As a predictive pipeline, SCRMshaw is subject to the usual tradeoffs of sensitivity versus specificity in generating results. Sensitivity is difficult to assess, as there is no “complete” known set of enhancers for any organism, and it is currently impossible to make an accurate estimate of what the yield should be for any of our training models. Similarly, without comprehensive <italic>in vivo</italic> testing using multiple conditions and methods, an accurate false-positive rate cannot be computed. However, we presented here several lines of evidence demonstrating that true positives significantly outweigh false positives: we non-randomly predict multiple enhancers for specific subsets of loci; we predict enhancers in orthologous loci at a rate significantly higher than random expectation; our predictions are highly enriched for regions of accessible chromatin; and we achieve a high rate of validation using <italic>in vivo</italic> reporter gene assays.</p>
<sec id="s3a">
<title>Validation success rates</title>
<p>The results from the <italic>in vivo</italic> validation experiments are summarized in <xref rid="tbl5" ref-type="table">Table 5</xref>. Overall, 17/22 (77%) of tested sequences revealed imaginal disc activity, consistent with our previous SCRMshaw success rates. 59% of these (10/17), or 45% of the total (10/22), had the correct target specificity of wing and/or haltere discs. This is again consistent with previous SCRMshaw experience, which shows that functional enhancers are predicted at a higher success rate than enhancers with specific targeted activity (<xref ref-type="bibr" rid="c13">13</xref>, <xref ref-type="bibr" rid="c14">14</xref>).</p>
<table-wrap id="tbl5" orientation="portrait" position="float">
<label>Table 5:</label>
<caption><title>Summary of <italic>in vivo</italic> validation results</title></caption>
<graphic xlink:href="576926v2_tbl5.tif" mime-subtype="tiff" mimetype="image"/>
</table-wrap>
<p>These numbers may in fact underestimate the success rate of our predictions, for several reasons. One, all of the testing was performed in transgenic <italic>D. melanogaster</italic>, despite the putative enhancers being from three additional species. Reduced efficiency of certain transcription factor or cofactor binding, reduced enhancer-promoter compatibility, or other species-specific differences may lead to elevated false negative results. Furthermore, the unique imaginal disc mode of adult epithelial development in <italic>D. melanogaster</italic> (<xref ref-type="bibr" rid="c56">56</xref>, <xref ref-type="bibr" rid="c57">57</xref>) might have prevented some enhancers of other species from working properly in <italic>D. melanogaster</italic> imaginal discs, likely producing additional false negative results. Evaluating enhancer activities in the native species will allow us to address the degree of false negatives produced by the cross-species setting. Two, predictions for <italic>ex, klu</italic>, and <italic>ush</italic> were made using a less-effective SCRMshaw post-processing algorithm and, for <italic>A. mellifera</italic>, a less-complete genome build. This may have led to suboptimal candidate enhancer selection. Three, it is possible that some of our predicted enhancers act as silencers rather than enhancers. Silencers, which attenuate rather than promote gene expression, are less well understood than enhancers, but in at least some instances have identical sequence characteristics (and in fact can act simultaneously as enhancers in some tissues and silencers in others) (<xref ref-type="bibr" rid="c58">58</xref>, <xref ref-type="bibr" rid="c59">59</xref>). Our reporter gene assay was not designed to detect silencers, which would therefore appear as false-positive predictions. Indeed, an intriguing possibility is that some of our sequences that had reporter gene activity in non-targeted discs (e.g., leg discs) act as enhancers in those tissues but as silencers in the targeted wing discs. Alternatively, identified enhancers that drove expression in multiple imaginal discs may simply represent pleiotropic enhancers.</p>
<p>Enhancer pleiotropy is not uncommon (<xref ref-type="bibr" rid="c60">60</xref>, <xref ref-type="bibr" rid="c61">61</xref>), and moreover, there is overlap in the enhancer content of our training sets for the various discs. There are many gene expression and developmental similarities among the discs, and likely significant regulatory overlap; chromatin profiling experiments have found that wing, leg, and eye discs share the majority of their regions of accessible chromatin (<xref ref-type="bibr" rid="c18">18</xref>, <xref ref-type="bibr" rid="c39">39</xref>). Additional experiments will be necessary to distinguish between pleiotropic enhancers, dual-function silencers/enhancers, and other possibilities.</p>
</sec>
<sec id="s3b">
<title>Redundant enhancers</title>
<p>Almost all of our training sets predicted multiple enhancers in certain loci at rates exceeding random expectation, consistent with the noted prevalence of redundant or “shadow” enhancers in numerous plant and animal species (reviewed by 34). The few training sets that did not may reflect types of tissues or regulatory processes for which having shadow enhancers is unusual, or may be indicative of poorly-performing training sets. Indeed, the sole training set that did not suggest the presence of shadow enhancers in either mosquito species, <italic>adult_PNS</italic>, also performed poorly by other metrics, such as failing to produce predictions in common over multiple species and scoring “poor” using our <italic>pCRM_eval</italic> training set evaluation tool (<xref ref-type="bibr" rid="c21">21</xref>).</p>
<p>Although the evolutionary origins of shadow enhancers are not well understood, several studies suggest that at least one important contribution of shadow enhancers is to provide phenotypic robustness during development, particularly during stress conditions (<xref ref-type="bibr" rid="c62">62</xref>–<xref ref-type="bibr" rid="c67">67</xref>). In some cases, a mechanistic basis for this robustness has been suggested by the finding that different groups of transcription factors act on individual members of a shadow enhancer set (<xref ref-type="bibr" rid="c68">68</xref>, <xref ref-type="bibr" rid="c69">69</xref>). Further support for this idea can be found in the fact that shadow enhancers appear to have limited sequence similarity (<xref ref-type="bibr" rid="c68">68</xref>, <xref ref-type="bibr" rid="c70">70</xref>). However, our successes with SCRMshaw have demonstrated that enhancers can have little overt sequence similarity but still rely on the same underlying subsequence (and presumably, transcription factor binding) model. As the potential shadow enhancers we identify are predicted using the same training set, they are likely to be functioning using similar, rather than independent, regulatory mechanisms. Thus, shadow enhancers appear to come in at least two flavors: those that use similar mechanisms, such as those identified here, and those that use different mechanisms, such as the <italic>Krüppel</italic> enhancers studied by (<xref ref-type="bibr" rid="c68">68</xref>).</p>
</sec>
<sec id="s3c">
<title>Enhancer evolution at large divergence distances</title>
<p>SCRMshaw’s ability to find putatively “orthologous” enhancers—i.e., enhancers in orthologous loci predicted using the same regulatory model—not only substantiates the non-random nature of our prediction approach, but also opens the door to exciting studies of enhancer evolution over large divergence ranges. We previously illustrated the power of such an approach with an analysis of a small number of enhancers in just 2-3 highly diverged species (<xref ref-type="bibr" rid="c17">17</xref>). However, the greatly increased number of species for which we now have regulatory predictions should allow us to follow the evolution of specific enhancers over their entire divergence range. Although the number of common loci tails off rapidly as the number of species considered increases, we suspect that this small (albeit statistically significant) number understates the true results. For one, only a limited number of species have been evaluated for common predictions so far, and these are not evenly distributed along the phylogenetic spectrum of sequenced species. Also, our analysis depends wholly on the presence of recognized <italic>D. melanogaster</italic> orthologous genes to define the common loci. This in turn is dependent on several factors, including the sensitivity and accuracy of our ortholog-calling pipeline, the reliability of the protein annotations we use in that pipeline, and on accurate target gene assignments. Each of these has known sources of error. For example, our ortholog-calling pipeline currently does not disambiguate multiple paralogs (see Methods), and ortholog detection has been shown to be sensitive to the method used for generating protein annotation, in particular when different annotation methods have been applied to different species in the comparison (<xref ref-type="bibr" rid="c71">71</xref>). Moreover, our target gene assignments are presently based solely on closest-gene relationships, a method which is known to mis-assign a fair number of enhancer-gene target pairings (e.g. 24, 25-27) and which does not take into account complexities in genome architecture such as nested genes, long non-coding RNAs, promoter competition, and the like (which are common in <italic>D. melanogaster</italic>, e.g. 72, 73). Developing effective and scalable solutions to these issues will be an important goal for future work.</p>
<p>As the number of species in our prediction database with fully mapped orthologs grows, it will become possible to ask increasingly sophisticated questions about the nature of enhancer evolution. For instance, we will be able to determine whether the numbers of putatively orthologous enhancer predictions follow phylogenetic relationships and degree of sequence divergence, and whether certain loci are only found in common for specific groups of species. Also in need of further investigation will be to determine whether sets of common enhancers are restricted to certain functional Gene Ontology categories, and how this might vary with phylogenetic grouping.</p>
</sec>
<sec id="s3d">
<title>Ongoing regulatory annotation</title>
<p>While effective, SCRMshaw still has various limitations and aspects in need of improvement. Errors in genome assembly and insufficient repeat masking both appear to contribute to overly long predictions (multiple kilobases) that are unlikely to represent individual single enhancers (HA and MSH, unpublished observations). Poor assembly, while not having major effects on SCRMshaw overall, can significantly affect results at specific loci, as can inaccurate gene annotation (<xref ref-type="bibr" rid="c21">21</xref>). Also requiring further investigation is how to best combine and weight the scores from the three individual SCRMshaw scoring methods of IMM, hexMCD, and PAC-rc. Individual SCRMshaw predictions should therefore be treated as just that—predictions—and appropriate validation experiments are recommended for any sequences of interest. The regulatory annotations presented here represent initial “1.0” versions of the regulatory genome. Like all genome annotations, these will continue to be revised as updated models become available, including addition of new training sets, confidence scores based on validation experiments, and improved genome builds. Nevertheless, the enrichment scores from our <italic>in silico</italic> experiments and the true-positive rates from our <italic>in vivo</italic> reporter gene experiments suggest that overall, true positive rates for most training sets are between 50-85%, with most likely having success rates exceeding 70%. Note that even this lower bound of a 50% true positive rate means that one out of every two predictions is correct—a more than acceptable rate to encourage follow-up experiments for predicted enhancers in organisms of interest. The catalog of predicted enhancers we introduce here, spanning 33 insect species and growing, should thus be useful resource for both large-scale and small-scale studies of the insect regulatory genome.</p>
</sec>
</sec>
<sec id="s4">
<title>Methods</title>
<sec id="s4a">
<title>Datasets</title>
<p>Sequence and annotation for <italic>D. melanogaster</italic> were obtained from FlyBase (<xref ref-type="bibr" rid="c74">74</xref>). For other species, wherever possible, genome sequence (FASTA) and annotation (GFF) files were downloaded from NCBI at <ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/datasets">www.ncbi.nlm.nih.gov/datasets</ext-link>. Otherwise, the genome sequences and annotations were obtained from the primary literature or directly from the data generators (see <xref rid="tbl1" ref-type="table">Table 1</xref> for references). Version numbers for genomes and annotations are provided in <xref rid="tbl1" ref-type="table">Table 1</xref>.</p>
</sec>
<sec id="s4b">
<title>Scripts</title>
<p>All scripts used for the analyses described in this paper can be found in the GitHub repository Asma_etal_2024_eLife at <ext-link ext-link-type="uri" xlink:href="https://github.com/HalfonLab/Asma_etal_2024_eLife">https://github.com/HalfonLab/Asma_etal_2024_eLife</ext-link>.</p>
</sec>
<sec id="s4c">
<title>SCRMshaw</title>
<p>A detailed SCRMshaw protocol can be found at Asma et al., 2024 (<xref ref-type="bibr" rid="c75">75</xref>).</p>
<p>Genome files were masked using Tandem Repeat Finder (<xref ref-type="bibr" rid="c22">22</xref>) using parameters 2 7 7 80 10 50 500 -m -h. Genome and annotation files were then assessed using our <italic>preflight</italic> script. Any sequence scaffolds not containing annotated genes were removed before passing the genome sequence to the main SCRMshaw program. SCRMshaw was run using the “HD” version, which scans the genome with a 500 bp window sliding in 10 bp increments, as described in Asma and Halfon (<xref ref-type="bibr" rid="c21">21</xref>). The <italic>--thiwt</italic> parameter was set to 5000, and <italic>--lb</italic> to [0,10,20,30,…,240] for each of 25 instances, respectively. All other parameters were kept at default values.</p>
<p>SCRMshaw makes use of three underlying statistical models, described briefly in (<xref ref-type="bibr" rid="c17">17</xref>) and in detail in (<xref ref-type="bibr" rid="c13">13</xref>, <xref ref-type="bibr" rid="c14">14</xref>). <italic>HexMCD</italic> trains a fifth-order Markov chain on all 6-mers and their one-away mismatches in the training and background sequence sets, respectively, while <italic>IMM</italic> trains an interpolated Markov model that combines Markov chains of all orders from 0-5. For both of these methods, the score for a sequence in the log-likelihood ratio of the model trained on the training sequences divided by model trained on the background sequences. <italic>PAC-rc</italic> quantifies the overrepresentation of 6-mers in the training versus background sets, assuming a Poisson distribution of word counts. Our previous work has shown that each method is effective, and each tends to have sensitivity for a different group of enhancers. Roughly three-quarters of all predicted enhancers were identified by just a single method (see Supplemental Table S2), although there is a strong correlation between high-ranking predictions and prediction by more than one method. The best results are obtained from combining the output of all three methods.</p>
<p>SCRMshaw can be downloaded at <ext-link ext-link-type="uri" xlink:href="https://github.com/HalfonLab/SCRMshaw_HD">https://github.com/HalfonLab/SCRMshaw_HD</ext-link>.</p>
</sec>
<sec id="s4d">
<title>Postprocessing</title>
<p>The top 5000 hits were extracted using scripts <italic>Generate_top_N_SCRM-hits.pl</italic> and <italic>concatenatenatingOffsetResults.sh</italic> and passed to <italic>postProcessingScrmshawPipeline.py</italic> with <italic>-num</italic> = “5000” and -<italic>topN</italic> = “Median.” Post-processing was performed essentially as previously described (<xref ref-type="bibr" rid="c15">15</xref>, <xref ref-type="bibr" rid="c21">21</xref>), with the following modifications (Supplemental Fig. S6): before evaluating each 10 bp region, scores from each of the 25 individual SCRMshaw instances were assessed and any 500 bp window whose score was below the value of the 5000<sup>th</sup> ranked score was eliminated by having its score reset to zero (Supplemental Fig. S6B). The “elbow” point of the SCRMshaw score curve of the 5000 top scores from each instance was then determined (Fig. S6C), and any scores below the elbow point were reset to zero (Supplemental Fig. S6D; gray boxes). This is a key modification to our previous SCRMshawHD protocol (<xref ref-type="bibr" rid="c21">21</xref>) and reduces the median prediction size by preventing concatenation of multiple adjacent low-scoring windows. Only after these two rounds of score evaluation were all windows grouped together (Fig. S6E, F) and subjected to peak calling on 10 bp intervals (Fig. S6G). Final “top predictions” were then any peaks with an amplitude above the selected amplitude threshold (elbow point of amplitude curve, represented by a red dot in Supplemental Fig. S6H), following the peak-calling step.</p>
<p>All scripts are available at <ext-link ext-link-type="uri" xlink:href="https://github.com/HalfonLab/">https://github.com/HalfonLab/</ext-link>.</p>
</sec>
<sec id="s4e">
<title>Orthology mapping</title>
<p>The final SCRMshaw output from the above steps was used as input to our orthology mapping pipeline. For each species, a FASTA-formatted file of all annotated proteins was downloaded from NCBI (<ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/datasets">www.ncbi.nlm.nih.gov/datasets</ext-link>). The annotated proteins from <italic>D. melanogaster</italic> plus each individual other species were used as input to <italic>Orthologer</italic> (<xref ref-type="bibr" rid="c31">31</xref>) to obtain the <italic>Drosophila</italic> ortholog for each protein (when existing). Our approach was designed to be minimally restrictive in that we did not enforce a one-to-one ortholog mapping; in cases of likely paralogs, we considered all of the paralogs as a potential result. Details on the orthology mapping protocol can be found in (<xref ref-type="bibr" rid="c75">75</xref>).</p>
</sec>
<sec id="s4f">
<title>Evaluating the number of predictions per locus</title>
<p>For each training set, BEDTools “merge” was used to remove any overlapping predictions (<xref ref-type="bibr" rid="c76">76</xref>). The results were then permuted 1000 times using BEDTools “shuffle”, with coding regions excluded. BEDTools “closest” was used to assign upstream and downstream flanking genes to each of the permuted predictions, with the following parameters: the ‘-io’ flag was enabled to ignore any overlaps between predictions and genes, and ‘-D ref-id’ and ‘-D ref-iu’ were used to obtain the closest 5’ and 3’ genes (with respect to chromosome coordinates), respectively. A custom Python script, <italic>checkSameLocus_ForSimulations.py</italic> (see <ext-link ext-link-type="uri" xlink:href="https://github.com/HalfonLab/Asma_etal_2024_eLife">https://github.com/HalfonLab/Asma_etal_2024_eLife</ext-link>) was used to calculate the number of predictions per locus for both the real and permuted results. For this purpose, “locus” was defined as the entire region between the left and right flanking genes for each SCRMshaw prediction; that is, for each prediction we take the region between the 3’ end of the upstream gene and the transcription start site of the downstream gene. If a prediction is within an intron, we define the locus as the entire span of the enclosing gene. If a prediction overlaps two genes, the locus is considered to be the entire span of the two genes. Average intergenic region sizes for the genomes were estimated using the distances between neighboring genes, discarding any nested or overlapping gene pairs, as provided by the SCRMshaw <italic>preflight</italic> script.</p>
<p>Significance was assessed by calculating the empirical <italic>P</italic>-value, defined as the number of real loci with a number of predictions greater than the maximum number obtained from the 1000 simulations, for each locus containing at least one SCRMshaw prediction.</p>
</sec>
<sec id="s4g">
<title>Evaluating the number of predictions in common across species</title>
<p>To determine the expected number of common predictions—i.e., predictions in orthologous loci—across species, we utilized SCRMshaw predictions for 16 species. Each set of predictions was sorted, merged, permuted, and mapped to new loci as described above for “Evaluating the number of predictions per locus.” The permuted predictions were used as input for the script “<italic>checkSameLocus_ForSimulations_crossSpecies.py</italic>” (see <ext-link ext-link-type="uri" xlink:href="https://github.com/HalfonLab/Asma_etal_2024_eLife">https://github.com/HalfonLab/Asma_etal_2024_eLife</ext-link>). This script identifies the <italic>Drosophila</italic> orthologs of the nearest flanking genes for each species and calculates the number of common loci flanking the simulated SCRMshaw predictions for 5, 10, 11, 12, 13, 14, 15, and 16 species. For each species, the process was repeated for a total of 360 permutations. The mean and standard deviation of the permuted results were then used to calculate a <italic>z</italic>-score for each training set. We considered training sets with <italic>z</italic>-score ≥ 1.645 to be significant (<italic>P</italic>&lt;0.05, not corrected for multiple testing).</p>
</sec>
<sec id="s4h">
<title>Overlap between SCRMshaw predictions and open chromatin regions</title>
<p>Open chromatin data were obtained from the following sources:</p>
<sec id="s4h1">
<title>D. melanogaster</title>
<p>Data for <italic>D. melanogaster</italic> were downloaded from GEO and consisted of ATAC-seq and FAIRE-seq data from accessions GSE101827, GSE38727 and GSE118240. These assays were performed using blastoderm embryos, eye-antennal discs, wing discs, haltere discs, leg discs, third instar central nervous system, and wing, leg, and haltere pharate appendages (<xref ref-type="bibr" rid="c38">38</xref>, <xref ref-type="bibr" rid="c39">39</xref>, <xref ref-type="bibr" rid="c77">77</xref>).</p>
</sec>
<sec id="s4h2">
<title>T. castaneum</title>
<p>FAIRE-seq data for three stages of embryogenesis, larval central nervous system, and larval second and third thoracic epidermal tissues were downloaded from GEO (GSE104495)(<xref ref-type="bibr" rid="c18">18</xref>). The FAIRE profiles were remapped to the version 5.2 of the <italic>T. castaneum</italic> genome (Tcas5.2) for this study. The remapped FAIRE profiles are available on iBeetle-Base (<ext-link ext-link-type="uri" xlink:href="https://ibeetle-base.uni-goettingen.de/)">https://ibeetle-base.uni-goettingen.de/)</ext-link>(<xref ref-type="bibr" rid="c78">78</xref>).</p>
</sec>
<sec id="s4h3">
<title>A. gambiae</title>
<p>ATAC-seq data for adult midgut and salivary gland were downloaded from GEO (GSE152924)(<xref ref-type="bibr" rid="c79">79</xref>).</p>
</sec>
<sec id="s4h4">
<title>D. plexippus, J. coenia, H. himera, and V. cardui</title>
<p>ATAC-seq data for larval forewing and hindwing tissues at stage M5 were provided by Anyi Mazo-Vargas and Robert Reed (<xref ref-type="bibr" rid="c40">40</xref>).</p>
<p>The overlap between SCRMshaw predictions and open chromatin regions was determined using BEDTools “intersect” with parameters -wa -u -f 0.1 such that sequences needed to overlap at least 10% of their length and were considered a single overlap in the event that more than one open chromatin peak overlapped a prediction.</p>
<p>To assess significance, the SCRMshaw predictions were permuted 500 times using BEDTools “shuffle” and the overlaps with open chromatin regions assessed as above. The mean and standard deviation of the permuted results were then used to calculate a <italic>z</italic>-score for each training set. We also calculated a “fold enrichment” score by dividing the observed number of overlapping regions by the expected number, to provide a sense of effect size in addition to statistical significance.</p>
</sec>
</sec>
<sec id="s4i">
<title>Reporter constructs and transgenic <italic>Drosophila</italic></title>
<p>Sequences for reporter gene analysis, including attL1 and attL2 sites, were synthesized <italic>de novo</italic> and cloned into pUC57 Kan-r (GenScript, Piscataway NJ) as entry vectors suitable for Gateway cloning (<xref ref-type="bibr" rid="c80">80</xref>). Gateway LR recombination was then used to move the sequences into piggyPhiGUGd (<italic>hth, Ubx,</italic> and <italic>psq</italic> lines) and piggyPhiGUGd-TomatoI (<italic>ex, klu,</italic> and <italic>ush</italic> lines) (<xref ref-type="bibr" rid="c41">41</xref>)(These reporter vectors are available from the <italic>Drosophila</italic> Genomics Resource Center (<ext-link ext-link-type="uri" xlink:href="https://dgrc.bio.indiana.edu">https://dgrc.bio.indiana.edu</ext-link>)). Transgenic flies were generated by BestGene (Chino Hills, CA) using PhiC31 recombination and the attP2 third chromosome insertion site. piggyPhiGUGd lines were subsequently crossed to G-TRACE for visualizing enhancer activities (<xref ref-type="bibr" rid="c42">42</xref>).</p>
<p>For each construct, imaginal discs from at least six larvae were dissected and mounted for direct fluorescence visualization using a Zeiss Axio Imager M2 microscope with ApoTome 2. Embryos were fixed and stained using standard <italic>Drosophila</italic> methods using anti-dsRed (Clontech) and visualized using the ABC-HRP kit (VectorLabs, Newark CA).</p>
</sec>
<sec id="s4j">
<title>Database implementation</title>
<p>The SCRMshaw results database is implemented as part of REDfly (RRID:SCR_006790) (<xref ref-type="bibr" rid="c16">16</xref>), a MariaDB-based database hosted on a private OpenStack cloud infrastructure maintained by the University at Buffalo Center for Computational Research. As part of an ongoing transition of REDfly to a more modern software architecture, backend functions are implemented in Node.JS and Python, while frontend components utilize React.JS and Next.js. GraphQL is used as the query language. REDfly is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives License v4 International (CC BY-NC-ND 4.0) and its underlying source code under a GNU General Public License v3 (GNU GPL 3.0).</p>
</sec>
<sec id="s4k">
<title>Image credits</title>
<p>The insect silhouettes in <xref rid="tbl5" ref-type="table">Table 5</xref> were obtained from The Noun Project (<ext-link ext-link-type="uri" xlink:href="http://thenounproject.com">thenounproject.com</ext-link>), artist Georgiana Ionescu, under a Creative Commons CC-BY 3.0 license.</p>
</sec>
</sec>
<sec id="s5">
<title>Data availability</title>
<p>All SCRMshaw output generated in this study has been deposited at Dryad at <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5061/dryad.3j9kd51t0">https://doi.org/10.5061/dryad.3j9kd51t0</ext-link> in both the standard SCRMshaw output format (modified BED) and in GFFv3 format. Data can also be obtained via the REDfly database at <ext-link ext-link-type="uri" xlink:href="http://redfly.ccr.buffalo.edu">http://redfly.ccr.buffalo.edu</ext-link>, as described above.</p>
</sec>
<sec id="d1e2081" sec-type="supplementary-material">
<title>Supporting information</title>
<supplementary-material id="d1e2225">
<label>Supplemental Figures S1-S6</label>
<media xlink:href="supplements/576926_file09.zip"/>
</supplementary-material>
<supplementary-material id="d1e2232">
<label>Supplemental Tables S1-S5</label>
<media xlink:href="supplements/576926_file10.zip"/>
</supplementary-material>
</sec>
</body>
<back>
<ack>
<title>Acknowledgments</title>
<p>We thank Bob Reed, Anyi Mazo-Vargas, and Tom Williams for sharing data and results, Jack Leatherbarrow for technical support, members of the Halfon and Tomoyasu labs for helpful discussion and advice, and Tom Williams for comments on the manuscript. SCRMshaw analyses were run using the resources of the University at Buffalo Center for Computational Research. The Center for Bioinformatics and Functional Genomics at Miami University provided instrumentation and technical support.</p>
</ack>
<sec id="s6">
<title>Funding</title>
<p>Miami University Faculty Research Grants Program (CFR) (to Y.T.), the National Science Foundation (NSF) (grant IOS1557936 to Y.T.), National Institutes of Health (NIH) (grant U24 GM142435 to M.S.H.), and the U.S. Department of Agriculture (USDA) (grant 2019-67013-29354 to M.S.H. and Y.T.).</p>
</sec>
<ref-list>
<title>References</title>
<ref id="c1"><label>1.</label><mixed-citation publication-type="web"><person-group person-group-type="author"><collab>NCBI. NCBI Datasets: Genome</collab>  [<collab>Available from</collab></person-group><year>2024</year> [: <ext-link ext-link-type="uri" xlink:href="https://www.ncbi.nlm.nih.gov/datasets/genome/?taxon=33208">https://www.ncbi.nlm.nih.gov/datasets/genome/?taxon=33208</ext-link>].</mixed-citation></ref>
<ref id="c2"><label>2.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Carroll</surname> <given-names>SB</given-names></string-name>, <string-name><surname>Grenier</surname> <given-names>JK</given-names></string-name> <string-name><surname>Weatherbee</surname>, <given-names>SD</given-names></string-name></person-group>, <source>From DNA to Diversity. Molecular Genetics and the Evolution of Animal Design</source>. <edition>2nd ed</edition>. <publisher-loc>Malden, MA</publisher-loc>: <publisher-name>Blackwell Publishing</publisher-name>; <year>2005</year>.</mixed-citation></ref>
<ref id="c3"><label>3.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Carroll</surname> <given-names>SB</given-names></string-name></person-group>. <article-title>Evo-devo and an expanding evolutionary synthesis: a genetic theory of morphological evolution</article-title>. <source>Cell</source>. <year>2008</year>;<volume>134</volume>(<issue>1</issue>):<fpage>25</fpage>–<lpage>36</lpage>.</mixed-citation></ref>
<ref id="c4"><label>4.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Claringbould</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zaugg</surname> <given-names>JB</given-names></string-name></person-group>. <article-title>Enhancers in disease: molecular basis and emerging treatment strategies</article-title>. <source>Trends Mol Med</source>. <year>2021</year>;<volume>27</volume>(<issue>11</issue>):<fpage>1060</fpage>–<lpage>73</lpage>.</mixed-citation></ref>
<ref id="c5"><label>5.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rickels</surname> <given-names>R</given-names></string-name>, <string-name><surname>Shilatifard</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Enhancer Logic and Mechanics in Development and Disease</article-title>. <source>Trends Cell Biol</source>. <year>2018</year>;<volume>28</volume>(<issue>8</issue>):<fpage>608</fpage>–<lpage>30</lpage>.</mixed-citation></ref>
<ref id="c6"><label>6.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Smith</surname> <given-names>E</given-names></string-name>, <string-name><surname>Shilatifard</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Enhancer biology and enhanceropathies</article-title>. <source>Nature structural &amp; molecular biology</source>. <year>2014</year>;<volume>21</volume>(<issue>3</issue>):<fpage>210</fpage>–<lpage>9</lpage>.</mixed-citation></ref>
<ref id="c7"><label>7.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Suryamohan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Halfon</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>Identifying transcriptional cis-regulatory modules in animal genomes</article-title>. <source>Wiley Interdisciplinary Reviews: Developmental Biology</source>. <year>2015</year>;<volume>4</volume>(<issue>2</issue>):<fpage>59</fpage>–<lpage>84</lpage>.</mixed-citation></ref>
<ref id="c8"><label>8.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Benton</surname> <given-names>ML</given-names></string-name>, <string-name><surname>Talipineni</surname> <given-names>SC</given-names></string-name>, <string-name><surname>Kostka</surname> <given-names>D</given-names></string-name>, <string-name><surname>Capra</surname> <given-names>JA</given-names></string-name></person-group>. <article-title>Genome-wide enhancer annotations differ significantly in genomic distribution, evolution, and function</article-title>. <source>BMC Genomics</source>. <year>2019</year>;<volume>20</volume>(<issue>1</issue>):<fpage>511</fpage>.</mixed-citation></ref>
<ref id="c9"><label>9.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Halfon</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>Studying Transcriptional Enhancers: The Founder Fallacy, Validation Creep, and Other Biases</article-title>. <source>Trends Genet</source>. <year>2019</year>;<volume>35</volume>(<issue>2</issue>):<fpage>93</fpage>–<lpage>103</lpage>.</mixed-citation></ref>
<ref id="c10"><label>10.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lindhorst</surname> <given-names>D</given-names></string-name>, <string-name><surname>Halfon</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>Reporter gene assays and chromatin-level assays define substantially non-overlapping sets of enhancer sequences</article-title>. <source>BMC Genomics</source>. <year>2023</year>;<volume>24</volume>(<issue>1</issue>):<fpage>17</fpage>.</mixed-citation></ref>
<ref id="c11"><label>11.</label><mixed-citation publication-type="web"><person-group person-group-type="author"><collab>IUCN. The IUCN list of threatened species</collab>  [<collab>Available from</collab></person-group><year>2022</year> [: <ext-link ext-link-type="uri" xlink:href="https://www.iucnredlist.org">https://www.iucnredlist.org</ext-link>.</mixed-citation></ref>
<ref id="c12"><label>12.</label><mixed-citation publication-type="web"><person-group person-group-type="author"><collab>Royal Entomological Society. Understanding Insects: Facts and figures St. Albans, UK</collab>  [<collab>Available from</collab></person-group><year>2023</year> [: <ext-link ext-link-type="uri" xlink:href="https://www.royensoc.co.uk/understanding-insects/facts-and-figures/">https://www.royensoc.co.uk/understanding-insects/facts-and-figures/</ext-link>.</mixed-citation></ref>
<ref id="c13"><label>13.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kantorovitz</surname> <given-names>MR</given-names></string-name>, <string-name><surname>Kazemian</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kinston</surname> <given-names>S</given-names></string-name>, <string-name><surname>Miranda-Saavedra</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Robinson</surname> <given-names>GE</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Motif-blind, genome-wide discovery of cis-regulatory modules in Drosophila and mouse</article-title>. <source>Dev Cell</source>. <year>2009</year>;<volume>17</volume>(<issue>4</issue>):<fpage>568</fpage>–<lpage>79</lpage>.</mixed-citation></ref>
<ref id="c14"><label>14.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kazemian</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Halfon</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Sinha</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Improved accuracy of supervised CRM discovery with interpolated Markov models and cross-species comparison</article-title>. <source>Nucleic Acids Res</source>. <year>2011</year>;<volume>39</volume>(<issue>22</issue>):<fpage>9463</fpage>–<lpage>72</lpage>.</mixed-citation></ref>
<ref id="c15"><label>15.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kazemian</surname> <given-names>M</given-names></string-name>, <string-name><surname>Halfon</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>CRM Discovery Beyond Model Insects</article-title>. <source>Methods Mol Biol</source>. <year>2019</year>;<volume>1858</volume>:<fpage>117</fpage>–<lpage>39</lpage>.</mixed-citation></ref>
<ref id="c16"><label>16.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Keränen</surname> <given-names>SVE</given-names></string-name>, <string-name><surname>Villahoz-Baleta</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bruno</surname> <given-names>AE</given-names></string-name>, <string-name><surname>Halfon</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>REDfly: An Integrated Knowledgebase for Insect Regulatory Genomics</article-title>. <source>Insects</source>. <year>2022</year>;<volume>13</volume>(<issue>7</issue>).</mixed-citation></ref>
<ref id="c17"><label>17.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kazemian</surname> <given-names>M</given-names></string-name>, <string-name><surname>Suryamohan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>JY</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Samee</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Halfon</surname> <given-names>MS</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Evidence for deep regulatory similarities in early developmental programs across highly diverged insects</article-title>. <source>Genome biology and evolution</source>. <year>2014</year>;<volume>6</volume>(<issue>9</issue>):<fpage>2301</fpage>–<lpage>20</lpage>.</mixed-citation></ref>
<ref id="c18"><label>18.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lai</surname> <given-names>YT</given-names></string-name>, <string-name><surname>Deem</surname> <given-names>KD</given-names></string-name>, <string-name><surname>Borras-Castells</surname> <given-names>F</given-names></string-name>, <string-name><surname>Sambrani</surname> <given-names>N</given-names></string-name>, <string-name><surname>Rudolf</surname> <given-names>H</given-names></string-name>, <string-name><surname>Suryamohan</surname> <given-names>K</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Enhancer identification and activity evaluation in the red flour beetle, Tribolium castaneum</article-title>. <source>Development</source>. <year>2018</year>;<volume>145</volume>(<issue>7</issue>).</mixed-citation></ref>
<ref id="c19"><label>19.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Schember</surname> <given-names>I</given-names></string-name>, <string-name><surname>Halfon</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>Identification of new Anopheles gambiae transcriptional enhancers using a cross-species prediction approach</article-title>. <source>Insect molecular biology</source>. <year>2021</year>;<volume>30</volume>(<issue>4</issue>):<fpage>410</fpage>–<lpage>9</lpage>.</mixed-citation></ref>
<ref id="c20"><label>20.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Suryamohan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Hanson</surname> <given-names>C</given-names></string-name>, <string-name><surname>Andrews</surname> <given-names>E</given-names></string-name>, <string-name><surname>Sinha</surname> <given-names>S</given-names></string-name>, <string-name><surname>Scheel</surname> <given-names>MD</given-names></string-name>, <string-name><surname>Halfon</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>Redeployment of a conserved gene regulatory network during Aedes aegypti development</article-title>. <source>Dev Biol</source>. <year>2016</year>;<volume>416</volume>(<issue>2</issue>):<fpage>402</fpage>–<lpage>13</lpage>.</mixed-citation></ref>
<ref id="c21"><label>21.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Asma</surname> <given-names>H</given-names></string-name>, <string-name><surname>Halfon</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>Computational enhancer prediction: evaluation and improvements</article-title>. <source>BMC bioinformatics</source>. <year>2019</year>;<volume>20</volume>(<issue>1</issue>):<fpage>174</fpage>.</mixed-citation></ref>
<ref id="c22"><label>22.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Benson</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Tandem repeats finder: a program to analyze DNA sequences</article-title>. <source>Nucleic Acids Res</source>. <year>1999</year>;<volume>27</volume>(<issue>2</issue>):<fpage>573</fpage>–<lpage>80</lpage>.</mixed-citation></ref>
<ref id="c23"><label>23.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>He</surname> <given-names>X</given-names></string-name>, <string-name><surname>Sinha</surname> <given-names>S</given-names></string-name>, <string-name><surname>Halfon</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Large-scale analysis of transcriptional cis-regulatory modules reveals both common features and distinct subclasses</article-title>. <source>Genome Biology</source>. <year>2007</year>;<volume>8</volume>(<issue>6</issue>):<fpage>R101</fpage>.</mixed-citation></ref>
<ref id="c24"><label>24.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sanyal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lajoie</surname> <given-names>BR</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>G</given-names></string-name>, <string-name><surname>Dekker</surname> <given-names>J</given-names></string-name></person-group>. <article-title>The long-range interaction landscape of gene promoters</article-title>. <source>Nature</source>. <year>2012</year>;<volume>489</volume>(<issue>7414</issue>):<fpage>109</fpage>–<lpage>13</lpage>.</mixed-citation></ref>
<ref id="c25"><label>25.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hafez</surname> <given-names>D</given-names></string-name>, <string-name><surname>Karabacak</surname> <given-names>A</given-names></string-name>, <string-name><surname>Krueger</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hwang</surname> <given-names>YC</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>LS</given-names></string-name>, <string-name><surname>Zinzen</surname> <given-names>RP</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>McEnhancer: predicting gene expression via semi-supervised assignment of enhancers to target genes</article-title>. <source>Genome Biol</source>. <year>2017</year>;<volume>18</volume>(<issue>1</issue>):<fpage>199</fpage>.</mixed-citation></ref>
<ref id="c26"><label>26.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chua</surname> <given-names>EHZ</given-names></string-name>, <string-name><surname>Yasar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Harmston</surname> <given-names>N</given-names></string-name></person-group>, <article-title>The importance of considering regulatory domains in genome-wide analyses - the nearest gene is often wrong!</article-title>. <source>Biol Open</source> <year>2022</year>;<volume>11</volume>(<issue>4</issue>).</mixed-citation></ref>
<ref id="c27"><label>27.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Qin</surname> <given-names>T</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Cavalcante</surname> <given-names>RG</given-names></string-name>, <string-name><surname>Orchard</surname> <given-names>P</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>H</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Comprehensive enhancer-target gene assignments improve gene set level interpretation of genome-wide regulatory data</article-title>. <source>Genome Biol</source>. <year>2022</year>;<volume>23</volume>(<issue>1</issue>):<fpage>105</fpage>.</mixed-citation></ref>
<ref id="c28"><label>28.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fishilevich</surname> <given-names>S</given-names></string-name>, <string-name><surname>Nudel</surname> <given-names>R</given-names></string-name>, <string-name><surname>Rappaport</surname> <given-names>N</given-names></string-name>, <string-name><surname>Hadar</surname> <given-names>R</given-names></string-name>, <string-name><surname>Plaschkes</surname> <given-names>I</given-names></string-name>, <string-name><surname>Iny Stein</surname> <given-names>T</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>GeneHancer: genome-wide integration of enhancers and target genes in GeneCards</article-title>. <source>Database: the journal of biological databases and curation</source>. <year>2017</year>;<fpage>2017</fpage>.</mixed-citation></ref>
<ref id="c29"><label>29.</label><mixed-citation publication-type="preprint"><person-group person-group-type="author"><string-name><surname>Gschwind</surname> <given-names>AR</given-names></string-name>, <string-name><surname>Mualim</surname> <given-names>KS</given-names></string-name>, <string-name><surname>Karbalayghareh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sheth</surname> <given-names>MU</given-names></string-name>, <string-name><surname>Dey</surname> <given-names>KK</given-names></string-name>, <string-name><surname>Jagoda</surname> <given-names>E</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>An encyclopedia of enhancer-gene regulatory interactions in the human genome</article-title>. <source>bioRxiv</source>. <year>2023</year>.</mixed-citation></ref>
<ref id="c30"><label>30.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Whalen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Truty</surname> <given-names>RM</given-names></string-name>, <string-name><surname>Pollard</surname> <given-names>KS</given-names></string-name></person-group>. <article-title>Enhancer-promoter interactions are encoded by complex genomic signatures on looping chromatin</article-title>. <source>Nat Genet</source>. <year>2016</year>;<volume>48</volume>(<issue>5</issue>):<fpage>488</fpage>–<lpage>96</lpage>.</mixed-citation></ref>
<ref id="c31"><label>31.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kuznetsov</surname> <given-names>D</given-names></string-name>, <string-name><surname>Tegenfeldt</surname> <given-names>F</given-names></string-name>, <string-name><surname>Manni</surname> <given-names>M</given-names></string-name>, <string-name><surname>Seppey</surname> <given-names>M</given-names></string-name>, <string-name><surname>Berkeley</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kriventseva</surname> <given-names>EV</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>OrthoDB v11: annotation of orthologs in the widest sampling of organismal diversity</article-title>. <source>Nucleic Acids Res</source>. <year>2023</year>;<volume>51</volume>(<issue>D1</issue>):<fpage>D445</fpage>–<lpage>D51</lpage>.</mixed-citation></ref>
<ref id="c32"><label>32.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Grosveld</surname> <given-names>F</given-names></string-name>, <string-name><surname>van Staalduinen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Stadhouders</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Transcriptional Regulation by (Super)Enhancers: From Discovery to Mechanisms</article-title>. <source>Annu Rev Genomics Hum Genet</source>. <year>2021</year>;<volume>22</volume>:<fpage>127</fpage>–<lpage>46</lpage>.</mixed-citation></ref>
<ref id="c33"><label>33.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Weinstein</surname> <given-names>ML</given-names></string-name>, <string-name><surname>Jaenke</surname> <given-names>CM</given-names></string-name>, <string-name><surname>Asma</surname> <given-names>H</given-names></string-name>, <string-name><surname>Spangler</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kohnen</surname> <given-names>KA</given-names></string-name>, <string-name><surname>Konys</surname> <given-names>CC</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>A novel role for trithorax in the gene regulatory network for a rapidly evolving fruit fly pigmentation trait</article-title>. <source>PLoS Genet</source>. <year>2023</year>;<volume>19</volume>(<issue>2</issue>):<fpage>e1010653</fpage>.</mixed-citation></ref>
<ref id="c34"><label>34.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kvon</surname> <given-names>EZ</given-names></string-name>, <string-name><surname>Waymack</surname> <given-names>R</given-names></string-name>, <string-name><surname>Gad</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wunderlich</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Enhancer redundancy in development and disease</article-title>. <source>Nat Rev Genet</source>. <year>2021</year>;<volume>22</volume>(<issue>5</issue>):<fpage>324</fpage>–<lpage>36</lpage>.</mixed-citation></ref>
<ref id="c35"><label>35.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Boyle</surname> <given-names>AP</given-names></string-name>, <string-name><surname>Davis</surname> <given-names>S</given-names></string-name>, <string-name><surname>Shulha</surname> <given-names>HP</given-names></string-name>, <string-name><surname>Meltzer</surname> <given-names>P</given-names></string-name>, <string-name><surname>Margulies</surname> <given-names>EH</given-names></string-name>, <string-name><surname>Weng</surname> <given-names>Z</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>High-resolution mapping and characterization of open chromatin across the genome</article-title>. <source>Cell</source>. <year>2008</year>;<volume>132</volume>(<issue>2</issue>):<fpage>311</fpage>–<lpage>22</lpage>.</mixed-citation></ref>
<ref id="c36"><label>36.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Buenrostro</surname> <given-names>JD</given-names></string-name>, <string-name><surname>Giresi</surname> <given-names>PG</given-names></string-name>, <string-name><surname>Zaba</surname> <given-names>LC</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>HY</given-names></string-name>, <string-name><surname>Greenleaf</surname> <given-names>WJ</given-names></string-name></person-group>. <article-title>Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position</article-title>. <source>Nature methods</source>. <year>2013</year>;<volume>10</volume>(<issue>12</issue>):<fpage>1213</fpage>–<lpage>8</lpage>.</mixed-citation></ref>
<ref id="c37"><label>37.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Giresi</surname> <given-names>PG</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>J</given-names></string-name>, <string-name><surname>McDaniell</surname> <given-names>RM</given-names></string-name>, <string-name><surname>Iyer</surname> <given-names>VR</given-names></string-name>, <string-name><surname>Lieb</surname> <given-names>JD</given-names></string-name></person-group>. <article-title>FAIRE (Formaldehyde-Assisted Isolation of Regulatory Elements) isolates active regulatory elements from human chromatin</article-title>. <source>Genome Res</source>. <year>2007</year>;<volume>17</volume>(<issue>6</issue>):<fpage>877</fpage>–<lpage>85</lpage>.</mixed-citation></ref>
<ref id="c38"><label>38.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bozek</surname> <given-names>M</given-names></string-name>, <string-name><surname>Cortini</surname> <given-names>R</given-names></string-name>, <string-name><surname>Storti</surname> <given-names>AE</given-names></string-name>, <string-name><surname>Unnerstall</surname> <given-names>U</given-names></string-name>, <string-name><surname>Gaul</surname> <given-names>U</given-names></string-name>, <string-name><surname>Gompel</surname> <given-names>N</given-names></string-name></person-group>. <article-title>ATAC-seq reveals regional differences in enhancer accessibility during the establishment of spatial coordinates in the Drosophila blastoderm</article-title>. <source>Genome Research</source>. <year>2019</year>;<volume>29</volume>(<issue>5</issue>):<fpage>771</fpage>–<lpage>83</lpage>.</mixed-citation></ref>
<ref id="c39"><label>39.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>McKay</surname> <given-names>DJ</given-names></string-name>, <string-name><surname>Lieb</surname> <given-names>JD</given-names></string-name></person-group>. <article-title>A Common Set of DNA Regulatory Elements Shapes Drosophila Appendages</article-title>. <source>Developmental Cell</source>. <year>2013</year>;<volume>27</volume>(<issue>3</issue>):<fpage>306</fpage>–<lpage>18</lpage>.</mixed-citation></ref>
<ref id="c40"><label>40.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mazo-Vargas</surname> <given-names>A</given-names></string-name>, <string-name><surname>Langmuller</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Wilder</surname> <given-names>A</given-names></string-name>, <string-name><surname>van der Burg</surname> <given-names>KRL</given-names></string-name>, <string-name><surname>Lewis</surname> <given-names>JJ</given-names></string-name>, <string-name><surname>Messer</surname> <given-names>PW</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Deep cis-regulatory homology of the butterfly wing pattern ground plan</article-title>. <source>Science</source>. <year>2022</year>;<volume>378</volume>(<issue>6617</issue>):<fpage>304</fpage>–<lpage>8</lpage>.</mixed-citation></ref>
<ref id="c41"><label>41.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Deem</surname> <given-names>KD</given-names></string-name>, <string-name><surname>Halfon</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Tomoyasu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A new suite of reporter vectors and a novel landing site survey system to study cis-regulatory elements in diverse insect species</article-title>. <source>Scientific reports</source>. <year>2024</year>;<volume>14</volume>(<issue>1</issue>):<fpage>10078</fpage>.</mixed-citation></ref>
<ref id="c42"><label>42.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Evans</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Olson</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Ngo</surname> <given-names>KT</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>E</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>NE</given-names></string-name>, <string-name><surname>Kuoy</surname> <given-names>E</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>G-TRACE: rapid Gal4-based cell lineage analysis in Drosophila</article-title>. <source>Nature methods</source>. <year>2009</year>;<volume>6</volume>(<issue>8</issue>):<fpage>603</fpage>–<lpage>5</lpage>.</mixed-citation></ref>
<ref id="c43"><label>43.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Boedigheimer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Laughon</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Expanded: a gene involved in the control of cell proliferation in imaginal discs</article-title>. <source>Development</source>. <year>1993</year>;<volume>118</volume>(<issue>4</issue>):<fpage>1291</fpage>–<lpage>301</lpage>.</mixed-citation></ref>
<ref id="c44"><label>44.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>LH</given-names></string-name>, <string-name><surname>Baker</surname> <given-names>NE</given-names></string-name></person-group>. <article-title>Salvador-Warts-Hippo pathway in a developmental checkpoint monitoring helix-loop-helix proteins</article-title>. <source>Dev Cell</source>. <year>2015</year>;<volume>32</volume>(<issue>2</issue>):<fpage>191</fpage>–<lpage>202</lpage>.</mixed-citation></ref>
<ref id="c45"><label>45.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>LH</given-names></string-name>, <string-name><surname>Baker</surname> <given-names>NE</given-names></string-name></person-group>. <article-title>Spatial regulation of expanded transcription in the Drosophila wing imaginal disc</article-title>. <source>PLoS One</source>. <year>2018</year>;<volume>13</volume>(<issue>7</issue>):<fpage>e0201317</fpage>.</mixed-citation></ref>
<ref id="c46"><label>46.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Klein</surname> <given-names>T</given-names></string-name> <string-name><surname>Campos-Ortega</surname><given-names>JA</given-names></string-name></person-group>, <article-title>klumpfuss, a Drosophila gene encoding a member of the EGR family of transcription factors, is involved in bristle and leg development</article-title>. <source>Development</source>. <year>1997</year>;<volume>124</volume>(<issue>16</issue>):<fpage>3123</fpage>–<lpage>34</lpage>.</mixed-citation></ref>
<ref id="c47"><label>47.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Buchberger</surname> <given-names>E</given-names></string-name>, <string-name><surname>Bilen</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ayaz</surname> <given-names>S</given-names></string-name>, <string-name><surname>Salamanca</surname> <given-names>D</given-names></string-name>, <string-name><surname>Matas de Las Heras</surname><given-names>C</given-names></string-name> <string-name><surname>Niksic</surname><given-names>A</given-names></string-name>, <etal>et al.</etal></person-group>, <article-title>Variation in Pleiotropic Hub Gene Expression Is Associated with Interspecific Differences in Head Shape and Eye Size in Drosophila</article-title>. <source>Mol Biol Evol</source> <year>2021</year>;<volume>38</volume>(<issue>5</issue>):<fpage>1924</fpage>–<lpage>42</lpage>.</mixed-citation></ref>
<ref id="c48"><label>48.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cubadda</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Heitzler</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ray</surname> <given-names>RP</given-names></string-name>, <string-name><surname>Bourouis</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ramain</surname> <given-names>P</given-names></string-name>, <string-name><surname>Gelbart</surname> <given-names>W</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>u-shaped encodes a zinc finger protein that regulates the proneural genes achaete and scute during the formation of bristles in Drosophila</article-title>. <source>Genes Dev</source>. <year>1997</year>;<volume>11</volume>(<issue>22</issue>):<fpage>3083</fpage>–<lpage>95</lpage>.</mixed-citation></ref>
<ref id="c49"><label>49.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tomoyasu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ueno</surname> <given-names>N</given-names></string-name>, <string-name><surname>Nakamura</surname> <given-names>M</given-names></string-name></person-group>. <article-title>The decapentaplegic morphogen gradient regulates the notal wingless expression through induction of pannier and u-shaped in Drosophila</article-title>. <source>Mech Dev</source>. <year>2000</year>;<volume>96</volume>(<issue>1</issue>):<fpage>37</fpage>–<lpage>49</lpage>.</mixed-citation></ref>
<ref id="c50"><label>50.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jory</surname> <given-names>A</given-names></string-name>, <string-name><surname>Estella</surname> <given-names>C</given-names></string-name>, <string-name><surname>Giorgianni</surname> <given-names>MW</given-names></string-name>, <string-name><surname>Slattery</surname> <given-names>M</given-names></string-name>, <string-name><surname>Laverty</surname> <given-names>TR</given-names></string-name>, <string-name><surname>Rubin</surname> <given-names>GM</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>A survey of 6,300 genomic fragments for cis-regulatory activity in the imaginal discs of Drosophila melanogaster</article-title>. <source>Cell reports</source>. <year>2012</year>;<volume>2</volume>(<issue>4</issue>):<fpage>1014</fpage>–<lpage>24</lpage>.</mixed-citation></ref>
<ref id="c51"><label>51.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Aldaz</surname> <given-names>S</given-names></string-name>, <string-name><surname>Morata</surname> <given-names>G</given-names></string-name>, <string-name><surname>Azpiazu</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Patterning function of homothorax/extradenticle in the thorax of Drosophila</article-title>. <source>Development</source>. <year>2005</year>;<volume>132</volume>(<issue>3</issue>):<fpage>439</fpage>–<lpage>46</lpage>.</mixed-citation></ref>
<ref id="c52"><label>52.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lewis</surname> <given-names>EB</given-names></string-name></person-group>. <article-title>A gene complex controlling segmentation in Drosophila</article-title>. <source>Nature</source>. <year>1978</year>;<volume>276</volume>(<issue>5688</issue>):<fpage>565</fpage>–<lpage>70</lpage>.</mixed-citation></ref>
<ref id="c53"><label>53.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Simon</surname> <given-names>J</given-names></string-name>, <string-name><surname>Peifer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bender</surname> <given-names>W</given-names></string-name>, <string-name><surname>O’Connor</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Regulatory elements of the bithorax complex that control expression along the anterior-posterior axis</article-title>. <source>EMBO J</source>. <year>1990</year>;<volume>9</volume>(<issue>12</issue>):<fpage>3945</fpage>–<lpage>56</lpage>.</mixed-citation></ref>
<ref id="c54"><label>54.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Prasad</surname> <given-names>N</given-names></string-name>, <string-name><surname>Tarikere</surname> <given-names>S</given-names></string-name>, <string-name><surname>Khanale</surname> <given-names>D</given-names></string-name>, <string-name><surname>Habib</surname> <given-names>F</given-names></string-name>, <string-name><surname>Shashidhara</surname> <given-names>LS</given-names></string-name></person-group>. <article-title>A comparative genomic analysis of targets of Hox protein Ultrabithorax amongst distant insect species</article-title>. <source>Scientific reports</source>. <year>2016</year>;<volume>6</volume>:<issue>27885</issue>.</mixed-citation></ref>
<ref id="c55"><label>55.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>DH</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>YL</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>CC</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>IC</given-names></string-name>, <string-name><surname>King</surname> <given-names>B</given-names></string-name></person-group>, <article-title>pipsqueak encodes a factor essential for sequence-specific targeting of a polycomb group protein complex</article-title>. <source>Mol Cell Biol</source>. <year>2002</year>;<volume>22</volume>(<issue>17</issue>):<fpage>6261</fpage>–<lpage>71</lpage>.</mixed-citation></ref>
<ref id="c56"><label>56.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Cohen</surname> <given-names>SM.</given-names></string-name> . In: <string-name><surname>Bate</surname> <given-names>M</given-names></string-name>, <string-name><surname>Martinez Arias</surname> <given-names>A</given-names></string-name></person-group><chapter-title>Imaginal disc development</chapter-title>. In: , , editors. <source>The development of Drosophila melanogaster</source>. <publisher-loc>Cold Spring Harbor, NY</publisher-loc>: <publisher-name>Cold Spring Harbor Laboratory Press</publisher-name>; <year>1993</year>.</mixed-citation></ref>
<ref id="c57"><label>57.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Svacha</surname> <given-names>P</given-names></string-name></person-group>. <article-title>What are and what are not imaginal discs: reevaluation of some basic concepts (Insecta Holometabola)</article-title>, <source>Dev Biol</source>. <year>1992</year>;<volume>154</volume>(<issue>1</issue>):<fpage>101</fpage>–<lpage>17</lpage>.</mixed-citation></ref>
<ref id="c58"><label>58.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Halfon</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>Silencers, Enhancers, and the Multifunctional Regulatory Genome</article-title>. <source>Trends Genet</source>. <year>2020</year>;<volume>36</volume>(<issue>3</issue>):<fpage>149</fpage>–<lpage>51</lpage>.</mixed-citation></ref>
<ref id="c59"><label>59.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Segert</surname> <given-names>JA</given-names></string-name>, <string-name><surname>Gisselbrecht</surname> <given-names>SS</given-names></string-name>, <string-name><surname>Bulyk</surname> <given-names>ML</given-names></string-name></person-group>. <article-title>Transcriptional Silencers: Driving Gene Expression with the Brakes On</article-title>. <source>Trends Genet</source>. <year>2021</year>;<volume>37</volume>(<issue>6</issue>):<fpage>514</fpage>–<lpage>27</lpage>.</mixed-citation></ref>
<ref id="c60"><label>60.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Laiker</surname> <given-names>I</given-names></string-name>, <string-name><surname>Frankel</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Pleiotropic Enhancers are Ubiquitous Regulatory Elements in the Human Genome</article-title>. <source>Genome biology and evolution</source>. <year>2022</year>;<volume>14</volume>(<issue>6</issue>).</mixed-citation></ref>
<ref id="c61"><label>61.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sabaris</surname> <given-names>G</given-names></string-name>, <string-name><surname>Laiker</surname> <given-names>I</given-names></string-name>, <string-name><surname>Preger-Ben Noon</surname> <given-names>E</given-names></string-name>, <string-name><surname>Frankel</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Actors with Multiple Roles: Pleiotropic Enhancers and the Paradigm of Enhancer Modularity</article-title>. <source>Trends Genet</source>. <year>2019</year>;<volume>35</volume>(<issue>6</issue>):<fpage>423</fpage>–<lpage>33</lpage>.</mixed-citation></ref>
<ref id="c62"><label>62.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Frankel</surname> <given-names>N</given-names></string-name>, <string-name><surname>Davis</surname> <given-names>GK</given-names></string-name>, <string-name><surname>Vargas</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Payre</surname> <given-names>F</given-names></string-name>, <string-name><surname>Stern</surname> <given-names>DL</given-names></string-name></person-group>. <article-title>Phenotypic robustness conferred by apparently redundant transcriptional enhancers</article-title>. <source>Nature</source>. <year>2010</year>;<volume>466</volume>(<issue>7305</issue>):<fpage>490</fpage>–<lpage>3</lpage>.</mixed-citation></ref>
<ref id="c63"><label>63.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Perry</surname> <given-names>MW</given-names></string-name>, <string-name><surname>Boettiger</surname> <given-names>AN</given-names></string-name>, <string-name><surname>Bothma</surname> <given-names>JP</given-names></string-name>, <string-name><surname>Levine</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Shadow enhancers foster robustness of Drosophila gastrulation</article-title>. <source>Curr Biol</source>. <year>2010</year>;<volume>20</volume>(<issue>17</issue>):<fpage>1562</fpage>–<lpage>7</lpage>.</mixed-citation></ref>
<ref id="c64"><label>64.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Perry</surname> <given-names>MW</given-names></string-name>, <string-name><surname>Boettiger</surname> <given-names>AN</given-names></string-name>, <string-name><surname>Levine</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Multiple enhancers ensure precision of gap gene-expression patterns in the Drosophila embryo</article-title>. <source>Proc Natl Acad Sci U S A</source>. <year>2011</year>;<volume>108</volume>(<issue>33</issue>):<fpage>13570</fpage>–<lpage>5</lpage>.</mixed-citation></ref>
<ref id="c65"><label>65.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Osterwalder</surname> <given-names>M</given-names></string-name>, <string-name><surname>Barozzi</surname> <given-names>I</given-names></string-name>, <string-name><surname>Tissieres</surname> <given-names>V</given-names></string-name>, <string-name><surname>Fukuda-Yuzawa</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Mannion</surname> <given-names>BJ</given-names></string-name>, <string-name><surname>Afzal</surname> <given-names>SY</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Enhancer redundancy provides phenotypic robustness in mammalian development</article-title>. <source>Nature</source>. <year>2018</year>;<volume>554</volume>(<issue>7691</issue>):<fpage>239</fpage>–<lpage>43</lpage>.</mixed-citation></ref>
<ref id="c66"><label>66.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Antosova</surname> <given-names>B</given-names></string-name>, <string-name><surname>Smolikova</surname> <given-names>J</given-names></string-name>, <string-name><surname>Klimova</surname> <given-names>L</given-names></string-name>, <string-name><surname>Lachova</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bendova</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kozmikova</surname> <given-names>I</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>The Gene Regulatory Network of Lens Induction Is Wired through Meis-Dependent Shadow Enhancers of Pax6</article-title>. <source>PLoS Genet</source>. <year>2016</year>;<volume>12</volume>(<issue>12</issue>):<fpage>e1006441</fpage>.</mixed-citation></ref>
<ref id="c67"><label>67.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sagai</surname> <given-names>T</given-names></string-name>, <string-name><surname>Amano</surname> <given-names>T</given-names></string-name>, <string-name><surname>Maeno</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kiyonari</surname> <given-names>H</given-names></string-name>, <string-name><surname>Seo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>SW</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>SHH signaling directed by two oral epithelium-specific enhancers controls tooth and oral development</article-title>. <source>Scientific reports</source>. <year>2017</year>;<volume>7</volume>(<issue>1</issue>):<fpage>13004</fpage>.</mixed-citation></ref>
<ref id="c68"><label>68.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Waymack</surname> <given-names>R</given-names></string-name>, <string-name><surname>Fletcher</surname> <given-names>A</given-names></string-name>, <string-name><surname>Enciso</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wunderlich</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Shadow enhancers can suppress input transcription factor noise through distinct regulatory logic</article-title>. <source>eLife</source>. <year>2020</year>;<volume>9</volume>.</mixed-citation></ref>
<ref id="c69"><label>69.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cannavo</surname> <given-names>E</given-names></string-name>, <string-name><surname>Khoueiry</surname> <given-names>P</given-names></string-name>, <string-name><surname>Garfield</surname> <given-names>DA</given-names></string-name>, <string-name><surname>Geeleher</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zichner</surname> <given-names>T</given-names></string-name>, <string-name><surname>Gustafson</surname> <given-names>EH</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Shadow Enhancers Are Pervasive Features of Developmental Regulatory Networks</article-title>. <source>Curr Biol</source>. <year>2016</year>;<volume>26</volume>(<issue>1</issue>):<fpage>38</fpage>–<lpage>51</lpage>.</mixed-citation></ref>
<ref id="c70"><label>70.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Barth</surname> <given-names>NKH</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Taher</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Independent Transposon Exaptation Is a Widespread Mechanism of Redundant Enhancer Evolution in the Mammalian Genome</article-title>. <source>Genome biology and evolution</source>. <year>2020</year>;<volume>12</volume>(<issue>3</issue>):<fpage>1</fpage>–<lpage>17</lpage>.</mixed-citation></ref>
<ref id="c71"><label>71.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Weisman</surname> <given-names>CM</given-names></string-name>, <string-name><surname>Murray</surname> <given-names>AW</given-names></string-name>, <string-name><surname>Eddy</surname> <given-names>SR</given-names></string-name></person-group>. <article-title>Mixing genome annotation methods in a comparative analysis inflates the apparent number of lineage-specific genes</article-title>. <source>Curr Biol</source>. <year>2022</year>;<volume>32</volume>(<issue>12</issue>):<fpage>2632</fpage>–<lpage>9 e2</lpage>.</mixed-citation></ref>
<ref id="c72"><label>72.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Crosby</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Gramates</surname> <given-names>LS</given-names></string-name>, <string-name><surname>Dos Santos</surname> <given-names>G</given-names></string-name>, <string-name><surname>Matthews</surname> <given-names>BB</given-names></string-name>, <string-name><surname>St Pierre</surname> <given-names>SE</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>P</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Gene Model Annotations for Drosophila melanogaster: The Rule-Benders</article-title>. <source>G3</source>. <year>2015</year>;<volume>5</volume>(<issue>8</issue>):<fpage>1737</fpage>-<lpage>49</lpage>.</mixed-citation></ref>
<ref id="c73"><label>73.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Matthews</surname> <given-names>BB</given-names></string-name>, <string-name><surname>Dos Santos</surname> <given-names>G</given-names></string-name>, <string-name><surname>Crosby</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Emmert</surname> <given-names>DB</given-names></string-name>, <string-name><surname>St Pierre</surname> <given-names>SE</given-names></string-name>, <string-name><surname>Gramates</surname> <given-names>LS</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Gene Model Annotations for Drosophila melanogaster: Impact of High-Throughput Data</article-title>. <source>G3</source>. <year>2015</year>;<volume>5</volume>(<issue>8</issue>):<fpage>1721</fpage>-<lpage>36</lpage>.</mixed-citation></ref>
<ref id="c74"><label>74.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gramates</surname> <given-names>LS</given-names></string-name>, <string-name><surname>Agapite</surname> <given-names>J</given-names></string-name>, <string-name><surname>Attrill</surname> <given-names>H</given-names></string-name>, <string-name><surname>Calvi</surname> <given-names>BR</given-names></string-name>, <string-name><surname>Crosby</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Dos Santos</surname> <given-names>G</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>FlyBase: a guided tour of highlighted features</article-title>. <source>Genetics</source>. <year>2022</year>;<volume>220</volume>(<issue>4</issue>).</mixed-citation></ref>
<ref id="c75"><label>75.</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Asma</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name> <string-name><surname>Halfon</surname><given-names>MS</given-names></string-name></person-group>. <article-title>SCRMshaw: supervised cis-regulatory module prediction for insect genomes</article-title>. <source>protocols.io</source> <year>2024</year>. <pub-id pub-id-type="doi">10.17504/protocols.io.e6nvw1129lmk/v2</pub-id>.</mixed-citation></ref>
<ref id="c76"><label>76.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Quinlan</surname> <given-names>AR</given-names></string-name>, <string-name><surname>Hall</surname> <given-names>IM</given-names></string-name></person-group>. <article-title>BEDTools: a flexible suite of utilities for comparing genomic features</article-title>. <source>Bioinformatics</source>. <year>2010</year>;<volume>26</volume>(<issue>6</issue>):<fpage>841</fpage>–<lpage>2</lpage>.</mixed-citation></ref>
<ref id="c77"><label>77.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jacobs</surname> <given-names>J</given-names></string-name>, <string-name><surname>Atkins</surname> <given-names>M</given-names></string-name>, <string-name><surname>Davie</surname> <given-names>K</given-names></string-name>, <string-name><surname>Imrichova</surname> <given-names>H</given-names></string-name>, <string-name><surname>Romanelli</surname> <given-names>L</given-names></string-name>, <string-name><surname>Christiaens</surname> <given-names>V</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>The transcription factor Grainy head primes epithelial enhancers for spatiotemporal activation by displacing nucleosomes</article-title>. <source>Nat Genet</source>. <year>2018</year>;<volume>50</volume>(<issue>7</issue>):<fpage>1011</fpage>–<lpage>20</lpage>.</mixed-citation></ref>
<ref id="c78"><label>78.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Donitz</surname> <given-names>J</given-names></string-name>, <string-name><surname>Gerischer</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hahnke</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pfeiffer</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bucher</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Expanded and updated data and a query pipeline for iBeetle-Base</article-title>. <source>Nucleic Acids Res</source>. <year>2018</year>;<volume>46</volume>(<issue>D1</issue>):<fpage>D831</fpage>–<lpage>D5</lpage>.</mixed-citation></ref>
<ref id="c79"><label>79.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ruiz</surname> <given-names>JL</given-names></string-name>, <string-name><surname>Ranford-Cartwright</surname> <given-names>LC</given-names></string-name>, <string-name><surname>Gomez-Diaz</surname> <given-names>E</given-names></string-name></person-group>. <article-title>The regulatory genome of the malaria vector Anopheles gambiae: integrating chromatin accessibility and gene expression</article-title>. <source>NAR Genom Bioinform</source>. <year>2021</year>;<volume>3</volume>(<issue>1</issue>):<elocation-id>lqaa113</elocation-id>.</mixed-citation></ref>
<ref id="c80"><label>80.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Katzen</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Gateway((R)) recombinational cloning: a biological operating system</article-title>. <source>Expert Opin Drug Discov</source>. <year>2007</year>;<volume>2</volume>(<issue>4</issue>):<fpage>571</fpage>–<lpage>89</lpage>.</mixed-citation></ref>
<ref id="c81"><label>81.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Asma</surname> <given-names>H</given-names></string-name>, <string-name><surname>Halfon</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>Annotating the Insect Regulatory Genome</article-title>. <source>Insects</source>. <year>2021</year>;<volume>12</volume>(<issue>7</issue>):<fpage>591</fpage>.</mixed-citation></ref>
</ref-list>
</back>
<sub-article id="sa0" article-type="editor-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.96738.2.sa3</article-id>
<title-group>
<article-title>eLife Assessment</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Dalal</surname>
<given-names>Yamini</given-names>
</name>
<role specific-use="editor">Reviewing Editor</role>
<aff>
<institution-wrap>
<institution>National Cancer Institute</institution>
</institution-wrap>
<city>Bethesda</city>
<country>United States of America</country>
</aff>
</contrib>
</contrib-group>
<kwd-group kwd-group-type="evidence-strength">
<kwd>Convincing</kwd>
</kwd-group>
<kwd-group kwd-group-type="claim-importance">
<kwd>Important</kwd>
</kwd-group>
</front-stub>
<body>
<p>In the revised version of this <bold>important</bold> study, the authors present a <bold>convincing</bold> pipeline for insect genome regulatory annotation across 33 insect genomes spanning 5 orders. Despite technical limitations in the field owing to the lack of comprehensive knowledge of enhancer content in any system, the authors employ several independent downstream analyses to support the validity of their enhancer predictions for a subset of these genomes. Taken together, the revised results suggest that this prediction pipeline may have uses in identifying functional enhancers across large phylogenetic distances. Reviewers note caveats that an experimental validation is not yet available in the field to validate a large class of newly identified enhancers across such evolutionary distances, and other pipelines might be of use to compare. This work will be of interest to the computational genomics, evolutionary biology, and gene regulation fields.</p>
</body>
</sub-article>
<sub-article id="sa1" article-type="referee-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.96738.2.sa2</article-id>
<title-group>
<article-title>Reviewer #1 (Public review):</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<anonymous/>
<role specific-use="referee">Reviewer</role>
</contrib>
</contrib-group>
</front-stub>
<body>
<p>Summary:</p>
<p>The authors provide an genome annotation resource of 33 insects using a motif-blind prediction methods for tissue-specific cis-regulatory modules. This is a welcome addition that may facilitate further research in new laboratory systems, and the approach seem to be relatively accurate, although it should be combined with other sources of evidence to be practical.</p>
<p>Strengths:</p>
<p>The paper clearly presents the resource, including the testing of candidate enhancers identified from various insects in Drosophila. This cross-species analysis, and the inherent suggestion that training datasets generated in flies can predict a cis-regulatory activity in distant insects, is interesting. While I can not be sure this approach will prevail in the future, for example with approaches that leverage the prediction of TF binding motifs, the SCRMShaw tool is certainly useful and worth of consideration for the large community of genome scientists working on insects.</p>
<p>Weaknesses from the previous version were appropriately corrected in this revision, as the authors improved data availability including with genome annotation resources.</p>
</body>
</sub-article>
<sub-article id="sa2" article-type="referee-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.96738.2.sa1</article-id>
<title-group>
<article-title>Reviewer #3 (Public review):</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<anonymous/>
<role specific-use="referee">Reviewer</role>
</contrib>
</contrib-group>
</front-stub>
<body>
<p>Summary:</p>
<p>In this ambitious paper, the authors develop an unparalleled community resource of insect genome regulatory annotations spanning five insect orders. They employ their previously-developed SCRMshaw method for computational cross-species enhancer prediction, drawing on available training datasets of validated enhancer sequence and expression from Drosophila melanogaster, which had been previously shown to perform well across select holometabolous insects (representing 160-345MY divergence). In this work they expand regulatory sequence annotation to 33 insect genomes spanning Holometabola and Hemiptera, which is even more distantly related to the fly model. They perform multiple downstream analyses of sets of predicted enhancers to assess the true-positive rate of predictions; the independent comparisons of real predictions with simulated predictions and with chromatin accessibility data, as well as the functional validation through reporter gene analysis strengthen their conclusions that their annotation pipeline achieves a high true-positive rate and can be used across long divergence times to computationally annotate regulatory genome regions, an ability that has been largely inaccessible for non-model insects and now is possible across the many newly-sequenced insect scaffold-level genomes.</p>
<p>Strengths:</p>
<p>This work fills a large gap in current methods and resources for predicting regulatory regions of the genome, a task that has long lagged behind that of coding region prediction and analysis.</p>
<p>Despite technical constraints in working outside of well-developed model insect systems, the authors creatively draw on existing resources to scaffold a pipeline and independently assess likelihood of prediction validity.</p>
<p>The established database will be a welcome community resource in its current state, and even more so as the authors continue to expand their annotations to more insect genomes as they indicate. Their available analysis pipeline itself will be useful to the community as well for research groups that may want to undertake their own regulatory genome annotation.</p>
<p>Weaknesses:</p>
<p>The work here is limited by the field-wide lack of an independently validated set of tissue specific enhancers that could be used to directly benchmark this pipeline. The prediction of true positive enhancer identification rates and in vivo reporter gene assays offer some insight into the rates of successful prediction, but the output of SCRMshaw regulatory annotation should be regarded as another prediction-generating tool.</p>
</body>
</sub-article>
<sub-article id="sa3" article-type="author-comment">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.96738.2.sa0</article-id>
<title-group>
<article-title>Author response:</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Asma</surname>
<given-names>Hasiba</given-names>
</name>
<role specific-use="author">Author</role>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Tieke</surname>
<given-names>Ellen</given-names>
</name>
<role specific-use="author">Author</role>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Deem</surname>
<given-names>Kevin D</given-names>
</name>
<role specific-use="author">Author</role>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Rahmat</surname>
<given-names>Jabale</given-names>
</name>
<role specific-use="author">Author</role>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Dong</surname>
<given-names>Tiffany</given-names>
</name>
<role specific-use="author">Author</role>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Huang</surname>
<given-names>Xinbo</given-names>
</name>
<role specific-use="author">Author</role>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Tomoyasu</surname>
<given-names>Yoshinori</given-names>
</name>
<role specific-use="author">Author</role>
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0001-9824-3454</contrib-id></contrib>
<contrib contrib-type="author">
<name>
<surname>Halfon</surname>
<given-names>Marc S</given-names>
</name>
<role specific-use="author">Author</role>
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0002-4149-2705</contrib-id></contrib>
</contrib-group>
</front-stub>
<body>
<p>The following is the authors’ response to the original reviews.</p>
<disp-quote content-type="editor-comment">
<p><bold>Public Reviews:</bold></p>
<p><bold>Reviewer #1 (Public Review):</bold></p>
<p>Strengths:</p>
<p>The paper clearly presents the resource, including the testing of candidate enhancers identified from various insects in Drosophila. This cross-species analysis, and the inherent suggestion that training datasets generated in flies can predict a cis-regulatory activity in distant insects, is interesting. While I can not be sure this approach will prevail in the future, for example with approaches that leverage the prediction of TF binding motifs, the SCRMShaw tool is certainly useful and worth consideration for the large community of genome scientists working on insects.</p>
</disp-quote>
<p>We thank the reviewer for the positive comments, and would just like to point out that we agree: while we cannot of course know if other methods will overtake SCRMshaw for enhancer prediction—we assume they will, at some point (although motif-based approaches have not fared as well in the past)—for now, SCRMshaw provides strong performance and is a useful part of the current toolkit.</p>
<disp-quote content-type="editor-comment">
<p>Weaknesses:</p>
<p>While the authors made the effort to provide access to the SCRMShaw annotations via the RedFly database, the usefulness of this resource is somewhat limited at the moment. First, it is possible to generate tables of annotated elements with coordinates, but it would be more useful to allow downloads of the 33 genome annotations in GFF (or equivalent) format, with SCRMshaw predictions appearing as a new feature. Also, I should note that unlike most species some annotations seem to have issues in the current RedFly implementation. For example, Vcar and Jcoen turn empty.</p>
</disp-quote>
<p>We have addressed these weaknesses in several ways:</p>
<p>(1) We have created GFF versions of the SCRMshaw predictions and provide them standalone and also merged into the available annotation GFFs for each of the 33 species</p>
<p>(2) We have made these GFF files, and also the original SCRMshaw output files, available for download in a Dryad repository linked to the publication (<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5061/dryad.3j9kd51t0">https://doi.org/10.5061/dryad.3j9kd51t0</ext-link>).</p>
<p>(3) We have added the inadvertently omitted species to the REDfly/SCRMshaw database.</p>
<p>We agree that the database functions are still somewhat limited, but note that database development is ongoing and we expect functionality to increase over time. In the meantime, the Dryad repository ensures that all results reported in this paper are directly available.</p>
<disp-quote content-type="editor-comment">
<p><bold>Reviewer #2 (Public Review):</bold></p>
<p>Summary:</p>
<p>… Upon identification of predicted enhancer regions, the authors perform post-processing step filtering and identify the most likely predicted enhancer candidates based on the proximity of an orthologous target gene. …</p>
</disp-quote>
<p>We respectfully point out a small misunderstanding here on the part of the reviewer. We stress that putative target gene assignments and identities have no impact at all on our prediction of regulatory sequences, i.e., they are not “based on the proximity of an orthologous target gene.” Predictions are solely based on sequence-dependent SCRMshaw scores, with no regard to the nature or identities of nearby annotated features. Putative target genes are mapped to Drosophila orthologs purely as a convenience to aid in interpreting and prioritizing the predicted regulatory elements. We have added language on page 8 (lines 189ff) to make this more clear in the text.</p>
<disp-quote content-type="editor-comment">
<p>Weaknesses:</p>
<p>This work provides predicted enhancer annotations across many insect species, with reporter gene analysis being conducted on selected regions to test the predictions. However, the code for the SCRMshaw analysis pipeline used in this work is not made available, making reproducibility of this work difficult. Additionally, while the authors claim the predicted enhancers are available within the REDfly database, the predicted enhancer coordinates are currently not downloadable as Supplementary Material or from a linked resource.</p>
</disp-quote>
<p>We have placed all the code for this paper into a GitHub repository “Asma_etal_2024_eLife” (<ext-link ext-link-type="uri" xlink:href="https://github.com/HalfonLab/Asma_etal_2024_eLife">https://github.com/HalfonLab/Asma_etal_2024_eLife</ext-link>) to address this concern. As described in our response to Reviewer 1, above, all results are now available in multiple formats in a linked Dryad repository in addition to the REDfly/SCRMshaw database.</p>
<disp-quote content-type="editor-comment">
<p>The authors do not validate or benchmark the application of SCRMshaw against other published methods, nor do they seek to apply SCRMshaw under a variety of conditions to confirm the robustness of the returned predicted enhancers across species. Since SCRMshaw relies on an established k-mer enrichment of the training loci, its performance is presumably highly sensitive to the selection of training regions as well as the statistical power of the given k-mer counts. The authors do not justify their selection of training regions by which they perform predictions.</p>
</disp-quote>
<p>Our objective in this study was not to provide proof-of-principle for the SCRMshaw method, as we have established the efficacy of the approach at this point in several previous publications. Rather, the objective here was to make use of SCRMshaw to provide an annotation resource for insect regulatory genomics. Note that the training regions we used here are the same as those we have used in earlier work. Naturally, we performed various assessments to establish that the method was working here, but we make no claims in this work about SCRMshaw’s relative efficiency compared to other methods. Some of our prior publications include assessments of the sort the reviewer references, which suggest that SCRMshaw is at least comparable to other enhancer discovery approaches. We note that benchmarking of such methods is in fact extremely complicated due to the fact that there are no established true positive/true negative data sets against which to benchmark (we have explored this in Asma et al. 2019 BMC Bioinformatics).</p>
<disp-quote content-type="editor-comment">
<p>While there is an attempt made to report and validate the annotated predicted enhancers using previously published data and tools, the validation lacks the depth to conclude with confidence that the predicted set of regions across each species is of high quality. In vivo, reporter assays were conducted to anecdotally confirm the validity of a few selected regions experimentally, but even these results are difficult to interpret. There is no large-scale attempt to assess the conservation of enhancer function across all annotated species.</p>
</disp-quote>
<p>We respectfully disagree that there is insufficient validation. We bring several different lines of evidence to bear suggesting that our results fall into the accuracy range—roughly 75%—established both here and in previous work. We are also clear about the fact that these are predictions only and need to be viewed as such (e.g. line 638). Although “large-scale” in vivo validation assays would certainly be both interesting and worthwhile, the necessary resources for such an assessment places it beyond our present capability.</p>
<disp-quote content-type="editor-comment">
<p>Lastly, it is suggested that predicted regions are derived from the shared presence of sequence features such as transcription factor binding motifs, detected through k-mer enrichment via SCRMshaw. This assumption has not been examined, although there are public motif discovery tools that would be appropriate to discover whether SCRMshaw is assigning predicted regions based on previously understood motif grammar, or due to other sequence patterns captured by k-mer count distributions. Understanding the sequence-derived nature of what drives predictions is within the scope of this work and would boost confidence in the predicted enhancers, even if it is limited to a few training examples for the sake of clarity of interpretation.</p>
</disp-quote>
<p>Again, we respectfully disagree that “this assumption has not been examined.” Although we did not undertake this analysis here, we have in the past, where we have shown that known TFBS motifs can be recovered from sets of SCRMshaw predictions (e.g., Kazemian et al. 2014 Genome Biology and Evolution). We return to this point when we address the Comments to Authors, below.</p>
<disp-quote content-type="editor-comment">
<p><bold>Reviewer #3 (Public Review):</bold></p>
<p>Weaknesses:</p>
<p>The rates of predicted true positive enhancer identification vary widely across the genomes included here based on the simulations and comparison to datasets of accessible chromatin in a manner that doesn't map neatly onto phylogenetic distance. At this point, it is unclear why these patterns may arise, although this may become more clear as regulatory annotation is undertaken for more genomes.</p>
</disp-quote>
<p>We agree that we do not see clear patterns with respect to phylogenetic distance in our results. However, we note that this initial data set is still fairly small, and not carefully phylogenetically distributed. We are hoping that, as the reviewer suggests, some of these questions become more clear as we add more genomes to our analysis. Fortunately, the list of available genomes with chromosome-level assembly is growing rapidly, and as we move ahead we should have much greater ability to choose informative species.</p>
<disp-quote content-type="editor-comment">
<p>Functional assessment of predicted enhancers was performed through reporter gene assays primarily in Drosophila melanogaster imaginal discs, a system amenable to transgenics. Unfortunately, this mode of canonical imaginal disc development is only representative of a subset of all holometabolous insects; therefore, it is difficult to interpret reporter gene expression in a fly imaginal disc as evidence of a true positive enhancer that would be active in its native species whose adult appendages develop differently through the larval stage (for example, Coleopteran and Lepidopteran legs). However, the reporter gene assays from other tissues do offer strong evidence of true positive enhancer detection, and constraints on transgenic experiments in other systems mean that this approach is the best available.</p>
</disp-quote>
<p>Please see an extensive discussion of this point in our response to Reviewer 3, below.</p>
<disp-quote content-type="editor-comment">
<p><bold>Recommendations for the authors:</bold></p>
<p><bold>Reviewer #2 (Recommendations For The Authors):</bold></p>
<p>Major Concerns:</p>
<p>(1) While the GitHub source code for SCRMshaw is provided, the authors do not provide a repository of manuscriptspecific code and scripts for readers. This is a barrier to reproducibility and the code used to perform the analysis should be made available. Additionally, links to available scripts do not work, see Line 690. Post-processing scripts point to a general lab folder, but again, no specific analysis or code is sourced for the work in this specific manuscript (e.g. Line 637).</p>
</disp-quote>
<p>As noted above, we have corrected this oversight and established a specific GitHub repository for this manuscript “Asma_etal_2024_eLife” (<ext-link ext-link-type="uri" xlink:href="https://github.com/HalfonLab/Asma_etal_2024_eLife">https://github.com/HalfonLab/Asma_etal_2024_eLife</ext-link>).</p>
<disp-quote content-type="editor-comment">
<p>(2) On lines 479-488, there is a discussion about the annotations being provided on REDfly, though no link is provided.</p>
</disp-quote>
<p>We have included a link in the text at this point (now line 515).</p>
<disp-quote content-type="editor-comment">
<p>Additionally, for transparency, it would be valuable to provide in Supplementary Table 1 the genomic coordinates of the original training sets in addition to their identity.</p>
</disp-quote>
<p>These coordinates have been added to Supplementary Table 1 as suggested.</p>
<disp-quote content-type="editor-comment">
<p>Also, it is suggested to provide genomic coordinates of the predicted enhancers for each training set across all species, perhaps with a column denoting a linked ID of one genomic coordinate in a species to another species (i.e. if there is a linked region found from D. melanogaster to J. coenia, labeling this column in both coordinate sets as blastoderm.mapping1_region1). Providing these annotations directly in the work enhances the transparency of the results.</p>
</disp-quote>
<p>We are unsure exactly what the reviewer means here by “a linked region.” It is critical to understanding our approach to recognize that the genome sequences have diverged to the point where there is no alignment of non-coding regions possible. Thus there is no way to directly “link” coordinates of a predicted enhancer from one species to those of a predicted enhancer in another species. The coordinates for each prediction are available on a per-species basis either through the database or in the files now available in the linked Dryad repository; these can be filtered for results from a specific training set. The database will allow users to select all results for a given orthologous locus, from any subset of species. More complex searches will continue to become available as we improve functionality of the database, an ongoing project in collaboration with the REDfly team.</p>
<disp-quote content-type="editor-comment">
<p>(3) Figure 2B: It is unclear what this figure shows. Are the No Fly Orthologs false positives, Orthology pipeline issues, or interesting biology?</p>
</disp-quote>
<p>We have clarified this in the Figure 2 legend. “No Mapped Fly Orthologs” indicates that our orthology mapping pipeline did not identify clear D. melanogaster orthologs. For any given gene, this could reflect either a true lack of a respective ortholog, or failure of our procedure to accurately identify an existing ortholog.</p>
<disp-quote content-type="editor-comment">
<p>(4) SCRMshaw appears to be a versatile tool, previously published in a variety of works. However, in this manuscript, there is little discussion of the sensitivity of SCRMshaw to different initial parameters, how the selection of training loci can impact outcomes, or how SCRMshaw k-mer discovery methods compare to other similar tools.</p>
<p>- This paper would be strengthened by addressing this weakness. Some specific suggestions below:</p>
<p>In order to strengthen confidence that SCRMshaw is a reliable predictor of enhancer regions in other species, it is suggested that you benchmark against other k-mer-derived methods to assign enhancers, such as GSK-SVM developed by the Beer Lab in 2016  (<ext-link ext-link-type="uri" xlink:href="https://www.beerlab.org/gkmsvm/">https://www.beerlab.org/gkmsvm/</ext-link>, <ext-link ext-link-type="uri" xlink:href="https://www.biorxiv.org/content/10.1101/2023.10.06.561128v1">https://www.biorxiv.org/content/10.1101/2023.10.06.561128v1</ext-link>).</p>
</disp-quote>
<p>We have established the effectiveness of SCRMshaw as an enhancer discovery method in previous work, and the main goal of this study was to make use of the established method to annotate numerous insect genomes as a community resource. Our claim here is that SCRMshaw works well for this purpose; we do not attempt a strong claim about whether other approaches may work equally well or marginally better (although we do not believe this is the case, based on prior work). Benchmarking enhancer discovery is challenging, as we point out in Asma et al. 2019 (BMC Bioinformatics), and, while important, best left for a dedicated comprehensive study. A major problem is that there are no independent objective “truth” sets for enhancers from the various species we interrogate here. Thus, while we could also run, e.g., GSK-SVM, what criteria would we use to establish which method had better accuracy for a given species? Note that the work from Beer’s lab took advantage of the ability to match human-mouse orthologous (or syntenic) regions and available open-chromatin data to assess whether conserved enhancers were discovered, but this is not possible given the degree of divergence, limited synteny, and relative lack of additional data for the insect genomes we are annotating.</p>
<disp-quote content-type="editor-comment">
<p>- In Table S1, we see that 7-146 regions are used as training sets, which is a huge variety. Does an increase in training set size provide a greater &quot;rate of return&quot; for predicted regions? Is the opposite true? Addressing this question would allow readers to understand if they wish to use SCRMshaw, a reasonable scope for their own training region selections.</p>
<p>- Within a training set, does subsampling provide the same outcomes in terms of prediction rates? There is no exploration of how &quot;brittle&quot; the training sets are, and whether the generalized k-mer count distributions that are established in a training set are consistent across randomly selected subgroups. Performing this analysis would raise confidence in the method applied and the resulting annotations.</p>
</disp-quote>
<p>These are interesting and important questions, but again we feel they are beyond the scope of this particular study, which is focused primarily on using SCRMshaw and not on optimizing various search parameters. That said, this is of course something we have investigated, although as with other aspects of enhancer discovery, the absence of a true gold standard enhancer set makes evaluation difficult. We have not found a clear correlation between training set size and performance beyond the very general finding that performance appears to be best when training set size is moderate, e.g. 20-40 initial enhancers. We suspect that larger training sets often contain too many members that don’t fit the core regulatory model and thus add noise, whereas sets that are too small may not contain enough signal for best performance (although small sets can still be useful, especially if used in an iterative cycle; see Weinstein et al. 2023 PLoS Genetics). However, establishing this rigorously is highly challenging given the limitations with assessing true and false positive rates at scale.</p>
<disp-quote content-type="editor-comment">
<p>(5) In Figure 2C, when plotting hexMCD, IMM, pacRC, and then the merged set, it is unclear whether the scorespecific bar allows coordinate redundancy, though this is implied. What might be more useful is a revision of this plot where the hexMCD/IMM/pac-RC-specific loci are plotted, with the merged set alongside as is currently reported. This would give the reader a clearer understanding of the variability between these scoring methods and why this variability occurs.</p>
</disp-quote>
<p>We have added the breakdowns between IMM, hexMCD, and pacRC in Supplementary Table S2, and made more complete reference to this in the text (lines 682ff). Both the database and the data files in the Dryad repository allow exploration of the overlap between the different methods and contain both separate and merged (for overlap and redundancy) results.</p>
<disp-quote content-type="editor-comment">
<p>Additionally, there is no information in the Methods section of these three SCRMshaw scores and what they represent, even colloquially. While SCRMshaw has been applied in several papers previously, it would help with scientific clarity to describe in a sentence or two what each score is meant to represent and why one is different from another.</p>
</disp-quote>
<p>We had chosen to err on the side of brevity given prior publication of the SCRMshaw methodology, but we recognize now that we went too far in that direction. We have added more complete descriptions of the methods in both the Results (lines 164-167) and the Methods (lines 667-681) sections.</p>
<disp-quote content-type="editor-comment">
<p>(6) When describing results in Figure 2, an important question arises: &quot;Is there an anti-correlation between the number of predicted regions and evolutionary distance?&quot; This would be an expected result that could complement Figure 4's point that shared orthology across 16 species is rarer than across 10 species. Visualizing and adding this to Figure 2 or Figure 4 would be a powerful statement that would boost confidence in the returned predicted enhancers and/or orthologous regions.</p>
</disp-quote>
<p>This is an important question and one in which we are very interested. Unfortunately, we do not have sufficient data at this time to address this proper statistical rigor. As we remarked above in response to Reviewer 3, “We agree that we do not see clear patterns with respect to phylogenetic distance in our results. However, we note that this initial data set is still fairly small, and not carefully phylogenetically distributed. We are hoping that, as the reviewer suggests, some of these questions become more clear as we add more genomes to our analysis. Fortunately, the list of available genomes with chromosome-level assembly is growing rapidly, and as we move ahead we should have much greater ability to choose informative species.”</p>
<disp-quote content-type="editor-comment">
<p>(7) In Figure 3, the authors seek to convey that SCRMshaw predicts enhancer regions that are mapped nearby one another, across different loci widths, and that this occurrence of nearby predicted regions occurs more than a randomly selected control. This is presumably meant to validate that SCRMshaw is not providing predictions with low specificity, but rather to highlight the possibility that SCRMshaw is identifying groups of shadow enhancers. However, these plots are extremely difficult to decipher and do not strongly support the claims due to the low resolution and difficult interpretability of the boxplot interquartile distributions.</p>
</disp-quote>
<p>Additionally, as the majority of predicted regions are around ~750bp, how does that address loci groups of &lt;1000bp? This suggests that predicted regions are overlapping, and therefore cannot be meaningfully interpreted as shadow enhancers. This plot should either be moved to the supplements or reworked to more effectively convey the point that &quot;SCRMshaw is detecting predicted regions that are proximal to one another and that this proximity is not due to chance&quot;.</p>
<disp-quote content-type="editor-comment">
<p>- A suggestion to rework this plot is to change this instead to a bar plot, where the y-axis instead represents &quot;number of predictions with at least 2 predicted regions proximal to one another&quot; divided by &quot;total number of predictions&quot;, separating bar color by simulated/observed values. The x-axis grouping can remain the same. Because this plot is a broad generalization of the statement you're trying to make above, knowing whether a few loci have 2 versus 4 proximal predicted enhancers doesn't enhance your point.</p>
</disp-quote>
<p>We agree with the reviewer that these are not the clearest plots, and thank them for the suggestions regarding revision. We tried many variations on visualizing these complex data, including those suggested by the reviewer, and have concluded that despite their weaknesses, these plots are still the best visualization. The main problem is that the observed data cluster heavily around zero, so that the box plots are very squat and mainly only the outlier large values are observed. The key point, however, is that the expected values almost never give values much greater than one, so that the observed outlier points are the only points seen in the upper ranges of the y-axis. This is true across the three species, across the bins of locus sizes, and across training sets (averaged into the box plots). The reviewer is correct as well about the bins where locus size is &lt; 1000. However, inspection of the data shows that this is not a large concern, as very few data points lie in this range and we never see multiple predicted enhancers there. Thus we believe while not the prettiest of graphs, Figure 3 does effectively support the claims made in the text. In keeping with our view that it is preferable to have data in the main paper whenever possible, we choose to keep the figure in place rather than move it to the Supplement.</p>
<disp-quote content-type="editor-comment">
<p>- Label the species for the reader's understanding of each subplot on the plot.</p>
</disp-quote>
<p>We apologize for this oversight and have now labeled each plot with its relevant species.</p>
<disp-quote content-type="editor-comment">
<p>(8) SCRMshaw operates on k-mer count distributions compared to a genomic background across different species, allowing it to assign predicted regions without prior knowledge of an organism's cis-regulatory sequences. This is powerful and boosts the versatility of the method. However, understanding the cis-regulatory origins of the kinds of kmers that are driving the detection of orthologous regions across species is crucial and absolutely within the scope of the paper, particularly for the justification of the provided annotations. Is SCRMshaw making use of enriched motifs within the training region set to assign regions in other species? One would presume so, but it is necessary to show this. There are many motif discovery tools that are readily available and require little up-front knowledge and little to no use of a CLI, such as MEMESuite (<ext-link ext-link-type="uri" xlink:href="https://meme-suite.org/meme/tools/meme">https://meme-suite.org/meme/tools/meme</ext-link>). It is highly recommended that, even for a few training pairs that are well understood (e.g. mesoderm.mapping1, dorsal_ectoderm.mapping1), assess the motif enrichment within the original sequence set, then see whether motif enrichments are reflected in the predicted enhancers. As evolutionary distance increases between D. melanogaster and the species of interest, is the assignment of enriched motifs more sparse? Is there a loss of a key motif? These are the kinds of questions that will allow readers to understand how these annotations are assigned as well as boost confidence in their usage.</p>
</disp-quote>
<p>This is a very important point and a subject of significant interest to us. We have demonstrated in earlier work (e.g., Kazemian et al. 2014 Genome Biol. Evol.) that SCRMshaw-predicted enhancers do contain expected TFBS motifs, across multiple species—and that even an overall arrangement of sites is sometimes conserved. Thus we have previously answered, in part, the reviewer’s question.</p>
<p>What we also learned from our previous work is that filtering out relevant motifs from the noise inherent in motif-finding is both arduous and challenging. As the reviewer is no doubt aware, while using motif discovery tools is simple, interpreting the output is much less so. In response to the reviewer’s comments, we revisited this issue with data from a small sample of training sets. We can discover motifs; we can see that the motif profiles are different between different training sets; and we can observe the presence of expected motifs based on the activity profile of the enhancers (e.g., Single-minded binding sites in our mesectoderm/midline training and result data). However, to do this cleanly and with appropriate statistical rigor is beyond what we feel would be practical for this paper. We hope to return to this important question in the future when we have a larger and phylogenetically more evenly-distributed set of species, and the time and resources to address it appropriately.</p>
<disp-quote content-type="editor-comment">
<p>(9) Figures 5-7 need to have better descriptions.</p>
</disp-quote>
<p>We have added to the figure 6 and 7 legends in response to this comment; please note as well that there is substantial detail provided in the text. If there are specific aspects of the figures that are not clear or which lack sufficient description, we are happy to make additional changes.</p>
<disp-quote content-type="editor-comment">
<p>Minor Concerns</p>
<p>(1)  In Figure 1A, it is implied that &quot;k-mer count distributions&quot; are actually only &quot;5-mer count distributions&quot;. However, in the published documentation of SCRMshaw, it is suggested that k-mers between 1-6 bp are involved in establishing sequence distributions. Please add a justification for the selection of these criteria. It would be helpful to understand the implications of using up to a 3-mer versus a 12-mer when assessing k-mer counts using SCRMshaw.</p>
</disp-quote>
<p>We have clarified in the Figure 1 legend that this is just an example, and the k-mers of different sizes are used in the IMM method; we have also increased the description of the basic method in the Methods section. To be clear, the hexMCD sub-method is 6-mer based (5th-order Markov chain), as is pacRC, while the IMM method considers Markov chains of orders 0-5.</p>
<disp-quote content-type="editor-comment">
<p>(2) Control the y-axis to remove white space from Figure 2D.</p>
</disp-quote>
<p>We have amended the figure as suggested.</p>
<disp-quote content-type="editor-comment">
<p>Additionally, expand in the manuscript on expected results from SCRMshaw. Given training regions of 750 bp, is the expectation that you return predicted enhancers of the same length? This is not explicitly stated, only a description of outliers.</p>
</disp-quote>
<p>The scoring is not dependent on the length of the training sequences, and there is no direct expectation of predicted enhancer length. Scores are calculated on 10-bp intervals, and a peak-calling algorithm is used to determine the endpoints of each prediction based on where the scores drop below a cutoff value. Thus there is no explicit minimum prediction length beyond the smallest possible length of 10-bp. That said, the initial scoring takes place over a 500-bp sequence window (for reasons of computational efficiency), which does influence scores away from the smaller end of the possible range. We correct for this in part by reducing scores below a certain threshold to zero, to prevent multiple low-scoring regions from combining to give a low but positive score over a long interval. Indeed, we found that in the original version of SCRMshawHD (Asma et al. 2019), multiple low-scoring but above-threshold intervals would get concatenated together in broad peaks, leading to an unrealistically large average prediction length. In the version used here, described in Supplementary Figure S6, low-scoring windows are now first reset to zero and a new threshold is calculated before overlapping scores are summed. This helps to prevent the broad peak problem, and we find that it results in a median prediction length ~750 bp, more in line with expected enhancer sizes.</p>
<disp-quote content-type="editor-comment">
<p><bold>Reviewer #3 (Recommendations For The Authors):</bold></p>
<p>Line 161: Given that the SCRMshaw HD method is the basis for the pipeline, the methodology deserves at least an &quot;in brief&quot; recapitulation in this manuscript.</p>
</disp-quote>
<p>As we remark in our response to Reviewer 2, above, “We had chosen to err on the side of brevity given prior publication of the SCRMshaw methodology, but we recognize now that we went too far in that direction. We have added more complete descriptions of the methods in both the Results (lines 164-167) and the Methods (lines 667-681) sections.”</p>
<disp-quote content-type="editor-comment">
<p>Line 219: Throughout the reporting of the results, there appeared to be a bit of inconsistency/potential typos regarding whether threshold or exact P values were reported. In lines 219, 222, 265, 696, and 811, the reported values seem to clearly be thresholds (&lt; a standard cutoff), while in lines 291,293, 297,300, values appear to be exact but are reported as thresholds (&lt;).</p>
</disp-quote>
<p>This is not an error but rather reflects two different types of analysis. The predictions per locus (originally lines 219, 222 etc) are evaluated using an empirical P-value based on 1000 permutations. As such, they are thresholded at 1/1000. The overlap with open chromatin regions, on the other hand, are based on a z-score with the P-values taken from a standard conversion of z-scores to P-values.</p>
<disp-quote content-type="editor-comment">
<p>Page 13/Table 2: At face value, it seems surprising that the overlap between Dmel SCRMshaw predictions with open chromatin is so much smaller than the overlap between predictions and open chromatin in other species, both in raw % (Tcas, D plexippus, H. himera) and fold enrichment (Tcas), given that the training sets for SCRMshaw are all derived from Dmel data. The discussion here does not touch on this aspect of the results, and the interpretation of this approach, in general, would be strengthened if the authors could comment on potential reasons why this pattern may be arising here, or at least acknowledge that this is an open question.</p>
</disp-quote>
<p>There are many variables at play here, as the data are from different species, from different tissues, and from different methods. Thus we think it is difficult to read too much into the precise results from these comparisons—the main take-home is really just that there is a significant amount of overlap. In acknowledgment of this, we have slightly modified the text in this section so that it now notes (line 302ff): “These comparisons are imperfect, as the tissues used to obtain the chromatin data do not precisely correspond to the training sequences used for SCRMshaw, and the data were obtained using a variety of methods.”</p>
<disp-quote content-type="editor-comment">
<p>Line 318-329: The inferences from the reporter gene assay deserve a more nuanced treatment than they are given here. The important nuance that was not addressed by the discussion here is that the imaginal disc mode of development in Drosophila is not broadly representative of the development of larval/adult epithelial tissues across Holometabola; thus, inference of a true positive validation becomes complicated in cases where predicted enhancers from a species were tested and shown to drive expression in a fly imaginal disc that the native species have no direct disc counterpart to. For example, in line 388 a Tcas enhancer is reported to drive expression in the eye-antennal disc, and in lines 404 and 423 additional Tcas enhancers were reported to drive expression in the leg discs; however, Tribolium larvae do not possess antennal discs or leg discs set aside during embryogenesis in the sense that flies do - instead the homologous epithelial tissues form larval antennae and larval legs external to the body wall that are actively used at this life stage and are starkly different in morphology than an internally invaginated epithelial disc, that will directly give rise to adult tissues in subsequent molts. Is the interpretation of an expression pattern driven in a fly disc as a true positive really as straightforward as it was presented here, when in the native species the expression pattern driven by the enhancer in question would be in the context of an extremely different tissue morphology? That said, I understand and am deeply sympathetic to the constraints on the authors in performing transgenic experiments outside of the model fly; but these divergent modes of development across Holometabola deserve a mention and nuance in the interpretation here.</p>
</disp-quote>
<p>This is indeed a very important point, and we greatly appreciate Reviewer 3 pointing out this caveat when interpreting the outcomes of our cross-species reporter assay. Reviewer 3 is correct that the imaginal disc mode of adult tissue (i.e. imaginal) development found in Diptera does not represent the imaginal development across Holometabola.</p>
<p>In fact, imaginal development is quite diverse among Holometabola. For instance, larval leg and antennal cells appear to directly develop into the adult legs and antennae in Coleoptera (i.e. primordial imaginal cells function as larval appendage cells), while some cells within the larval legs and antennae are set aside during larval development specifically for adult appendages in Lepidopteran species (i.e. imaginal cells exist within the larval appendages but do not contribute to the formation of larval appendages). In contrast, an almost entire set of cells that develop into adult epithelia are set aside as imaginal discs during embryogenesis in Diptera. Furthermore, the imaginal disc mode of development appears to have evolved independently in</p>
<p>Hymenoptera. Therefore, determining how imaginal primordial tissues correspond to each other among Holometabola has been a challenging task and a topic of high interest within the evo-devo and entomology communities.</p>
<p>Nevertheless, despite these differences in mode of imaginal development, decades of evo-devo studies suggest that the gene regulatory networks (GRNs) operating in imaginal primordial tissues appear to be fairly well conserved among holometabolan species (for example, see Tomoyasu et al. 2009 regarding wing development and Angelini et al. 2012 regarding leg development between flies and beetles). These outcomes imply that a significant portion of the transcriptional landscape might be conserved across different modes of imaginal development. Therefore, an enhancer functioning in the Tribolium larval leg tissue (which also functions as adult leg primordium) could be active even in the leg imaginal disc of Drosophila, if the trans factors essential for the activation of the enhancer are conserved between the two imaginal tissues.</p>
<p>That being said, we fully expect there to be both false negative and false positive results in our cross-species reporter assay. We are optimistic about the biological relevance of the positive outcomes of our crossspecies reporter assay, especially when the enhancer activity recapitulates the expression of the corresponding gene in Drosophila (for example, Am_ex Fig6B and Tc_hth Fig7B). Nonetheless, the biological relevance of these enhancer activities needs to be further verified in the native species through reporter assays, enhancer knock-outs, or similar experiments.</p>
<p>In recognition of the Reviewer’s important point, we added the following caveat in our Discussion (lines 549553): “Furthermore, the unique imaginal disc mode of adult epithelial development in D. melanogaster  might have prevented some enhancers of other species from working properly in D. melanogaster imaginal discs, likely producing additional false negative results. Evaluating enhancer activities in the native species will allow us to address the degree of false negatives produced by the cross-species setting.” We moreover mention this caveat in the Results section when we first introduce the reporter assays (line 342).</p>
<disp-quote content-type="editor-comment">
<p>Line 580: This is the first time that the weakness of the closest-gene pairing approach is mentioned. This deserves mention earlier in the manuscript, as unfortunately, this is one of the major bottlenecks to this and any other approaches to investigating enhancer function. Could the authors address this earlier, perhaps pages 7-8, and provide citations for current understanding in the field of how often closest-gene pairing approaches correctly match enhancers to target genes?</p>
</disp-quote>
<p>We have added text as suggested on p.7-8 acknowledging the shortcomings of the closest-gene approach. We also clarify at the end of that section (lines 173-181) that target gene assignments, while useful for interpretation, have no bearing on the enhancer predictions themselves (which are generated prior to the target gene assignment steps).</p>
</body>
</sub-article>
</article>