Sunday, March 21, 2010

MicroRNAs and metazoan phylogeny: big trees from little genes

Understanding the evolution of a clade, from either a morphological or genomic perspective, first and foremost requires a correct phylogenetic tree top- ology. This allows for the polarization of traits so that synapomorphies (innovations) can be distin- guished from plesiomorphies and homoplasies. Metazoan phylogeny was originally formulated on the basis of morphological similarity, and in some areas of the tree was robustly supported by molecu- lar analyses, whereas in others it was strongly repu- diated. Nonetheless, some areas of the tree still remain largely unknown, despite decades, if not centuries, of research. This lack of consensus may be largely due to apomorphic body plans combined with apomorphic sequences. Here, we propose that microRNAs (miRNAs) may represent a new data set that can unequivocally resolve many relationships in metazoan phylogeny, ranging from the interre- lationships among genera to the interrelationships among phyla. miRNAs, small non-coding regula- tory genes, show three properties that make them excellent candidates for phylogenetic markers: (1) new miRNA families are continually being incorpo- rated into metazoan genomes through time; (2) they show very low homoplasy, with only rare instances of secondary loss, and only rare instances of substi- tutions occurring in the mature gene sequence; and (3) they are almost impossible to evolve convergently. Because of these three properties, we propose that miRNAs are a novel type of data that can be applied to virtually any area of the metazoan tree, to test among competing hypotheses or to forge new ones, and to help finally resolve the correct topology of the metazoan tree.
15.1 Introduction

Since the dawn of molecular phylogenetics, the relationships between animal groups, from spe- cies to the deepest nodes in Metazoa, have been the domain of ribosomal, mitochondrial, and nuclear protein-coding genes. Orthologous genes are amp- lified and sequenced, the sequences are aligned, and the alignment is analysed with increasingly sophisticated phylogenetic algorithms to gain an estimate of relationships. It is unarguable that our understanding of metazoan phylogeny has pro- gressed through the use of these genes and the application of standard phylogenetic methods. Many relationships originally proposed on mor- phological grounds have been confirmed, while others, such as the grouping of annelids and arthropods as the Articulata, have been strongly refuted, leading to a new understanding of mor- phological evolution (Eernisse and Peterson, 2004; Halanych, 2004). However, many areas of the metazoan tree have remained recalcitrant, yield- ing trees with low statistical support or with little resemblance to any credible scenario of morpho- logical evolution. It has often been assumed that these problems would disappear as more data (i.e. more genes and/or more taxa) were applied to the questions at hand. A sampling of the literature on multigene phylogenetics demonstrates that this has not been the case. Indeed, despite the fact that the amount of sequence data in public data bases such as NCBI’s GenBank doubles every 10 months, many phylogenetic questions remain as intractable today as they were before the advent of molecular


157

158 AN I M AL EV O L UTI O N



systematics. As just one example, from one of our parochial areas of interest, the interrelationships among three lophotrochozoan phyla, the nemer teans, annelids, and molluscs, still remain effect- ively unknown. This is despite a number of mul- tigene studies that have recently been published including complete 18S + 28S ribosomal RNA genes (Passamaneck and Halanych, 2006), complete mitochondrial genomes (Yokobori et al., 2008), mul- tiple PCR-amplified nuclear housekeeping genes (Helmkampf et al., 2008a; Peterson et al., 2008), and expressed sequence tag (EST) studies (Dunn et al.
2008; Struck and Fisse, 2008), with all three possible arrangements of these three phyla being advocated by at least one data set (Figure 15.1).
Three problems have always plagued (and will forever plague) the field of molecular phylogenet- ics: differential rates of molecular evolution, long internodes caused by a recent origin of the crown group, and fast, deep radiations. Indeed, in our reading of the metazoan phylogenetic literature, a large number of questions were robustly answered in the first or second pass using 18S rDNA, and were then largely confirmed using other types of data and/or algorithms. But the remaining nodes, which are usually hampered by at least one of these three problems, have remained largely intractable despite the ever-increasing number of taxa and genes being applied. Because of this, we believe that it is not more of the same data that will
ultimately answer these questions, but new types of data.
It was originally hoped that large-scale gen- omic changes, such as gene rearrangements in mitochondrial genomes (Boore et al., 1998) or insertion–deletion events, retroposon integrations, or gene duplications in nuclear genomes (Rokas and Holland, 2000) would provide this new data set, and provide a complementary approach to sequence-based phylogenetic estimation. These sources have provided robust support to topologies previously identified in sequence-based phylogen- etic studies, notably the placement of phoronids and brachiopods within the Protostomia (Helfenbein and Boore, 2004) in the case of mitochondrial gene order, or resolution of the whale–hippo clade by ret- roposon analysis (Shimamura et al., 1997; Nikaido et al., 1999) in the case of nuclear genome changes. Ultimately, however, these structural changes have not been the panacea it was hoped they would be. The most comprehensive coding of mitochondrial gene order demonstrated that in some clades, such as the vertebrates, rearrangement was too slow or non-existent, leading to a polytomy, whereas in other clades, such as the molluscs, rearrangement was too fast, leading to a nonsensical tree (Fritzsch et al., 2006). Mutational decay of the flanking regions surrounding retroposons makes their detection, at best, difficult in taxa that diverged more than about c. 50 million years ago (Ma), restricting their utility



(a) (b) (c)










Figure 15.1 The three proposed hypotheses for the interrelationships among nemerteans, annelids, and molluscs with respect to arthropods. (a) The Neotrochozoa hypothesis (Peterson and Eernisse, 2001) posits that annelids and molluscs are each other’s closest relatives with respect to nemerteans, found using morphological characters (Peterson and Eernisse, 2001) as well as an analysis using concatenated amino acid sequences of nuclear protein-coding genes (Peterson et al. 2008). (b) Mollusca + Nemertea was found with concatenated
amino acid sequences of nuclear protein-coding genes (Hausdorf et al., 2007; Helmkampf et al., 2008a; Struck and Fisse, 2008), as well as concatenated amino acid sequences of mitochondrial protein-coding genes (Yokobori et al., 2008). (c) Annelida + Nemertea was found using combined 18S rDNA + 28S rDNA by Passamaneck and Halanych (2006) and concatenated amino acid sequences of nuclear protein-coding genes (Dunn et al., 2008).

MIC R ORN A S 159



to the post-Mesozoic portions of the phylogeny (Luo, 2000; Rokas and Holland, 2000). Moreover, the utility of these rare events has been hampered by their very rarity; although in some fortuitous examples a strong synapomorphy is captured, they are simply not present in sufficient numbers such that investigators can reliably base a research pro- gram around using them to test hypotheses con- cerning metazoan interrelationships.
The ultimate problem in resolving evolutionary relationships is homoplasy: similarity caused not by shared ancestry but by convergent evolution, or loss and reversion to the primitive condition. Homoplasy occurs in every data set, from examples like the torpedo shape of fish, whales, and ich- thyosaurs, to the rapid gene rearrangements in molluscan mitochondrial genomes, to multiple sub- stitutions or convergent changes in gene sequences that are limited to only four nucleotide or 20 amino acid character states. Both elevated sequence evolu- tion in some taxa and long internodes cause homo- plasy in molecular data sets, causing informative synapomophies to be eroded to misleading homo- plasies. The key to resolving intractable nodes will lie in data sets that minimize homoplasy as much as possible, but whose characters arise or change at a high enough rate that they record the divergences in question. In this paper, we propose that short, highly conserved genes within the non-coding por- tion of the genome, specifically miRNA, may be one such data set. Not only is it almost impossible to evolve the same miRNA twice independently, but miRNAs are continually being added to metazoan genomes through time; only rarely are they sec- ondarily lost, and nucleotide substitutions to the mature gene sequence are infrequent. Importantly, ascertaining the miRNA complement of a taxon does not require any prior knowledge of the miRNA sequences themselves, greatly facilitating their util- ity for attacking phylogenetic questions at all scales of the animal tree, from phyla to species.


15.2 Background

Briefly, miRNAs are small, c. 22 nucleotides, non-coding genes that negatively regulate pro- tein-coding genes by binding, with imperfect complementarity, to sites in their 3c untranslated
regions (UTRs), thereby subjecting the transcript to cleavage or to blockage of its translation (Zhao and Srivastava, 2007; Filipowicz et al., 2008; Hobert
2008; Stefani and Slack, 2008). The first retrospect- ively recognized animal miRNA, lin-4, was discov- ered in the nematode worm Caenorhabditis elegans, where it is involved in regulating the timing of cell cycle division in the larval worm by binding to tar- get sites in the protein-coding gene lin-14 and pre- venting its translation (Lee et al., 1993; Wightman et al., 1993). Because lin-4 could not be found out- side of nematodes, it was considered to be a quirk of the nematode developmental process. The wider significance of this discovery came when it was shown that this type of gene regulation exists in other systems, particularly vertebrates (Ruvkun et al., 2004; Wickens and Takayama, 1994). This occurred with the finding of a second miRNA, let-7, which was originally discovered again in C. elegans (Reinhart et al., 2000), but was soon found in numerous other taxa including fruitflies and ver- tebrates (Pasquinelli et al., 2000), and quickly led to the discovery of many small regulatory RNAs subsequently named miRNAs (Lagos-Quintana et al., 2001; Lau et al., 2001; Lee and Ambros, 2001). let-7 had three intriguing characteristics that held promise for a future role in phylogenetic recon- struction for these small RNA genes (Pasquinelli et al., 2000). First, the mature gene product of let-7 is unchanged in sequence between nematodes, humans, and Drosophila, despite a total of almost
2000 million years of independent evolution in these three taxa. Second, let-7 was found in every protostome and deuterostome analysed, with no suggestions of secondary loss. Third, the gene was not present in any non-metazoan genome and was not detectable by Northern analysis in sponges or cnidarians, suggesting that the gene arose within Eumetazoa at the base of the nephrozoan triploblasts (i.e. protostomes and deuterostomes). Subsequent studies confirmed this pattern for let-7 as the gene was found in, for example, chaetog- naths, nemerteans, and polyclad and triclad flat- worms, but not in acoel flatworms or ctenophores (Pasquinelli et al., 2003).
miRNAs are defined by their mode of biogen- esis, which is intimately related to their unique hairpin secondary structure (Figure 15.2) and not

160 AN I M AL EV O L UTI O N




Dme Dpu Isc Csp
consensus
10 20 30 40 50 60 70 80


Star Loop Mature






Figure 15.2 Alignment and secondary structure of representative sequences of the miRNA bantam. The mature sequence of the Drosophila melanogaster (Dme) bantam gene (miRBase) was used as a query against the trace archive sequences of Daphnia pulux (Dpu), Ixodes scapularis (Isc), and Capitella sp. (Csp) using the default settings (see Wheeler et al., 2009). About 85 nucleotides of the best hits were then folded using mfold (Zuker et al., 1999). Shown at the top is an alignment of these best hits using the default settings of ClustalW (MacVector, version 9.5.2), and shown below are the structures of two of these sequences, D. melanogaster (left) and Capitella (right) as determined by mfold. The mature and star sequences are shown.





by their specific nucleotide sequence. There are two components to a miRNA, the mature gene product, which is what binds to the 3c UTRs of tar- get genes, and the star sequence, the complement of the mature sequence, which is often degraded but is sometimes used as a gene product as well. miRNAs, which can be located either in intergenic regions or in introns, are transcribed as long pri- mary transcripts that are capped and polyade- nylated in typical Pol II fashion. However, because of the complementarity and spacing of this com- plementarity, the primary miRNA transcript folds into a hairpin structure, which is recognized by an enzyme complex involving at least two proteins, Drosha and Pasha, which cleave the pro-RNA into a c. 70 nucleotide precursor miRNA (Kim, 2005). This pre-miRNA is then exported into the cytoplasm where it is further processed by another RNAse enzyme called Dicer, and the mature gene product is then incorporated into an RNA–protein moiety that serves as the repressive entity with respect to messenger RNA translation and/or stability. Hence, miRNA biogenesis relies solely on miRNA struc- ture and not on miRNA sequence per se, greatly facilitating their utility for phylogenetics because it obviates the need for a researcher to know any particular miRNA sequence (see below).
miRNAs are named in sequential order of dis- covery, with identical or near identical mature sequences in the same or different organism given the same number (Ambros et al., 2003). miRNAs
given different numbers have different primary sequences and are assumed to have arisen independ- ently of other named miRNAs. This can be shown using a standard maximum parsimony analysis. If the first 20 miRNAs listed for both the fly Drosophila melanogaster and the human are aligned (Figure
15.3, left), and analysed using bootstrap analysis (Figure 15.3, right; see legend for details) one can easily see the orthology between similarly named miRNAs in the fly and human (e.g. let7, miR-1), and the paralogy of similarly named miRNAs in each taxon (e.g. miR-2). Further, the unique nature of each numbered miRNA or groups of miRNAs is readily apparent as they share virtually no similar- ity with any other miRNA in the alignment (Figure
15.3, left) and do not cluster together in the boot- strap analysis (Figure 15.3, right).
However, phylogenetic analyses are rarely, if ever, used to help name miRNAs, and thus nomen- clature problems can and do arise. For example, the two copies of miR-13 group with miR-2 (Figure 15.3, right), not unexpected given a cursory look at the alignment (Figure 15.3, left), and hence there are five, not three, copies of miR-2 in the fly genome. Even worse is when the same miRNA is given dif- ferent names in different organisms. For example, Sempere et al. (2006) reconstructed the protostome- specific set of miRNAs to include miR-8, and the deuterostome-specific set of miRNAs to include miR-141 and miR-200, and part of the reason for this was that the seed sequences of these genes, which

MIC R ORN A S 161




Dme bantam
Dme let7
Dme miR-1
Dme miR-2a
Dme miR-2b
Dme miR-2c
Dme miR-3
Dme miR-4
Dme miR-5
Dme miR-6
Dme miR-7
Dme miR-8
Dme miR-9a
Dme miR-9b
Dme miR-9c
Dme miR-10
Dme miR-11
Dme miR-12
Dme miR-13a
Dme miR-13b
Hsa let7a
Hsa let7b
Hsa let7c
Hsa let7d
Hsa let7e
Hsa let7f1
Hsa let7f2
Hsa let7g
Hsa let7i
Hsa miR-1-1
Hsa miR-1-2
Hsa miR-7-1
Hsa miR-7-2
Hsa miR-7-3
Hsa miR-9-1
Hsa miR-9-2
Hsa miR-9-3
Hsa miR-10a
Hsa miR-10b
Hsa miR-15a
consensus
10 20
Dme bantam
Dme let7
Hsa let7a
Hsa let7b
Hsa let7c
73 Hsa let7d
Hsa let7e
Hsa let7f1
Hsa let7f2
Hsa let7g
Hsa let7i
89 Dme miR-1
Hsa miR-1-1
Hsa miR-1-2
Dme miR-2a
Dme miR-2b
96 Dme miR-2c
Dme miR-13a
Dme miR-13b
Dme miR-3
Dme miR-4
Dme miR-5
Dme miR-6
Dme miR-7
79 Hsa miR-7-1
Hsa miR-7-2
Hsa miR-7-3
Dme miR-8
Dme miR-9a
Hsa miR-9-1
73 Hsa miR-9-2
Hsa miR-9-3
Dme miR-9b
Dme miR-9c
Dme miR-10
Hsa miR-10a
Hsa miR-10b
Dme miR-11
Dme miR-12
Hsa miR-15a

Figure 15.3 Alignment (left) and phylogenetic analysis (right) of the first 20 miRNAs listed for the fly Drosophila melanogaster (Dme) and the human Homo sapiens (Has) in miRBase. The alignment used the same parameters as Figure 15.2. Right: The 70% bootstrap tree. Sequences were analysed by PAUP, version 4.0b10 (Swofford, 2002) using maximum parsimony. Nodes found less than 70% of the time (1000 replications) were collapsed into polytomies. Note that similarly named miRNAs in the two systems cluster together, as do obvious paralogues in each system (e.g. let-7, miR-1). Note also that some differently numbered miRNAs (e.g. miR-2 and miR-13) group together as well, and as such constitute miRNA families, similar to, for example, the let-7 family or miR-1 family.



are positions 2–8 of the mature gene products, were slightly different (Figure 15.4, left). And because the seed sequence is the most important area of the mature gene sequence for target recognition it is primarily used for family-level classification (Filipowicz et al., 2008). But the use of only the seed sequence to name (and hence classify) miRNAs is a functional rather than a phylogenetic distinction, and in this case it is clear from a bootstrap analysis (Figure 15.4, right) that this is the same gene fam- ily, with the protostome versions called miR-8, and deuterostome versions called miR-141/200.


15.3 miRNAs as phylogenetic characters

miRNAs show three characteristics that make them outstanding candidates to arbitrate among
competing phylogenetic hypotheses and to forge new ones: (1) new miRNA families are continu- ously being added to metazoan genomes through time; (2) once incorporated into a gene regulatory network, there are only rare instances of secondary gene loss and they show only rare nucleotide sub- stitutions to the mature gene product; and (3) there is an infinitesimally small chance that miRNAs with the same mature sequence will evolve more than once.


15.3.1 Continuous addition of miRNA
families to metazoan genomes

Sempere et al. (2006) showed that the miRNA rep- ertoires of both fly and human were added sequen- tially through time such that each node leading to the fly, or to the human, could be characterized by

162 AN I M AL EV O L UTI O N




Hsa miR-141
Hsa miR-200a
Hsa miR-200b
Hsa miR-200c
Bfl1
Bfl2
Bfl3
Sko
Spu
Dme miR-8
Aga
Ame
Bmo Dpu Isc Csp Lgi
10 20




Deuterostomes



72



Protostomes

Figure 15.4 Homologous, but differently numbered, miRNAs in protostomes and deuterostomes. miR-8 in protostomes is clearly the same miRNA as miR-141 and miR-200 in deuterostomes, and indeed is supported as such in a 70% bootstrap analysis (right), despite some chordate paralogues possessing changes in the seed sequence (nucleotides 2–8) with respect to miR-8. Specifically, the human miR-141 and miR-200a, and the third paralogue found in the genomic traces of the cephalochordate Branchiostoma floridae (Bf13), have T to C changes in position 4 (left). Other abbreviations: Sko, Saccoglossus kowalevskii; Spu, Strongylocentrotus purpuratus; Aga, Anopheles gambiae; Ame, Apis mellifera; Bmo, Bombyx mori; Lgi, Lottia gigantea.




a distinctive miRNA or set of miRNAs. To explore this further, we again traced the phylogenetic his- tory of 132 uniquely numbered D. melanogaster miRNAs (see Heimberg et al., 2008, for vertebrates), but this time used many more genomes com- bined with 454 sequencing of small RNA librar- ies (Wheeler et al., 2009). We chose the arthropod example for three reasons. First, the miRNA rep- ertoire of D. melanogaster is the most extensively studied of any model organism (Ruby et al., 2007; Stark et al., 2007a,b), and we can be confident we are examining almost every miRNA in the organ- ism. Second, there are a large number of arthro- pod genomes available, including 12 from the genus Drosophila alone, allowing us to trace the phylogenetic acquisition over a range of taxonomic scales. And third, and most importantly for test- ing this data set for phylogenetic utility, there is an accepted phylogeny, allowing us to map miRNA gain (and losses, see below) against a known top- ology (Stark et al., 2007a).
Figure 15.5 shows the phylogenetic history of all
132 D. melanogaster miRNAs considered, as well
as the ancient miRNAs that should be present in
Drosophila, as determined by Wheeler et al., (2009),
but have been secondarily lost. Where we identify the gain of a new miRNA, it is shown in black under the node with paralogues of previously existing miRNA families underlined (see figure legend for details). Importantly, every node since the diver- gence between D. melanogaster and demosponges, where a sequenced genome is available and/or where a miRNA library has been constructed and sequenced (e.g. Priapulida; Wheeler et al., 2009), is characterized by the addition of at least one novel miRNA. This often involves the innovation of new families (e.g. miR-2 at the base of Protostomia), but sometimes additionally involves the generation of a paralogue from an existing gene (e.g. miR-13 at the base of Ecdysozoa). Hence, miRNAs could be used to resolve the interrelationships of taxa at vir- tually every level in the taxonomic hierarchy, from species to phyla.
We emphasize that we are showing the phylo- genetic history of the D. melanogaster miRNAs because they are well known and because the large number of genomes available enables such a study through bioinformatics alone. This does not imply that the other terminal tips will not have a similar number of miRNAs; groups such

MIC R ORN A S 163







100





1
31
79
92
124
219
252







let7
7
8
9
10
33
34
71
125
133
137
153
184
190
193
210
242
263
278
281
283
285
315
365
375
980
2001








Bantam
2
12
76
87
277
279
307
317
750
1175
1993








–242
–365
13
993










–1993 iab-4-3p
iab-4-5p
275
276












–2001
965













–153
14
277
282
286
305
927
929
932
970
988
989
998
1000















–750
11
306
308
316
957
996
999


















–71
3
4
5
6
274
280
284
287
288
289
304
309
314
318





















959
960
964
974
978
986
1003























1014

























961
968
975
1011
1015



























310
311
Demospongia Cnidaria Acoela Deuterostomia Eutrochozoa Priapulus
Ixodes Daphnia Tribolium Aedes
D. virilis
D. mojavensis D. grimshawi D. willistoni D. persimilis
D. pseudoobscura
D. ananassae
D. yakuba
D. erecta
D. sechellia

139 gains
7 known losses
955
956
962
963
969
971
976
987
994
995
1006
1007
1010
312
973
977
982*
985*
990
991*
992
1001*
1002
1004
1005
1008
1009
1012
1013
1016

313
966
967
983


303
954
972
979
984
D. simulans
D. melanogaster

Figure 15.5 Gains and secondary losses of 132 differently numbered miRNAs in Drosophila melanogaster. Gains are shown in black below
the node, and the seven secondary losses are shown above the node in grey (and where they were originally acquired are shown boxed below the node). Underlined miRNAs are paralogues of previously acquired miRNAs; those that are starred have slightly different seed sequences in other species of Drosophila, but fold properly and hence are considered gains at that point on the tree. This figure only considers gains and losses on
the lineage leading to the single terminal D. melanogaster and does not consider those leading to other terminals, although groups like beetles or chelicerates will clearly have their own sets of clade-specific miRNAs. Results from Demospongia to Daphnia are taken from Wheeler et al. (2009); the trace archives of the remaining insects were searched using all D. melanogaster miRNAs as query sequences. Potential hits were folded using the program mfold and assessed using standard structural criteria (see Wheeler et al., 2009, for materials and methods).



as mosquitoes or chelicerates will have their own clade-specific set of miRNAs that (we suspect) can (and hopefully will) be used to ascertain their internal phylogenetics. Indeed, Wheeler et al.,
(2009) showed that each major lineage of meta- zoans, except for Deuterostomia, could be char- acterized by the acquisition of at least one novel miRNA family. For example, ambulacrarians were

164 AN I M AL EV O L UTI O N



characterized by the addition of five novel miRNA families, and eleutherozoan echinoderms were characterized by the addition of 10 novel miRNA families. Even cnidarians have a novel miRNA family found only in Hydra and Nematostella and not anywhere else in the animal kingdom. And within these groups, the hemichordate worm Saccoglossus kowalevskii has at least an additional
34 miRNAs not found in the two echinoderms analysed, the sea urchin Strongylocentrotus purpu- ratus and the starfish Henricia sanguinolenta, and the hydrozoan cnidarian Hydra has at least an additional 17 miRNAs not found in Nematostella (Peterson et al., unpublished data). These novel miRNAs could then be used as phylogenetic markers to explore hemichordate and hydrozoan interrelationships, respectively, assuming they show low homoplasy. Fortunately, if they are similar to virtually all other known miRNAs, this will indeed be the case.


15.3.2 Minimal secondary gene loss and rare substitutions to the mature
sequence

miRNA homoplasy results from the possible com- bination of two factors, the first being the conver- gent evolution of the same miRNA in two taxa. The second is either complete loss from the gen- ome, or nucleotide substitutions in the mature sequence that destroy the ability to recognize its true orthology. As argued below, independent evolution of miRNAs is extremely limited, but sec- ondary gene loss and substitutions to the mature sequence can and do occur, and could obscure not only the interrelationships among the miRNAs but among the animal taxa as well. Nonetheless, in the Drosophila example discussed above (Figure 15.5), there are only seven secondary losses in D. mela- nogaster as compared with 139 gains—these losses are shown in grey in Figure 15.5 with their point of origin shown below the node and their inferred location of loss above the node. Note that loss can occur at any point in the evolutionary history— two of the genes not present in the fly were lost at the base of Ecdysozoa (miR-242 and miR-365), whereas one gene (miR-71) was lost in Drosophila after this lineage split from Aedes but before the
diversification of the 12 species under consider- ation (Figure 15.5).
If it could be shown that for most metazoan taxa miRNA gains outnumber miRNA losses by over an order of magnitude, as they do in this example, then their utility as phylogenetic markers would be unsurpassed, assuming that the mature sequence does not degrade over time. Sempere et al. (2006) argued that this was indeed the case, otherwise it would not be possible to map the origin of these
139 miRNAs with such minimal numbers of sec- ondary losses, as in this example (Figure 15.5). Further, Sempere et al. (2006) showed that miRNAs were some of the most, if not the most, conserved genetic elements in the genome, with most fly and eutherian mammal miRNAs showing no substi- tutions to the mature sequence. But because their focus was necessarily on flies and vertebrates, it could be argued these evolutionary patterns were particular to flies and vertebrates. Subsequently, Wheeler et al. (2009) quantified the number and position of substitutions of all 93 shared miRNAs across 14 nephrozoan taxa, and because this study relied primarily on isolating mature sequences in small RNA libraries it was not biased towards finding only conserved miRNAs. These authors analysed 16,729 nucleotides and showed that the substitution rate of all known and novel miRNAs across these 14 taxa, whose independent evolution- ary history spans over 7800 million years, is only
3.5% (567 total substitutions). When compared with
18S rDNA, one of the most conserved genes in the
metazoan genome, this rate is impressively slow:
aligning 18S rDNA from the same 14 taxa and
removing the unalignable regions using Gblocks
resulted in a substitution rate of 7.3% (Wheeler
et al., 2009). Hence, miRNAs evolve more than twice
as slowly as the most conserved positions in a gene
that is often used for reconstructing the deepest
nodes in the tree of life.


15.3.3 Exceedingly small probability of the independent evolution of
the same miRNA

In terms of convergent evolution, each unique
22-nucleotide sequence occurs by chance once
for every 1.76 × 1013 nucleotides (422), or once for

MIC R ORN A S 165



every 5864 human-genome-sized chunks of DNA queried. However, this is not an accurate estimate of the chances of two miRNAs evolving twice independently. For example, we took the (arbitrar- ily chosen) protostome-specific bantam miRNA gene (see Figure 15.2) from D. melanogaster and searched both protostomes and deuterostome genomes for this sequence in taxa that diverged from one another at least 500 Ma (Figure 15.7). In no case was the very same 23-nucleotide sequence found in any of these genomes (and aside from hits to D. melanogaster it is not found in the nucleotide data base deposited at GenBank, which consisted of
24,006,283,182 letters as of June 2008). Nonetheless,
23-nucleotide sequences were found in the two
arthropods, the water flea Daphnia and the tick
Ixodes, that are identical to each other but that differ
from that of D. melanogaster by a single nucleotide
at position 11. A single sequence was also found in
the genomic traces of the sea urchin S. purpuratus
that differs from the fly bantam sequence by a single
nucleotide, at position 13; the best hits in all of the
remaining deuterostomes have numerous differ-
ences, many of which are distributed in positions
2–6. The putative orthologue of bantam in the poly-
chaete annelid Capitella shares the same nucleotide
at position 11 as the water flea and the tick, but dif-
fers from all of the arthropods at positions 17, 20,
and 23. Because clearly orthologous miRNAs often
differ by two or three nucleotides, rather than com-
puting the probability for 23 nucleotides, a more
appropriate calculation is for the occurrence of a
stretch of 19 nucleotides, which is expected every
2.75 × 1011 bases, or once in every 91 human- genome equivalents, with the possibility of a few nucleotide substitutions (see below).
On the other hand, these numbers are decep- tively low because there are more constraints on a miRNA than the mature sequence of 22 nucle- otides; it must also fold with a free energy value lower than about –20 kcal/mol and often lower than –25 kcal/mol. In addition, the spacing has to be such that the mature sequence, which has to be located in one of the two hairpin arms, occurs within about two nucleotides from the loop, with the entire pre-miRNA generally being from 60–80 nucleotides long. Further, the structure cannot
contain large, and in particular asymmetrical, internal loops or bulges (Ambros et al., 2003). Thus, if one compares the two bantam miRNA sequences from Drosophila and the annelid Capitella it is obvi- ous that these are real miRNA genes; they have the requisite free energy values and structure to be processed and thus function as bona fide miRNA genes (Figures 15.2 and 15.6). But when the deu- terostome sequences are folded in silico, it is readily apparent that none of these are miRNAs, let alone orthologues of bantam. In the hemichordate (Sko), amphioxus (Bfl), and lamprey (Pma, see Figure 15.6) the free energy of these sequences are extremely high, c. –8 kcal/mol. In both the zebrafish (Dre, Figure 15.6) and ascidian (Cin) they have relatively low free energy values (c. –22 kcal/mol), but in both cases large and asymmetrical bulges and loops are present. Finally, in the sea urchin (Spu), which has the highest nucleotide similarity with the proto- stome sequences, not only is the free energy too high (–15 kcal/mol), but it too has large and asym- metrical bulges (Figure 15.6).
These non-folds are consistent with the observed substitution profile of the mature miRNA sequence as revealed by Wheeler et al. (2009). These authors found that most substitutions occurred at the 3c end of the mature sequence, but other regions of the gene, especially nucleotide 1 and nucleotide
10, showed a relatively high percentage of sub- stitutions. Importantly, the two most infrequent places for substitutions to occur are the seed region (positions 2–8) and the 3c complementar- ity region spanning nucleotides 13–16, especially position 15, in concordance with the hypothesized importance of these two regions for base pairing with the 3c UTR of targets (Filipowicz et al., 2008). Thus, unlike the protostome substitutions, which occur in statistically likely places (positions 11,
17, 20, and 23, see Figure 15.2), in deuterostomes, differences occur in the most conserved areas of miRNAs, positions 2–8 and 13–15. Conservation of sequence of orthologous miRNAs is explained by the constraints governing not only folding but base-pairing with targets, and these structural considerations also explain why the same miRNA gene sequence evolving twice independently is highly unlikely.





Dme Dpu Isc
Csp
Spu
Sko
Bfl
Cin
Pma
Dre
Consensus
10 20 30 40 50 60 70 80 90



Dme–bantam initial dG = –25.50






Csp–bantam initial dG = –29.80







Pma initial dG = –8.80
Spu initial dG = –15.70












Dre initial dG = –22.90








Figure 15.6 Alignment of the bantam gene taken from Figure 15.2 with the best hits from six different deuterostome genomes including the lamprey Petromyzon marinus (Pma) and the zebrafish Danio rerio (Dre) (other abbreviations are listed in Figure 15.4). Note that although some similarity is found in the mature sequence, especially with the sea urchin Strongylocentrotus purpuratus (Spu), there is no similarity in the star region, and, contra the protostome sequences (Dme and Csp), the deuterostome sequences do not show canonical folds (bottom), highlighting the improbability of evolving two miRNAs with the same mature sequence twice independently.

MIC R ORN A S 167


Symsagittifera roscoffensis Caenorhabditis elegans Ciona intestinalis

Cnidarians
Lophotrochozoans
10 10
Cnidarians
Lophotrochozoans
10
Cnidarians
Lophotrochozoans


let 7 133
1 137
7 153
8 184
9 190
22 193
29 210
31 216
33 219
Bm 67 277 750
2 76 279 1175
12 87 317 1993


217




126
135
155
Ecdysozoans Ambulacrarians Cephalochorates Vertebrates

Let7 133
1 137
7 153
8 184
9 190
22 193
29 210
31 216
33 219
34 242
71 252
79 278
92 281
96 315
124 365
125 375
2001
Bm 67 277 750
2 76 279 1175
12 87 317 1993


217




126
135
155
Ecdysozoans Ambulacrarians Cephalochorates Vertebrates

Let7 133
1 137
7 153
8 184
9 190
22 193
29 210
31 216
33 219
34 242
71 252
79 278
92 281
96 315
124 365
125 375
2001
Bm 67 277 750
2 76 279 1175
12 87 317 1993


217




126
135
155
Ecdysozoans Ambulacrarians Cephalochorates Vertebrates

Figure 15.7 Primitive repertoire versus secondary loss of miRNAs. Shown are three taxa with reduced complements of miRNAs, the acoel flatworm Symsagittifera roscoffensis (Sempere et al., 2007; Wheeler et al., 2009); the nematode Caenorhabditis elegans (Ruby et al., 2006), and the ascidian urochordate Ciona intestinalis (Norden-Krichmaer et al., 2007). The miRNAs found in each taxon are shown in a grey box; the ones in black are known to characterize that particular node based on extensive comparative analyses (Wheeler et al., 2009). The 37 miRNA families known to characterize vertebrates (Heimberg et al., 2008) are not shown; only the three shared between vertebrates and
ascidians. Note that, unlike the acoel flatworm, the nematode and the ascidian, while missing many primitive miRNAs, do possess protostome or chordate-specific miRNAs, respectively. In fact, the ascidian is grouped as the sister taxon to the vertebrates given that it shares three miRNA families with them (Heimberg et al., 2008). The nematode is clearly a protostome, but cannot (at the moment) be allied with the ecdysozoans based on miRNAs, as hypothesized by numerous other data sets.




15.4 miRNAs in organisms with fast molecular evolution and frequent gene loss

We wish to take a moment to emphasize that miRNAs are not immutable, but are components of genomes that will experience some of the same processes that affect other components, especially when there is a high rate of gene loss and/or a high substitution rate. Nonetheless, the pattern that emerges from such instances will still allow an investigator to draw accurate if imprecise conclu- sions concerning the taxon’s phylogenetic position. Ascidian urochordates, nematode worms, and acoel flatworms are all characterized by high rates of molecular evolution, and both nematodes (Copley et al., 2004) and ascidians (Hughes and Friedman,
2005) are further characterized by large amounts of secondary gene loss. Both nematodes and ascidians have taken a phylogenetic ‘bump up’ recently with nematodes going from basal bilaterians to near relatives of arthropods (Aguinaldo et al., 1997), and ascidians moving from basal chordates to the sis- ter taxon of vertebrates (Delsuc et al., 2006). Acoels, on the other hand, have followed the opposite
phylogenetic trajectory—they were originally included within the Platyhelminthes, but mul- tiple studies with different genes have suggested that acoels and nemertodermatids form a grade or clade at the base of Bilateria (recently reviewed in Baguñà et al., 2008).
Because all three of these hypotheses are con- troversial, they serve as useful test cases for the utility of miRNAs. Figure 15.7 shows the phylogen- etic distribution of miRNAs in the acoel flatworm Symsagittifera roscoffensis (Wheeler et al., 2009), the nematode C. elegans (Ruby et al., 2006), and the ascidian Ciona intestinalis (Norden-Krichmaer et al.,
2007). In stark contrast to both the nematode and the ascidian, the acoel possesses only a subset of the bilaterian set of miRNAs, and no miRNAs that characterize higher clades, namely Protostomia or Platyhelminthes (Figure 15.7, left; see also Sempere et al., 2007). Clearly, if acoels were in fact Platyhelminthes, or even within the protostomes or deuterostomes as suggested by recent EST studies (Philippe et al., 2007; Dunn et al., 2008) and have sec- ondarily lost miRNAs so as to artificially appear to be basal bilaterians, then they have lost miRNAs in an extremely unlikely pattern. The position

168 AN I M AL EV O L UTI O N



recovered by Dunn et al. (2008), which also corres- ponds to the traditional morphological hypothesis that allies them with the Platyhelminthes, would require the loss of 26 nephrozoan-specific and
12 protostome-specific miRNAs, in addition to some unknown set of platyzoan-specific miRNAs. The miRNA complement of C. elegans is well
known from deep 454 sequencing (Ruby et al.,
2006), and it is clear that it has lost a number of
miRNAs (Figure 15.7, centre), as it possesses just
over half of the reconstructed repertoire of the
ancestral bilaterian miRNA-family complement
(19 of 34). Nonetheless, they are clearly not basal
bilaterians as they also have over half of the
protostome-specific miRNAs as well (6 of 12). The
ascidian C. intestinalis has also lost many miRNA
families (Figure 15.8, right), as it only possesses
14 of the 34 miRNAs families present in the last
common ancestor of protostomes and deuteros-
tomes, but it also has the chordate-specific miRNA
miR-217, and three miRNAs otherwise found only
in vertebrates (Heimberg et al., 2008), which is
entirely consistent with the hypothesis that they,
and not cephalochordates, are the sister taxon to
the vertebrates (Delsuc et al., 2006). Thus, it is these
mosaic patterns of miRNA gene loss that charac-
terize a secondary reduction in terms of miRNA
content from primary absence and distinguishes
nematodes and ascidians from acoels (Sempere
et al., 2007).
There is a good reason for suspecting high gene
loss and/or high rates of sequence evolution in
miRNA genes in these two taxa. The presence and
sequence constraint of miRNA is probably dictated,
to a large degree, by targeting numerous mes-
senger RNA gene products. Several studies have
shown that miRNAs can regulate up to hundreds
of protein-coding genes (Lim et al., 2005; Baek et al.,
2008; Selbach et al., 2008), and because miRNAs
regulate so many different transcripts, and must
functionally interact with the 3c UTR of all targets,
it is difficult to lose the gene or change the pri-
mary sequence. Nematodes and ascidians have
both lost a considerable fraction of their protein-
coding genome, and consequently, in nematodes at
least, each miRNA probably regulates, at best, only
one or a few protein-coding genes (Ambros and
Chen, 2007). This allows for individual miRNAs
to be more easily lost when their target messenger RNA is lost, or else to track the target site without being constrained by other targets. Thus, if a taxon is known to have a high proportion of secondary gene losses, it is likely that there will also be a rela- tively high number of missing and/or unrecog- nizable miRNAs. The pattern should nevertheless be both mosaic and random still allowing for an accurate (but possibly imprecise) placement on the metazoan tree of life.


15.5 Returning to the lophotrochozoan problem . . .

Ultimately, taxa such as nematodes are exceptions—most organisms have not experienced drastic gene losses. We hypothesize that lophotro- chozoans in particular, which show little secondary gene loss and very little secondary modifications to their genomes (Tessmar-Raible and Arendt, 2003; Raible et al., 2005), will make a near-perfect test case for miRNA phylogenetics. Returning to the problem introduced earlier in the paper, the inter- relationships among nemerteans, annelids, and molluscs with respect to arthropods (Figure 15.1), the miRNAs are unequivocal. Both Eutrochozoa and Neotrochozoa are monophyletic, as nemerte- ans share with annelids and molluscs three unique miRNA families, one of which is the star sequence of miR-958, and annelids and molluscs share two miRNA families not found in nemerteans or any other taxon, one of which is the star sequence of an ancient miRNA family miR-133 (Wheeler et al.,
2009). Further, nemerteans do not share any miR- NAs with either the annelids or the molluscs to the exclusion of the other, nor do they share with annelids and molluscs second copies of miR-10 and miR-22. Thus, among the three possible arrange- ments of these three taxa (Figure 15.1) miRNAs support the topology derived from morphological and embryological considerations (Figure 15.1a) (Peterson and Eernisse, 2001).


15.6 Methodology for miRNA
phylogenetics

Because the primary sequence of the mature sequences of miRNAs is so fundamentally

MIC R ORN A S 169



conserved, miRNA phylogenetics is essentially a binary system, involving simply the presence or absence of given miRNAs in different organ- isms. miRNAs can be identified as present in an organism by bioinformatic searches in genomes, Northern analysis, or sequencing of libraries tar- geting the products of Dicer cleavage, usually with new high-throughput sequencing technologies (e.g. Wheeler et al., 2009). The literature on discov- ery and validation of miRNAs is too large to cover in this chapter, and we point the reader towards the general reviews cited above as a starting point to this literature, as well as to Ambros et al. (2003), which explains the requirements for annotation in miRBase, the online miRNA repository (Griffiths- Jones et al., 2006).
With the exception of the phylogenetic position of acoels (Sempere et al., 2007; Wheeler et al., 2009), the utility of miRNAs as phylogenetic characters has not been fully tested. The main problem for miRNA-based phylogenetics is in positively dem- onstrating absence, which (needless to say) is far more difficult than demonstrating presence, espe- cially in organisms without sequenced genomes. miRNAs with low expression levels will be hard to detect with both Northern analysis and librar- ies. As an extreme example, the miRNA lys-6, which was discovered in C. elegans by genetic screens (Johnston and Hobert, 2003) is expressed in fewer than 10 cells, and consequently has yet to be found in small RNA libraries even by extremely deep sequencing (Ruby et al., 2006). Furthermore, reaction kinetics for Northern analyses indicate that only a few base changes will result in non- detection (Sempere et al., 2006; Pierce et al., 2008), so a negative result could be the result of a few nucleotide changes as opposed to an absence of the gene product.
Although the absence of a gene can never be proved, there are several relatively straightforward ways strongly to suggest absence. First, studies should strive to sequence libraries deeply enough to provide some confidence that the absence is real. The cost of next-generation sequencing tech- nology is dropping quickly and as this happens the ability to sequence many organisms deeply and obtain a near complete understanding of their
miRNA complement will be possible. Second, until the methodology of miRNA-based phylogenetics is more fully developed, studies should focus on understanding the relationships of non-genomic organisms to genomic organisms, rather than tack- ling questions where there are no genomes avail- able at all. Genomes are never completed in the true sense of the word, but absence in both a finished genome and in a small RNA library is extremely unlikely to be a false negative. Working in a con- text where at least some of the organisms have sequenced genomes allows for the demonstration that the putative miRNA folds correctly in at least some of the organisms, precluding the possibility that the shared library reads are a degraded and highly conserved fragment of another gene. For non-genomic organisms, it is experimentally feas- ible to amplify miRNA loci using genome walk- ing from the taxon of interest to demonstrate the necessary structural features (Wheeler et al., 2009). The third, and probably most important, approach, much like any other form of phylogenetic infer- ence, is taxon sampling. Studying more than one organism per clade of interest (especially for library construction) has the benefit of sampling two dif- ferent transcriptomes which are likely to differ in their miRNA expression levels but which will help establish the polarity of individual characters and distinguish synapomorphies from plesiomor- phies. As an example, miR-750 was reconstructed as a lophotrochozoan-specific miRNA by Sempere et al., (2007), but the presence of miR-750 in the small RNA library of the priapulid Priapulus caudautus (Wheeler et al., 2009) demonstrated that this was in fact a protostome-specific miRNA that had been lost at the base of Insecta (Wheeler et al., 2009; see Figure 15.5).


15.7 Conclusions

We see two main advantages to miRNA-based phylogeny. First, as demonstrated in Sempere et al., (2006) and also in our Figure 15.5, they are applic- able over a wide range of phylogenetic scales, from species divergences within a genus to phylum-level relationships at the base of Bilateria. Second, the constraints on miRNA structure means that any

170 AN I M AL EV O L UTI O N



taxon can be queried for its complement of miRNAs, without having prior knowledge of a single miRNA sequence, simply by building a small RNA library. Given that miRNAs are continually added over time, rarely change in primary sequence, and are only rarely secondarily lost, they are potentially the near homoplasy-free data set that systematists have long wished for, and one that can be used to resolve the interrelationships among eumetazoan taxa at virtually any hierarchical level.
15.8 Acknowledgements

KJP would like to thank the National Science Foundation for funding and T. Littlewood and M. Telford for the invitation to contribute to the symposium. EAS would like to thank the Lerner- Gray fellowship from the American Museum of Natural History, the Systematics Research Fund of the Systematics Association and the Yale Enders Fund for funding. We thank D. Pisani and S. Smith for helpful suggestions on the manuscript.

The animal in the genome: comparative genomics and evolution

Comparisons between completely sequenced metazoan genomes have generally emphasized how similar their encoded protein content is, even when the comparison is between phyla. Given the manifest differences between phyla and, in par- ticular, intuitive notions that some animals are more complex than others, this creates something of a paradox. Simplistic explanations have included arguments such as increased numbers of genes, greater numbers of protein products produced through alternative splicing, increased numbers of regulatory non-coding RNAs, and increased com- plexity of the cis-regulatory code. An obvious value of complete genome sequences lies in their ability to provide us with inventories of such components. Here I examine progress being made in linking genome content to the pattern of animal evolu- tion, and argue that the gap between genome and phenotypic complexity can only be understood through the totality of interacting components.

Deus ex machina: ‘A power, event, person, or thing that comes in the nick of time to solve a difficulty; providen- tial interposition . . .’
Oxford English Dictionary


14.1 Introduction

Complete genome sequences provide limits to our imaginations. Even just a few years before the human genome was available in rough draft form, it was widely believed to encode at least 50,000 genes (Fields et al., 1994; Editorial, 2000). In contrast, the initial publications estimated 25,000–40,000
protein-coding genes (Lander et al., 2001; Venter et al., 2001), and since then estimates have generally carried a downward momentum, most recently approaching 20,000 (Goodstadt and Ponting, 2006; Pennisi, 2007). Although this number is higher than the 16,000 or so found in invertebrate chord- ates (Dehal et al., 2002) it is basically the same total as in the nematode worm Caenorhabditis elegans (Hillier et al., 2005). Whether or not these low num- bers of protein-coding genes for vertebrates stand the test of time, the sense of unease surrounding the lack of correlation between organismal com- plexity (often measured in number of distinct cell types) and protein-coding gene count is evident from the framing of the ‘g-value paradox’ by Hahn and Wray (2002), and the various explanations that have been put forward to ease it, including, for example, miRNAs (Sempere et al., 2006, Heimberg et al., 2008), non-protein-coding DNA (Taft et al.,
2007), and alternative splicing (Kim et al., 2007). Similar gene counts are, of course, a crude meas-
ure of biological complexity. There is no reason why two genomes should not encode very different sets of protein-coding genes, but still have similar over- all totals. Within the field of animal evolution and the evolution of development (evo-devo), however, the g-value paradox has a particular resonance. Studies in different animal phyla have repeatedly shown the reuse of a core set of developmental genes, the so-called ‘toolkit’ (Carroll et al., 2005), with the HOX genes, in particular, taking on an iconic significance. Broadly, toolkit genes come from a handful of transcription factor families, defined by the presence of particular structural domains


148

CO M P A R A T I V E G E N O M I C S 149



such as the helix–turn–helix (HTH the class that includes, the homeobox genes), zinc fingers (ZnF), leucine zippers, and the helix–loop–helix (HLH). As well as transcription factors there are seven well- conserved pathways responsible for intercellular signalling (Pires-daSilva and Sommer, 2003), many of which appear to be present in sponges, the earli- est branching clade of animals (Nichols et al., 2006). An extreme interpretation of these data is provided by Davidson (2006): ‘if we focus explicitly on the genes encoding transcription factors, and [ . . . ] sig- nalling systems required for developmental spatial regulation, there is almost no qualitative variation among the genomes of bilaterians’.
Given all this, where in the genome do the phenotypic differences between animal taxa arise? The undoubted conservation of the protein-coding developmental genes has, particularly in the evo- devo field with its morphological concerns, focused attention on cis-enhancer elements affecting tran- scription (Carroll et al., 2005; Davidson, 2006; Wray,
2007; Simpson, 2007), although there are alternative views emphasizing the importance of different kinds of regulatory elements (Alonso and Wilkins,
2005) and different protein classes, such as struc- tural genes (Hoekstra and Coyne, 2007). Below I outline some major themes being developed by large-scale genome comparisons, principally of nematodes, insects, and vertebrates. My aim is not to present an exhaustive account, but to highlight areas where functionally relevant species-specific differences may arise, within apparently conserved systems. Although I concentrate on the evolution of the systems regulating animal development, this is not to lose sight of the things being regu- lated: the proteins involved in making nematode cuticles, or asynchronous flight muscles in insects, or the human brain and adaptive immune system, to name but a few, are what made it necessary to evolve those systems.


14.2 Gene duplication

Usefully summarizing the differences and similarities between more than 10,000 protein- coding genes from several species at once is not necessarily straightforward. Although pair- wise similarities between sequences are easy to
compute, they suffer from the imposition of arbi- trary cut-offs and are less easy to interpret than measures that explicitly reflect phylogeny. Genes in different species are most obviously compared by grouping into sets of orthologues (that is, genes related by speciation events) and paralogues (genes related by intragenome duplication events). Closely related species share large numbers of orthologues: 93% of dog (Canis familiaris) and 82% of the marsupial Monodelphis domesticus gene pre- dictions have orthologues in human (Goodstadt et al., 2007). The Linnean hierarchy, however, is not necessarily a good guide of genomic relatedness by this definition of similarity. Within the nematodes
65% of C. elegans genes share an orthologue with Caenorhabditis briggsae, despite their being from the same genus (Stein et al., 2003). For more distantly related genomes, orthologue counts can drop rap- idly. This may be as much a sign of difficulties in reliably assigning gene phylogenies on a large scale as a real indication of the extents of the con- served cores.
Paralogues often arise via tandem duplication of genes, giving rise to localized clusters of func- tionally related genes. As these are the regions where gene content is evolving most rapidly between closely related species, the functions of these genes are of special interest for understand- ing animal-specific differences. For the most part, for any two closely related vertebrate genomes the functional classes of genes duplicated in this way are similar—olfaction and chemosensation, repro- duction, and effectors of the immune response— although the duplications have occurred independently in each lineage (Emes et al., 2003). These large groups of paralogues often show evi- dence of adaptive evolution in their amino acid sequences, suggesting that new functions have been selected for (Emes et al., 2004a,b).
The recurrent nature of duplications within particular functional classes, coupled with the observed diversifying selection, suggests that they are a standard adaptive genomic response to environmental challenges. Does similar rapid duplication occur in the kinds of genes, such as transcription factors, that might be implicated in development? A growing number of examples are known. Perhaps most dramatically, in mice a set

150 AN I M AL EV O L UTI O N



of 32 tandemly duplicated homeoboxes have arisen from apparently one or two genes in the common ancestor of humans and rodents; they are believed to play a role in germ cell development and embry- onic stem cell differentiation (Maclean et al., 2005; Jackson et al., 2006).
Zinc finger-containing transcription factors have undergone independent rounds of gene duplication in insects and tetrapods. In insects a set of zinc fingers are found to co-occur with a zinc finger associated domain (ZAD) (Chung et al., 2007); this ZAD class is found in around
100 copies in Drosophila melanogaster and and 150 copies in the mosquito Anopheles gambiae; there is only a single copy in vertebrates (Chung et al.,
2007). In D. melanogaster, many are expressed in the female germline, suggesting a role in oocyte development or embryogenesis (Chung et al.,
2007). An analogous story is found with Krüppel associated box (KRAB) containing zinc fingers in tetrapods. Successive independent tandem dupli- cation events have occurred in different mam- malian lineages, leading to more than 400 copies in the human genome (Huntley et al., 2006). The KRAB domain itself appears to have been co- opted from a progenitor sequence conserved throughout eukaryotes (Birtle and Ponting, 2006); however, it has evolved so much as to make this similarity difficult to detect; clearly identifi- able KRAB domains are specific to tetrapods. Their functions are largely unknown, and have not been tied to any general aspects of tetrapod biology. As such, why the family as a whole has expanded is a puzzle.
Nematodes too exhibit lineage-specific expansions of particular transcription factor families, most not- ably the nuclear hormone receptors (NHRs). The C. elegans genome encodes 284, many more than the 48 in human and 21 in D. melanogaster. The bulk of these (> 200) have arisen from an apparently nematode- specific expansion of a unique gene (Lander et al.,
2001, Robinson-Rechavi et al., 2005). Once more, the reasons for such a dramatic lineage-specific expan- sion of a particular transcription factor family, and any links to taxon-specific biology, are obscure, although it has been speculated that C. elegans relies less on combinatorial reuse of different transcrip- tion factors (Antebi, 2006). A less dramatic lineage-
specific expansion occurs in the case of the T-box-containing transcription factors: there are 21 in C. elegans, with 17 arising from a lineage-specific expansion when compared with D. melanogaster and humans.Ascertainingwhenandinwhichcladesthese C. elegans duplications took place is currently frus- trated by a lack of relevant genome sequences. As a set these T-box genes map to several genomic loca- tions, suggesting that they have arisen over a more protracted timescale than the examples discussed above; some, at least, have known roles in the devel- opment of C. elegans (Poole and Hobert, 2006).


14.3 The ‘invention’ of new genes

A number of gene families appear to be meta- zoan novelties, with no clear sequence similarity to other genes outside the Metazoa, but present in the more basal animal phyla such as cnidarians and sponges. These include key families involved in animal development, like T-box and SMAD transcription factors, and signalling molecules such as WNTs and FGFs (Putnam et al., 2007). The most closely related non-metazoan eukaryote sequenced to date, the choanoflagellate Monosiga brevicollis, was reported to be missing true HOX, ETS, NHR, POU, and T-box class transcription fac- tors, strongly suggesting their origin was co-inci- dent with that of the metazoans (King et al., 2008). Analysis of preliminary data from the sponge genome indicates that, although present, many of these gene families were much smaller in number prior to the divergence of sponges and cnidarians (Larroux et al., 2008).
Was the invention of such families a prerequisite for the evolution of the Metazoa, and were analo- gous protein inventions required for the evolution of particular taxa, such as insects and vertebrates? At the level of three-dimensional structures (i.e. the protein fold itself), there is some reason to be sceptical that this is the case. In many cases, exam- ination of similarities in three-dimensional pro- tein structures shows that these genes have distant homologues in non-metazoan genomes. The MH1 (DNA-binding) domain of SMADs, for instance, is probably homologous to a family of homing endo- nucleases found in all kingdoms of life (Grishin,
2001); the T-box shares structural similarities

CO M P A R A T I V E G E N O M I C S 151



indicative of homology with a variety of other transcription factors, such as STAT DNA-binding domains, which are found in other eukaryotes (Murzin et al., 1995; Soler-Lopez et al., 2004); and the signalling domain of metazoan hedgehog proteins shares detailed similarities with members of a fam- ily of bacterial peptidases, suggesting that they too are likely to be homologous (Murzin et al., 1995). In these cases the novel families are likely to be cases of rapid sequence evolution, accompanying func- tional shifts, within stem lineages leading to the Metazoa. Sparse sequence sampling of non-fungal and metazoan eukaryotic genomes may contribute to the apparent co-origin of these protein domains with the animals.
As this type of domain evolution is occurring from pre-existing domain types, the process fits within a standard framework of accelerated point mutation and selection for new functions. The invention of the domain type is not a key innov- ation in itself; rather, it can be seen as the exten- sion of functional diversification of subfamilies of the kind that is apparent when comparing more closely related species. The fact that so many new domain types are found to be co-incident with the origin of metazoans suggests that the selective
pressures giving rise to this kind of accelerated sequence evolution were greater in the metazoan stem lineage.
An example of a more recent domain innovation is found in the Drosophila gene brinker, which plays a key role in the establishment of dorsoventral pat- terning. Although the protein-coding sequence of its DNA-binding domain is well conserved in insects, using current sequence data bases it shows no significant sequence similarity to proteins from any other taxa (Figure 14.1 and Plate 9), although there is weak (non-significant) similarity to pogo- like transposases, and the structure, which is only folded when complexed with DNA, suggests simi- larity to various transcription factors (Cordier et al.,
2006).


14.4 Evolution of transcription factors:
the animal in the orthologue

Lineage-specific duplication followed by sequence divergence provides one route to species-specific biology, but what scope is there for lineage-specific functional shifts within orthologous genes? In the absence of gene duplication, it is hard to imagine how the DNA specificity of a particular factor



a. b.
A.pisum
B mori
A mellifera
N.vitripennis
P.humanus
D mojavensis
D melanogaster
D.pseudoobscura
D.ananassae
D.erecta
D.yakuba
D.sechellia
D.simulans
D.grimshawi
D.virilis
T.castaneum
C.pipiens
A.aegypti
A.gambiae CCLHKTYHAHSLLSVLDSYRQDSDCQGNQRATARKYGIHRRQIQKWLQTE AGSRRIFPPQFKLQVLEAYRRDSQCRGNQRATARKFGIHRRQIQKWLQAE MGSRRIFAPAFKLKVLDSYRNDIDCRGNQRATARKYGIHRRQIQKWLQCE MGSRRIFAPAFKLKVLDSYRKDIDCRGNQRATARKYGIHRRQIQKWLQCE VGSRRIFSPHFKLQVLDSYRYDADCRGNQRATARKYNIHRRQIQKWLQCE MGSRRIFTPQFKLQVLESYRHDNDCKGNQRATARKYNIHRRQIQKWLQCE MGSRRIFTPHFKLQVLESYRNDNDCKGNQRATARKYNIHRRQIQKWLQCE MGSRRIFTPHFKLQVLESYRNDNDCKGNQRATARKYNIHRRQIQKWLQCE MGSRRIFTPHFKLQVLESYRNDNDCKGNQRATARKYNIHRRQIQKWLQCE MGSRRIFTPHFKLQVLESYRNDNDCKGNQRATARKYNIHRRQIQKWLQCE MGSRRIFTPHFKLQVLESYRNDNDCKGNQRATARKYNIHRRQIQKWLQCE MGSRRIFTPHFKLQVLESYRNDNDCKGNQRATARKYNIHRRQIQKWLQCE MGSRRIFTPHFKLQVLESYRNDNDCKGNQRATARKYNIHRRQIQKWLQCE MGSRRIFTPQFKLQVLESYRNDNDCKGNQRATARKYNIHRRQIQKWLQCE MGSRRIFTPQFKLQVLESYRNDNDCKGNQRATARKYNIHRRQIQKWLQCE IGSRRIFAPHFKLQVLDSYRNDADCKGNQRATARKYGIHRRQIQKWLQVE MGSRRIFTPQFKLQVLDSYRNDSDCKGNQRATARKYGIHRRQIQKWLQVE MGSRRIFTPQFKLQVLDSYRNDSDCKGNQRATARKYGIHRRQIQKWLQVE MGSRRIFTAQFKLQVLDSYRNDGDCKGNQRATARKYGIHRRQIQKWLQVE
Consensus/90% hGSRRIFss.FKLpVL-SYRpD.DC+GNQRATARKYsIHRRQIQKWLQsE

Figure 14.1 The DNA-binding domain of brinker is conserved within insects, but has no significantly similar sequences in other taxa.
(a) The alignment shows the conserved core from a selection of insect species. Sequences of Drosophila species were taken from the UCSC
web browser (http://genome.ucsc.edu/), Anopheles and Aedes from ENSEMBL (http://www.ensembl.org/), other predictions were made
from sequences at the NCBI. GI accession numbers: N. vitripennis 146253130; T. castaneum 73486274; C. pipiens 145464888; P. humanus
145365328; A. mellifera 63051942; B. mori 91842977; A. pisum 47522326. (b) The three-dimensional structure of the aligned region when binding DNA. The structure was taken from the PDB file 2glo. (See also Plate 9.)

152 AN I M AL EV O L UTI O N



might be significantly changed, in such a way that it targets new genes, without deleterious con- sequences. The modular structure of proteins, however, suggests that other routes of functional evolution are available. A protein may have pleio- tropic effects, but that is not the same as saying that every amino acid in the protein will be dir- ectly involved in all those effects. A recent illustra- tive example from the hox gene Ultrabithorax, is of an insect-specific ‘QA’ protein motif, found outside the homeodomain. The region is involved in limb repression; the effects of deleting the motif are strong in some tissues but close to undetectable in others (Hittinger et al., 2005). Clearly, changes in the protein-coding sequences of transcription fac- tors, apart from their more obvious DNA-binding residues, must be integrated into our understand- ing of the evolution of developmental regulation (Wagner and Lynch, 2008).
The majority of residues in metazoan transcrip- tion factors do not fall within regions of well-defined globular structure, with many belonging to so- called ‘intrinsically disordered’ regions—regions that may form a structure when complexed with other macromolecules (J. Liu et al., 2006, Minezaki et al., 2006). The specific sequences of these regions are typically not obviously conserved between par- alogues; because they are unique to particular fam- ilies they are not covered in domain data bases such as SMART and Pfam (Finn et al., 2006; Letunic et al.,
2006). The lack of extreme conservation between distant species has sometimes masked the fact that, within closely related species, these regions are conserved. Comparisons of orthologous sequences from closely related genomes (e.g. vertebrates or drosophilids) often show that substantial propor- tions of these non-domain sequences are undergo- ing strong purifying selection—they accumulate many more synonymous nucleotide changes than non-synonymous changes—and are thus func- tional. For the large part, precisely what these bio- logical functions are is unknown; two possibilities, however, suggest themselves. Firstly, they may have relatively uninteresting non-specific effects, such as facilitating folding of the major domain (for instance by reducing aggregation) or acting as spacers between globular domains. Secondly, and more interestingly from the point of view of
animal evolution, they may include short linear peptide motifs that mediate protein–protein inter- actions (Dyson and Wright, 2005; Neduva et al.,
2005; Neduva and Russell, 2005).
There are numerous examples of regula-
tory motifs found outside of transcription factor
domains. Many hox proteins include a YPWM-like
hexapeptide motif that interacts with other home-
odomain-containing proteins (In der Rieden et al.,
2004); Drosophila fushi tarazu ( ftz) orthologues have
lost this motif but acquired an LXXLL motif cou-
pled to a new role in segmentation (Lohr and Pick,
2005); and an N-terminal SSYF-like motif believed
to be involved in transcriptional activation is con-
served across hox orthologues and paralogues
from different phyla (Tour et al., 2005). Interaction
motifs can be coupled with signalling pathways to
create cell-type specificity. They can, for instance,
be regulated by phosphorylation, such that the
phosphorylation status governs what interactions
can be made (e.g. Sapkota et al., 2007), or alternative
splicing can result in protein–protein interaction
motifs being included or excluded from particular
cell types, providing additional layers of regula-
tory complexity that are likely to be species specific
(Neduva and Russell, 2005).
The challenge of identifying small regulatory
motifs means that their species distributions, and
how their presence might produce taxon-specific
differences in protein functions, have not been
well studied. Examples that tie cleanly to one taxo-
nomic group are less common, but an interesting
case has been proposed in bilaterian orthologues
of the Brachyury gene. These possess an N-terminal
motif that is not found in non-bilaterian Metazoa
(Marcellini, 2006), which instead have a well-defined
EH1-like motif (Copley, 2005). The bilaterian motif
is believed to be responsible for an interaction with
Smad1, and hence to link gastrulation to bilateral
pattern formation (Marcellini, 2006).


14.5 Enhancers: transcription factor binding sites and ultraconserved regions

Theoretical considerations have led to an intense focus on transcription factor binding sites (TFBSs) as a major molecular source of morphological novelty

CO M P A R A T I V E G E N O M I C S 153



(Wray et al., 2003; Carroll et al. 2005, Davidson, 2006; Wray, 2007; although see Hoekstra and Coyne, 2007, for a critique). Individual TFBSs show rapid turn- over in comparisons of closely related genomes, with many being lineage specific (Dermitzakis and Clark, 2002; Moses et al., 2006). This dynamic nature may not be revealed in the phenotype— patterns of gene expression may be conserved even though regulatory sequences change at the molecular level (Ludwig et al., 2000; Romano and Wray, 2003; Fisher et al., 2006). On the other hand, the gain and loss of individual TFBSs has been implicated in several recent cases of morphological evolution, in both vertebrates and invertebrates (reviewed in Wray, 2007, and Simpson, 2007). The relationship between individual TFBSs and enhan- cer function is clearly not straightforward, beyond the fact that clustering of individual binding sites can identify some enhancer regions (Markstein et al., 2002). Cases of functional linkages between particular transcription factors have been pro- posed, for example, between Dorsal, twist, Su(H), and an unidentified motif in neurogenic ectoderm formation in Diptera (Markstein et al., 2004), and even a coupling originating prior to the origin of Bilateria, of hairy and E(spl), promoting neural cell fate (Rebeiz et al., 2005).
Comparisons of vertebrate genomes have revealed large regions (more than 100 nucleotides) of extreme conservation of non-coding sequences (conserved non-coding elements, CNEs) (Bejerano et al., 2004). These regions are often found near transcription factors and other developmental genes (Sandelin et al., 2004). Outside of the verte- brates, there is evidence for similar regions occur- ring near developmental genes in flies (Glazov et al., 2005) and nematodes (Vavouri et al., 2007). Although in many cases the conserved regions are even found near orthologous genes, there is no evidence that they are homologous; they appear to have evolved independently in each of the phyla (Vavouri et al., 2007). Experimental evidence from vertebrates shows that many instances have roles as tissue-specific enhancer elements (Woolfe et al.,
2005, Pennacchio et al., 2006).
The length, and lack of interphyla conserva-
tion of CNEs is in contrast to individual TFBSs.
The DNA specificity of orthologous transcription
factors is usually well conserved over large phylo- genetic distances, but typical TFBSs are short, of the order of six to ten nucleotides. An obvious pos- sibility is that longer CNEs are composed of over- lapping or adjacent TFBSs. This would suggest a tight packing of transcription factor proteins on the genomic DNA of these CNEs. There is direct evidence for this: some fragments of highly con- served non-coding sequences are present in crys- tal structures of transcription factor complexes. An atomic model based on known crystal structures of the interferon-E enhancer, for example, shows
50 consecutive nucleotides in contact with eight different proteins; these nucleotides are well con- served in mammalian species (Panne et al., 2007; see Figure 14.2 (also Plate 10) for another example). Given that such structures exist, it is not such a leap to imagine 16 proteins binding to 100 nucle- otides, or even bigger complexes. This suggests a model where CNE enhancer regions controlling orthologous genes in different phyla are controlled by multiple transcription factor binding sites, although not necessarily the same transcription factors or in the same orientation. Moreover, the tight packing of transcription factors on the gen- omic DNA suggests that the proteins themselves may be co-adapted to interact with each other and aid the cooperative formation of enhancer com- plexes. Previously, Ruvinsky and Ruvkun (2003) have presented experimental evidence that enhanc- ers and transcription factors co-evolve in this way, with neuronal and muscle-specific enhancer elem- ents from D. melanogaster failing to drive expres- sion in homologous tissue types in C. elegans, and Dover and co-workers (McGregor et al., 2001; Shaw et al., 2002) have argued for co-evolution of bicoid protein and hunchback regulatory regions. Wagner (2007) has proposed that the protein–protein inter- actions of co-adapted transcription factors may form the underpinnings of ‘character identity net- works’; that is, the gene regulatory networks that control the development of homologous morpho- logical characters.
If protein–protein interactions between transcrip- tion factors are often required for the formation of enhancer complexes, close analysis of transcription factor sequence and structure may reveal evidence for co-adaptation of proteins, such as the HOX

154 AN I M AL EV O L UTI O N



CEBP

BRLZ




CEBP

BRLZ




PFAM: Runt
RUNX1

Pfam: RunxI














Human GCAACCACAGAGTTTGGAAATCTT Chimp ........................ Rhesus ........................ Mouse G...T.....A.........A.. Rat G...T...............A.. Rabbit ........................ Dog ........................ Cow ........................ Elephant ..........T..........-.. Tenrec ..........G.............

Figure 14.2 Adjacent transcription factor binding sites cause extended regions of DNA sequence conservation. Structure of CEBPE homodimer and RUNX1 (Tahirov et al., 2001). Three transcription factors (2× CEBPE and RUNX1) bind in a region of 25 nucleotides conserved throughout placental mammals. The DNA-binding domains represented as three-dimensional structures are boxed and colour- coded in the schematic representation of the proteins. In each case, the majority of the protein is not represented in the structure; these regions could interact with other transcription factors, activators, and repressors. The human sequence coordinates are chromosome 5, bases
149,446,373–149,446,396 of the NCBI build 36. The alignment is taken from the UCSC web browser (http://genome.ucsc.edu/). (See also
Plate 10.)




hexapeptide motif, through which homeotic pro- teins form complexes with TALE class homeodo- mains (LaRonde-LeBlanc and Wolberger, 2003). We might expect instances of co-adapted transcription factor combinations to be taxon specific, to match the taxon specificity of enhancer sequences.


14.6 Alternative splicing

Not all CNEs are associated with enhancer regions.
There is good evidence that many are involved
in regulating alternative splicing events, includ- ing the alternative splicing of mRNAs of proteins which themselves regulate alternative splicing (Lareau et al., 2007; Ni et al., 2007). The presence of highly conserved control elements to regulate alternative splicing indicates that the functional consequences are of importance. Although large very conserved elements may be the exception rather than the rule, detailed comparative analyses have identified smaller conserved motifs regulat- ing alternative splicing, for instance in nematodes

CO M P A R A T I V E G E N O M I C S 155



(Kabat et al., 2006) and vertebrates (Sorek and Ast,
2003; Yeo et al., 2005).
Alternative splicing is often touted as a mechan-
ism by which proteomic complexity is increased.
Although early reports suggested that levels of
alternative splicing were comparable in vertebrates
and invertebrates (Brett et al., 2002), more recent
studies suggest that there is indeed more alterna-
tive splicing of transcripts in vertebrates (Kim et al.,
2007), suggesting a link with increased phenotypic
complexity. How relevant is alternative splicing for
species-specific biology and morphological differ-
ences? Quantitatively, the gene products that appear
to be most affected by alternative splicing are typic-
ally involved in the functioning of the nervous and
immune systems (Modrek et al., 2001). There are,
however, ample examples of alternatively spliced
transcription factors—as many as 63% of mouse
transcription factors have variant exons (Taneri
et al., 2004). Although the differences in molecu-
lar roles of the alternatively spliced products are
often unknown, the genes themselves include
developmental classics such as members of Hox,
SMAD, and T-box families (Noro et al., 2006; Fan
et al., 2004; Dunn et al., 2005;), although they do not
necessarily present obvious morphological corre-
lates (Yoder and Carroll, 2006). Alternative splicing
of modular proteins is an obvious route through
which functions can be changed by including or
excluding particular combinations of domains. In
this regard, it is interesting that alternative splicing
often affects intrinsically disordered regions out-
side known protein domains (Romero et al., 2006)—
this again points to a critical role for finely tuned
protein–protein interactions among transcriptional
regulators.
There are few known cases of distant conserva-
tion of alternative splice variants of transcription
factors; typically, examples are conserved within
phyla at best. Widening the search to other classes
of gene again suggests that splice variants are not
conserved over long periods, although it should
be remembered that transcript coverage of most
species, from which evidence of alternative spli-
cing is obtained, is very restricted. Perhaps the
best example is currently that of fibroblast growth
factor receptor 2 (FGFR2), where an exon configur-
ation diagnostic of mutually exclusive alternative
splicing is found in both vertebrates and the sea
urchin Strongylocentrotus purpuratus (Mistry et al.,
2003). Examples of orthologous ion channel encod-
ing genes showing similar alternative splicing pat-
terns in D. melanogaster, C. elegans, and humans are
likely to be cases of parallel evolution (Copley,
2004). The shared ability of vertebrates and at
least insects and C. elegans to produce alternative
transcripts in a regulated manner, alongside the
absence of large numbers of conserved alterna-
tive splicing between protostomes and deuteros-
tomes, suggests that gene products have become
alternatively spliced in parallel between different
lineages, while at the same time hinting that the
functions performed by alternative splice variants
may, over time, be replaced by different genomic
solutions.


14.7 Summary

Key genetic innovations, such as alternative spli- cing, the invention of hox genes or the advent of micro-RNAs, have held a strong appeal for those seeking to explain animal evolution in terms of genomes. Without denying the importance of such phenomena, a more nuanced outlook is preferable. Much of the molecular complexity found in ani- mals could have its origins in non-adaptive proc- esses attributable to small population sizes (Lynch,
2007a,b), but this complexity may then be exploited in the service of phenotypic adaptation (Lynch,
2007a) within a framework of point mutation and
selection.
Although most major classes of protein involved
in animal development may be conserved through-
out the Metazoa, detailed comparative analysis of
these gene types reveals a more dynamic picture,
with frequent gene duplication, gene loss, coup-
lings with new motifs, and other processes such as
alternative splicing and regulation by micro-RNAs,
all of which are likely to be important for a full
understanding of function. Cis-regulatory vari-
ation may well be revealed to be quantitatively the
most common form of variation between species,
but it seems likely that the cumulative effects of
multiple cis-regulatory changes will have required
that protein networks evolve to accommodate and
correctly regulate changed enhancer structures.

156 AN I M AL EV O L UTI O N



Our knowledge of animal evolution and the pic- ture presented here is currently based on a very small sampling of almost exclusively nematode, insect, and vertebrate genomes. Although this situ- ation is beginning to change, the fact that many important functional regions, especially those that
do not encode proteins, are only revealed by hav- ing sets of closely related genome sequences, and that there are 35 or so animal phyla, gives some idea of the huge scale of the challenges ahead. The rapidly falling costs of genome sequencing do, however, give grounds for optimism.

Beyond linear sequence comparisons: the use of genome-level characters for phylogenetic reconstruction

The first whole genomes to be compared for phylogenetic inference were those of mitochondria, which provided the first sets of genome-level char- acters for phylogenetic reconstruction. Most power- ful among these characters have been comparisons of the relative arrangements of genes, which have convincingly resolved numerous branching points, including some that had remained recalcitrant even to very large molecular sequence compari- sons. Now the world faces a tsunami of complete nuclear genome sequences. In addition to the tre- mendous amount of DNA sequence that is becom- ing available for comparison, there is also the potential for many more genome-level characters to be developed, including the relative positions of introns, the domain structures of proteins, gene family membership, the presence of particular bio- chemical pathways, aspects of DNA replication or transcription, and many others. These characters can be especially convincing because of their low likelihood of reverting to a primitive condition or occurring independently in separate lineages, so reducing the occurrence of homoplasy. The com- parisons of organelle genomes pioneered the way for the use of such features for phylogenetic recon- structions, and it is almost certainly true, as ever more genomic sequence becomes available, that further use of genome-level characters will play a big role in outlining the relationships among major animal groups.
13.1 Why do we need anything other than molecular sequence comparisons?

Over the past few decades, the comparison of nucleotide and amino acid sequences has revolutionized our understanding of evolution- ary relationships for many groups of organisms. The broader field of systematics has been reinvig- orated and a generation of evolutionary biologists have come to accept that molecular sequence com- parisons are an essential component for inferring the phylogeny of any group. These studies have led to extensive revision of animal systematics and to the overturning of previous reliance on features of the coelom and segmentation (Adoutte et al., 1999).
In the 1980s, when comparing molecular sequences for phylogenetic inference was first becoming common, some asserted with great con- fidence that all evolutionary relationships would soon be convincingly resolved solely with this type of data, leading to much consternation. However, some of the relationships that were equivocal in early molecular studies have remained highly recalcitrant even with many more DNA sequence data in hand. There are several potential explan- ations, including:

1. Multiple nucleotide or amino acid substitutions may have occurred at a single site, obscuring any accumulated signal.


139

140 AN I M AL EV O L UTI O N



2. Convergent or parallel substitutions may have occurred among different lineages due to having only four (for nucleotides) or 20 (for amino acids) possible character states, exacerbated by conver- gent biases in base composition (Naylor and Brown,
1998), which may even cause ever-increasing con- fidence measures for incorrect associations with ever larger data sets (Phillips et al., 2004).
3. The analysis may show artefactual association
of the more rapidly changing lineages (Felsenstein,
1978), including the attraction of long branches
to the base of the ingroup in association with the
outgroup (which is almost always a long branch;
Philippe and Laurent, 1998).
4. In some cases, non-orthologous gene copies may be inadvertently compared among various lineages due to ancestral gene duplications followed by dif- ferential losses, or due to incomplete sampling.
5. Differing views of scientists on alignments, exclusion sets, and weighting schemes frequently cannot be arbitrated based on objective criteria and can lead to radically different phylogenetic reconstructions.
6. The most difficult problems are when the time of shared ancestry is short relative to the subse- quent time of divergence, where there has been little opportunity to accumulate signal and ample time for it to have been erased.

Molecular sequence comparison is now a mature field that has influenced the culture of systematics. Many have come to expect that the future of sys- tematics will be dominated by creating ever more sophisticated methods for teasing a weak signal from noisy data. This causes concern that differing preferences for various methods will ensure that no consensus on many evolutionary relationships will ever be reached.
However, an alternative is possible, that there may be other, less explored, types of characters that could be powerful for resolving these conten- tious relationships. There is no doubt that com- parisons of some characters have identified certain robust synapomorphies (shared and derived char- acter states) that have supported long-standing, little-contested evolutionary relationships, such as the monophyly of mammals, tetrapods, and
echinoderms. These synapomorphies are sub- jectively judged to be of characters so unlikely to revert to an earlier condition or to occur mul- tiple times in parallel that they could only have arisen once in the common ancestor of the group. Can new sets of characters be found that would meet these criteria to provide confident resolution of some problematic evolutionary relationships? Although there is a broad range of character types to explore, we will focus here specifically on com- parison of features of genomes.

13.2 Comparisons of mitochondrial genomes have laid the foundation

Sequences from mitochondrial genes and genomes have been used extensively for phylogenetic infer- ence, with complete mtDNA sequences being pub- licly available for more than 1000 animal species. (For a summary of the characteristics of animal mtDNAs, see Boore, 1999.) It has been long-argued (e.g. Boore and Brown, 1998) that the relative arrange- ment of the (normally) 37 genes in animal mitochon- drial genomes constitutes an especially powerful type of character for phylogenetic inference, and so constitutes the first set of genome-level features to be used extensively for animal phylogeny. Briefly summarized, these genes are present in nearly all animal groups, are unambiguously homologous, and can potentially be rearranged into an enor- mous number of states such that convergent rear- rangements are very unlikely (and demonstrated to be uncommon). All genes on each strand are tran- scribed together in cases where it has been studied (Clayton, 1992), so selection on gene arrangements is expected to be minimal. A summary of the evo- lutionary relationships convincingly demonstrated by these types of data (and in many cases left unre- solved by all other studies) is found in Boore (2006), but here are a few of the more significant conclu- sions of deep-branch phylogenetic relationships: (1) the superphylum Eutrochozoa includes cestode platyhelminths (von Nickisch-Rosenegk et al., 2001) and the phylum Phoronida (Helfenbein and Boore,
2004); (2) Sipuncula is closely related to Annelida rather than to Mollusca (Boore and Staton, 2002); (3) Annelida is more closely related to Mollusca than to Arthropoda (Boore and Brown, 2000);

G E N O M E S A N D P H Y L O G EN Y 141


Table 13.1 URLs for the largest public DNA sequencing centres

DNA sequencing centre website


Wellcome Trust Sanger Institute http://www.sanger.ac.uk/ DOE Joint Genome Institute http://www.jgi.doe.gov/
Mammals 10 40
Birds 1 1

Washington University Genome
Sequencing Center
http://genome.wustl.edu/
Reptiles 1 1
Amphibians 1 1
Coelacanths 0 2

Broad Institute http://www.broad.mit.edu/
Bony fish 5 6

Baylor College of Medicine
Genome Center
http://www.hgsc.bcm.tmc.edu/
Cartilaginous fish 0 2
Jawless fish 0 2
Cephalochordates 1 0

Beijing Genomics Institute http://www.genomics.org.cn/
en/index.php
Riken Genomic Sciences Center http://www.gsc.riken.go.jp/ J. Craig Venter Institute http://www.jcvi.org/ Genoscope http://www.genoscope.cns.
fr/spip/
Urochordates 2 1
Hemichordates 0 1
Echinoderms 1 0
Mollusks 1 2
Flatworms 0 3
Annelids 2 0
Arthropods 22 23
Priapulids 0 1
Tardigrades 0 1

Nematodes
3 21


(4) Arthropoda is monophyletic and, within this phylum, Crustacea is united with Hexapoda to the exclusion of Myriapoda and Onychophora (Boore et al., 1995, 1998); (5) Pentastomida is not a phy- lum, but rather a type of crustacean, and joins with Cephalocarida and Maxillopoda to the exclusion of other major crustacean groups (Lavrov et al., 2004).


13.3 Nuclear genomes, a treasure-trove of phylogenetic characters

By a great margin, more DNA sequence is being generated than ever before. Facilities built and techniques developed for sequencing the human genome are now focusing on many other organ- isms. As recently as a year ago, the nine largest genome sequencing centres (see Table 13.1) collect- ively produced well over 170 billion nucleotides of DNA sequence per year: approximately 57-fold the coverage of the human genome. With next-genera- tion sequencing platforms now in regular use, that number is exploding. Imminently there will be complete genomes of at least draft quality for many dozens of animals representing a phylogenetically diverse sample and including several equivocally placed lineages (Figure 13.1, Table 13.2).
In these genomic data are many higher-order features, beyond the linear sequences, that consti- tute genome-level characters that are potentially
Cnidarians 1 1
Placozoans 1 0
Poriferans 0 1

Figure 13.1 This reconstruction of the major branches of
animal evolution is used to plot the numbers of taxa with complete genome sequences done and under way. The taxonomic ranks shown are arbitrary, split for illustration, but not meant to be consistent among the major groups, and taxa listed do not comprehensively cover all of life. Branch length holds no meaning. While opinions may differ on particular genomes as to whether
they are complete versus needing more work, and whether they
are well enough along to consider them ‘under way’, it is clear that there will soon be a large and phylogenetically broad sampling of genome sequences.


useful for phylogenetic reconstruction, including: (1) gene content, including components of multiu- nit complexes such as the ribosome, spliceosome, DNA replication machinery, or oxidative phosphor- ylation enzymes and the presence versus absence of particular biochemical pathways (e.g. de Rosa et al., 1999; Fitz-Gibbon and House, 1999; House and Fitz-Gibbon, 2002; Huson and Steel, 2004; Snel et al., 1999, 2005); (2) the relative arrangements of genes (Boore and Brown, 1998); (3) movements of genes among intracellular compartments (i.e. plastid, mitochondrion, nucleus) (e.g. Nugent and Palmer, 1991); (4) insertions of segments of DNA, including transposons and numts (Fukuda et al.,
1985; Richly and Leister, 2004); (5) variation in intron positions (e.g. Qiu et al., 1998); (6) secondary structures of rRNAs or tRNAs (e.g. Murrell et al.,

142 AN I M AL EV O L UTI O N


Table 13.2 Complete nuclear genome sequencing projects completely drafted (i.e. not necessarily having every gap closed) or under way as summarized in Figure 13.1. Some of the taxa listed as under way are currently funded to only low (generally 2X) coverage. There are many other taxa not listed here whose genomes are being investigated at even lower levels of coverage.

Taxonomy Organism Common description

COMPLETE GENOMES
Chordata, Mammalia Bos taurus Cow Callithrix jacchus Marmoset Canis familiaris Dog
Mus musculus Mouse
Homo sapiens Human
Macaca mulatta Rhesus macaque Monodelphis domestica Opossum Ornithorhynchus anatinus Duck-billed platypus Pan troglodytes Chimpanzee
Pongo pygmaeus abelii Orangutan Chordata, Aves Gallus gallus Red jungle fowl Chordata, Sauria Anolis carolinensis Anole lizard Chordata, Amphibia Xenopus tropicalis Western clawed frog Chordata, Teleostei Danio rerio Zebrafish
Gasterosteus aculeatus Stickleback Oryzias latipes Medakafish Takifugu rubripes Japanese pufferfish
Tetraodon nigroviridis Green spotted pufferfish
Chordata, Cephalochordata Branchiostoma floridae Lancelet Chordata, Urochordata Ciona intestinalis, C. savignyi Sea squirt Echinodermata, Echinozoa Strongylocentrotus purpuratus Purple sea urchin Mollusca, Bivalvia Lottia gigantea Owl limpet Annelida, Oligochaeta Helobdella robusta Leech
Annelida, Polychaeta Capitella capitata None Arthropoda, Coleoptera Tribolium castaneum Red flour beetle Arthropoda, Diptera Aedes aegypti Yellow fever mosquito
Anopheles gambiae Malaria mosquito
Culex pipiens House mosquito

Drosophila ananassae, D. erecta, D. grimshawi, D. melanogaster,
D. mojavensis, D. persimilis, D. pseudoobscura, D. sechellia,
D. simulans (8), D. virilis, D. willistoni, D. yakuba
Fruit flies
Arthropoda, Hemiptera Pediculus humanus corporis Louse
Arthropoda, Hymenoptera Apis mellifera Honeybee
Nasonia vitripennis Parasitic wasp Arthropoda, Lepidoptera Bombyx mori Silkworm Arthropoda, Crustacea Daphnia pulex Water flea Arthropoda, Chelicerata Ixodes scapularis Deer tick Nematoda, Chromadorea Caenorhabditis briggsae, C. elegans, C. remanei Roundworms
Pristionchus pacificus None
Meloidogyne incognita Root-knot nematode

G E N O M E S A N D P H Y L O G EN Y 143


Table 13.2 (Continued.)

Taxonomy Organism Common description

Cnidaria, Anthozoa Nematostella vectensis Sea anemone Placozoa Trichoplax adhaerens Tablet animal GENOMES IN PROGRESS
Chordata, Mammalia Cavia porcellus Guinea pig Choloepus hoffmanni Two-toed sloth Cryptomys sp. Mole Cynocephalus volans Flying lemur
Dasypus novemcinctus Nine-banded armadillo
Dipodomys panamintinus Kangaroo rat
Echinops telfairi Lesser hedgehog tenrec
Elephantulus sp. Elephant shrew
Equus caballus Horse
Erinaceus europaeus Western European hedgehog
Felis catus Cat Gorilla gorilla Gorilla Loxodonta africana African elephant
Macaca fascicularis Crab-eating macaque
Macropus eugenii Wallaby
Manis pentadactyla Chinese pangolin Microcebus murinus Mouse lemur Mustela putorius furo Ferret
Myotis lucifugus Little brown bat Nomascus leucogenys Gibbon Ochotona princeps Pika
Oryctolagus cuniculus European rabbit
Otolemur garnettii Bushbaby Pan paniscus Bonobo Papio anubis Baboon

Peromyscus californicus, P. leucopus,
P. maniculatus,
P. polionotus
Mice
Procavia capensis Hyrax
Pteropus vampyrus Flying fox
Saimiri sp. Squirrel monkey
Sorex araneus European common shrew
Spermophilus tridecemlineatus Ground squirrel
Sus scrofa Pig
Tarsius syrichta Tarsier
Tenrec ecaudatus Common tenrec Tupaia belangeri Tree shrew Tursiops truncatus Dolphin
Vicugna pacos Alpaca

144 AN I M AL EV O L UTI O N


Table 13.2 (Continued.)

Taxonomy Organism Common description

Chordata, Aves Taeniopygia guttata Zebra finch Chordata, Testudines Chrysemys picta Painted turtle Chordata, Amphibia Xenopus laevis African clawed frog Chordata, Coelocanthiformes Latimeria chalumnae Indonesian coelacanth
Latimeria menadoensis South African coelacanth
Chordata, Teleostei Astatotilapia burtoni Tilapia Lepisosteus oculatus Spotted gar Metriaclima zebra Tilapia Oreochromis niloticus Tilapia Paralibidichromis chilotes Tilapia
Salmo salar Atlantic salmon
Chordata, Chondrichthys Callorhinchus milii Elephant shark
Raja erinacea Skate Chordata, Hyperotreti Eptatretus burgeri Hagfish Chordata, Hyperoartia Petromyzon marinus Sea lamprey Chordata, Urochordata Oikopleura dioica Tunicate Hemichordata, Enteropneusta Saccoglossus kowalevskii Acorn worm Mollusca, Gastropoda Aplysia californica Sea hare
Biomphalaria glabrata Snail
Platyhelminthes, Cestoda Echinococcus multilocularis Tapeworm
Taenia solium Pork tapeworm
Platyhelminthes, Turbellaria Schmidtea mediterranea Flatworm
Platyhelminthes, Trematoda Schistosoma mansoni, S. japonicum Blood flukes (schistosomes)

Arthropoda, Diptera Drosophila americana, D. auraria, D. equinoxialis, D. hydei, D. littoralis, D. mercatorum, D. mimica, D. miranda, D. novamexicana, D. repleta,
D. silvestris
Fruit flies
Glossina morsitans Tsetse fly Lutzomyia longipalpis Sand fly Phlebotomus papatasi Sand fly
Arthropoda, Hemiptera Acyrthosiphon pisum Pea aphid
Rhodnius prolixus Kissing bug Arthropoda, Hymenoptera Nasonia giraulti, N. longicornis Parasitic wasps Arthropoda, Crustacea Jassa slatteryi Amphipod
Parhyale hawaiensis Amphipod
Arthropoda, Chelicerata Limulus polyphemus Horseshoe crab
Tetranychus urticae Spider mite Arthropoda, Myriapod Strigamia maritima Centipede Priapula Priapulus caudatus Priapulid worm Tardigrada Hypsibius dujardini Water bear

G E N O M E S A N D P H Y L O G EN Y 145


Table 13.2 (Continued.)
Taxonomy Organism Common description
Nematoda, Chromadorea Ancylostoma caninum Canine hookworm
Ascaris lumbricoides Human intestinal roundworm
Brugia malayi Filarial roundworm
Caenorhabditis brenneri, C. japonica None
Cooperia oncophora Intestinal worm
Dictyocaulus viviparus Bovine lungworm
Haemonchus contortus Barber pole worm
Heterorhabditis bacteriophora None
Necator americanus New World hookworm
Nematodirus battus Thread necked worm
Nippostrongylus brasiliensis Rat intestinal nematode
Oesophagostomum dentatum Nodule worm
Onchocerca volvulus River blindness roundworm
Ostertagia ostertagi Stomach worm
Strongyloides ratti Threadworm
Teladorsagia circumcincta Brown stomach worm
Trichostrongylus vitrinus Black scour worm
Nematoda, Enoplia Trichinella spiralis Trichina worm
Trichuris muris Whipworm
Cnidaria, Hydrozoa Hydra magnipapillata Hydra
Porifera, Demosponge Reniera sp. None




2003); (7) details of genome-level processes, such as the rearrangements that generate antibody diver- sity (Frieder et al., 2006); and (8) deviations from the
‘universal’ genetic code (Telford et al., 2000; Santos,
2004). Many others are likely to be found.
Of course, the reliability of these features can
only be assessed by study of their consistency
with other characters, and several are already
suspect. Convergent gene losses, for example, may
be common as organisms independently evolve
smaller genomes or no longer experience selection
for maintaining a particular biochemical path-
way; in contrast, convergent gain of genes seems
much less likely. Independent evolution of smaller
genomes may also lead to parallel losses of the
most expendable structures in RNA or protein
genes. There is a certain time-horizon that limits
the usefulness of any particular type of charac-
ter; for example, once retro-elements degrade in
sequence beyond the point where the insertion can be reliably inferred to be of single origin, the insertion is no longer useful as a phylogenetic character. Certain changes in the genetic code and in tRNA secondary structures of mitochondria are known to have occurred convergently (although occasional homoplasy has not disqualified the use of either morphological characters or molecular sequence comparisons). There is also the problem in the case of closely spaced sequential internodes where random partitioning of polymorphisms, including those of genome-level characters, can lead to incorrect inference of phylogeny (e.g. Salem et al., 2003). See Boore (2006) for additional caveats and precautions.
Already there have been important insights gained from comparing such features, including: (1) tarsiers have been shown to be the sister group to the clade of monkeys and apes rather than the

146 AN I M AL EV O L UTI O N



prosimians based on patterns of short interspersed nuclear element (SINE) integration (Schmitz et al., 2001); (2) patterns of SINE and long inter- spersed nuclear element (LINE) insertions have also supported the monophyly of toothed plus baleen whales, that hippopotamuses are the sister group to cetaceans, that camels are the most basal cetartiodactyls (Nikaido et al., 1999), and that river dolphins are paraphyletic (Nikaido et al., 2001); (3) animal interphylum relationships have been clari- fied by comparisons of the gene membership within Hox clusters (de Rosa et al., 1999); and (4) a study of the presence of spliceosomal introns supports the monophyly of Actinopterygia and clarifies sev- eral relationships within the group, including the basal position of bichirs (Venkatesh et al., 1999). For further discussion see Murphy et al. (2004), Okada et al. (2004), and Boore (2006).

(a)
















(b)




Taxon 1




Taxon 2




Taxon 3






Taxon1


13.4 What are the advantages of using these genome-level characters?

In general, these types of features would be expected to change in a saltatory, non-clocklike manner. This may seem, at first, to be wrong- headed, since great effort has been expended in many studies to identify clocklike characters, to enable accurate molecular clock estimates of time of divergence. But it is this aspect that makes these genome-level characters especially useful for addressing the most difficult branch points, those with a short time of shared history followed by a long period of divergence, as mentioned above. It is for resolving these relationships that clocklike behaviour guarantees failure, since the ratio of signal to noise will closely match the ratio of the two time periods. Rather it is the least clocklike of characters that are expected to prevail, where an occasional and abrupt change may have occurred and then remain (Figure 13.2). Admittedly, the con- comitant disadvantage is that many such charac- ters must typically be examined in order to find those that happened to have changed during the period of shared ancestry and so marking the rela- tionship (see Boore, 2006, for further analysis and discussion).



Taxon 2




Taxon 3

Figure 13.2 Illustration of why clocklike characters (a) may be less informative than non-clocklike characters (b) when the internode between subsequent lineage splits is short. Each of the four shapes is meant to be a character with states indicated by patterning. In (a) the circle and triangle are not informative and the square and pentagon are homoplasious. The two changes accumulated in the common ancestor of taxon 1 and 2 (for the pentagram and circle), that were at one point synapomorphies, have been erased by subsequent changes. In (b), the changes are
rarer and saltatory. The pentagram and triangle are not informative and the circle is constant, but the square is informative for uniting taxa 1 and 2.




13.5 What about clades without representative genome sequences?

This enormous data set provides a new class of characters that could lead to definitive resolution of some branches of the tree of life, not only for these taxa but also for others where targeted searches for

G E N O M E S A N D P H Y L O G EN Y 147



identified characters could be fruitful. As shown in Figure 13.1, whole-genome sampling will include many major lineages, but not all. Fortunately, we can use genomes in hand to identify sets of genome-level characters that can be diagnostic for the relationships of related groups without gen- ome projects. One could then determine gene order by using Southern hybridization, for example, or probe a large DNA insert library (i.e. in bacterial artificial chromosome (BAC) or fosmid vectors) to find a clone to sequence for the region of interest of the genome. Gene rearrangements, losses, and duplications can also be identified using compara- tive genomic hybridization (CGH) chips with tiled large-insert clones, as has been done for a sampling of diverse human populations (Sharp et al., 2005) and more broadly across the great apes (Locke et al., 2003) or by using arrays of oligonucleotides (representational oligonucleotide microarray analysis (ROMA); Sebat et al., 2004).


13.6 What are the main challenges before us?

First, we must increase the representation of under- studied groups of animals for large-scale genomic sequencing. There is no reason to believe that taxa that have been traditionally studied intensively, i.e.
those with higher species richness, greater breadth of niche occupation, more important roles in patho- genesis, or amenability to laboratory experimenta- tion, will be more informative toward the goals of understanding broad patterns of the evolution of animals and their genomes. Next, we need to have a codification of nomenclature for genes that is based on assessment of orthology (Dehal and Boore, 2006). The renaming of genes to indicate orthology is not feasible because it would ren- der large bodies of literature difficult to interpret and because scientists who study model organ- isms (and who have largely done the naming) are invested in their parochial nomenclature. Thus, the solution must be a lexicon superimposed on these names already in place. Third, a system must be devised for codifying the genome-level char- acters themselves for entry into data bases and matrices for broad comparisons. Lastly, we need for the community to devise standards of inter- pretation and analysis, such as the use of cladistic reasoning rather than associating taxa by similar- ity alone (Boore, 2006). Then it seems likely that genome-level characters will provide the best data set for convincingly reconstructing relationships for some of the most hotly contended nodes in the tree of life and for establishing a framework for all organismal relationships.