Thursday, December 6, 2012

Parallel evolution in adaptive phenotypes: the case of the threespine stickleback


ResearchBlogging.orgHow do adaptive phenotypes evolve? This question, despite the increasing availability of genomic and other molecular data, remains still largely unanswered. Among the different aspects investigated, a major point of discussion in this topic is the extent of the contribution of coding versus non-coding variation in the evolution of new traits. Although many research groups suggested that non-coding mutations might play a pivotal role because might avoid pleiotropic effects, still few examples are available to discard a potential major contribution of coding variants in adaptive evolution.
The paper from Jones et al. we discussed tried to answer this question by looking at the differences between distinct populations of threespine sticklebacks (Gasterosteus aculeatus). This species, originally found in marine habitats, colonized the freshwater environment evolving specific phenotypic traits, but still maintaining the ability to hybridize with the marine individuals. An important feature of this species, already known from previous studies, is the presence of shared genomic variants in geographically unrelated populations distinguishing the marine from the freshwater populations. This finding suggested the possibility of a parallel adaptive evolution of phenotypic traits due to the reuse of standing genetic variation. To test this hypothesis, Jones et al. generated a reference genomic assembly of a female freshwater stickleback (Sanger sequencing, 9.0x coverage, total gapped size: 463Mb). This reference genome provided the basis to analyze genomic differences in marine and freshwater populations collected in several locations around the world (Europe, North America and Japan). For this purpose, a total number of 20 individuals (classified in clearly marine and clearly freshwater based on multiple phenotypic features) were sequenced at 2.3x average coverage and genome-wide single nucleotide polymorphisms (SNPs) identified.
The data collected were analyzed using three different approaches with the aim of finding regions in the genomes showing a high similarity among the freshwater individuals and differing from the corresponding loci in the marine samples. The first approach consisted in a self-organizing map-based iterative Hidden Markov Model (SOM/HMM), used to reconstruct common relationships (trees) among the individuals. Although most of the phylogenies recapitulated the geographical relationships among the samples, four of them separated most of the marine from most of the freshwater individuals, identifying genomic loci putatively involved in the differentiation of the ecotypes. The second and the third approaches used a sliding window analysis to detect the divergence between the two populations. The second consisted in the calculation of a cluster separation score (CSS) to quantify the distance between the marine and the freshwater clusters; the third consisted in an unguided Bayesian model-based data-driven clustering (DDC) to calculate for each window a maximum number of clusters to which assign the individual samples. In total, 242 genomic regions showing a shared marine-freshwater divergence were identified by either method (0.5% of the genome). Testing of these approaches on a genomic location known to have evolved adaptively in the distinct species (EDA gene, fig.1) revealed the reliability of the three and the power of their complementary usage to spot putatively adaptively evolving loci.


Figure 1: Parallel divergence signals at known armour plate locus. a) Ensembl gene models around EDA. b) Visual genotypes for sequenced fish (homozygous sites for most frequent allele in marine fish (red); homozygous for alternative allele (blue); heterozygous (yellow), or non-variable/missing/repeat- masked data (white)). c) DDC cluster assignments for marine (red) and freshwater populations (blue). Most fish are assigned to cluster k1, except in the boxed region, where freshwater fish are assigned to a distinct cluster (k2). d) SOM/HMM analysis supports patterns of divergence with a marine– freshwater-like tree topology in the centre, but not edges, of the window (trees a–d). e, f) Similar support is shown by CSS analysis (e) and its associated P-value (f). The combined analyses define a consensus 16-kb region shared in freshwater fish (vertical shaded box), matching the minimal haplotype known to control repeated low armour evolution in sticklebacks. 

To test the extent of parallel reuse of these regions in adaptation to the freshwater environment in contrast to newly evolved adaptive loci, an independent sample of a pair of marine and freshwater individuals from the same geographical zone (River Tyne, Scotland) was subjected to sequencing and SNP analysis. The experiment showed that, within the most highly divergent windows of the genome, only a part (35.3% of the 0.1% most divergent windows) contained the globally shared loci (fig. 2). The result indicates that part of the divergence between the two phenotypes actually derives from shared standing variation, but also that new population specific mutation can play a role in the determination of the specific traits.


Figure 2: How much of local marine–freshwater adaptation occurs by reuse of global variants? a) Classic marine and freshwater ecotypes are maintained in downstream and upstream locations of the River Tyne, Scotland, despite extensive hybridization at intermediate sites16. b) Pairwise sequence comparisons identify many genomic regions that show high divergence between upstream and downstream fish (x axis). Many, but not all, of these regions also show high global marine–freshwater divergence (y axis; red points indicate significant CSS FDR , 0.05), indicating that both global and local variants contribute to formation and reproductive isolation of a marine– freshwater species pair.


Interestingly, the group found three loci showing clear marine-freshwater divergence within regions involved in chromosomal inversions (chromosomes I, XI and XXI, fig. 3). The finding supported the hypothesis that molecular mechanisms, such as chromosomal inversions, suppressing recombination between adaptive loci can be favored by selection for the maintenance of contrasting ecotypes in hybridizing populations.


Figure 3: Genome-wide distribution of marine–freshwater divergence regions. Whole-genome profiles of SOM/HMM and CSS analyses reveal many loci distributed on multiple chromosomes (plus unlinked scaffolds, here grouped as ‘ChrUn’). Extended regions of marine–freshwater divergence on chromosomes I, XI and XXI correspond to inversions (red arrows). Marine–freshwater divergent regions detected by CSS are shown as grey peaks with grey points above chromosomes indicating regions of significant marine– freshwater divergence (FDR , 0.05). Genomic regions with marine– freshwater-like tree topologies detected by SOM/HMM are shown as green points below chromosomes.

Finally, the analysis of the 64 genomic regions showing the strongest evidence of parallel evolution were investigated to determine the contribution of coding and non-coding variation to the adaptation to a different environment. Only 17% of them could be classified as coding based on the presence of non-synonymous substitutions, while the remaining part could be attributed to regulatory or probably regulatory changes (fig. 4a). To actually test if any regulatory change could be linked to these regions, a whole-genome microarray expression analysis was performed on tissues from a marine and a freshwater sample. The results obtained from genes mapping within or close to the loci identified by either method show a general divergence in the expression levels in different tissues between the two ecotypes (fig. 4b).


Figure 4: Contributions of coding and regulatory changes to parallel marine–freshwater stickleback adaptation. a) A genome-wide set of marine– freshwater divergent loci recovered by both SOM/HMM and CSS analyses includes regions with consistent amino acid substitutions between marine and freshwater ecotypes (yellow sector); regions with no predicted coding sequence (green sector); and regions with both coding and non-coding sequences, but no consistent marine–freshwater amino acid substitutions (grey). b) Genome-wide expression analysis shows that marine–freshwater regions identified by SOM/ HMM or CSS analyses are enriched for genes showing significant expression differences in 6 out of 7 tissues between marine LITC and freshwater FTC fish (observed, grey bars; expected, white bars; *P , 0.01, **P , 0.001, ***P , 0.0001, ****P = 0.00001), consistent with a role for regulatory changes in marine–freshwater evolution.

In summary, the work provided strong evidence that in threespine sticklebacks many loci implicated in adaptation from marine to freshwater environment were reused independently in several distinct populations, suggesting parallel evolution had a deep impact in the adaptation process. Furthermore, several putatively adaptive loci were found to be involved in chromosomal inversions, supporting the idea that genomic rearrangements can hamper recombination in these genomic locations thus preserving the features of the different ecotypes. Finally, the results obtained suggest regulatory mutation might have had a major role in the evolutionary processes leading to the adaptation of this species to a new environment.

Jones FC, Grabherr MG, Chan YF, Russell P, Mauceli E, Johnson J, Swofford R, Pirun M, Zody MC, White S, Birney E, Searle S, Schmutz J, Grimwood J, Dickson MC, Myers RM, Miller CT, Summers BR, Knecht AK, Brady SD, Zhang H, Pollen AA, Howes T, Amemiya C, Broad Institute Genome Sequencing Platform & Whole Genome Assembly Team, Baldwin J, Bloom T, Jaffe DB, Nicol R, Wilkinson J, Lander ES, Di Palma F, Lindblad-Toh K, & Kingsley DM (2012). The genomic basis of adaptive evolution in threespine sticklebacks. Nature, 484 (7392), 55-61 PMID: 22481358

Thursday, November 22, 2012

Ecological success of recently emerged bacterial hybrids living in the wild

ResearchBlogging.orgMicrobial species are one of the most ubiquitous living group on Earth's biosphere, showing incredible ability to thrive even in ambient conditions to the limit of human endurance. By virtue of their rapid growth, bacteria are ideal for unraveling the molecular mechanisms of many evolutionary processes. Their rapidity to respond to changes has been associated to the combined effect of evolutionary processes, species composition or gene expression shifts. Most of the studies have focused so far on isolation and comparison of cultured bacterial population, while very few data are available concerning free-living bacteria. Therefore, it is still controversial how quickly, to which extent and by which mechanism microorganisms evolve in their natural environment. 

Two researchers of the University of California, Denef and Banfield, have tried to answer this question, as described in a paper recently published in Science. Their work report evolutionary rate estimates from bacterial populations living in a really challenging site, the hot, humid, low-pH, metal-rich and low-oxygen acid mine drainage in the Richmond Mine (Iron Mountain, CA), over the course of 9 years. This would seem not the ideal model system for conducting such kind of study, for the low accessibility at the sites only in limited periods of the year, but it perfectly fits the requirement of a discrete, reproducible and simple microbial community meeting very restricted input from other regions. In fact, the air-solution interface biofilm community consists in few organisms types, normally four to six. In this case Leptospirillum group II, which comprises iron-oxidizing bacteria that can live in sulfuric acid, dominates. The authors tried to trace back a lineage history of the group, starting by using the data of previous metagenomic studies Simmons et al. from the same lab published on Plos Biology in 2008. These led to the reconstruction of Leptospirillum group II type I and VI “reference” genomes, which share about 94% average nucleotide identity. As shown in figure 1, the two populations were sampled in 2002 and 2005 at 5-way and UBA locations, respectively.


Figure 1. Adapted from Denef et al. Richmond Mine schematic map, with pie charts indicating genotype proportions in 24 samples, estimated on the basis of read recruitment. 

Already assembled genomes were compared against total population DNA, thus allowing the authors to find other four distinct Leptospirillum genotypes (types II to V). As indicated by proteomics-based results published by the same group in 2009, some particular sites, like C75, were clearly dominated by Leptospirillum type III genotype, which appeared to be a recombinant hybrid of genotypes I and VI. The recombination points, which were found by identifying discontinuities in reads alignments with the reference genomes, were all located within genes. The population sampled at C75 was ideal to be used to calculate the substitution rate, because of low level of variation within population and across space. A high rate of substitution of 1.4 × 10−9 per nucleotide per generation was calculated, if compared to previous estimates of bacterial genome-wide short term substitution rates, which have ranged from 7.2 × 10−11 to 4.0 × 10−9. Many reasons can account for these unexpected results, which anyway match with universal mutations-per-genome size predicted by Drake; for example, they used a unique approach combining proteomics inferred genotyping dataset, population genomic time series and, for the first time, cultivation-independent population genotypes, but also they sampled a unique model system, where human and natural perturbations, combined with the low biological complexity, can both affect the evolutionary rate estimates. 

Population genomic analyses suggested that the six Leptospirillum genotypes consist of a mosaic of type I and type VI genome blocks tens to hundreds of kb in length, probably recombined in a single cell and fixed in its descendants, as shown by the recurrence of the same transition points. Each genotype’s fixed mutation were used to construct the phylogenetic tree showed in figure 2, which suggests that the six Leptospirillum genotypes diverged from a common ancestor in a matter of decades (time of coalescence estimated between 2 and 44 years).



Figure 2. Adapted from Denef et al. Evolutionary history of the sampled genotypes, based on the variant loci inferred using the maximum parsimony method. The percentage of replicate trees in which the associated taxa clustered together in the bootstrap test is shown next to the branches. Dotted arrows indicate recombination events; circle schematics represent the regions affected. The timeline indicates the calculated time ranges of recombination events as well as historical events. Branch D* is presented as two strains, each assigned half the total number of UBA 06/05 SNPs, because low incidence of SNPs precluded their linkage. BP, years before the present. 

Nicely, evidence for positive selection for hybrid genotypes was given by the finding of fixed non-synonymous substitutions in high number of genes involved in signal transduction, transcriptional regulation and global regulators. This study suggest that the evolution of Leptospirillum consisted of a mosaic of different events, comprising homologous recombination, fixation and selective sweeps that generated the different genotypes that can be currently observed. Some limits related to the paper can be found in the final author’s speculation that states selection between genotypes as due to genotypes divergence by only a few nucleotides. In fact it’s likely that this result was biased by the approach they chose, that only relied on the comparison of genes present in their reference genomes, which can produce incomplete or erroneous interpretation if the genome assemblies are not corrected. 


Denef, V., & Banfield, J. (2012). In Situ Evolutionary Rate Measurements Show Ecological Success of Recently Emerged Bacterial Hybrids Science, 336 (6080), 462-466 DOI: 10.1126/science.1218389

Wednesday, October 10, 2012

What could our genomes actually tell about disease risk?

ResearchBlogging.org
Despite the recent advances in whole-genome sequencing, two recent studies let us think that we are far from uncovering the genetic basis of common diseases risk. In fact, information relevant to complex diseases might hide within rare or even private genome variations, often too scarce to be studied statistically. We might thus have to change radically our way of thinking of genes-diseases associations to make a step forward and make the DNA talk.

Whereas a few, usually rare and severe “genetic disorders” can be traced to variations at one or two locations, or “loci”, in the DNA sequence, most common diseases are the result of complex interactions between protein-coding genes, non-coding DNA and environmental effects. These well-named “complex diseases” include cardiovascular, metabolic, neurologic and psychiatric conditions of great concern to health policies, such as early-onset stroke, myocardial infarction, diabetes, dyslipemia, Alzeihmer's, bipolar disorder or schizophrenia.

Some of these complex diseases have a high heritability, which means that a great part of individual differences in the probability to develop the disease can be explained by differences in genomes. For example, the heritability of early-onset myocardial infarction is about 60% [1]: genomes are more important than environment in explaining the differences in early-onset infarction between individuals. Thus a lot of work has been going into identifying the changes in DNA sequences involved in complex disease heritability. Especially, the development of new sequencing technologies has allowed for comparison of hundreds of individual sequences and their mapping to various symptoms, a method known as “genome-wide association studies”. Hundreds of disease-related genetic variations have been identified this way. However they explain only a very small fraction of the heritability: in the case of early-onset myocardial infarction, only 2.8% of the heritability has already been linked to particular genes [2].

To explain the low power of association studies to identify genetic variants contributing to complex diseases, it was hypothesized that most variation in disease predisposition were due to “high risk” variants, that have a strong negative impact on health, but remain rare in a population because they are counter-selected [3]. In consequence, we would only need to increase sample size and therefore our power to detect rare variants to better explain the genetic basis of common diseases. In that scope, two studies published in the July issue of Science have used large datasets (respectively 2 440  and 14 002 genomes) to investigate the potential role of rare variants, defined when one of the variants at one locus is present in less than 0.5% of the individuals sampled. The large sample sizes allowed for detection of lots of previously unknown variants, thus highlighting the limits of previous smaller-scale studies: 90% of rare variants, but only 5% of common variants, found in 202 drug-target genes were novel, and estimates of discovery rates showed that lots of new variants are still to discover (Fig. 1). 

Nelson et al., An Abundance of Rare Functional Variants in 202 Drug Target Genes Sequenced in 14,002 People, Science 337, 2012. 

Fig. 1 Number of variants discovered per kilobase of sequence with sample sizes increasing to 5000 people for multiple populations.


The studies also confirmed that variants with an potential impact on health remained rare: the proportion of non-synonymous variants, which result in an alteration of the protein synthesized, was higher in rare than in common variants (Fig. 2).

Nelson et al., An Abundance of Rare Functional Variants in 202 Drug Target Genes Sequenced in 14,002 People, Science 337, 2012. 
Fig. 2 Expected ratios of non-synonymous to synonymous variants in the absence of selection and observed ratios for rare to common alleles, from left to right. MAF (Minor Allele Frequency) is the frequency of the rarest version of a variant.

However, rare variants were found to be more numerous than previously thought: around 90% of variants were rare. Interestingly, individuals of African ancestry exhibited less rare variants, but more variants of intermediate frequency than those of European ancestry. Moreover, most rare variants were population-specific (Fig. 3 and 4) and about 60% of all variants were only present in one individual. 


Casals and Bertranpetit, Human Genetic Variation, Shared and Private,  Science 337, 2012. Data from Tennessen et al., Science 2012.
Fig. 3 Proportion of shared and unshared (private) variants between the African-American and the European-American populations.


Nelson et al., An Abundance of Rare Functional Variants in 202 Drug Target Genes Sequenced in 14,002 People, Science 2012. 
Fig. 4 Allele sharing and variant abundance. (A-C) The average allele sharing between pairs of populations for rare (A), intermediate (B) and common (C) variants computed as the frequency in the pooled population pair. (D) The number of variants per kilobase found in population samples of 2,500 individuals.

Such figures are in contradiction with the current estimates derived from recent population growth. In fact, human demography is currently described by the “Out-of-Africa” model, that posits an emergence of European and Asian populations from a small population in Africa about 60,000 years ago [4]. As the individuals that migrated represented only a small fraction of the ancestral African population, some genetic variants, especially the rarest one, were lost. Then African and European populations were supposed to increase regularly, all the while acquiring new population-specific variants by mutation that would increase in frequency only if they are not deleterious, or else eventually disappear. Such “bottleneck effect” can be observed in the Finns that have less variants, but more population specific variants than other Europeans (Fig. 4). But the overall excess in rare variants in Europeans does not fit the model: most of these variants should have either disappeared or increased in frequency over such a time-scale. Such pattern can however be explained by accelerated population growth in the last thousands year, during which lots of new mutations could occur in a short time (Fig. 5). Therefore, rare variants provide a precious insight into recent demography.


Tennessen et al., Evolution and Functional Impact of Rare Coding Variation from Deep Sequencing of Human Exomes, Science 2012.
Fig. 5 Schematic representation (not to scale) of the inferred demographic model. kya, thousand years ago.

These findings are not very good news for complex disease research. Of course, the rare variants discovered in protein-coding genes are numerous and often deleterious, and could therefore play an important role in disease risk. However they rarity makes it difficult to actually test that role: over thousands of genomes, less than 5% protein-coding genes afforded a sufficient power to detect the effect of rare variations on disease risk, even when that effect is relatively strong. Consistently, no significant association was found for 202 drug-target genes. Moreover, most variants are population- or even individual-specific. Thus association studies should be at least replicated across populations, with a careful determination of ancestry, to be universal and avoid false associations between population-specific traits and variants.

The high number of individual-specific variants and the low power of association for other rare variants highlight the importance of genome-wide functional studies to accurately estimate disease risk where association studies fail. Functional predictions might be the next prevailing tool in the study of genes and disease association. Ideally, such studies would directly estimate the functional impact of a given variant, but the methods currently implemented are rather inconsistent and have a high false-positive rate, that is they often detect a functional impact where there is none. Such caveats make them still unsuitable for applied uses in medical diagnosis. A better knowledge of molecular biology and its link to physiology seems still necessary to assess the actual impact of rare variants on complex diseases.



1       Nora J.J. et al. (1980). Genetic-epidemiologic study of early onset ischemic heart disease. Circulation 61:503 – 508.
2       Myocardial Infarction Genetics Consortium (2009). Genome-wide association of early-onset myocardial infarction with single-nucleotide polymorphisms and copy number variants. Nature Genetics 41:334-341.
3       Manolio T.A. et al. (2009). Finding the missing heritability of complex diseases. Nature 461:747-753.
4       Laval G. et al. (2010). Formulating a historical and demographic model of recent human evolution based on resequencing data from noncoding regions. PLoS ONE 5(4): e10284.

Nelson et al. (2012). An Abundance of Rare Functional Variants in 202 Drug Target Genes Sequenced in 14,002 People. Science, 337, 100-104 DOI: 10.1126/science.1217876

Tennessen et al. (2012). Evolution and Functional Impact of Rare Coding Variation from Deep Sequencing of Human Exomes. Science, 337, 64-69 DOI: 10.1126/science.1219240

Casals and Bertranpetit (2012). Human Genetic Variation, Shared and Private Science, 337, 39-40 DOI: 10.1126/science.1224528

Wednesday, October 3, 2012

Evolutionary consequences of sex: It's not about what you're doing, but who you're doing it with...

ResearchBlogging.org
Bacteria are one of the most ubiquitous living group and exhibit finely tuned adaptations to a wide range of habitats, even the most inhospitable ones. Their ability to evolve rapidly is at the roots of many public health issues, such as the development of resistances to antibiotics or the rapid evolution of seasonal diseases, but can also be of great help to humans by creating new metabolic pathways to transform human-made pollutants and harmful substances. In the early 20th century, new bacterial genomes were still thought to be the result of mutations only, and to be then transmitted vertically within a clonal strain. In the 40’s, the discovery of bacterial DNA recombination through transformation (Avery, MacLeod and McCarty experiment in 1944) or conjugation (Lederberg and Tatum experiment in 1946) shed light on the processes responsible for the rapid ecological differentiation of bacterial strains: an individual can acquire new genes or alleles through recombination that allow it to stand new ecological conditions.

In Eucaryotes, genetic exchange and recombination through sexual reproduction is considered the basis of gene-specific transmission and selection among a population. However, the importance of genetic exchange between bacteria in uncoupling selection processes between different genes remains a controversial issue. In fact, contradictory observations have elicited two models of selection:
  1. On one hand, the ecological clustering of bacterial biodiversity in genetically consistent ecotypes support the traditional view that adaptive mutations are selected through whole-genome clonal selection. Moreover the low measured levels of recombination are insufficient to unlink a gene from the rest of the genome.
  2. On the other hand, the existence of environment specific genes and alleles suggests that recombination can unlink parts of the genome. Moreover, some loci exhibit low nucleotidic diversities compared to the rest of the genome, with suggest purifying selection on these regions. Thus adaptive mutations seems to be selected quite independently of the rest of the genome.
To disentangle those apparently incompatible observations and assess the degree of gene uncoupling in bacteria, researchers of the MIT examined in a recently published study[i] the genomes of 20 strains representing two ecotypes in the marine species Vibrio cyclitrophicus. As the genomes of these ecotypes are extremely similar, they can be considered the result of recent ecological differentiation, thus giving us a snapshot of this evolutionary process. Based on the comparison of these sequences, the authors claim that gene-specific sweeps do occur and can lead to environment specific-genes on a short time scale, but also to ecological clustering, through preferential within-habitat recombination, on a longer time scale.

In fact, different parts of the genome have different evolutionary histories. Especially, ecotype-specific SNPs are only found on a few locations in the genome, whereas the rest of the polymorphic genome supports a genetic intermingling between the ecotypes. Moreover, the two chromosomes of V. cyclitrophicus support different phylogenies, with chromosome 1 grouping the ecotypes, whereas chromosome 2 splits one ecotype into two groups. The phylogeny within one of these two groups is strongly supported by chr2 but not by chr1. Thus, habitat-specific genes are evolving quite independently and do not drive genomewide selective sweeps, an observation consistent with the environment specific genes and alleles that have already been documented[ii].

These results highlight the need for high quality sequencing data and fine grained analysis to understand the evolutionary histories of different parts of the genome. In fact, the authors show that a few loci with consistent phylogeny, such as the ecotype-specific SNPs here, are sufficient to drive the whole-genome phylogeny, if the signal of clonal ancestry in the rest of the genome has been blurred by homologous recombination (Fig. 1). Therefore, the ecotype theory might be based on phylogenies biased toward the history of a few loci under purifying habitat-driven selection rather than on neutral loci with inconsistent histories accounting for most of the genome.

Fig. 1: A. Maximum-likelihood phylogeny for the core genome (genes presents in all strains) of chromosome I in V. cyclitrophicus. Scale is substitution per site. All nodes have a 100% bootstrap support unless indicated. B. Genome regions with uninterrupted support for (black points) or against (grey points) the ecological split. ML trees for three major regions are shown. Adapted from Shapiro et al. 2012.

The most important point made by the authors remains however their evidences for preferential within-ecotype recombination. When examining recombination events affecting recently diverged pairs of strains, recombination rates were found to be higher within than between habitats. The authors make here an essential point toward a unified theory of bacterial genomes evolution. Such preferential recombination indeed provides an explanation for the development of ecotypes, usually considered an evidence of genomewide genetic sweeps, from gene-specific sweeps. Even if the mechanisms involved in genes transmission are quite different between Eubacteria and Eucaryotes, they seem to converge in allowing gene-specific selective sweeps and in restricting genetic exchange between habitat. This study show a more universal picture of selective pressures on evolutionary mechanisms than previously thought between Eucaryotes and Eubacteria (Fig. 2).


Fig. 2: Model of ecological differentiation between bacterial ecotypes (from Shapiro et al. 2012). Thin grey (resp. black) arrows represent recombination within (resp. between) ecologically associated populations. Thick coloured arrows represent acquisition of adaptive alleles for red or green habitat.

Eventually, this new insight into bacterial evolution plead for a new structure of bacterial diversity. Whether and how a species-level could be defined in Eubacteria has been a long standing controversy. The current species criterion dates back to 1987: it defines a species as a group of clonal strains characterized by at least one phenotypic trait and 70% DNA–DNA hybridization[iii]. However this definition was often criticized for grouping within one species extremely diverse phenotypes[iv]. Moreover, the very idea of bacterial species was sometimes rejected on the basis of common genetic exchanges between morphologically and ecologically distinct bacteria[v]. That last assertion is here contradicted by demonstrating the existence of barriers to gene flow between habitats. If recombination, that is “bacterial sex”, is more frequent within than between habitats, ecotypes are quite close to the definition of a species in Eucaryotes, an “ecotype-hypothesis” already put forward about ten years ago[vi]. Therefore this study provides a first step toward the “solid understanding of the genetic basis of the ecological distinctiveness of the ecotype” advocated by Konstantinidis et al. in 2006[iv], although it does not exclude the possibility that some ecotypes be defined by gene expression rather than gene content.

New paths for research are here opened, especially to define the barriers to gene flow that could explain habitat-specific recombination. The bacteria studied here are not known well enough to define the ecological differentiation observed as sympatric or allopatric: if their habitats within sea water are distinct enough, a physical barrier could be considered. If the differentiation can be regarded as sympatric, the mechanisms that prevent recombination between ecotypes are still to be investigated.




[i] Shapiro B.J. et al. (2012). Population genomics of early events in the ecological differentiation of bacteria. Science 336:48-51.
[ii] Coleman M.L. and Chisholm S.W. (2010). Ecosystem-specific selection pressures revealed through comparative population genomics. Proc. Natl. Acad. Sci. U.S.A. 107(43):18634-18639.
[iii] Wayne L.G. et al. (1987). Report of the ad hoc committee on reconciliation of approaches to bacterial systematics. Int. J. Syst. Bacteriol. 37:463-464.
[iv] Konstantinidis K.T., Ramette A. and Tiedje J.M. (2006). The bacterial species definition in the genomic era. Philos Trans R Soc Lond B Biol Sci. 361(1475): 1929-1940.
[v] Margulis L. and Sagan D. (1997). Microcosmos: Four Billion Years of Microbial Evolution. University of California Press.
[vi] Cohan F.M. (2002). What are bacterial species? Annu Rev Microbiol. 56:457-487.
 
Shapiro BJ, Friedman J, Cordero OX, Preheim SP, Timberlake SC, Szabó G, Polz MF, & Alm EJ (2012). Population genomics of early events in the ecological differentiation of bacteria. Science (New York, N.Y.), 336 (6077), 48-51 PMID: 22491847

Friday, September 21, 2012

The yak genome and adaptation to life at high altitude

ResearchBlogging.org

The domestic yak (Bos grunniens) is an important domesticated species for Tibetans. Domestic yaks provide meat and other basic resources of necessity. The analysis of yak genome provides important insights into adaptation to a high altitude. Here discussed study was published in Nature Genetics.

The study compares the yak genome with the genome of taurine cattle (B. taurus). Yak and cattle are cross-fertile, that means that they are genetically very similar. However the cattle suffer from hypertension when living in the yak habitat, thus, comparing this two species can provide the information about evolutionary adaptation to high altitude.

In the study, researches sequenced genome of a female yak. They found three genes that help the animal to deal with a low concentration of oxygen that is typical for high altitude. Five further genes provide a better nutritional assimilation, as a consequence of the limited herbal resources available in the mountains where they live.

Fig.1 Qiu et al., The yak genome and adaptation to life at high altitude., Nature Genetics 44, 2012
Venn diagram showing unique and shared gene families between the yak, cattle, dog and human genoms.

One the Fig.1 unique and shared gene families from four different species (yak, cattle, human and dog) are shown. Gene family is a set of genes that show sequence similarity, and generally (not always) have similar functions. We discussed the choice of the species on the Venn diagram. The two other species for comparison, beside of yak and cattle, were chosen probably because of the good annotated genome. It would be also nice to see the comparison to chimp, to see the relationships yak-cattle and human-chimp together. However the authors mention that yak and cattle have diverged approximately 4.9 million years ago, which is comparable to the time at which humans and chimpanzees diverged. Why the comparison is made to mammals and not other species? The comparison was done to show influence of adaptation, and for this purpose it is better to take more related species.


Fig.2 Qiu et al., The yak genome and adaptation to life at high altitude., Nature Genetics 44, 2012
Gene expasion and contraction in the yak genome

In the Fig.2 a neighbor-joining tree of mammalian Hig domain sequences is presented. The Hig domain is known to play role in hypoxia, so adaptation to high altitude. Hig domain sequences group in different clusters, but yak and cattle sequences seem to be similar. Some brunches of the tree are long, that means that the sequences were very diverged.

Authors also show that yak has three more positively selected genes then cattle, and it is more then the rest of positive selected genes. The paper also compares Ka/Ks ratio of GO categories of yak and cattle. Ka/Ks ratio is a proportion of synonymous to nonsynonymous substitutions. It is assumed that synonymous substitutions represent the background and the selection influence nonsynonymous substitutions.


For the further studies it would be interesting to compare the results from other species, such as goats in mountains and fields, or even wild species. We wondered why there is now comparison of the results of the current study with already known result from the studies of human genome on adaptation to high altitude?

Summary
The study present analysis for the genome adapted to high altitude. The authors sequenced genome of yak using a whole-genome shotgun strategy and the Illumina platform. The scientist hope that this results can help in research of hypoxia-related diseases in humans.


Qiu Q, Zhang G, Ma T, Qian W, Wang J, Ye Z, Cao C, Hu Q, Kim J, Larkin DM, Auvil L, Capitanu B, Ma J, Lewin HA, Qian X, Lang Y, Zhou R, Wang L, Wang K, Xia J, Liao S, Pan S, Lu X, Hou H, Wang Y, Zang X, Yin Y, Ma H, Zhang J, Wang Z, Zhang Y, Zhang D, Yonezawa T, Hasegawa M, Zhong Y, Liu W, Zhang Y, Huang Z, Zhang S, Long R, Yang H, Wang J, Lenstra JA, Cooper DN, Wu Y, Wang J, Shi P, Wang J, & Liu J (2012). The yak genome and adaptation to life at high altitude. Nature genetics, 44 (8), 946-9 PMID: 22751099