"Because" versus "so that".

I want to make a quick point about how evolution works and how it does not. The reason is that two stories about non-coding DNA posted today include a major misconception about evolution. Unfortunately, this is a misconception attributed in the articles to biologists, so I can only imagine what the state of comprehension is among non-scientists.

The distinction is between “because” and “so that”. In evolution, things evolve “because,” meaning that there are causes and effects that can be identified. Why are some strains of bacteria resistant to antibiotics? Because a mutation that occurred that happened to be beneficial under the conditions of antibiotic treatment became common in the population over the course of several generations. By contrast, things do not evolve “so that”. Bacteria do not experience mutations so that they will become resistant to antibiotic agents.

Why is there so much non-coding DNA? Because transposable elements spread, or because there are accidental duplications that are not eliminated by selection, or because of the interaction of some other mutational processes and their consequences (or lack thereof). So much non-coding DNA did not evolve so that it might someday be useful, or so that it could be coopted when needed, or so that evolution would have more potential in the form of genetic raw materials.

So why, then, do we see quotes like these?

Wired One Scientist’s Junk Is a Creationist’s Treasure:

“I’ve stopped using the term [‘junk’],” Collins said. “Think about it the way you think about stuff you keep in your basement. Stuff you might need some time. Go down, rummage around, pull it out if you might need it.”

Reuters Human instruction book not so simple: studies:

“It is not the sort of clutter that you get rid of without consequences because you might need it. Evolution may need it,” [Collins] said.

That little extra padding might be just what an animal needs to adapt to some unforeseen circumstance, the researchers said. “They may become useful in the future,” Birney said.

The latter quote by Ewan Birney illustrates the problem that can arise when a detailed, nuanced discussion is summarized into a short soundbite. I know this from experience, and I suspect that this is what has happened here, given how his very reasonable interpretation is paraphrased in New Scientist ‘Junk’ DNA makes compulsive reading:

Birney says that the additional switches may be mutations that appear by accident and then generate new slugs of RNA, but because they are produced randomly, most are evolutionarily neutral ‘passengers’ in the genome. There might be rare occasions, however, when a new RNA does confer an advantage.

Collins, on the other hand, seems to have said his bit to two different reporters, so I strain to give him the benefit of the doubt on this one. When I began this blog, I did not think I would be pointing out obvious misconceptions about evolution, genomes, and DNA as propagated by the likes of Collins or Nature. But here we are.


Junk DNA gets Wired.

There is a new article on the Wired website about junk DNA [One Scientist’s Junk Is a Creationist’s Treasure]. I make a very brief appearance in it, and I just want to clarify what I meant by the statement cited (I’m still learning that even an hour-long interview might result in only a short blurb).

My quote is “Function at the organism level is something that requires evidence”. I make this statement because there are several different sorts of DNA sequences in the genome whose presence can be explained even if they do not benefit (and indeed, even if they slightly harm) the organism carrying them. Pseudogenes, satellite DNA, transposable elements (45% of our genome), and other non-coding sequences may or may not be functional — that requires evidence — and some may exist as a result of accidental duplication or even due to selection at the level of the elements themselves (by “intragenomic selection”). The old assumption that all non-coding DNA must be beneficial to the organism or it would have been deleted by now ignores genome-specific processes by which non-coding DNA evolves.

As I have discussed previously, both hardcore adaptationists (if any exist anymore) and creationists have a vested interest in having all non-coding DNA be functional. I believe that real-world variability in genome size argues strongly against such a prospect, but of course it is possible, and this is the point that people like Ohno, Doolittle, Orgel, and Crick made in the 1980s. The important point is that yes, some non-coding DNA is functional at the organism level (as opposed to existing for its own sake or because there is no strong selection against it). And certainly, non-coding DNA has effects at the organism level. But current evidence suggests that about 5% of the human genome is functional, and even the least conservative ENCODE participants (whose primary, and important, objective is to identify the functional elements and their features) are betting that 20% is functional.

In the end, it is obvious that non-coding DNA is the product of evolution whether it all turns out to be functional or not. The cases in which former parasites (transposons) have taken on function at the organism level are a perfect illustration of cooption, which is the same basic process that allows explanations for the evolution of complex structures like eyes or flagella. The research into function of non-coding DNA, which the creationists are eager to cite, can be carried out only under an evolutionary framework — it is meaningless to talk about “conserved non-coding DNA sequences” otherwise.

Finally, let me say one thing about Francis Collins’s quote: “Think about it the way you think about stuff you keep in your basement. Stuff you might need some time. Go down, rummage around, pull it out if you might need it.” With all due respect (which is considerable, given his contribution to the Human Genome Project), it makes no sense to explain the existence of non-coding DNA because it might someday prove useful. Evolution does not work that way. Elements might be coopted, but maintaining this option explains neither the origin nor the persistence of non-coding sequences.

As to what the creationists have to say, well, I leave that to others with more (or less?) patience to attend to.

____________

Updates:


Decoding the blueprint. Sigh.

The results of the proof-of-principle phase of ENCODE, the Encyclopedia of DNA Elements Project, appear in the June 14 issue of Nature. It’s a very interesting project, and it has revealed a few more surprises (or at least, added evidence in favour of previously surprising observations). I will probably post more about it soon, but for the time being let me just offer a brief apology to the science writers out there whom I have given a hard time about invoking sloppy language to describe non-coding DNA, sequencing, and genomes (recent example, but one I will leave alone, ‘Junk’ DNA makes compulsive reading online at New Scientist).

The reason I am sorry is that I simply cannot hold you to a higher standard than is maintained by one of the most prestigious journals on planet Earth. You see, Nature has decided to depict the ENCODE project on the cover as “Decoding the Blueprint”. Needless to say (again), genomes are not blueprints (as the ENCODE project shows!) and no one is decoding anything at this point.

I have said all this before, and even I am getting tired of my complaints about it. Thus, I will focus only on the interesting science in a later post.

Sigh.


Am I a MacGregor?

The name “Gregory” is used as both a first name and a surname, and I wish I had a nickel for every time someone said “No, your last name” after I told them my name was “Gregory”. Jokes about having two (actually, three) “first” names have been a staple in my life as well.

There have been 16 popes with the name “Gregory”, including Pope Gregory I (“Gregory the Great”, which, had it not been taken, would have been a nickname I would have aspired to myself; he can keep “Saint Gregory”). Think “Gregorian calendar” (Pope Gregory XIII) or “Gregorian chants” (though these are probably not actually a product of Pope Gregory I). Readers with a snarkier side may consider this blog an example of “Gregorian rants” if they so desire.

There are many derivatives of the name “Gregor”, of which “Gregory” is one. It appears to date back to the Latin “Gregorious” and the Greek “Gregorios”, meaning “alert, watchful, or vigilant”. When my father and stepmother were in Greece, they were often told that they had a “very good Greek name”. Other languages have their own versions as well.

When I was living in the west end of London (specifically, the “London Borough of Richmond-Upon-Thames“), I would have my hair cut by a fantastic old-school barber, an ex-merchant marine who lived in a long boat on the Thames and who did the final trim on one’s neck with a straight razor. On my first visit, he remarked that I “must have Scottish blood”. The reason, apparently, had to do with my thick hair and reddish goatee. “What’s your surname?” he asked. “Gregory,” I replied. “Well there you go,” he said.

You see, the other, more circuitous origin of the name “Gregory” is via the Scottish Clan MacGregor (meaning “son of Gregor”, and thus linked back to the Latin/Greek origin). It seems the MacGregors ran afoul of King James VI, who made bearing the MacGregor name a capital offence in 1603. You may be familiar with subsequent adventure involving the “Scottish Robin Hood”, Rob Roy MacGregor, as portrayed on screen by Liam Neeson (who is not a Scottish folk hero at all, but a Northern Irish Jedi).

When given the choice between changing their names or being executed, most MacGregors opted for the former. The resulting names, which numbered more than 100 and of which Gregory was one of the more obvious, became septs of Clan MacGregor. The ban on the name MacGregor was lifted in 1774, but the division into different septs remains.

Today, my fellow DNA Network member Blaine Bettinger of The Genetic Genealogist reports on an effort by the Clan Gregor Society to use DNA to reunite the Clan MacGregor.

The idea of the MacGregor DNA Project is to draw comparisons to a genetic profile from a known descendant of the chief’s line (known only as “kit 2124”). Anyone who shares 31 out of the 37 DNA markers with this individual will be given full membership in the Clan Gregor Society, regardless of current surname. Gregory is one of a few surnames focused on explicitly as part of the project.

The project primarily is making use of Y-chromosome loci, which would mean that only descendants related through their father’s side would register. It appears that some mitochondrial DNA analysis is also being conducted, which would identify individuals related through descent on their mother’s side.

As per the old tradition in our society, I received my surname from my father and, as per the old tradition in biology, I also received my Y chromosome from him. In other words, it would be perfectly feasible for me to take the test and see if my red beard is homologous to that of Rob Roy.

But really, what’s the point? I am not Scottish, I am Canadian, and I am perfectly happy with that identity. Moreover, like a great many North Americans, I represent a mixture of many different families: Gregory, Davis, Sager, MacKenzie, and who knows what else (though I confess that the ingredients here are pretty limited in their variety, coming as they all do from the British Isles). It’s really only because of a quirk of our culture that I associate almost exclusively with Gregory.

Still, it would be pretty cool to wear an official clan tartan…


Genomes large and small.

The past few years have witnessed the discovery of both very large and small genomes in different groups of organisms. Here are some highlights from this research.

The first represents the largest genome so far reported for a crustacean, in the Arctic-dwelling amphipod Ampelisca macrocephala. The genome of this small invertebrate is a whopping 63.2 billion base pairs, or about 20 times larger than the human genome (Rees et al. 2007). Again, this sort of observation should dispel the notion that all non-coding DNA is functional for protecting against mutagens or some such thing.

The second interesting finding is of the largest viral genome so far discovered. The virus, dubbed Mimivirus, was sufficiently odd that it was originally assumed to be a bacterium when first observed, but on closer examination was found to be a virus. Its genome size is estimated as 1.2 million bases, which is larger than the genome of many bacteria (Raoult et al 2007). So, now there is overlap in reported genome sizes between viruses and bacteria, which goes along with the known overlap between the genome sizes of bacteria and eukaryotes (Gregory 2005).

And now for some small genomes. More specifically, the smallest flowering plant genome, that of Genlisea margaretae at a mere 63 million base pairs, less than half the size of the previous record holder, Arabidopsis thaliana at about 157 million base pairs. This increases the range in angiosperm genome sizes to more than 2,000-fold. (In animals the total range is about 3,300-fold; Gregory et al. 2007).

The smallest insect genome so far estimated was reported fairly recently as well. It belongs to Caenocholax fenyesi, a twisted-wing parasite, and is a mere 108 million base pairs (Johnston et al. 2004). Not to spoil the fun, but my lab has also found genome sizes this small in other groups, though these have not yet been published. The largest insect genome size known is found in the mountain grasshopper Podisma pedestris at 16.6 billion base pairs (Westerman et al. 1987).

The smallest eukaryotic genome known to date is that of the protist Encephalitozoon intestinalis, a parasitic microsporidian with a genome size of only 2.3 million base pairs, which is smaller than that of many bacteria (Vivarès and Méténier 2000). The smallest free-living eukaryote genome size is found in Ostreococcus tauri at 12.6 million base pairs (Derelle et al. 2006). The largest reliable protozoan genome size estimate reported to date is 97.8 billion base pairs in the dinoflagellate Gonyaulax polyedra (Shuter et al. 1983). That is a more than 33,000-fold range among protists.

It should be pointed out that the largest published eukaryote genome size estimate is 1,400 billion base pairs (400 times larger than human) in the free-living amoeba Chaos chaos (Friz 1968), although the largest genome size is often attributed to Amoeba dubia at 700 billion base pairs based on the same study. These data are not generally considered reliable, for several reasons. First, these values for amoebae were based on rough biochemical measurements of total cellular DNA content, which probably includes a significant fraction of mitochondrial DNA. Second, Friz’s (1968) value of 300pg for Amoeba proteus is an order of magnitude higher than those reported in subsequent studies (Byers 1986). Third, some amoebae (e.g., A. proteus) contain 500-1000 small chromosomes and are quite possibly highly polyploid (Byers 1986), in which case these values would be inappropriate for a comparison of haploid genome sizes among eukaryotes.

Finally, the smallest genome so far known for any cellular organism also was discovered recently — that of the endosymbiotic bacterium Carsonella ruddii at a miniscule 159,662 base pairs (Nakabachi et al. 2006). This species resides within specialized cells inside the body of psyllid insect hosts. The genome is so small, and the insect and bacterium so mutually dependent, that this species blurs the lines between bacteria and organelles, and probably is similar in some ways to an intermediate stage in the evolution of other obligate intracellular symbionts turned organelles like mitochondria and chloroplasts.

The old assumption, still often repeated, that viruses have smaller genomes than bacteria which have smaller genomes than single-celled eukaryotes which have smaller genomes than multicellular eukaryotes is beginning to wear thin. The pattern remains in a general sense, but focusing only on such a coarse scale overlooks a significant amount of diversity within, and increasingly apparent overlap between, groups of life.

___________

Readers interested in exploring genome size data can check out the various online databases for more.

References

Byers, T.J. 1986. Molecular biology of DNA in Acanthamoeba, Amoeba, Entamoeba, and Naegleria. International Review of Cytology 99: 311-341.

Derelle, E., C. Ferraz, S. Rombauts, P. Rouzé, A.Z. Worden, S. Robbens, F. Partensky, S. Degroeve, S. Echeynié, R. Cooke, Y. Saeys, J. Wuyts, K. Jabbari, C. Bowler, O. Panaud, B. Piégu, S.G. Ball, J.P. Ral, F.Y. Bouget, G. Piganeau, B. De Baets, A. Picard, M. Delseny, J. Demaille, Y. Van de Peer, H. Moreau. 2006. Genome analysis of the smallest free-living eukaryote Ostreococcus tauri unveils many unique features. Proceedings of the National Academy of Sciences of the USA 103: 11647-11652.

Friz, C.T. 1968. The biochemical composition of the free-living amoebae Chaos chaos, Amoeba dubia, and Amoeba proteus. Comparative Biochemistry and Physiology 26: 81-90.

Gregory, T.R. 2005. Synergy between sequence and size in large-scale genomics. Nature Reviews Genetics 6: 699-708.

Gregory, T.R., J.A. Nicol, H. Tamm, B. Kullman, K. Kullman, I.J. Leitch, B.G. Murray, D.F. Kapraun, J. Greilhuber, and M.D. Bennett. 2007. Eukaryotic genome size databases. Nucleic Acids Research 35 (Suppl. 1): D332-D338.

Greilhuber, J., T. Borsch, K. Müller, A. Worberg, S. Porembski, W. Barthlott. 2006. Smallest angiosperm genomes found in lentibulariaceae, with chromosomes of bacterial size. Plant Biology 8: 770-777.

Johnston, J.S., L.D. Ross, L. Beani, D.P. Hughes, and J. Kathirithamby. Tiny genomes and endoreduplication in Strepsiptera. Insect Molecular Biology 13: 851-585.

Nakabachi A, A. Yamashita, H. Toh, H. Ishikawa, H.E. Dunbar, N.A. Moran, and M. Hattori. 2006. The 160-kilobase genome of the bacterial endosymbiont Carsonella. Science 314: 267.

Raoult, D., B. La Scola, and R. Birtles. 2007. The discovery and characterization of Mimivirus, the largest known virus and putative pneumonia agent. Clinical Infectious Diseases 45: 95-102.

Rees, D.J., F. Dufresne, H. Glémet, and C. Belzile. 2007. Amphipod genome sizes: first estimates for Arctic species reveal genomic giants. Genome 50: 151-158.

Vivarès, C.P. and G. Méténier 2000. Towards the minimal eukaryotic parasitic genome. Current Opinion in Microbiology 3: 463–467.

Westerman, M., N.H. Barton, and G.M. Hewitt (1987). Differences in DNA content between two chromosomal races of the grasshopper Podisma pedestris. Heredity 58: 221-228


Two-for-one misconceptions about genomes from the New York Times.

To date, two identified human beings have had their genomes sequenced: J. Craig Venter and James D. Watson. Venter’s was completed in draft form in 2001 and the final version was completed recently. Watson received his genome sequence on disk (a hard drive, not a DVD as reported) from Jonathan Rothberg, founder of 454 Life Sciences, at Baylor College of Medicine yesterday. You can watch the presentation here.

The notion that individual people can have their genomes sequenced (still for about $2 million, but the cost will fall precipitously in the future) is sure to elicit some interesting discussions about medical applications, ethical implications, and intriguing research into human variation. Certainly, the completion of Watson’s genome sequence has already gained media attention. Unfortunately, the same old catchphrases and errors abound. Apparently, even the mighty combined forces of Genomicron, Evolgen, and Sandwalk are insufficient to stop this.

Today, both RPM of Evolgen and Jonathan Badger at T. taxus take aim at the New York Times, who not only confuse sequencing with “deciphering”, but think that Watson discovered DNA in 1953 (Genome of DNA Discoverer Is Deciphered by Nicholas Wade).

To clarify, DNA (“nuclein”) was discovered by Friedrich Miescher in 1869. Watson and Crick elucidated the double helix structure of DNA in the 1950s, based on the results of decades of work on the chemical properties of the molecule by a large number of researchers.

I give full credit to Watson and Crick for their monumental contribution, which rightly garnered them the 1962 Nobel Prize. But credit is also due to Miescher and the countless others whose work was integral to the subsequent rise of molecular genetics and genome sequencing.

Here are two headlines announcing the same story, one inaccurate and the other fine:

Genome of DNA Discoverer is Deciphered (New York Times)

Nobel Laureate James Watson Receives Personal Genome (ScienceDaily)

Is one less catchy than the other? It seems to me that getting the history and the science right would be relatively simple and would only add to the strength of a story.

____________

Updates:

The Genetic Genealogist mentions the story and argues that Nicholas Wade may not be responsible for the headline. Fair enough — my criticism is about the entire presentation, whether that be the fault of the author, editor, or other. It does bear noting, however, that Wade has used this terminology several times previously, including describing it in the main text as the “project to sequence, or decode, the genome.”

Sandwalk has opened a discussion about whether readers would (or, like Larry, would not) want to have their genomes sequenced.

DNADirectTalk repeats the standard inaccuracies.

I don’t think we’re going to be rid of the “decoding” analogy any time soon, especially since sequencers themselves use it. Venter has a book coming out in October, with the unfortunate title A Life Decoded: My Genome: My Life. (Wouldn’t The Sequence of My Life or My Life’s Sequence have been catchier anyway?). The US Department of Energy (which financed much of the Human Genome Project) still has it on their website Human Genome Research: Decoding DNA also. To be fair to science writers, we can’t hold them to a higher standard of terminological accuracy than applies to scientists. In other words, we need to clean it up on our side first and then, hopefully, reporters will follow our lead.