Hype vs. content.

I think I should clarify my position on the human evolution acceleration issue, because I don’t want my comments about nonsense in the press to be misconstrued as a rejection of the study. The basic theoretical arguments make good sense, and I am eagerly awaiting (peer-reviewed) commentary regarding the particular method and dataset. As far as whether I accept the possibility of recent selection among humans, allow me to start by showing three slides that I use in my section on human history. These were presented three weeks ago, before I knew anything about the Hawks et al. paper.



If the Hawks et al. study holds up by next fall, I will add some slides about it specifically, because I think it nicely brings together ideas about adaptive peaks, population size, and selection.

My problem, as noted, is with the hype in the press, most of it direct quotes of the authors. Over at evolgen, Hawks suggests that “Most [bloggers] seem to be reacting viscerally to the idea that evolution could ‘accelerate’ by as much as we claim. That’s a deficiency of theory and/or knowledge about human history, which we’re trying as hard as we can to make a dent in.”

Maybe so. But that’s not the case with me, as should be obvious from above. But let’s evaluate how well these authors are succeeding in clarifying misconceptions about human evolution.

“History looks more and more like a science fiction novel in which mutants repeatedly arose and displaced normal humans – sometimes quietly, by surviving starvation and disease better, sometimes as a conquering horde. And we are those mutants.”

Now imagine you’re not a biologist. Your concept of “mutant” is based on what you know from science fiction. And a scientist tells you that human history really did include stampeding hordes of mutants, right out of science fiction. Does this help you to understand that it is an *allele* that is a mutant, and what really happened is that some of us, though more than in the past, are descendants who carry these alleles?

“Five thousand years is such a small sliver of time – it’s 100 to 200 generations ago,” he says. “That’s how long it’s been since some of these genes originated, and today they are in 30 or 40 percent of people because they’ve had such an advantage. It’s like ‘invasion of the body snatchers.’”

Again with the science fiction. I have absolutely no idea how allele frequencies changing *over many generations* by normal vertical inheritance is anything remotely like body snatchers.

“We found very many human genes undergoing selection,” says anthropologist Gregory Cochran of the University of Utah, a member of the team that analyzed the 3.9 million genes showing the most variation. “Most are very recent, so much so that the rate of human evolution over the past few thousand years is far greater than it has been over the past few million years.”

A few million years ago there we no humans. Six (or so) million years ago, there would have been one species that eventually branched into the lineage of which humans are currently the only representative, and the one of which chimps and bonobos are the only extant examples. The *rate* of allele frequency change may be the highest it has ever been, but these statements are unbelievably misleading.

“In the last 40,000 years humans have changed as much as they did in the previous 2 million years.”

Modern humans, as a species, have existed for about 200,000 years or so. More than 90% of the history covered in this claim includes Homo habilis and Homo erectus. Allele frequency changes *may* be faster and more abundant now than during all that time, but this is hardly as significant under the normal conception of “evolutionary change” as the actual origin of two new species, including our own.

Their paper may be great. What they have been stating about it is irresponsible.

Evolution with a bullet.

There is a lot of buzz about the recent (but still unavailable [update: link]) PNAS paper by John Hawks et al. reporting an accelerated rate of natural selection in humans. This time, I am not going to pick on the media who, predictably, are selling this as a conclusive finding when those of us in the scientific community have not even had a chance to read the paper yet, let alone for anyone to try to critically assess it. It may be fantastic work, and if it holds up I will certainly give it a significant place in my lecture on human evolutionary history. My complaint is about what the authors themselves have been telling the press.

“Ten thousand years ago, no one on planet Earth had blue eyes,” Hawks notes, because that gene—OCA2—had not yet developed. “We are different from people who lived only 400 generations ago in ways that are very obvious; that you can see with your eyes.”

Interesting idea. But I suppose at no time in history could there have been another variant that caused blue eyes? Is this really the sort of thing that can only evolve once and in one way?

“We aren’t the same as people even 1,000 or 2,000 years ago,” he says, which may explain, for example, part of the difference between Viking invaders and their peaceful Swedish descendants.

Um, ok. I wonder if blue eyes make you peaceful?

Harpending says genetic differences among different human populations “cannot be used to justify discrimination. Rights in the Constitution aren’t predicated on utter equality. People have rights and should have opportunities whatever their group.”

And, by implication, the groups are not utterly equal. Care to speculate on which groups are more equal than which others?

The new study comes from two of the same University of Utah scientists – Harpending and Cochran – who created a stir in 2005 when they published a study arguing that above-average intelligence in Ashkenazi Jews – those of northern European heritage – resulted from natural selection in medieval Europe, where they were pressured into jobs as financiers, traders, managers and tax collectors. Those who were smarter succeeded, grew wealthy and had bigger families to pass on their genes. Yet that intelligence also is linked to genetic diseases such as Tay-Sachs and Gaucher in Jews.

No comment.

“History looks more and more like a science fiction novel in which mutants repeatedly arose and displaced normal humans – sometimes quietly, by surviving starvation and disease better, sometimes as a conquering horde. And we are those mutants.”

Michael Crichton’s latest: LACTASE, the harrowing story of a small mutation that conferred a slightly better ability to digest milk and reached a higher frequency in some human populations. Expect the movie in summer 2010.

“Five thousand years is such a small sliver of time – it’s 100 to 200 generations ago,” he says. “That’s how long it’s been since some of these genes originated, and today they are in 30 or 40 percent of people because they’ve had such an advantage. It’s like ‘invasion of the body snatchers.’”

Genes, alleles. Tomayto, tomahto. Either way, they’re out to take us over!

“We are always trying to outrun disease.”

And body snatchers.

“Natural selection cares about how many children you have. People will have kids younger and younger.”

Where’s Bart Simpson when you need him? Natural selection is not conscious. Natural selection is not conscious. Natural selection is not conscious. Natural selection is n…

“Genetic engineering will make all this irrelevant. If people want green-haired kids they will go to the doctor and get them in 100 years.”

No they won’t, because people will be marrying robots by then.

“We are more different genetically from people living 5,000 years ago than they were different from Neanderthals.”

MNSdfnklcn. Oops, sorry… that was Coke sprayed all over my keyboard.

“In the last 40,000 years humans have changed as much as they did in the previous 2 million years.”

Nxjbjbecbc. Dammit… again!

“We found very many human genes undergoing selection,” says anthropologist Gregory Cochran of the University of Utah, a member of the team that analyzed the 3.9 million genes showing the most variation. “Most are very recent, so much so that the rate of human evolution over the past few thousand years is far greater than it has been over the past few million years.” [emphasis added]

Really? A few million years ago there were no humans at all.


Authors use inappropriate terminology in "lower" paper.

I know that many medically-oriented geneticists don’t understand even the basics of evolution, but do they have to make it so painfully clear?

Schlegel A, Stainier DYR. (2007). Lessons from “lower” organisms: what worms, flies, and zebrafish can teach us about human energy metabolism. PLoS Genet 3(11): e199 doi:10.1371/journal.pgen.0030199

Some tidbits:

Recent studies using the worm Caenorhabditis elegans, the fly Drosophila melanogaster, and the zebrafish Danio rerio indicate that these “lower” metazoans possess unique attributes that should help in identifying, investigating, and even validating new pharmaceutical targets for these diseases.


As will be discussed below, unbiased methods have been used to identify more genes whose mutation in lower metazoans leads to phenotypes that are comparable to human syndromes of altered energy homeostasis like obesity.

Rather, studies on energy homeostasis in C. elegans, Drosophila, and zebrafish are proving that genetically tractable lower organisms can alter our understanding of the relationship of metabolic processes underlying obesity and its related illnesses (atherosclerotic vascular disease and type 2 diabetes mellitus).


Genome size, code bloat, and proof-by-analogy.

I recently did an interview with New Scientist for what, I am happy to say, was one of the most reasonable popular reviews of “junk DNA” that has appeared in recent times (Pearson 2007). My small section appeared in a box entitled “Survival of the fattest”, in which most of the discussion related to diversity in genome size and its causes and consequences. It even included mention of “the onion test“, which I proposed as a tonic for anyone who thinks they have discovered “the” functional explanation for the existence of vast amounts of non-coding DNA within eukaryotic genomes. Also thrown in, though not because I said anything about it, was a brief analogy to computer code: “Computer scientists who use a technique called genetic programming to ‘evolve’ software also find their pieces of code grow ever larger — a phenomenon called code bloat or ‘survival of the fattest'”.

I do not follow the literature of computer science, though I am aware that “genetic algorithms” (i.e., program evolution by mutation and selection) is a useful approach to solving complex puzzles. When I read the line about code bloat, my impression was that it probably gave other readers an interesting, though obviously tangential, analogy by which to understand the fact that streamlined efficiency of any coding system, genetic or computational, is not a given when it is the product of a messy process like evolution.

More recently, I have been made aware of an electronic article published in the (non-peer-reviewed) online repository known as arXiv (pr. “archive”; the “X” is really “chi”) that takes this analogy to an entirely different level. Indeed, the authors of the paper (Feverati and Musso 2007) claim to use a computer model to provide insights into how some eukaryotic genomes become so bloated. That is, instead of applying biological observations (i.e., naturally evolving genomes can become large) to a computational phenomenon (i.e., programs evolved in silico can become large, too), the authors flipped the situation around and decided that a computer model could provide substantive information about how genomes evolve in nature.

I will state up front that I am rarely (read: never) convinced by proof-by-analogy studies. Yes, modeling can be helpful if it provides a simplified way to test the influence of individual parameters in complex systems, but only insofar as the conclusions are then compared against reality. When it comes to something like genome size evolution, which applies to millions of species (billions if you consider that every species that has ever lived, about 99% of which are extinct, had a genome) and billions of years, one should be very skeptical of a model that involves only a handful of simplified parameters. This is especially true if no effort is made to test the model in the one way that counts: by asking if it conforms to known facts about the real world.

The abstract of the Feverati and Musso (2007) article says the following:

The development of a large non-coding fraction in eukaryotic DNA and the phenomenon of the code-bloat in the field of evolutionary computations show a striking similarity. This seems to suggest that (in the presence of mechanisms of code growth) the evolution of a complex code can’t be attained without maintaining a large inactive fraction. To test this hypothesis we performed computer simulations of an evolutionary toy model for Turing machines, studying the relations among fitness and coding/non-coding ratio while varying mutation and code growth rates. The results suggest that, in our model, having a large reservoir of non-coding states constitutes a great (long term) evolutionary advantage.

I will not embarrass myself by trying to address the validity of the computer model itself — I am but a layman in this area, and I am happy to assume for the sake of argument that it is the single greatest evolutionary toy model for Turing machines ever developed. It does not follow, however, that the authors are correct in their assertion that they “have developed an abstract model mimicking biological evolution”.

As I understand it, the simulation is based on devising a pre-defined “goal” sequence, similarity to which forms the basis of selecting among randomly varying algorithms. As algorithms undergo evolution by selection, they tend to accumulate more non-coding elements, and the ones that reach the goal most effectively turn out to be those with an “optimal coding/non-coding ratio” which, in this case, was less than 2%. The implication, not surprisingly, is that genomes evolve to become larger because this improves long-term evolvability by providing fodder for the emergence of new genes.

Before discussing this conclusion, it is worth considering the assumptions that were built into the model. The authors note that:

For the sake of simplicity, we imposed various restrictions on our model that can be relinquished to make the model more realistic from a biological point of view. In particular we decided that:

  1. non-coding states accumulate at a constant rate (determined by the state-increase rate pi) without any deletion mechanism [this is actually two distinct claims rolled into one],
  2. there is no selective disadvantage associated with the accumulation of both coding and non-coding states,
  3. the only mutation mechanism is given by point mutation and it also occurs at a constant rate (determined by the mutation rate pm),
  4. there is a unique ecological niche (defined by the target tape),
  5. population is constant,
  6. reproduction is asexual.

As noted, I am fine with considering this a fantastic computer simulation — it just isn’t a simulation that has any resemblance to the biological systems that it purports to mimic. Consider the following:

  • Although some authors have suggested that non-coding DNA accumulates at a constant rate (e.g., Martin and Gordon 1995), this is clearly not generally true. All extant lineages can trace their ancestries back to a single common ancestor, and thus all living lineages (though not necessarily all taxonomic groups) have existed for exactly the same amount of time. And yet the amount of non-coding DNA varies dramatically among lineages, even among closely related ones. Ergo, the rate of accumulation of non-coding DNA differs among lineages. Premise 1 is rejected.
  • The insertion of non-coding elements can be selectively relevant not only in terms of effects on protein-coding genes (many transposable elements are, after all, disease-causing mutagens), but also in terms of bulk effects on cell division, cell size, and associated organism-level traits (Gregory 2005). Premise 2 is rejected.
  • The accumulation of non-coding DNA in eukaryotes does not occur by point mutation, except in the sense that genes that are duplicated may become pseudogenized by this mechanism. Indeed, the model seems only to involve a switch between coding and non-coding elements without the addition of new “nucleotides”, which makes it even more distant from true genomes. Moreover, the primary mechanisms of DNA insertion, including gene duplication and inactivation, transposable element insertion, and replication and recombination errors, do not occur at a constant rate. In fact, the presence of some non-coding DNA can have a feedback effect in which the likelihood of additional change is increased, be it by insertions (e.g., into non-coding regions, such that mutational consequences are minimized) or deletions (e.g., illegitimate recombination among LTR elements) or both (e.g., unequal crossing over or replication slippage enhanced by the presence of repetitive sequences). Premise 3 is rejected.
  • Evolution does not have a pre-defined goal. Evolutionary change occurs along trajectories that are channeled by constraints and history, but not by foresight. As long as a given combination of features allows an organism to fill some niche better than alternatives, it will persist. Not only this, but models like the one being discussed are inherently limited in that they include only one evolutionary process: adaptation. Evolution in the biological world also occurs by non-adaptive processes, and this is perhaps particularly true of the evolution of non-coding DNA. It is on these points that the analogy between evolutionary computation and biological evolution fundamentally breaks down. Premise 4 is rejected in the strongest possible terms.
  • Real populations of organisms are not constant in size, though one could argue that in some cases they are held close to the carrying capacity of an available niche. However, this assumes the existence of only one conceivable niche. Real populations can evolve to exploit different niches. Premise 5 is rejected.
  • With a few exceptions (e.g., DNA transposons), transposable elements are sexually transmitted parasites of the genome, and these elements make up the single largest portion of eukaryotic genomes (roughly half of the human genome, for example). Ignoring this fact makes the model inapplicable to the very question it seeks to address. Premise 6 is rejected.

The main problem with proofs-by-analogy such as this is that they disregard most of the characteristics that make biological questions complex in the first place. Non-coding DNA evolves not as part of a simple, goal-directed, constant-rate process, but one typified by the influence of non-adaptive processes (e.g., gene duplication and pseudogenization), selection at multiple levels (e.g, both intragenomic and organismal), and open-ended trajectories. An “evolutionary” simulation this may be, but a model of biological evolution it is not.

Finally, it is essential to note that “non-coding elements make future evolution possible” explanations, though invoked by an alarming number of genome biologists, contradict basic evolutionary principles. Natural selection cannot favour a feature, especially a potentially costly one such as the presence of large amounts of non-coding DNA, because it may be useful down the line. Selection occurs in the here and now, and is based on reproductive success relative to competing alternatives. Long-term consequences are not part of the equation except in artificial situations where there is a pre-determined finish line to which variants are made to race.

That said, there can be long-term consequences in which inter-lineage sorting plays a role. In terms of processes such as alternative splicing and exon shuffling, which rely on the existence of non-coding introns, an effect on evolvability is plausible and may help to explain why lineages of eukaryotes with introns are so common (Doolittle 1987; Patthy 1999; Carroll 2002). However, this is not necessarily linked to total non-coding DNA amount. For a process of inter-lineage sorting to affect genome size more generally, large amounts of non-coding DNA would have to be insufficiently detrimental in the short term to be removed by organism-level selection, and would have to improve lineage survival and/or enhance speciation rates, such that over time one would observe a world dominated by lineages with huge genomes. In principle, this would be compatible with the conclusions of the model under discussion, at least in broad outline. In practice, however, this is undone by evidence that lineages with exorbitant genomes are restricted to narrower habitats (e.g., Knight et al. 2005), are less speciose (e.g., Olmo 2006), and may be more prone to extinction (e.g., Vinogradov 2003) than those with smaller genomes.

Non-coding DNA does not accumulate “so that” it will result in longer-term evolutionary advantage. And even if this explanation made sense from an evolutionary standpoint, it is not the effect that is observed in any case. No computer simulation changes this.

__________

References

Carroll, R.L. 2002. Evolution of the capacity to evolve. Journal of Evolutionary Biology 15: 911-921.

Doolittle, W.F. 1987. What introns have to tell us: hierarchy in genome evolution. Cold Spring Harbor Symposia on Quantitative Biology 52: 907-913.

Feverati, G. and F. Musso. 2007. An evolutionary model with Turing machines. arXiv.0711.3580v1.

Gregory, T.R. 2005. Genome size evolution in animals. In: The Evolution of the Genome (edited by T.R. Gregory). Elsevier, San Diego, pp. 3-87.

Knight, C.A., N.A. Molinari, and D.A. Petrov. 2005. The large genome constraint hypothesis: evolution, ecology and phenotype. Annals of Botany 95: 177-190.

Martin, C.C. and R. Gordon. 1995. Differentiation trees, a junk DNA molecular clock, and the evolution of neoteny in salamanders. Journal of Evolutionary Biology 8: 339-354.

Olmo, E. 2006. Genome size and evolutionary diversification in vertebrates. Italian Journal of Zoology 73: 167-171.

Patthy, L. 1999. Genome evolution and the evolution of exon shuffling — a review. Gene 238: 103-114.

Pearson, A, 2007. Junking the genome. New Scienist 14 July: 42-45.

Vinogradov, A.E. 2003. Selfish DNA is maladaptive: evidence from the plant Red List. Trends in Genetics 19: 609-614.

___________

Update: The author’s responses are posted and addressed here.