Hits.

I began this blog largely as an experiment in public outreach. It is my belief that many people are interested in science, and that they would like an opportunity to interact with practicing scientists in a blog format. In this regard, I have been very happy to come across the blogs of other front line researchers, such as Jonathan Eisen’s The Tree of Life, Rosie Redfield’s RRResearch, John Dennehy’s The Evilutionary Biologist, Rod Page’s iPhylo, and John Logsdon’s Sex, Genes, and Evolution (Best. Title. Ever.), along with those of quite a number of active grad students.

The question was, would anyone visit my blog, in light of established (and unabashedly political and correspondingly popular) options such as PZ Myers’s Pharyngula (like anyone still needs a link) and Larry Moran’s Sandwalk, or the excellent science reporting of Carl Zimmer’s The Loom?

Well, after just under three weeks in the blogosphere, Genomicron has received over 1,250 hits from roughly 850 unique visitors. Not bad for an upstart. At least, the null hypothesis that a blog is not useful for outreach has taken a bit of a thrashing. Thanks to those who have stopped by, and I hope to see you again soon.


Comments on "Noncoding DNA and Junk DNA" (re-post).

The following is a re-post of my comments on the recently posted Noncoding DNA and Junk DNA at Sandwalk. Needless to say, I am quite pleased to see such active discussion about non-coding DNA. Passages in italics are excerpts from the original article.

TR Gregory said…

Ryan Gregory has serious doubts about the usefulness of the term as he explains in his excellent article A word about “junk DNA”.

Just to clarify, I think the term could be useful — indeed, it was useful when Ohno coined it. The problem is that it is seldom used in an appropriate way. If the meaning were specified explicitly to be “regions strongly suspected of being non-functional with evidence to back it up” (which, incidentally, is not the original definition according to Ohno (1972) or Comings (1972)), and if people used it only in this way, then I would not have a problem with this. But given the difficulty that people seem to have in accepting that some DNA may truly not have a function at the organism level, I don’t know if we could ever get it to be used with such precision.

…a new term, Junctional DNA, to describe DNA that probably has a function but that function isn’t known… think we don’t need to go there. It’s sufficient to remind people that lots of DNA outside of genes has a function and these functions have been known for decades.

That neologism was suggested in response to Minkel’s appeal for a term that would “make the distinction between functional and nonfunctional noncoding DNA clear to a popular audience”. My main suggestion was to call DNA by what it is known to be, if at all possible, by function (“regulatory DNA”, “structural DNA”) or by type (“pseudogene”, “transposable element”, “intron”). Your definition of “junk DNA” is also more precise than most usages, meaning that you specify that the term only be applied to sequences for which there is evidence (not just assumption) of non-function. That leaves us with something in between for journalists to talk about with a catchy buzzword. “Junctional DNA” lets them specify that we’re not talking about “junk DNA” or “functional DNA” — i.e., there is some evidence for function (e.g., being conserved) but no evidence of what that function is. The main utility would be to stop the very frustrating leap that gets made from “this 1% of the genome may have a function, so the whole thing must have this function” kind of reporting. Now they could say “another 1% has moved into the category of ‘junctional DNA'”. I think that would be considerably less misleading than current wording.

Note that I’m avoiding the term “noncoding” DNA here. This is because to me the term “coding DNA” only refers to the coding region of a gene that encodes a protein … there are many genes for RNAs that are not properly called coding regions so they would fall into the noncoding DNA category … introns in eukaryotic genomes would be “noncoding DNA” as far as I’m concerned. I think that Ryan Gregory and others use the term “noncoding DNA” to refer to all DNA that’s not part of a gene instead of all DNA that’s not part of the coding region of a protein encoding gene. I’m not certain of this.

By definition, non-coding DNA is, and always has been, everything other than exons. The reason this is relevant is that early work in genome biology assumed that there should be a 1 to 1 correspondence between DNA content and protein-coding gene number. This is work that occurred for at least two decades before the discovery of introns, pseudogenes, and other non-coding DNA. Now we have more descriptive names for the categories of DNA that are not the genes, all the genes, and nothing but the genes. I actually don’t know of anyone else who would have a problem calling introns, pseudogenes, and regulatory regions “non-coding DNA”. Certainly, Ohno, Crick, and many others have historically put introns in the same non-protein-coding grouping as pseudogenes. It’s just a category — you also have more specific subcategories to apply to each of the types of non-coding DNA. Perhaps your objection relates to an undue emphasis on the distinction between exons and everything else — well, that’s the history of the past half century of this field, so it should be no surprise that the terminology reflects this.

Read Gregory’s article for the short concise version of this dispute. What it means is that junk DNA threatens the worldviews of both Dembski and Dawkins!

Not quite. What you’re leaving out of this is the possibility of multiple levels of selection. In the original edition of The Selfish Gene (1976, p.76), Dawkins argued that “the simplest way to explain the surplus DNA is to suppose that it is a parasite, or at best a harmless but useless passenger, hitching a ride in the survival machines created by the other DNA”. Cavalier-Smith (1977) drew a similar conclusion (before he had read Dawkins), and Doolittle and Sapienza (1980) and Orgel and Crick (1980) [yes, that Crick] independently developed the concept of “selfish DNA” a few years later. This is an explicitly multi-level selection approach because it specifies that non-coding DNA can be present due to selection within the genome rather than exclusively on the organism (or gene, in Dawkins’s case) (see, e.g., Gregory 2004, 2005). (Incidentally, this idea of parasitic DNA dates back at least to 1945, when Gunnar Östergren characterized B chromosomes in this fashion). Of course, they tended to do what Ohno did and applied this one idea to all non-coding DNA, which is too ambitious. The modern view is more pluralistic (see, e.g., Pagel and Johnstone 1992 vs. Gregory 2003). Some non-coding DNA is just accumulated “junk” (in the definition of evidence-supported non-function that you espouse). Some (perhaps most) is “selfish” or “parasitic” and persists because there is selection within the genome as well as on organisms (in fact, an argument could be, and has been, made that “selfish DNA” would be a much more accurate term than “junk DNA” for most non-coding DNA). Some non-coding DNA is clearly functional at the organism level, including regulatory regions and chromosome structure components. Some of these latter functional non-coding DNA sequences are derived from elements that originally were of one of the first two types, most notably transposable elements that take on a regulatory function through co-option (or, in another manner of thinking, that undergo a shift in level of selection).

Junk DNA is not noncoding DNA and anyone who claims otherwise just doesn’t know what they’re talking about.

I’m afraid I don’t follow what you mean here. By your definition, “junk DNA” is any non-functional sequence of DNA, including pseudogenes (i.e., the original meaning). Those sequences do not encode proteins. Hence, your version of junk DNA is non-coding. I think this reflects the confusion that is imposed by the term “junk DNA”, which is why I generally think it is more obfuscating than enlightening.

________

References

Cavalier-Smith, T. 1977. Visualising jumping genes. Nature 270: 10-12.

Comings, D.E. 1972. The structure and function of chromatin. Advances in Human Genetics 3: 237-431.

Dawkins, R. 1976. The Selfish Gene. Oxford University Press, Oxford.

Doolittle, W.F. and C. Sapienza. 1980. Selfish genes, the phenotype paradigm and genome evolution. Nature 284: 601-603.

Gregory, T.R. 2003. Variation across amphibian species in the size of the nuclear genome supports a pluralistic, hierarchical approach to the C-value enigma. Biological Journal of the Linnean Society 79: 329-339.

Gregory, T.R. 2004. Macroevolution, hierarchy theory, and the C-value enigma. Paleobiology 30: 179-202.

Gregory, T.R. 2005. Macroevolution and the genome. In The Evolution of the Genome (ed. T.R. Gregory), pp. 679-729. Elsevier, San Diego.

Ohno, S. 1972. So much “junk” DNA in our genome. In Evolution of Genetic Systems (ed. H.H. Smith), pp. 366-370. Gordon and Breach, New York.

Orgel, L.E. and F.H.C. Crick. 1980. Selfish DNA: the ultimate parasite. Nature 284: 604-607.

Östergren, G. 1945. Parasitic nature of extra fragment chromosomes. Botaniska Notiser 2: 157-163.

Pagel, M. and R.A. Johnstone. 1992. Variation across species in the size of the nuclear genome supports the junk-DNA explanantion for the C-value paradox. Proceedings of the Royal Society of London, Series B: Biological Sciences 249: 119-124.


Peer review.

John Dennehy has posted an interesting summary on The Evilutionary Biologist about professional peer review1. He notes, along with Marc Hauser and Ernst Fehr, that delays imposed by slow reviewers can be a significant source of frustration with the peer review process. The suggestion by Hauser and Fehr (2007) is to institute a system of punishments and rewards to get reviewers to submit reviews on schedule. Interesting idea, though I strongly oppose intentionally subjecting anyone’s work to delay as punishment, no matter how dawdling they are as reviewers. The scientific community at large should not be held back in order to punish specific individuals. There is also the obvious difficulty that reviewers may begin to substitute speed for quality in their review of manuscripts. I am currently reviewing four papers for four different journals. It will take time to get through them, and I hope to get them all in on time, but rushing them to meet a deadline won’t help the peer review process.

Long turnaround times are a real issue, and I have my own stories (one paper took over a year to show up in print). But my complaint regarding peer review comes as a reviewer rather than as an author. One of the biggest frustrations comes when one reviews a paper carefully, provides detailed comments, points out significant problems with the data, analysis, or interpretation, and recommends that the paper be rejected in its present format — and then it shows up in one’s mailbox again, unaltered, after simply having been submitted to a different journal, or worse, appears in print in another journal with none of the errors corrected. If anything shakes my confidence in the efficacy of peer review, it is this.

I understand full well the pressure to publish, but something has to be done about the tendency to submit a rejected paper — sometimes without even fixing typos that have been pointed out — to journal after journal (my current record is reviewing the same paper three times for three journals) until it gets through reviewers who are willing to let the mistakes slide or who lack the expertise to recognize the problems.

My suggested solution is that authors should be required to submit all previous reviews to any new journal to which they are sending the same paper. They should be required to show the editor that changes have been made or to justify why they have not. Otherwise, the peer review process is undermined, the quality of the science suffers, and the reviewers’ time is completely wasted.

End rant.

_________

Notes

1Not to be confused with the spectacle currently going on with regard to the flagellum paper.

References

Hauser M, Fehr E (2007) An incentive solution to the peer review problem. PLoS Biology 5: e107.


Genome size databases.

In case anyone is unaware of their existence, here are the links to the available genome size databases.

For a summary of the databases, see Gregory et al. (2007).

For a discussion about units of measurement in genome size, see here.

A summary of genome size ranges in various animals is available here.

A much smaller database of genome sizes that also includes some taxa besides animals, plants, and fungi is posted here.

For bacterial and archaeal (“prokaryote”) genome size data, see here and here and here.

For a list of completed and ongoing genome sequencing initiatives, see the Genomes OnLine Database (GOLD).

For vertebrate red blood cell sizes, see here.


Junctional DNA.

JR Minkel at the Scientific American blog has responded to the post on Evolgen about his earlier story regarding “junk DNA” (did you catch all that?). At the end of the post, he asks:

Scientists and scientist bloggers: Again, do you care [if journalists call it junk DNA]? If so, what term would you propose instead, or how would you make the distinction between functional and nonfunctional noncoding DNA clear to a popular audience?

Yes, I care, and here are my suggestions. If you mean the general category without any speculation either way about function, then it is simply and accurately “noncoding DNA”. If it has a function, then you specify what that function is: “regulatory DNA” or “structural DNA” or what have you. If the type of sequence is known, then you can use that as well or instead: “transposable elements” or “mobile DNA” or “pseudogenes” or “introns”. Maybe readers won’t know what those terms mean. This is a good opportunity to inform them.

What is missing is a term to describe a given collection of noncoding DNA for which there is thought to be some function, but for which that function and/or the type of sequence is unknown. This would reside somewhere between “junk DNA” (in the vernacular sense) and “functional DNA” (to which specific names can be applied). I therefore suggest the neologism “junctional DNA” to encompass this category. Note that Petsko (2003) suggested “funk DNA” to represent “functionally unknown DNA”, but I think “junctional DNA” is a little less, uh, funky.

Let me be even more specific. The proposed term “junctional DNA” derives from a dual etymology: 1) a simple portmanteau of “junk” and “functional”; 2) an indication that the sequences so described reside at the crossroads between DNA with no evident function and that with a clear function.

Two terms in one day — “the onion test” and “junctional DNA” — how ’bout that.

Incidentally, my annoyance with such reports has less to do with the terminology than with the fact that the highly conserved sequences in question make up about 5% of the total genome. To jump from this to imply that all noncoding DNA is recognized as functional is inappropriate and misleading. I also wish they would cite the source papers they reference; some of us would like to look up the primary material when we see a summary in a news story.

_______________

Update: Other bloggers (RPM of Evolgen in personal correspondence, Sandwalk) seem to think this term is not needed. I point out that this post was given in direct response to Minkel’s appeal for a term that would “make the distinction between functional and nonfunctional noncoding DNA clear to a popular audience”. In light of the fact that a journalist sees the need for such a term, and that it was coined in response to that need, I think ‘junctional DNA’ could be a useful term.


The onion test.

I am not sure how official this is, but here is a term I would like to coin right here on my blog: “The onion test”.

The onion test is a simple reality check for anyone who thinks they have come up with a universal function for non-coding DNA1. Whatever your proposed function, ask yourself this question: Can I explain why an onion needs about five times more non-coding DNA for this function than a human?

The onion, Allium cepa, is a diploid (2n = 16) plant with a haploid genome size of about 17 pg. Human, Homo sapiens, is a diploid (2n = 46) animal with a haploid genome size of about 3.5 pg. This comparison is chosen more or less arbitrarily (there are far bigger genomes than onion, and far smaller ones than human), but it makes the problem of universal function for non-coding DNA clear2.

Further, if you think perhaps onions are somehow special, consider that members of the genus Allium range in genome size from 7 pg to 31.5 pg. So why can A. altyncolicum make do with one fifth as much regulation, structural maintenance, protection against mutagens, or [insert preferred universal function] as A. ursinum?

Left, A. altyncolicum (7 pg); centre, A. cepa (17 pg); right, A. ursinum (31.5 pg).


There you have it. The onion test. To be applied to any ambitious claims that a universal function has been found for non-coding DNA.

____________

1 I do not endorse the use of the term “junk DNA”, which I think has deviated far too much from its original meaning and is now little more than a loaded buzzword; the descriptive term “non-coding DNA” is what I use to refer to the majority of eukaryotic sequences (of various types) that do not encode protein products.

2 Some non-coding DNA certainly has a function at the organismal level, but this does not justify a huge leap from “this bit of non-coding DNA [usually less than 5% of the genome] is functional” to “ergo, all non-coding DNA is functional”.



Genome size is good for you.

I imagine that every practicing scientist has experienced, in one form or another, the tendency of many non-scientists to expect all research to be directly beneficial to human health and well-being. I used to respond facetiously to these kinds of expectations when expressed by friends or family members, with something along the lines of “My work has absolutely no practical applications to human welfare whatsoever”.

Of course, this is not true. Genome size is becoming very relevant to fields of inquiry that are likely to have major significance for medicine. Notably, genome size data provide an important indication of the cost and difficulty of sequencing a given genome, and thus represent a prime criterion in the choice of sequencing targets. As an example, I performed a genome size estimate for Biomphalaria glabrata, a planorbid snail that serves as an intermediate host for the trematode flatworm Schistosoma mansoni which causes the debilitating disease known as schistosomiasis. The genome of B. glabrata is one of the smallest so far reported for a gastropod, and is now being sequenced (along with S. mansoni).

More recently, Jenner and Wills (2007) made explicit mention of genome size as an important factor in deciding on the next set of models for evo-devo studies. Discoveries regarding the fundamental genetic underpinnings of development have obvious implications for medical science and here, too, genome size is becoming increasingly seen as important. As they put it,

Whole-genome sequences are an increasingly important resource for many biological disciplines, including evo–devo15, 49, 50. However, financial and technical constraints mean that there is currently a preference for species with small genomes. This compounds the bias that is already introduced by the big six. First, putatively general conclusions about genome evolution might actually be specific to those smaller genomes that have been fully sequenced. For example, when focusing only on sequenced genomes, a close correspondence between genome size and gene number in eukaryotes is observed. The C-value paradox becomes apparent only when genome-size data from non-sequenced genomes is included51. Second, there are important genetic, morphological, physiological and ecological correlates of genome size in a range of animals and plants51, 52. Some correlates seem ubiquitous in animals and plants, such as those between genome size and cell size, body size and the inverse of developmental rate52. Others are group specific: genome size correlates mostly with metabolic rate in homeotherms, but with developmental type and ecology in amphibians53, and is positively correlated with egg size in copepods, plethodontid salamanders and fishes51, 52, 54. Studying these correlated traits in phylogenetically disparate taxa could illuminate the relationships between small genome size and rapid development, as well as the evolution of strongly cell-lineage-dependent development in taxa such as tunicates and nematodes, and the partial fragmentation of their Hox clusters55, 56.

References 51, 52, and 53 in that paragraph are papers of mine, so again I am forced to admit that my work may have some practical application after all.

My main focus is on genome size diversity in eukaryotes, which mostly means differences among species in the abundance of noncoding DNA. In bacteria, most of the genome is composed of protein-coding genes, so unlike in eukaryotes there is a very strong correlation between genome size and gene number. Genome size is generally small in parasites and endosymbionts and larger in free-living species (probably because population bottlenecks and relaxed selection on gene function result in gene loss by deletion bias in bacteria associated with hosts [Mira et al. 2001]).

But this observation is not the link between genome size and human health that I had in mind for this post. In this month’s issue of Antimicrobial Agents and Chemotherapy, Steven Projan argues that genome size is associated with the evolution of antibiotic resistance in bacteria. In Dr. Projan’s own words,

It is observed here that the ability of a given bacterium to evolve toward a multidrug resistance phenotype is a function of genome size. In Table 1, a number of examples are provided, but even an expanded analysis shows that this observation holds true. That is, the larger the genome the greater the propensity of a bacterium to display multidrug resistance phenotypes and the smaller the genome the less likely it is that antibacterial resistance will emerge and disseminate within that species. What is proposed here is that, just as there is a continuum of genome sizes among bacteria, there is a continuum in the ability or propensity of a bacterium to become “multidrug resistant” and that continuum is reflected in the size of the genome. This is not to say that we do not observe resistance to certain agents even in organisms with the smallest genomes (macrolide resistance appears in virtually every pathogen at some level). There is probably a solid biological reason for this observation; organisms with larger genomes are more adaptable to environmental changes because they have more (genetic) information to draw upon. It appears that organisms with smaller genomes have become more “specialized,” residing in particular environmental niches (Treponema pallidum and the Chlamydiae are cases in point), and their lack of versatility in adapting to different environments is also manifest in an inability to develop mechanisms for coping with antibiotics. Indeed, we have learned that virtually each and every time a bacterium either acquires a novel resistance determinant or a mutant strain arises with decreased susceptibility to an antibacterial drug, the bacterium experiences a “fitness burden.” With time, compensatory mutations are selected in which the bacterium accumulates mutations that allow for something like wild-type growth in a strain that is now phenotypically resistant (e.g., topA mutations in gyrB mutant strains). Bacteria with larger genomes simply have a greater opportunity to develop these compensatory mutations. It must be emphasized that it does not matter whether we are discussing the acquisition of a novel resistance gene as opposed to a mutation that alters the target or results in up-regulation of an efflux pump. The accumulating evidence tells us that all require some form of adaptation. Another consequence of this phenomenon is that antibiotic cycling in health care settings is unlikely to result in a reversion of the local microflora to susceptibility as the compensatory mutations “lock in” the resistance phenotype.

He continues by noting, “I and several of those I have discussed this observation with were perplexed that it had not previously been articulated. Although to be fair, others have suggested it is a trivial, if not nonsensical, observation and worthy only of cocktail party conversation… in fact, I believe that this is an important guide as to where and which organisms we actually need novel antibacterial agents for.” Projan blames an overemphasis on individual organisms with small genomes for the overlooking of this potentially important pattern. In other words, it is the sort of thing that can only be applied to human health research if one takes a broad view of genomic diversity.

As much fun as it is to study genome size for purely academic reasons, it seems it actually may be good for us too.


More interest in genome size.

The buzz on a few blogs today is the pending release of new books. Sort of the academic blogger equivalent to summer blockbusters, I suppose. In any case, it’s great to see that two of the eagerly anticipated items, Darwinian Detectives by Norman Johnson and The Origins of Genome Architecture by Michael Lynch, will both include significant space devoted to the topic of genome size. Not having read either book, it would not be prudent for me to recommend them to anyone (and it is no secret that I have problems with Lynch’s model, which is not the first and probably not the last one-dimensional explanation), but I do suggest that eyes be kept open for their arrival in June.

On another practical note, genome size is no longer just considered an important criterion for choosing genome sequencing targets, it has also been mentioned as directly relevant in the selection of the next wave of evo-devo models. It is also an interesting and important subject of investigation in its own right, of course.

So, while I may not ascribe to some of the explanations for genome size diversity that have been put forth of late, I am very glad to see that this is an active area of discussion that is gaining more attention every day.

__________

References

Evans, J.D. and D. Gundersen-Rindal. 2003. Beenomes to Bombyx: future directions in applied insect genomics. Genome Biology 4: 107.101-107.104.

Gregory, T.R. 2005. Synergy between sequence and size in large-scale genomics. Nature Reviews Genetics 6: 699-708.

Jenner, R.A. and M.A. Wills. 2007. The choice of model organisms in evo-devo. Nature Reviews Genetics 8: 311-319.

Pryer, K.M., H. Schneider, E.A. Zimmer, and J.A. Banks. 2002. Deciding among green plants for whole genome studies. Trends in Plant Sciences 7: 550-554.


Genomics, evolution, and health: comparisons of avian flu genomes.

An article by Steven Sternberg and colleagues is set to appear in the May issue of the journal Emerging Infectious Diseases. In it, the authors describe the results of complete genome sequence comparisons for 36 recent isolates of the avian flu virus (influenza H5N1). Their results “clearly depict the lineages now infecting wild and domestic birds in Europe and Africa and show the relationships among these isolates and other strains affecting both birds and humans”. More specifically,

The isolates fall into 3 distinct lineages, 1 of which contains all known non-Asian isolates. This new Euro-African lineage, which was the cause of several recent (2006) fatal human infections in Egypt and Iraq, has been introduced at least 3 times into the European-African region and has split into 3 distinct, independently evolving sublineages.


Figure 1. Phylogenetic tree of hemagglutinin (HA) segments from 36 avian influenza samples. A 2001 strain (A/duck/Anyang/AVL-1/2001) is used as an outgroup at top. Clade V1 comprises the 5 Vietnamese isolates at the bottom of the tree, and clade V2 comprises the 9 Vietnamese isolates near the top of the tree. The European-Middle Eastern-African (EMA) clade contains the remaining 22 isolates sequenced in this study; the 3 subclades are indicated by red, blue, and purple lines. The reassortant strain, A/chicken/Nigeria/1047–62/2006, is highlighted in red.

This is a study in phylogenetics — that is, it reconstructs evolutionary relationships among viral strains using the same tools that many evolutionary biologists use to study the relationships among species. It is well known that viruses evolve very rapidly, and tracking their their past changes contributes to the ability to predict future ones. As the authors conclude,

These findings show how whole-genome analysis of influenza (H5N1) viruses is instrumental to the better understanding of the evolution and epidemiology of this infection, which is now present in the 3 continents that contain most of the world’s population. This and related analyses, facilitated by global initiatives on sharing influenza data, will help us understand the dynamics of infection between wild and domesticated bird populations, which in turn should promote the development of control and prevention strategies.

Evolution is not something that only happened to the myriad fossil specimens housed in museum drawers, and evolutionary biology is not merely relevant to academics tucked away in research labs. Evolution is both an ongoing process and an active and exciting area of research. More than ever, an understanding of the processes involved is relevant to the well-being of people from all regions of the world.


Darwin’s death.

Today, April 19th, is the anniversary of Charles Darwin‘s death in 1882. I refer you to an excellent post by PZ Myers on Pharyngula about the details of Darwin’s passing [The Death of Darwin].

Darwin is buried at Westminster Abbey in London, within a few yards of Sir Isaac Newton. There is a bronze bust of Darwin as part of a memorial to several scholars near the grave that was installed by his family in 1888. The grave itself is very understated, a simple marble slab in the floor marking his name and the dates of his birth and death.


There is also a memorial to Darwin in Kent, where Down House is located, in the form of a sundial on the side of the local church.


Charles Robert Darwin, 12 February 1809 – 19 April 1882.