How science works.

I am exhausted by the constant necessity of defending discovery and exploration science as being as legitimate as (and, in my opinion, far more influential than) “the” scientific method of testing very focused hypotheses. I am especially tired of having to reassure my graduate students that, despite what other faculty insist on implying, what they are doing is good science. As you can imagine, I am thrilled to see that the excellent resource site Understanding Science (a follow-up to their equally wonderful Understanding Evolution) agrees with me.

The alternative that they propose is pretty complex, but so is the real world.


If you’re one of those people who thinks there is one scientific method and that it involves only testing hypotheses, then you owe it to yourself and to your (and my) students to read this resource.

Flaws of the fudge factor.

A nice property of blogs is that they provide an opportunity to share ideas that, for one reason or another, never made it to publication. Here is a letter to Science that I wrote a few years ago which was bounced. It is in response to a piece called “Testing hypotheses: prediction or prejudice” by Peter Lipton (Science 307: 219-221; see replies by several others in Science 308: 1409-1412).

Flaws of the fudge factor

Lipton’s (1) analysis of common explanations for the intuitive notion that prediction is more convincing than accommodation is interesting and apt. His own contribution to the discussion – namely that differences in the opportunity for “fudging” in support of a favoured hypothesis drive this disparity – is less compelling. A few things to consider in this regard:

1) Who can fudge? Lipton (1) argues that “fudging” is more problematic in accommodation than in prediction, but this appears to be based on the assumption that the predictions are made by one scientist and tested independently by another. Certainly, this pertains to some fields of physics, but in a great many instances in the life sciences, the predictor and tester are one and the same. Thus, there is ample opportunity for (unintentional) “fudging” of predictive tests in support of a preferred hypothesis at all stages, from experimental design to data collection to interpretation. Moreover, one could make the argument that “tests” based on accommodation, which use data generated independently by many other authors and in which errors are random with respect to the hypothesis being tested, are more – not less – objective. It also should be self-evident that if an author ignores contradictory data during an accommodation-based test, then other members of the scientific community are likely to expose this omission.

2) A question of scale. If the comparison being made is in regard to a single datum, then indeed prediction may be far more convincing than accommodation. However, if the scales differ, as they often do in real life, such that the prediction relates to a tractable and therefore simple test but the accommodation deals with multiple, independent types of information, then the latter may be considered much more convincing. Therefore, it matters a great deal that prediction often tests one aspect of an issue whereas accommodation can include a much broader “consilience of inductions” (2).

3) Prediction as confirmation. The history of science is replete with examples in which predictive testing, while undoubtedly particularly convincing, has served mainly to confirm ideas that had been developed primarily on the basis of accommodation. This has included some of the most revolutionary breakthroughs in science, including relativity, natural selection, atomic theory, and genetics (both Mendelian and molecular). The question of which is the superior way to advance science is a false dichotomy.

4) An alternative hypothesis: x + y > x. Perhaps the most important point overlooked by Lipton (1) is that hypotheses that are not compatible with existing data – i.e., which have not already achieved accommodation – are rejected before ever being tested by prediction. In this sense, a more parsimonious explanation for the perception that prediction is more powerful than accommodation is that the former is a second-order test that only occurs after the first has been completed. In other words, prediction seems more powerful because its very existence implies that a more fundamental criterion, accommodation, has already been met. It is a mere truism that two tests will be more convincing than one, especially if the second test indicates inherently that the first has been passed.

References
1. Lipton, P. Testing hypotheses: predictions and prejudice. Science 307: 219-221, 2005.

2. Whewell, W. The Philosophy of the Inductive Sciences, Founded Upon Their History, 1840.

What would I do with more research support? Part One: Background.

One of the great joys of being a scientist is that we get to spend our lives exploring the aspects of the natural world that most intrigue and excite us. However, the equally great frustration of being a researcher is that our curiosity and passion invariably outstrip the resources available for our explorations. It often feels like we spend the bulk of our creative energy begging for money, and when this is declined — as it often is — it can be crushing. What keeps us going is the conviction that what we are doing, and what we have not yet found a way to do, is interesting and important and worth pursuing.

The primary focus of my research is the evolution of genome size in animals. Genome size is the amount of DNA in one copy of the chromosome set of a species, generally measured in terms of the number of base pairs (bp) or in mass (in picograms, or 10-12g). What makes this an intriguing topic of research is the enormous variability that exists across species: in animals, genome sizes range more than 7,000-fold. Think about that for a moment. Some animals have 7,000 times more DNA in their cells than others. Even within vertebrates, there is huge diversity at the genomic level: the largest (lungfish) is 350 times larger than the smallest (pufferfish). Or consider amphibians, which range about 120-fold from the smallest in some frogs to the largest in a few aquatic salamanders.

The human genome contains about 3.2 billion base pairs. In the simplest terms, one might expect this to be the largest genome of all — humans are the most complicated organisms (right?) and that should require the most genes (right?) which in turn means more DNA (right?). This was indeed the assumption when researchers began assessing genome sizes in the late 1940s — before the structure of DNA was elucidated, and even before it had been established that DNA is the hereditary molecule. At this time it was reported that the amount of DNA in a species’ cells is mostly constant (thus, genome size is also called “C-value”). This itself was suggested to indicate that DNA, and not protein, serves as the molecular basis of inheritance. However, it was also obvious by 1951 that the amount of DNA varies dramatically among species, and that the “complexity” of an animal and its genome size are decoupled. There are, it was discovered, salamanders with 40x more DNA per genome than in humans. This made no sense. DNA amount is constant within species because it is what genes are made of, and yet more complicated organisms (which presumably require more genes) may have substantially less DNA in their genomes than simpler organisms. This became known as the “C-value paradox” in the early 1970s.

It was not long before the apparent “paradox” was resolved: most DNA in animal and plant genomes is not genes (it is “non-coding DNA”). This means that genome size need not be related to the number of protein-coding genes, and that there is no reason to expect more complex animals to have more DNA in their genomes. However, this raised many new questions: What is this non-coding DNA? Where does it come from? How does it increase or decrease in amount in different genomes? Does it have any effect on the organism? Does it have any function? Why do some species have so much of it and others so little?

Despite several decades of research, most of these questions remain at best only partially answered. This is where my lab’s research comes in. We are interested in genome size diversity across all animals, in its effects on organism biology, and in the factors ranging in scale from individual DNA elements to ecological properties that accentuate or constrain amounts of DNA in the genomes of different species.

One thing that has become clear over the past several decades is that genome size is not randomly distributed across taxa. Some, like birds, all seem to have relatively small genomes. Others, like salamanders, all have large genomes. The quantity of DNA also relates to important features such as cell size and cell division rate, such that large genomes are found in cells that are big and divide slowly. Because all animals are made of cells, this means that any feature relating to cell size or cell division rate could be indirectly related to genome size. Body size is an obvious possibility, at least when cell numbers are held mostly constant. Metabolic rate is another possibility, because the larger a cell gets, the lower its relative surface area is, and this can influence gas exchange. Developmental rate is yet another, because slower individual cell divisions can add up to protracted development overall.

We have found that body size is correlated with genome size not only in some invertebrates like flatworms and copepod crustaceans, but also within specific groups of vertebrates like rodents, bats, and birds. Inverse relationships between genome size and metabolic rate have been reported in both mammals and birds, and in particular it has been argued that flight imposes a constraint on genome size due to its high metabolic demands. This latter idea has been around for several years, but it has recently become the subject of renewed interest and some intriguing new discoveries. For example, my colleague Chris Organ has used fossil cell size measurements to reveal that theropod dinosaurs (the lineage from which birds evolved) already had somewhat reduced genome sizes relative to other lineages before birds evolved, and that pterosaurs (the first vertebrates to evolve flight) also had small genomes. One of my students has been working on flight in birds, and showed that wing parameters associated with flight ability are related to genome size as well. We have also found recently that hummingbirds have the smallest genomes among birds (this isn’t published yet, but we’re writing the paper as we speak).

In terms of development, we have found in insects like lady beetles and vinegar flies that larger genomes are associated with slower overall development. Similar correlations have been known for some time in amphibians. What is more interesting is the pattern that we see with regard to metamorphosis, which represents a period of rapid and extreme physical reorganization. Groups with intensive metamorphosis, like frogs living in deserts that complete their life cycle quickly during wet seasons, have very small genomes (smaller than birds). Others, like aquatic salamanders that have lost the ability to metamorphose, have some of the largest genomes among animals. This also seems to apply to the major lineages of insects. Orders exhibiting complete metamorphosis (“holometabolous development”) appear almost never to exceed about 2 billion base pairs, whereas some without complete metamorphosis (“hemimetabolous development”) can be very large — there are grasshoppers with 5x more DNA than in humans.

Although genome size has been investigated for more than 60 years, some of these trends are only now coming to light. One reason is that we are focusing on the “big picture” now. Another reason is that we have technology that allows us to estimate genome sizes for large numbers of species. To give one example, an undergraduate student and I produced new data for more than 300 species of moths last summer alone. Previously, only 50 moth species had been analyzed (almost all of them in a pilot study I did a few years ago). Of course, this is a miniscule fraction of the 180,000 or so described species in the order, but it’s infinitely better than no information at all. Various students of mine have begun filling other major gaps, including in mammals, birds, insects, worms, and molluscs, but a huge amount of work remains just to get a basic picture of genomic diversity and its significance.

Over the upcoming series of posts, I will highlight some of the projects that I am very interested in undertaking, but which are on indefinite hold due to lack of funds. (It’s not that I haven’t tried — but granting agencies tend not to like this kind of large-scale “discovery” science as compared to the testing of very focused hypotheses). There are several reasons why I think it is worth doing this. First, most members of the public get only snippets of what goes on in research labs, most often provided by news reports. The raw curiosity that drives basic research is not often conveyed, particularly when projects are first conceived (vs. once they’re completed and published). Second, this is the stuff that gets me out of bed in the morning, and I hope that others can share in the excitement that my students and I feel when we think about, and try to answer, these fundamental questions about the diversity of life. Third, I believe it is useful for people to grasp the frustration that every scientist lives with when he or she feels that there are great ideas collecting dust for simple lack of funds. Finally, it provides an opportunity to talk about some intriguing animal groups from a perspective that most people haven’t considered. In that sense, it should be an interesting exercise in thinking about the wondrous biological diversity that surrounds us.

In the meantime, you are welcome to explore the Animal Genome Size Database to get a sense of the tremendous diversity — and glaring gaps in our knowledge — that drive my research program.

Adopt-a-lab.

While I am still daydreaming, how ’bout this? Once the economy is in better shape, have a program set up where individuals/companies can sponsor a specific science lab. Someone could put together a list of labs and what they do, and if anything sparks the interest of a sponsor, he or she could fund it. I don’t think people are aware just how little money many labs in Canada receive, and therefore just how much of a difference some outside assistance could make. The average starting grant from NSERC in many panels is $25K-$30K per year. So, the number of students that a lab can support and the amount of work they can do could pretty much be doubled for the change in your average corporate couch. Basic science does not have to be totally dependent on grants — maybe it’s time to become more creative.

Private foundations could save Canadian science.

Having spent time at a large museum in the US, I am familiar with the common occurrence of privately endowed “[generous person or foundation] lab of [awesome research]” initiatives. The culture of funding basic research is not as well developed in Canada as it is in the US, so these are much less commonly found up here, at least in terms of support for specific researchers working on basic science. Given the extreme challenges facing scientists in this nation, especially those who are working on basic research (vs. technology, health, environment, etc.), this could be an area where private funding could really help. Obviously, this would be contingent on foundations with an interest in basic science, but it’s something that seems to me increasingly needed as grants shrink, frustration builds, and brains drain. Funding agencies tend to shy away from high-risk (what we would call interesting and novel) projects, and are focusing more and more on projects that have commercial potential or some other applied use. It is very rare to be able to explore without being strongly constrained in some way.

Quotes of interest – ERVs.

It has been quite some time since the last update to the Quotes of interest series on junk DNA. Most of the posts have sought to demonstrate that the exhausting cliché that scientists dismissed possible functions for non-coding DNA until recently is false. Therefore, I have provided many quotes indicating that many (if not most) biologists continued to consider possible functions for various non-coding elements throughout the mythical period of neglect. This time, I want to discuss an example in which a particular kind of non-coding sequence was considered as probably non-functional — but because of knowledge about its biology, not because no function could be imagined.

The elements under discussion are endogenous retroviruses (ERVs) which, as the name suggests, are viral-like sequences that exist within the genome. Depending on who you ask, they are either very similar to or are interchangeable with long terminal repeat (LTR) retrotransposons. ERVs make up approximately 8% of the human genome, while LTR elements account for 50-80% of the maize genome.

Endogenous retroviruses were discovered in the 1960s and 1970s (see Weiss 2006), but were first dubbed “endogenous viruses” by David Baltimore in 1974 (published in Baltimore 1975).

Here is how Baltimore (1975) explained their origin:

Evidence has accumulated that viruses have entered the germ line at various times during the ancestry of different species. For convenience, two different cases can be considered: acquisition of viral genomes during inbreeding or domestication and acquisition of viral genomes during the evolution of a species. In principle, viruses could have become part of an animal genome at any stage of evolution and still be detectable now.

Baltimore (1975) discussed the fact that these “endogenous viruses” generally do not grow well in the species in which they had been identified, and that they often show signs of degrading by mutation. Moreover, being clearly similar to viruses and sometimes causing diseases, it seemed very unlikely that they were maintained because they conferred some functional advantage to their hosts. As he concluded,

It is my guess that these viruses have no positive function to play in the life of the animals in which they are resident. Rather, there is an evolutionary equilibrium balancing their acquisition and loss. The viruses are being inserted into the germ line at very low frequency, after which they require many thousands or millions of years to be mutated away because they have little or no detrimental effect on the animal in which they are resident. Viruses that did have a detrimental effect would be lost rapidly and might never come to our attention.

It is worth noting that Baltimore (1975) does not cite Ohno (1972), makes no reference to “junk DNA”, and reaches a tentative conclusion about lack of function from his consideration of the origin and properties of the elements.

Today, some examples are known of ERVs with beneficial effects, such as in placental development (Mi et al. 2000) and p53 binding sites (Wang et al. 2007). (Note, however, that only 0.5% of identified ERVs are associated with binding sites). As Weiss (2006) summarized the present situation:

As Mendelian elements, retroviruses must be subject to host selection. However, with the exception of enrolling env genes in placental differentiation, ERV appear to be parasitic DNA sequences for which the host has little use, other than to protect against further retrovirus infection. Potentially, ERV can damage the host by mutational insertion and by homologous recombination. But despite a tendency to implicate ERV in many ‘non-infectious’ diseases in humans, there is scant evidence that they play a significant role. There are only rare examples where a recessive single gene disorder in a family lineage is caused by an endogenous retroviral insertion disrupting gene function.

Seems like Baltimore’s (1975) assessment was largely correct with regard to mammalian ERV sequences.

_________

Baltimore, D. (1975). Tumor viruses: 1974. Cold Spring Harbor Symposia on Quantitative Biology 39: 1187-1200.

Mi, S. et al. (2000). Syncytin is a captive retroviral envelope protein involved in human placental morphogenesis. Nature 403: 785-789.

Wang, T. et al. (2007). Species-specific endogenous retroviruses shape the transcriptional network of the human tumor suppressor protein p53. Proceedings of the National Academy of Sciences USA 104: 18613-18618.

Weiss, R.A. (2006). The discovery of endogenous retroviruses. Retrovirology 3: 67.

____________

Part of the Quotes of interest series.