Showing posts with label diversification. Show all posts
Showing posts with label diversification. Show all posts

Saturday, December 28, 2013

Maximum marginal likelihood

In my last post, I mentioned the possibility of comparing alternative diversification models, conceived of as alternative priors on divergence times in a molecular dating context, using Bayes factors.

As an alternative to Bayes factors, one could integrate the likelihood over rates and times and maximize the resulting integrated (marginal) likelihood as a function of the diversification parameters (e.g. speciation and extinction rates, let us call them $\theta$). We can then use standard likelihood ratio tests to compare diversification models. Numerically maximizing and calculating the likelihood is a bit tricky, but this is just a technical problem, not a conceptual one.

Finally, for a given diversification model, and once an estimate $\hat \theta$ of its parameters has been obtained by maximum likelihood, divergence times can be inferred based on the plug-in posterior distribution, i.e. the posterior distribution on divergence times obtained by fixing the diversification parameters at $\hat \theta$.

This type of approach is sometimes called maximum marginal likelihood, or ML-II in the literature (Berger, 1985). It may also be called empirical Bayes, thus emphasizing the fact that the estimation of divergence times is still Bayesian in spirit, although now based on an empirically determined plug-in prior. However, for me, "empirical Bayes" has a more general meaning -- after all, the prior on divergence times is also empirically determined in the fully Bayesian approach. For that reason, I prefer the "maximum marginal likelihood" terminology.

Maximum marginal likelihood or full Bayes, which is best ? I am not really sure.

Conceptually speaking, the maximum marginal likelihood approach does not invoke any second stage prior on $\theta$, which can be seen as an advantage if one is not too sure about the choice of the prior nor about its philosophical meaning (subjective? uninformative?, etc). Maximum marginal likelihood may also be better for comparing models, as Bayes factors tend to be sensitive to the prior.

Computationally speaking, on the other hand, I find it easier to put a prior on $\theta$, because this makes the MCMC easier to implement.

Finally, there is a classical argument saying that the fully Bayesian approach has the advantage of integrating uncertainty about $\theta$. However, this again depends on the second-stage prior, a point that certainly requires further discussion, but that's for another day.

---

As for now, we can draw an interesting parallel here. Indeed, all this looks very much like another likelihood approach used in population genetics for estimating the scaled mutation rate (let's call it again $\theta = 4 N_e u$) using intra-specific sequence data (Kuhner et al, 1995). In that case, the sequence data are assumed to be generated by a mutation process along a genealogy, which is itself described by Kingman's coalescent. The probability of the sequence data given the genealogy is integrated over the coalescent and then maximized as a function of the scaled mutation rate $\theta$ (all this by Monte Carlo).

Here also, once $\theta$ has been estimated, one can infer the age of the ancestor of the sample by calculating the plug-in posterior distribution over this age (Thomson et al, 2000). Here also, instead of maximizing with respect to $\theta$, one can decide to put a second-stage prior on $\theta$ and then sample from the joint posterior over $\theta$ and the coalescence times (this is what is done in Beast, for instance, Drummond and Rambaut, 2007).

Altogether, there is a very close parallel between the two situations:
species / individuals
substitutions / mutations
phylogeny / genealogy
prior on divergence times / coalescent
diversification parameters / scaled mutation rate
divergence time of the last common ancestor / age of the ancestor of the sample.

The hierarchical structure of the two models is virtually the same. In both cases, we can either maximize the marginal likelihood with respect to $\theta$, or put a second-stage prior on $\theta$ and sample from the joint posterior over $\theta$ and divergence or coalescence times.

I find this interesting because the two situations are so similar, and yet we tend to consider them as essentially different. In particular, we call the birth-death process a prior on divergence times, and we consider divergence times as parameters of the model. On the other hand, coalescence times are not usually considered as parameters, and I doubt that we would call the coalescent a prior.

I do not think that the status of the intermediate level of the model (the distribution over divergence or coalescence times) should depend on wether we use a maximum likelihood or a Bayesian method for estimating the upper level of the model (the parameter $\theta$). Instead, I believe that these distinctions are merely cultural. They are just a historical accident: one of the two methods is originally a Bayesian approach, which could be modified and turned into a maximum marginal likelihood procedure. The other is a frequentist maximum likelihood method that has only secondarily been endowed with a second-stage prior and implemented in a Bayesian software program.

In any case, all this shows that, in everyday life, there is much more overlap between Bayesian inference and maximum likelihood than the traditional vision in terms of a fundamental opposition between frequentists versus subjectivists would seem to suggest. At the end, the differences are rather subtle.

Finally, this little comparison allows me to anticipate on another problem that will be the subject of a future post: how do you determine the number of parameters of a model?

Berger, J. O. (1985). Statistical Decision Theory and Bayesian Analysis (1985 ed.). New-York: Springer-Verlag.

Drummond, A. J., & Rambaut, A. (2007). BEAST: Bayesian evolutionary analysis by sampling trees. BMC Evolutionary Biology, 7, 214.

Kuhner, M. K., Yamato, J., & Felsenstein, J. (1995). Estimating effective population size and mutation rate from sequence data using Metropolis-Hastings sampling. Genetics, 140(4), 1421–1430.

Thomson, R., Pritchard, J. K., Shen, P., Oefner, P. J., & Feldman, M. W. (2000). Recent common ancestry of human Y chromosomes: evidence from DNA sequence data. Proceedings of the National Academy of Sciences of the United States of America, 97(13), 7360–7365.

Friday, December 27, 2013

Two sides of the same coin

One of the things that people want to do with phylogenies is to estimate parameters or test hypotheses about species diversification processes (Nee et al, 1992). The idea has been revisited recently, for testing for the presence of diversity dependence (Etienne et al 2012), or for estimating variation in speciation or extinction rates over time (Morlon et al, 2011, Stadler 2011).

In order to study diversification processes, however, one should first estimate a time-calibrated phylogeny.

Time-calibrated phylogenies are usually estimated using Bayesian methods.

However, a Bayesian method requires a prior on divergence times. The choice of the prior on divergence times has a strong impact on divergence times estimation. It will therefore potentially also have a strong impact on the outcome of the test of alternative diversification processes.

More fundamentally, some priors on divergence times have themselves an interpretation in terms of an underlying diversification process: the birth-death prior (Yang and Rannala, 1996), for instance, amounts to assuming constant speciation and extinction rates. Thus, testing diversification models based on a time-calibrated phylogeny that has itself been estimated assuming a given diversification model is either circular (if the two diversification models are identical) or contradictory (if the models are different).

So, what should we do ?

Obviously, what we could do here is, use the alternative diversification models that we want to test as alternative priors on divergence times. We can then compare the resulting alternative models, for instance using Bayes factors. By doing so, we will simultaneously (1) integrate the uncertainty about divergence times in our comparison of alternative diversification models (while avoiding the circularity issues mentioned above) and (2) infer divergence times under several diversification models, typically deciding to keep those obtained under the best fitting model.

Priors which can be interpreted in terms of macro-evolutionary processes are what I would call mechanistic priors. They are not meant to be uninformative priors, like the uniform prior on divergence times (although I am not totally sure that the uniform prior on divergence times is really uninformative), nor subjective priors.

Trying to derive mechanistic priors is, I think, an interesting and constructive answer to prior sensitivity issues. It is a risky business, because mechanistic priors tend to make strong assumptions and therefore potentially lack robustness. On the other hand, doing this is potentially more insightful in the long term.

Also, it naturally leads to an integration of different levels of macro-evolutionary studies. In the present case, what we obtain is an elegant statistical formalization of the idea that molecular dating and diversification studies are in fact two sides of the same coin.

--

Etienne, R. S., Haegeman, B., Stadler, T., Aze, T., Pearson, P. N., Purvis, A., & Phillimore, A. B. (2012). Diversity-dependence brings molecular phylogenies closer to agreement with the fossil record. Proceedings of the Royal Society B: Biological Sciences, 279:1300–1309.

Morlon, H., Parsons, T. L., & Plotkin, J. B. (2011). Reconciling molecular phylogenies with the fossil record. Proceedings of the National Academy of Sciences, 108:16327–16332.

Nee, S., Mooers, A. O., & Harvey, P. H. (1992). Tempo and mode of evolution revealed from molecular phylogenies. Proceedings of the National Academy of Sciences of the United States of America, 89:8322–8326.

Rannala, B., & Yang, Z. (1996). Probability distribution of molecular evolutionary trees: a new method of phylogenetic inference. Journal of Molecular Evolution, 43:304–311.

Stadler, T. (2011). Mammalian phylogeny reveals recent diversification rate shifts. Proceedings of the National Academy of Sciences, 108:6187–6192.