Skip to content

R Inference & Modeling: Bayesian Statistics

Abish Pius edited this page Mar 6, 2020 · 1 revision

Bayesian Statistics Introduction

  • In the urn model, it does not make sense to talk about the probability of p being greater than a certain value because pis a fixed value.
  • With Bayesian statistics, we assume that p is in fact random, which allows us to calculate probabilities related to p.
  • Hierarchical models describe variability at different levels and incorporate all these levels into a model for estimating p.

Bayes' Theorem

  • Bayes' Theorem states that the probability of event A happening given event B is equal to the probability of both A and B divided by the probability of event B:
    Pr(A|B) = {Pr(B|A)Pr(A)}/Pr(B)
  • Bayes' Theorem shows that a test for a very rare disease will have a high percentage of false positives even if the accuracy of the test is high.
  • Examples can be found here.

Code: Monte Carlo Bayes' Theorem

prev <- 0.00025    # disease prevalence
N <- 100000    # number of tests
outcome <- sample(c("Disease", "Healthy"), N, replace = TRUE, prob = c(prev, 1-prev))

N_D <- sum(outcome == "Disease")    # number with disease
N_H <- sum(outcome == "Healthy")    # number healthy

# for each person, randomly determine if test is + or -
accuracy <- 0.99
test <- vector("character", N)
test[outcome == "Disease"] <- sample(c("+", "-"), N_D, replace=TRUE, prob = c(accuracy, 1-accuracy))
test[outcome == "Healthy"] <- sample(c("-", "+"), N_H, replace=TRUE, prob = c(accuracy, 1-accuracy))

table(outcome, test)

Bayes Theorem in Practice

  • The techniques we have used up until now are referred to as frequentist statistics as they consider only the frequency of outcomes in a dataset and do not include any outside information. Frequentist statistics allow us to compute confidence intervals and p-values.
  • Frequentist statistics can have problems when sample sizes are small and when the data are extreme compared to historical results.
  • Bayesian statistics allows prior knowledge to modify observed results, which alters our conclusions about event probabilities.

The Hierarchical Model

  • Hierarchical models use multiple levels of variability to model results. They are hierarchical because values in the lower levels of the model are computed using values from higher levels of the model.
  • We model baseball player batting average using a hierarchical model with two levels of variability:
    • p~N(μ,τ) describes player-to-player variability in natural ability to hit, which has a mean μ and standard deviation τ.
    • Y|p~N(p,σ) describes a player's observed batting average given their ability p which has a mean p and standard deviation σ = sqrt(p(1-p)/N). This represents variability due to luck.
    • In Bayesian hierarchical models, the first level is called the prior distribution and the second level is called the sampling distribution.
  • The posterior distribution allows us to compute the probability distribution of p given that we have observed data Y
  • By the continuous version of Bayes' rule, the expected value of the posterior distribution p given Y=y is a weighted average between the prior mean μ and the observed data Y:
    E(p|y) = Bμ +(1-B)Y where B = σ^2/(σ^2+τ^2)
  • The standard error of the posterior distribution SE(p|Y)^2 = 1/(1/σ^2 +1/τ^2). Note that you will need to take the square root of both sides to solve for the standard error.
  • This Bayesian approach is also known as shrinking. When σ is large, B is close to 1 and our prediction of p shrinks towards the mean μ. When σ is small, B is close to 0 and our prediction of p is more weighted towards the observed data Y.

Clone this wiki locally