Skip to content

R Stats: Random Variables

Abish Pius edited this page Feb 23, 2020 · 1 revision

Random Variables

  • Random variables are numeric outcomes resulting from random processes.
  • Statistical inference offers a framework for quantifying uncertainty due to randomness.

Code: Random Variables Examples

# define random variable x to be 1 if blue, 0 otherwise
beads <- rep(c("red", "blue"), times = c(2, 3))
x <- ifelse(sample(beads, 1) == "blue", 1, 0)

# demonstrate that the random variable is different every time
ifelse(sample(beads, 1) == "blue", 1, 0)
ifelse(sample(beads, 1) == "blue", 1, 0)
ifelse(sample(beads, 1) == "blue", 1, 0)

Sampling Models

  • A sampling model models the random behavior of a process as the sampling of draws from an urn.
  • The probability distribution of a random variable is the probability of the observed value falling in any given interval.
  • We can define a CDF F(a)=Pr(S≤a) to answer questions related to the probability of S being in any interval.
  • The average of many draws of a random variable is called its expected value.
  • The standard deviation of many draws of a random variable is called its standard error.

Code: Sampling Model Example: Monte Carlo Losing Money on Roulette Wheel

# sampling model 1: define urn, then sample
color <- rep(c("Black", "Red", "Green"), c(18, 18, 2)) # define the urn for the sampling model
n <- 1000
X <- sample(ifelse(color == "Red", -1, 1), n, replace = TRUE)
X[1:10]

# sampling model 2: define urn inside sample function by noting probabilities
x <- sample(c(-1, 1), n, replace = TRUE, prob = c(9/19, 10/19))    # 1000 independent draws
S <- sum(x)    # total winnings = sum of draws
S
n <- 1000    # number of roulette players
B <- 10000    # number of Monte Carlo experiments
S <- replicate(B, {
    X <- sample(c(-1,1), n, replace = TRUE, prob = c(9/19, 10/19))    # simulate 1000 spins
    sum(X)    # determine total profit
})

mean(S < 0)    # probability of the casino losing money
library(tidyverse)
s <- seq(min(S), max(S), length = 100)    # sequence of 100 values across range of S
normal_density <- data.frame(s = s, f = dnorm(s, mean(S), sd(S))) # generate normal density for S
data.frame (S = S) %>%    # make data frame of S for histogram
    ggplot(aes(S, ..density..)) +
    geom_histogram(color = "black", binwidth = 10) +
    ylab("Probability") +
    geom_line(data = normal_density, mapping = aes(s, f), color = "blue")

Distributions vs Probability Distributions

  • A random variable X has a probability distribution function F(a) that defines Pr(X≤a) over all values of a.
  • Any list of numbers has a distribution. The probability distribution function of a random variable is defined mathematically and does not depend on a list of numbers.
  • The results of a Monte Carlo simulation with a large enough number of observations will approximate the probability distribution of X.
  • If a random variable is defined as draws from an urn:
    • The probability distribution function of the random variable is defined as the distribution of the list of values in the urn.
    • The expected value of the random variable is the average of values in the urn.
    • The standard error of one draw of the random variable is the standard deviation of the values of the urn.

Random Variable Notation

  • Capital letters denote random variables (X) and lowercase letters denote observed values (x).
  • In the notation Pr(X=x), we are asking how frequently the random variable X is equal to the value x. For example, if x=6, this statement becomes Pr(X=6).

Central Limit Theorem

  • The Central Limit Theorem (CLT) says that the distribution of the sum of a random variable is approximated by a normal distribution.
  • The expected value of a random variable, E[X]=μ, is the average of the values in the urn. This represents the expectation of one draw.
  • The standard error of one draw of a random variable is the standard deviation of the values in the urn.
  • The expected value of the sum of draws is the number of draws times the expected value of the random variable.
  • The standard error of the sum of independent draws of a random variable is the square root of the number of draws times the standard deviation of the urn.

Equations

These equations apply to the case where there are only two outcomes, a and b with proportions p and 1−p respectively. The general principles above also apply to random variables with more than two outcomes.

Expected value of a random variable:
ap+b(1−p)

Expected value of the sum of n draws of a random variable:
n×(ap+b(1−p))

Standard deviation of an urn with two values:
∣b–a∣*sqrt{p(1−p)}

Standard error of the sum of n draws of a random variable:
sqrt{n} *∣b–a∣ * sqrt{p(1−p)}

Clone this wiki locally