Repository navigation
R Stats: Random Variables
Abish Pius edited this page Feb 23, 2020
·
1 revision
- Random variables are numeric outcomes resulting from random processes.
- Statistical inference offers a framework for quantifying uncertainty due to randomness.
# define random variable x to be 1 if blue, 0 otherwise
beads <- rep(c("red", "blue"), times = c(2, 3))
x <- ifelse(sample(beads, 1) == "blue", 1, 0)
# demonstrate that the random variable is different every time
ifelse(sample(beads, 1) == "blue", 1, 0)
ifelse(sample(beads, 1) == "blue", 1, 0)
ifelse(sample(beads, 1) == "blue", 1, 0)- A sampling model models the random behavior of a process as the sampling of draws from an urn.
- The probability distribution of a random variable is the probability of the observed value falling in any given interval.
- We can define a CDF F(a)=Pr(S≤a) to answer questions related to the probability of S being in any interval.
- The average of many draws of a random variable is called its expected value.
- The standard deviation of many draws of a random variable is called its standard error.
# sampling model 1: define urn, then sample
color <- rep(c("Black", "Red", "Green"), c(18, 18, 2)) # define the urn for the sampling model
n <- 1000
X <- sample(ifelse(color == "Red", -1, 1), n, replace = TRUE)
X[1:10]
# sampling model 2: define urn inside sample function by noting probabilities
x <- sample(c(-1, 1), n, replace = TRUE, prob = c(9/19, 10/19)) # 1000 independent draws
S <- sum(x) # total winnings = sum of draws
S
n <- 1000 # number of roulette players
B <- 10000 # number of Monte Carlo experiments
S <- replicate(B, {
X <- sample(c(-1,1), n, replace = TRUE, prob = c(9/19, 10/19)) # simulate 1000 spins
sum(X) # determine total profit
})
mean(S < 0) # probability of the casino losing money
library(tidyverse)
s <- seq(min(S), max(S), length = 100) # sequence of 100 values across range of S
normal_density <- data.frame(s = s, f = dnorm(s, mean(S), sd(S))) # generate normal density for S
data.frame (S = S) %>% # make data frame of S for histogram
ggplot(aes(S, ..density..)) +
geom_histogram(color = "black", binwidth = 10) +
ylab("Probability") +
geom_line(data = normal_density, mapping = aes(s, f), color = "blue")- A random variable X has a probability distribution function F(a) that defines Pr(X≤a) over all values of a.
- Any list of numbers has a distribution. The probability distribution function of a random variable is defined mathematically and does not depend on a list of numbers.
- The results of a Monte Carlo simulation with a large enough number of observations will approximate the probability distribution of X.
- If a random variable is defined as draws from an urn:
-
- The probability distribution function of the random variable is defined as the distribution of the list of values in the urn.
-
- The expected value of the random variable is the average of values in the urn.
-
- The standard error of one draw of the random variable is the standard deviation of the values of the urn.
- Capital letters denote random variables (X) and lowercase letters denote observed values (x).
- In the notation Pr(X=x), we are asking how frequently the random variable X is equal to the value x. For example, if x=6, this statement becomes Pr(X=6).
- The Central Limit Theorem (CLT) says that the distribution of the sum of a random variable is approximated by a normal distribution.
- The expected value of a random variable, E[X]=μ, is the average of the values in the urn. This represents the expectation of one draw.
- The standard error of one draw of a random variable is the standard deviation of the values in the urn.
- The expected value of the sum of draws is the number of draws times the expected value of the random variable.
- The standard error of the sum of independent draws of a random variable is the square root of the number of draws times the standard deviation of the urn.
These equations apply to the case where there are only two outcomes, a and b with proportions p and 1−p respectively. The general principles above also apply to random variables with more than two outcomes.
Expected value of a random variable:
ap+b(1−p)
Expected value of the sum of n draws of a random variable:
n×(ap+b(1−p))
Standard deviation of an urn with two values:
∣b–a∣*sqrt{p(1−p)}
Standard error of the sum of n draws of a random variable:
sqrt{n} *∣b–a∣ * sqrt{p(1−p)}