You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Abish Pius edited this page Apr 1, 2020
·
1 revision
Conditional Probabilities
Conditional probabilities for each class: pk(x)=Pr(Y=k|X=x),fork=1,...,K
In machine learning, this is referred to as Bayes' Rule. This is a theoretical rule because in practice we don't know
p(x). Having a good estimate of the p(x) will suffice for us to build optimal prediction models, since we can control the balance between specificity and sensitivity however we wish. In fact, estimating these conditional probabilities can be thought of as the main challenge of machine learning.
Conditional Expectations and Loss Function
Due to the connection between conditional probabilities and conditional expectations: pk(x)=Pr(Y=k|X=x),fork=1,...,K
we often only use the expectation to denote both the conditional probability and conditional expectation.
For continuous outcomes, we define a loss function to evaluate the model. The most commonly used one is MSE (Mean Squared Error). The reason why we care about the conditional expectation in machine learning is that the expected value minimizes the MSE: Y^=E(Y|X=x)minimizesE{(Y^−Y)^2|X=x}
Due to this property, a succinct description of the main task of machine learning is that we use data to estimate for any set of features. The main way in which competing machine learning algorithms differ is in their approach to estimating this expectation.