Abstract:
Forecasts are a central goal of machine learning. The current relevance of machine learning
in technological development and society motivates a renewed examination of fundamental
questions: What is the meaning or “goodness” of a prediction? What warrants the use
of a quantified and/or statistical prediction? I argue that forecasts are speech acts, which
as such do something. For example, a weather forecast warns of extreme precipitation.
Accordingly, forecasts should be measured by how felicitous they are, i.e., whether they
are successful in what they do. In this thesis, I examine felicity conditions for statistical
forecasts in relation to the aforementioned questions. I will focus on two, possibly over-
lapping, types of prevalent statistical forecasts. The first type is based on modeling an
aggregate which is relevant to the subject of the forecast. I show that notions of measura-
bility and randomness are not universally given model components. Instead, these aspects
are context-dependent modeling choices that determine the success of the prediction. The
second type of forecast is based on the individual assignment of numerical values that are
evaluated on an aggregate. I investigate the landscape of evaluation metrics with a focus
on calibration and its relationship to regret. I demonstrate that there are two accounts of
calibration currently conflated in literature. Furthermore, most evaluation metrics used in
machine learning, including certain types of calibration and regret, can be expressed as the
capital of a gambler betting against the forecaster. Beyond the aforementioned facets of fe-
licity conditions, I point out commonly overlooked, open questions, particularly with regard
to randomness, individuality, performativity, and communication of forecasts. A tabular
questionnaire designed to help understand, create, and analyze (in)felicitous forecasts on a
practical level concludes the work.