Machine learning, health, and the end of theory

This paper deals with the possibility and conditions for using machine learning for epistemic purposes and for generating knowledge about the world based on the statistical correlations found in data. Some supporters of this view, referred to as the “end of theory” view, argue that the correlations and patterns discovered in data make human-made causal models and explanatory theories unnecessary for scientific progress. Without seeking to exhaust the topic, the paper engages with it by discussing the concept of health and juxtaposing machine learning with a well-discussed data-driven theory of health that seeks to define what health is, i.e., the Bio-statistical theory of health. The paper explores a series of problems with the “end of theory” view. Prior theory, subjectivity, and values blight a naturalistic effort solely based on data. The necessary labelling required for training data in supervised ML systems introduces an element of circularity that is inadmissible from a naturalistic point of view. At the same time, assessing the appropriateness of a reference class determined by unsupervised machine learning and profiling techniques also requires prior theoretical conceptions of health. Hypothetical cases are used to support the claim that while machine learning could certainly aid scientific discovery this cannot happen without previous theories and normative concepts that guide the exploration. These normative concepts, I argue, cannot be meaningfully obtained by machine learning alone. Defining fraught and value-laden notions such as health is an ongoing societal dialogue where the discussion is not only about what is, but also about what should be. This dialogical engagement is all about politics, ethics, and epistemic sovereignty, not about positive correlations and mathematical calculations.