Is Machine Learning a silver bullet for naturalism about health?

The aim of this paper is to explore whether machine learning can contribute to strengthening naturalistic theories of health. To structure the discussion, Christopher Boorse’s “biostatistical” theory is used, in particular the notion of reference class, which is an important locus of criticism (especially by Elselijn Kingma) against the theory. The paper briefly explores the biostatistical theory of health and introduces in a non-technical way the main notions of machine learning. The thrust is the exploration of a number of possibilities and problems that arise when applying machine learning to the formulation of a naturalistic theory of health. The central topic that is explored is whether machine learning can contribute to determining the reference classes that Boorse’s biostatistical theory necessitates for evaluative analysis, while, at the same time, avoiding to violate the principle of non-normative justification that the theory prescribes. Five main problems related to elements of normativity, which creep in during the application of machine learning, are highlighted and discussed. The paper concludes that both supervised and unsupervised machine learning inevitably carry elements of normativity and subjectivity that make the application of machine learning unfeasible if the goal is to obtain reference classes that are non-normative and suitable for a naturalistic theory in the strict sense. These problems do not imply, however, that machine learning is invalidated for making a contribution to the theoretical analysis of health. Once the appropriate reference classes have been defined, machine learning can make significant contributions, albeit not purely naturalistic ones.