159Evaluating Evaluations of Innovation Policy: Exploring Reliability,. . .
Background: Evaluation as a Practice
Evaluating public policy is a somewhat difficult operation. Any society will likely succumb to public waste without any evaluative elements making sure that public resources are not wasted or misused (Furubo et al., 2002). Yet, it is easy to imagine how too close and frequent control of public servants or policy quickly becomes absurd. Having a grade school teacher being monitored in detail during daily classes or having every agency’s decision double-checked by another auditing agency would not only prove costly but also, most likely, quite futile. Hence, a balance between the two is necessary—societies need both trust and evaluation in order to work. The term evaluation is often used in a rather general and arbitrary way. In a broader sense, evaluation is distinguished from similar practices like auditing or reviewing through the fact it features judgment. An evaluation is not just a display of numbers or opinions but includes some sort of judgment of the studied practice in relation to a predesignated norm or goal (Scriven, 1991; Pollitt, 2003; Knill & Tosun, 2012). Based on this definition, a multitude of evaluation practices exist. Among these no specific practice can be distinguished as superior to the others. Different practices rather serve different purposes. As with scientific methodology in general, the choice of evaluation method and practice depends on the value or goals of interest to the evaluator. The trend toward the large-scale evaluations we see today started in the United States in the 1950s and 1960s. Great hope was then invested in various social and political scientists, who, with the help of quantitative and objective methods, were to scientifically find the best ways to govern society. Subsequent evaluators would question this evaluation practice in favor of what can be described as a more constructivist approach. Greater emphasis was put on experiences from public officials and the people targeted by the studied political intervention. Today, both traditions live on and are present in many Western countries (Dahler-Larsen, 2007; Bovens et al., 2008). Since the late 1900s, evaluation activities in society have increased exponentially, noted not least by Power (1997) in The Audit Society. The huge increase in public scrutiny can be attributed to an expanded public interest in such activities, an increased focus on goal and result management and several of the various governance practices that are referred to under the name New Public Management (NPM)—in part replacing the preceding Weberian public servant model predominating in Western democracies throughout the twentieth century.1
Other factors driving the trend toward more evaluation are organizations such as the European Union and the World Bank putting external pressure on countries to
1Named after the German sociologist Max Weber, Weberian bureaucracy is seen as a system of public administration in which bureaucrats in hierarchal organizations executed political decisions in accordance with predetermined and exact rules and equality before the law.