When Algorithms Meet Institutions: Evaluation Ecologies in Public Sector Machine Learning
Abstract
This article introduces the concept of âevaluation ecologiesâ to theorize how machine learning (ML) systems are assessed in public sector contexts. Drawing on Science and Technology Studies (STS) and critical AI scholarship, and supported by case studies of ML deployments in Danish educational institutions and Dutch psychiatric clinics, we show that ML evaluation is not a discrete technical procedure but an ongoing, contested process in which multiple registers of judgment interact. Building on Halpern and Mitchellâs (2023) account of experimental governance and Amooreâs (2020) analysis of cloud ethics, the notion of evaluation ecologies captures how technical metrics, professional discretion, legal frameworks, ethical norms, and political objectives coexist and compete in shaping algorithmic outcomes. The Danish case demonstrates how legal and ethical registers can terminate an algorithm despite technical robustness, while the Dutch case shows how perceptions of operational utility and policy alignment may override professional skepticism about algorithmic predictions. Together, these cases reveal that ML evaluation in public institutions is fundamentally relational and context-dependent, structured by power, expertise, and accountability rather than by performance metrics alone. By foregrounding evaluation as an ecological and political process, the article contributes to debates on algorithmic governance by showing how local contexts and professional judgment remain decisive in the age of public-sector AI.
1 Introduction
Machine learning has entered public administration amid both enthusiasm for innovation and fierce controversy (Carboni et al., 2025a; Dencik et al., 2019; Hoeyer 2023; Ranchordas, 2021; Ratner & Schrøder, 2024; Ravn et al., 2026; Thylstrup et al., 2022; Vonk & Ranchordås, 2025). As ML systems take on consequential roles in education, healthcare, and social services, their evaluation has become a contested political terrain. Who gets to judge whether an algorithm works? What counts as evidence of success or failure? How do technical metrics relate to professional judgment, ethical principles, or legal requirements?
This article develops the concept of âevaluation ecologiesâ to analyze how different modes of assessment including technical, professional, ethical, legal, interact and compete when ML systems are deployed in public institutions. Traditionally, evaluation in machine learning has been the province of data scientists wielding quantitative metrics: accuracy, precision, efficiency. But as algorithms move from laboratories into consequential public services, new dimensions of assessment emerge, often generating friction between competing stakeholders and value systems. We argue that ML evaluation in the public sector operates as ecologies: dynamic systems where diverse actors such as data scientists, frontline professionals, citizens, legal authorities, policymakers assess algorithms according to different registers, each carrying distinct logics and authority claims. These registers do not harmonize. They collide, compete, and reshape one another. Technical validity may be overridden by ethical concerns; professional skepticism may yield to policy imperatives; legal frameworks may terminate systems that perform well by computational standards.
Through case studies of ML deployments in Danish educational institutions and Dutch psychiatric clinics, we trace how these evaluation ecologies unfold in practice. The Danish case shows how legal and ethical registers terminated an algorithm despite technical validity. The Dutch case reveals how operational utility and policy alignment overrode professional reservations about algorithmic predictions. Together, they demonstrate that successful algorithmic governance requires attending not just to technical performance but to the broader ecology of evaluation: understanding how different registers gain or lose influence, which actors can mobilize which forms of authority, and how local contexts shape what counts as legitimate assessment.
Our contribution is both conceptual and empirical. Conceptually, we offer âevaluation ecologiesâ as a framework for understanding the inherently political nature of ML assessment in public institutions. Empirically, we demonstrate how this framework illuminates otherwise puzzling outcomes: why technically sound algorithms fail, why professionally contested systems succeed, and why evaluation is never settled but continuously renegotiated across multiple registers of judgment.
1.1 Engaging with Evaluation in the Public Sector
The concept of evaluation has a long, complex history. It first appeared in the English language in 1755 in an insurance context to denote an act of appraisal. In the same century, certain sciences adopted it to describe âthe action of evaluating or determining the value of (a mathematical expression, a physical quantity, etc.) or of estimating the force of (probabilities, evidence, etc.)â (OED). This later meaning evolved in the 19th century, as described by Ian Hacking (1990), as the state and data-rich organizations sought to âtame chanceâ through probabilistic methods. By the 20th century, evaluation practices solidified into a scientific field closely aligned with public sector growth, establishing systematic evaluation methods to assess program effectiveness (Hogan, 2007). Private sector management techniques subsequently influenced the public sector, shaping not only evaluation practices but also the very organization of governmental bureaucracy, shifting from a bureaucratic ethos to a business-oriented one, focused on outcomes, performance indicators, and accountability (Hood, 1991). As Dahler-Larsen (2011) describes in The Evaluation Society, this new ethos introduced a technological regime in which âevaluation, accreditation, auditing, benchmarking, performance management, quality assurance, and similar documentation practices produce datascapes as an important dimension in social life.â Thus, evaluation as an idea and practice has a robust genealogy in public administration, extending well beyond the recent development of ML and its performance evaluations.
Our concept of âecologies of evaluationâ brings together these diverse registers, emphasizing that machine learning assessment does not occur within a closed epistemic environment - particularly when ML models are deployed in complex, socially significant settings.Footnote 1 While various contexts naturally involve multiple registers of machine learning evaluation (cf. Neyland, 2016), public service delivery offers a uniquely pertinent lens for examining these registers. Firstly, the public sector relies on a high degree of professional discretion, where professionals like nurses, social workers, or teachers use their judgment to address citizensâ unique needs (Evans, 2016; Lipsky, 2010). This discretion allows professionals to adapt policies and interventions to complex real-world contexts that highly standardized, automated systems cannot accommodate. In algorithmic decision-support contexts, professional discretion and human judgment thus often entail an evaluation of algorithmic outputs. Secondly, the public sector is under extensive scrutiny, subject to traditional jurisdictional accountability and to the attention of interest groups and public media (Ratner & Elmholdt, 2023). Such âpublicsâ constitute a crucial site for evaluating emerging ML technologies. In our analysis, then, we examine three key registers of evaluation that emerge in public sector ML deployments: (1) technical evaluation, where data scientists assess algorithmic performance through metrics like accuracy and reliability; (2) legal and ethical evaluation, where stakeholders examine compliance with regulations and social norms; and (3) professional evaluation, where domain experts like nurses and teachers apply situated knowledge to assess ML systemsâ practical value. Through our case studies, we demonstrate how these registers interact, compete, and sometimes fundamentally conflict with one another, creating dynamic evaluation ecologies.
This article argues that understanding the politics of evaluation in machine learning requires examining not only the internal practices of AI labs but also the broader ecologies within which these practices are embedded. While recent STS and media studies scholarship has illuminated the norms, practices, and conditions of machine learning evaluation within laboratory and broader data science contexts (Jaton, 2021; Engdahl, 2024; Luitse & Hansen, 2024; Raji et al., 2021; Schjøtt & Blanke, 2026; Thylstrup & Hartley, 2025), we demonstrate that ML initiatives encounter multiple forms of assessment throughout their lifecycles. Beyond formal evaluations like benchmarking, these technologies face ongoing scrutiny from professional and public stakeholders, creating a complex web of interacting and mutually influencing evaluation registers (Ratner & Thylstrup, 2025; Wenzelburger et al., 2024). To capture this complexity, we propose conceptualizing evaluations as ecologies, emphasizing their multidimensional, domain-crossing, and dynamic character. This ecological framework reveals how different forms of evaluation, including technical, professional, and public, coexist and shape one another in ways that traditional, lab-centric approaches might overlook.
The structure is as follows: first, we develop the notion of ecologies of evaluation, linking diverse genealogies of evaluation in ML and the public sector; second, we analyze two distinct sites and registers of ML evaluation: (1) the tension between data scientists evaluating a modelâs precision for predicting student dropouts vs. stakeholdersâ focus on legal compliance; (2) nursesâ assessment of the relevance of a model predicting inpatient violence, based on their professional knowledge. Finally, the article discusses the politics of evaluation in datafied welfare regimes (Dencik & Kaun, 2020), considering how the increasing integration of AI in the public sector expands the sites and registers of ML evaluation.
2 Theorizing Evaluation Ecologies in the Age of AI
Recent scholarship has identified a new politics of testing and experimentation emerging around computational governance. Orit Halpern and Robert Mitchell (2023) develop the concept of âsmartnessâ to describe âan emerging form of technical rationality whose major goal is the management of an uncertain future through a constant deferral of future resultsâ and âperpetual and unending evaluation through a continuous mode of self-referential data collection.â Their historical analysis traces how machine learning experiments are assessed through computational methods that fundamentally reshape what counts as efficient governance. Louise Amoore (2020) extends this argument by conceptualizing cloud computing as an âexperimental chamberâ where technologies engage in continuous trial-and-error rather than bounded design periods. Machine learning models, operating on vast and varied data inputs, exemplify this shift toward open-ended experimentation. Amoore further argues that cloud-based systems, which leverage both public and classified data in black-box mechanisms, introduce profound accountability challenges by suggesting âwe donât need to see the data,â thereby complicating traditional evaluation methods (Amoore, 2023). The shift from transparent decision trees to opaque large language models intensifies these complications. If existing evaluation frameworks remain largely rules-based, how should assessment evolve in the dissolution of the situatedness of data (Campolo & Schwerzmann, 2023) and a parallel renegotiation of expertise in the implementation and finetuning of AI as general purpose technologies into sensitive and contested domains (Rella et al., 2025)?
This article addresses these critical questions through situated studies of ML evaluation practices within public sector institutions. We argue that examining concrete evaluation encounters reveals how ML experimentation multiplies and diversifies assessment sites and logics, generating friction between competing forms of expertise and authority. While our analysis confirms the broader politics identified by Halpern and Mitchell and Amoore, i.e. the deferral of results, perpetual testing, opacity as governance mode, we also find that these macro-political imperatives play out through micro-political contestations. Situated evaluation practices are shaped not only by efficiency mandates and computational opacity but also by professional resistance, ethical intervention, legal constraint, and organizational negotiation.
Our approach aligns with STS scholarship on evaluation and valuation (Felt, 2017, 2025; Heuts & Mol, 2013) that attends to practices rather than fixed qualities. As Heuts and Mol (2013) argue in their study of tomato valuation: âwe shifted from talking about âworthâ (a quality) to foregrounding âvaluingâ (an activity) and from âeconomiesâ (that come with a single gradient each) to âregistersâ (that indicate a shared relevance, while what is or isnât good in relation to this relevance may differ from one situation to another).â Following their shift from worth to valuing, we move from evaluation frameworks to registers of evaluation by examining how ML systems in the public sector are judged as âgoodâ or âbadâ through multiple, often competing logics. We pay particular attention to how MLâs technical registers encounter public accountability requirements and domain expertsâ professional and ethical commitments. This focus reveals how evaluation ecologies contain differential power relations: certain registers achieve resonance within deploying institutions while others barely register at all.
Building on Clarke and Starâs (2008) social worlds framework, we conceptualize evaluation ecologies as relentlessly ecological phenomena encompassing diverse social worlds each with distinct perspectives, commitments, and evaluation registers. These ecologies operate across three dimensions. First, they are inherently relational: multiple collective actors (data scientists, domain experts, policymakers, citizens) engage in evaluation practices that mutually influence one another despite distinct epistemological foundations. Second, they are materially situated: evaluation unfolds within specific âinfrastructural ecologiesâ (AUTHOR 2024) that shape what can be evaluated and how. Third, they are dynamically contested: characterized by ongoing power struggles over which evaluation registers should take precedence. ML systems thus function as what Star and Griesemer (1989) term âboundary objectsâ, i.e. entities existing at intersections of multiple social worlds, interpreted differently by each yet serving as coordination sites. More specifically, following Suchman et al. (2002), they operate as âworking artifactsâ: socio-material configurations aligning interests and practices across heterogeneous worlds through âboundary-crossing activitiesâ that create situations âfor the meeting of different partial knowledgesâ (Suchman et al., 2002). Their meaning and value are not predetermined but discovered through collaborative evaluation-in-use, where different assessment registers continually inform and transform one another. By focusing on ecologies rather than individual evaluation frameworks, we thus highlight the inherent multiplicity and contestation involved in ML assessment. This ecological view reveals how ML evaluations in public institutions are never merely technical exercises but always already social, political, and ethical endeavors embedded in broader contexts of power, expertise, and accountability.
Our ecological perspective on ML evaluation extends Lucy Suchmanâs foundational work on technology in practice, drawing on both her early theorization of situated action (1987) and her later development of âlocated accountabilities in technology productionâ (2002). Together, these frameworks reconceptualize ML evaluation in public sector contexts. Suchmanâs Plans and Situated Actions (1987) demonstrated that human-computer interaction cannot be predetermined by abstract plans but emerges through specific contexts. We extend this insight to show that ML evaluation in public institutions similarly resists reduction to predetermined technical metrics, instead unfolding within particular institutional, professional, and political contexts. Building on this foundation, Suchmanâs concept of âlocated accountabilityâ (2002) offers an explicit alternative to âdesign from nowhereâ (with its illusion of universal objectivity) by acknowledging the specific sociomaterial networks within which both designers and users operate. As Suchman states, âlocated accountability is built on what Haraway terms âpartial, locatable, critical knowledgesââ (2002). This concept anchors our evaluation ecologies framework in three ways. First, where Suchman argued that responsible technology design requires recognizing oneâs position within âextended networks of sociomaterial relationsâ (2002), we apply this insight to ML evaluation practices in public institutions. The Danish case exemplifies located accountability: external stakeholders successfully contested the algorithm by making visible the ethical and legal relationships that developers had obscured, revealing how technical evaluation registers are always already entangled with legal, ethical, and professional accountabilities. Second, we extend Suchmanâs critique of âdetached intimacyâ by showing how ML evaluators develop close relationships within technical communities while maintaining distance from other stakeholders. In the Dutch case, the data scientist operated with primarily technical accountability metrics (predictive accuracy), while nurses practiced located accountability grounded in professional care ethics and situated patient knowledge. This tension illustrates Suchmanâs observation that âbecoming a participant in the worlds of technology production necessarily involves finding a relation to professional designâ (2002), with professionals from different worlds bringing divergent accountabilities to evaluation. Third, our concept of evaluation ecologies advances Suchmanâs vision of âartful integrationsâ over technological hegemonies. Where she proposes that âdesign success rests on the extent and efficacy of oneâs analysis of specific environments of devices and working practicesâ (2002), we show that ML evaluation in public contexts requires integrating multiple evaluation registers across professional boundaries.
3 Briefly on the Empirical Material Informing the Article
European public sectors increasingly employ algorithmically supported decision-making. In this article we focus on a suite of semi-formalized projects experimenting with AI in public service. Our analysis focuses on two cases in particular that illuminate different dimensions of ML evaluation ecologies. The first case is a 2014 initiative by the Danish company MaCom to introduce a dropout prediction algorithm within their Lectio learning management system, used across Danish secondary educational institutions. This case, exploring ML evaluationâs intersection with public accountability, draws on publicly accessible materials including research publications, media coverage, and legal documentation. The Danish context proves particularly relevant given the nationâs strategic emphasis on public sector AI, exemplified by substantial investments in what the government calls AI âsignature projectsâ and ambitions of achieving global leadership in governmental AI deployment (Danish Government, 2018). The second case is a 2022 pilot program predicting violence risk in two Dutch acute psychiatric clinics. This study, revealing everyday evaluation practices rooted in professional expertise and ethics, emerges from three months of direct observation as clinical staff and IT personnel grappled with implementing the algorithm. Initiated by the organizationâs IT department and swiftly approved by management, the pilot exemplifies the enthusiasm with which Dutch healthcare institutions have embraced experimental ML deployments. Together, these cases offer complementary insights into how evaluation practices unfold across different institutional settings, stakeholder groups, and regulatory environments. Though distinct in their specific contexts, both reveal the complex interplay between technical assessment and broader social, ethical, and professional considerations that shape ML implementation in public services.
4 Machine Learning Evaluation Meets Public Accountability
Compared to other countries, Denmark has limited experience with the deployment of predictive analytics in education. In 2014, the private learning management platform provider, MaCom, presented a dropout prediction algorithm as part of its digital platform Lectio, used by the majority of Danish secondary educational institutions (HHX/gymnasium). The algorithm had been developed through a partnership with University of Copenhagen by using machine learning on data from 70,000 students that had used Lectio. With an average dropout of 20%, the company assumed that the taximeter funded educational institutions were interested in identifying students at risk for dropping out (Mølsted, 2014). About one week after the predictive function had been launched, the association âDanske Gymnasierâ requested to have it turned off. They were concerned about the algorithm, especially the questions it raised in relation to privacy issues and lack of consent (Møllerhøj 2015a). As we will outline in the following, despite the algorithm receiving a favorable evaluation within the register of data science, especially in relation to its efficiency and precision, it also elicited a critical register of evaluation related to public accountability, which in the end turned off the algorithmic model before it had been put to operational use.
4.1 Data Scientistsâ Registers of Machine Learning Evaluation
In the context of developing the predictive ML model, the data scientists working on the project mobilized different registers of evaluation for assessing its performance, robustness and utility. These registers of evaluation are common among data scientists and document how the âgoodâ machine learning model is understood and which methods or ideas are mobilized to conduct the evaluation. First and foremost, the data scientists mobilize a register of evaluation emphasizing the volume of data as a parameter of quality and research contribution. The data was extracted from the Lectio study administration system. As the data scientists working on the project write: âIn contrast to existing studies that were based on only a few hundred students, we considered a considerably larger sample (âŚ) We queried the MaCom Lectio database for students enrolled after 2009 and extracted 72598 pupils, 55259 of which graduated and 17339 dropped out, giving a dropout rate of 23.8%, which is close to the Danish averageâ (Sara et al., 2015). In an interview with the tech media Version2, the director of Macom describes the dataset as a âholy grail,â highlighting its unparalleled scale, i.e. data collected from hundreds of thousands of students across a decade. Comparing the model to similar models built on smaller datasets used in other countries further emphasizes this, portraying the sheer amount of data as making the system âfar more bulletproofâ than others (Mølsted, 2014).
Here, the magnitude of the dataset is not only a technical register for evaluation but also becomes a rhetorical tool for asserting the reliability of the model. This register reflects more mainstream notions in data scientific discourses that equate large datasets with more accurate, objective, and trustworthy insights (DâIgnazio & Klein, 2018). This massive dataset also reflects a political register of evaluation, where scale plays into the institutional logic of governance. When such a comprehensive amount of data is collected and centralized, it signals a form of state capacity such as its ability to monitor, manage, and intervene in public education systems.
The data scientistsâ process of ârandomly splitting the data equally into a training and test set with 36299 samples eachâ (Sara et al., 2015) serves as another critical register of evaluation. The splitting of data is a common methodological practice in machine learning, wherein the ability to train a model and then test its performance on a separate, unseen dataset is used to determine its generalizability. This register of evaluation is rooted in ideas in data science of robustness and the avoidance of overfitting models. That the ratio of graduates to dropouts in the data set matches the national dropout rate further supports the idea that this split preserves the datasetâs external validity. Emphasizing that the test data mirrors real-world conditions (here, the national average in drop-out rate) aligns the modelâs outputs with institutional expectations. If the model is trained and tested on a ârepresentativeâ dataset, the idea is that it can be trusted for decision-making in educational contexts.
Finally, the data scientists also compared different machine learning models, introducing a register of model performance focusing on different machine learning modelsâ accuracy and Receiver Operating Characteristic Curves (ROC). Accuracy and ROC curves are standard metrics in data science, used to evaluate how well different models perform in predicting outcomes. In this evaluative register, emphasis is placed on optimizing predictive performance through comparative analysis. One study, for example, assessed several machine learning models (random forest, support vector machines (SVM), classification and regression trees (CART), and naĂŻve Bayes classifiers) and found that random forest achieved the highest predictive accuracy, with other models performing at slightly lower levels (Sara et al. 2015) (p. 322). This approach reflects a model of evaluation grounded in statistical optimization and the prioritization of predictive precision.
4.2 Stakeholdersâ Registers of Evaluation
The development of the drop-out algorithm emerged from a collaboration between MaCom and a University of Copenhagen masterâs thesis student, exemplifying MaComâs established practice of partnering with students from the University of Copenhagen and the Technical University of Denmark. As MaComâs owner (in Møllerhøj, 2015) notes, these collaborations were traditionally viewed as âmutual benefits for both the university students, who base their theses on a real-world problem, and my client, who gains insight into issues they otherwise wouldnât have the resources to explore.â However, this particular collaboration sparked significant controversy when two influential stakeholder organization, the Danish association for secondary education (âDanske Gymnasierâ) and the Danish association for vocational education (âDanske Erhvervsskolerâ), challenged the legality of the algorithmâs development. Their opposition centered on privacy and data security concerns, which can be analyzed through three distinct registers of evaluation.
The first register emerged during a pivotal meeting between Danske Gymnasier, MaCom, and representatives from the Danish Agency for IT and Learning. As Møllerhøj (2015c) reported, the primary contention centered on MaComâs unauthorized sharing of educational data with a master student from the University of Copenhagen, the third-party researcher: âThe theme of the meeting was that Macom had shared data with a third party, Nicolae-Bogdan Čara, without specific agreements in place with the individual educational institutions, and that Macom had, in this context, violated the data processing agreement made with the schools.â Møllerhøj (2015c) further noted that âit was also emphasized at the meeting that it is the schoolâs management that must decide who should have access to which data, including information that could indicate a dropout risk and thus allegedly stigmatize the students.â This unauthorized data sharing register of evaluation highlighted not just procedural violations but deeper ethical concerns about data control and access.
The second register, also evident in the above quote, focused on studentsâ risk and stigmatization. Stakeholders expressed concern that the algorithmâs predictive capabilities could prematurely label students, potentially creating bias and negatively impacting their educational experiences. This register moved beyond purely technical or legal considerations to address the broader social implications of implementing predictive analytics in educational settings.
The third register of evaluation emerged through Danske Erhvervsskolerâs engagement of the legal firm Bech-Bruun. Their assessment focused specifically on data governance and legal compliance, as evidenced in their formal communication with MaCom. As cited in Møllerhøj (2015b), Bech-Bruun stated:
As is well-known, both the individual vocational school, which is the data controller, and any data processors have a legal responsibility to ensure that both parties handle the personal data of students, staff, and other individuals connected to the respective vocational school appropriately. [âŚ] We further request that we also receive, no later than October 27, 2014, a report on MaCom A/Sâs handling of the information that the listed vocational schools have entrusted to MaCom A/S in connection with the use of Lectio, including the use that has taken place in relation to external persons (for example, computer science students at the University of Copenhagen) participating in the development of an IT tool to predict student dropout.
Bech-Bruunâs involvement elevated the scrutiny of MaComâs data handling practices, particularly regarding external researchersâ access to sensitive student information. This legal register of evaluation emphasized the importance of proper data governance frameworks and compliance with data protection laws, to encompass broader regulatory requirements.
Significantly, these externally mobilized registers of evaluation, which focused on legal compliance, ethical data handling, and potential stigmatization, proved more influential than technical considerations about the algorithmâs performance (Møllerhøj 2015a, c). While data scientists might prioritize model accuracy and efficiency, these stakeholder organizations successfully shifted the evaluation framework toward legal and ethical considerations specific to public sector applications. The primacy of these registers ultimately led to the algorithmâs termination after only five days, demonstrating how legal and ethical considerations can override technical achievement in public sector machine learning applications.
This case illustrates how external stakeholders can effectively mobilize alternative registers of evaluation that prioritize compliance, ethics, and social impact over technical performance metrics. It also highlights the unique challenges of implementing machine learning systems in public institutions, where legal adherence and ethical treatment of personal data may carry more weight than algorithmic efficiency or innovation.
5 Machine Learning Evaluation Meets Domain Expertise and Ethics
The implementation of machine learning technologies in Dutch healthcare has largely remained âpiecemeal and small-scaleâ (KPMG, 2020), as exemplified by the pilot program in two inpatient, acute psychiatric clinics that we examine in this section. This pilot centered on an algorithm designed to predict violent episodes among patients through an analysis of the words staff used in patient records. The algorithm was initially developed by a PhD student for a different hospital, and subsequently adapted for the two clinics in question. The initiativeâs rollout was facilitated by several factors. First, the local IT department had access to the algorithm. Second, the algorithm itself aligned with national and international policies aimed at reducing coercive care in psychiatry (Smith et al., 2023), particularly Dutch initiatives to phase out patient seclusion (Steinert et al., 2014). Third, the algorithm had the potential to automate the violence risk assessments manually conducted by nurses.
5.1 Data Scientistsâ Registers of Machine Learning Evaluation
The pilotâs rollout revealed distinct registers of evaluation among different stakeholders. The lead data scientist, working within the organizationâs IT department, immediately modified the algorithm to predict violence within a 24-hour window rather than the original two-week timeframe. As she explained during an interview:
We didnât like the outcomes from [the original algorithm]. There was a low recall. So, you could not really predict most of the incidents. ⌠For me, I guess it was logical and [closer to] the needs of the staff to predict an aggression incident for one day and not for the next two weeks. [Two weeks is] a little abstract, I guess.
This adjustment reflected a primary register of evaluation that balanced predictive accuracy with practical utility, characteristic of applied data science. However, her perspective was further complicated by data quality concerns. Violence episodes represented what she termed âa stupid outcome indicator,â with only 10% of incidents being reported in the designated digital platform. This limited reporting stemmed from both contested definitions of violence among staff, and nursesâ habituation to volatile environments amid heavy workloads. She expressed frustration with the disconnect between recorded and actual incidents, and with its repercussion on the algorithmâs predictive accuracy:
Itâs not doing what it should do when we donât have [data on] the real incidents and maybe even the âalmost-incident.â Itâs fuzzy. The model says âOkay, [in the patient files] I see words like âaggressiveâ and âthrowingâ or âhitâ ⌠And there is no incident. Huh? What do I do? Itâs weird.â And it is weird, because there was an incident, but it was not reported.
In her view, this could be solved through a stricter definition of violence incidents (i.e., physical attacks to objects or people, or clear verbal attacks), and thus more consistent reporting on the part of nurses. In turn, this would enable her to âre-train the algorithm not only with the aggression incident-or actually, with the real aggression incidents and not only with the 10% that is reported.â This resonates with modalities of evaluation of machine learning as an ongoing, experimental enterprise, always amenable to update and adjustment (Amoore, 2020; Halpern & Mitchell, 2023).
Finally, at the evaluation meeting concluding the pilot, the data scientist reported that during the pilotâs three months, the algorithm had flagged as at risk of violence more than 500 cases that nurses had not identified. This data was interpreted by both the data scientist and the clinicsâ management as evidence of the algorithmâs superior risk assessment capabilities. The data scientistâs evaluative framework thus encompassed three key elements: the practical utility of algorithmic predictions; dataset quality linked to predictive accuracy; and a specific focus on minimizing false negatives.
5.2 Nursesâ Registers of Evaluation
While nurses officially deferred to the data scientistâs technical evaluation about whether the algorithm was âworking,â they simultanously conducted ongoing informal assessments based on their professional judgment. Their evaluative register centered on relational and affective complexity that they believed the algorithm failed to capture. This emerged clearly in daily practice. As noted during field observations:
Josh ⌠showed me the file with the risk scores for that day, comparing it with the printout from the handover. He saw the name of the patient scored as the highest risk: âThis one, for example, doesnât make any sense. This is a patient who is asking to be isolated from the group. That has nothing to do with violence!
Nursesâ concerns manifested in three primary areas: the algorithmâs capacity to predict complex violent behavior based solely on patient files; the importance of contextualizing behavioral signs within broader care trajectories; the relational nature of violence in psychiatric settings. As captured in field notes:
When, at the end of the handover, I asked the nurse and the doctor about how they deal with violence, they explained that what counts as a âsignificant episodeâ is very complicated: âIf someone hits the window, like today, is it violence to objects, or is it just that they got a bit angry?â said Josh, a nurse. Karin, a doctor, agreed: âDo you come in and pump them full of lorazepam?â âExactly, or put them in a straitjacket?â
The nursesâ experiential register of evaluation frequently clashed with the data scientistâs pragmatic approach. While nurses proposed including contextual variables, such as staff composition and experience levels, the data scientist prioritized maintaining a âsimpleâ algorithm. This tension highlighted the fundamental difference between data-driven and experience-based evaluation registers.
To sum up, their register of evaluation appears to be rooted in nursesâ humanistic professional ethos, scaffolded by care ethics that emphasize relationality in all aspects of care encounters (Carboni et al., 2025a, b). Moreover, this ethos resonates fully with goals of moving psychiatry away from coercive care. In contrast to the data scientistâs focus on false negatives, nurses repeatedly expressed concerns with false positives and with the algorithmâs predictions turning into what they referred to as âself-fulfilling propheciesâ that would lead to more coercive care. Interestingly, although their concerns were heard during evaluations, they had no concrete effect on the algorithm itself or on the ultimate evaluation of the pilot, which was considered a success and continued beyond the ethnography reported here.
6 Discussion: Contingency, Contestation, and Power
Understanding the politics of machine learning evaluation in public sector contexts requires examining how different registers of evaluation interact within broader evaluation ecologies. Our comparative analysis reveals that these evaluation processes are relational, i.e. constituted through interactions among multiple actors whose assessments mutually shape one another, as well as contingent, producing divergent outcomes despite technical similarities. The Danish and Dutch cases, both deploying predictive algorithms to improve welfare services, met different fates because evaluation registers gained influence through specific socio-political configurations rather than predetermined hierarchies.
First, the Danish and Dutch cases demonstrate the contingent nature of evaluation hierarchies. Which registers gain influence depends not on inherent authority but on how specific socio-political contexts configure their interactions. In Denmark, a seemingly technically robust dropout prediction algorithm was discontinued when stakeholders raised privacy and data governance concerns. Legal and ethical registers overrode technical considerations. This was not because they inherently outrank technical assessment, but because the relational dynamics among educational administrators, data protection authorities, and civil society organizations elevated these registers within this particular context. The Dutch case presents a contrasting configuration, demonstrating how the same types of registers (technical, ethical, professional) interact differently when institutional power aligns with policy objectives. Despite significant concerns about data quality and false positives, the algorithmâs perceived practical utility and alignment with national policies to reduce coercive care became dominant forces. Here, relationality operated through different actor configurations: psychiatric administrators prioritizing policy alignment, national authorities seeking efficiency metrics, and frontline nurses whose concerns achieved less traction.
Second, these cases reveal the inherently relational constitution of evaluation hierarchies. No single register possesses intrinsic authority; rather, influence emerges through interactions among actors who mobilize different registers. The Danish context demonstrated legal and ethical registers achieving hegemonic status through alliances between external stakeholders, educational professionals, and data protection experts who collectively challenged the algorithm. The Dutch case exemplifies a contested hierarchy where practical utility and policy alignment took precedence through alignment between psychiatric administrators and national policymakers, potentially marginalizing nursesâ professional concerns that lacked institutional backing. This relationality explains why technically similar systems meet different fates: evaluation outcomes depend on which actors can successfully mobilize which registers within specific institutional configurations.
Third, both cases foreground how professional discretion operates relationally within evaluation ecologies, serving as counterweight to technical authority when institutional conditions allow. In Denmark, educational professionals actively questioned the dropout prediction algorithmâs utility based on their expertise. Crucially, their concerns gained traction because they aligned with legal and ethical registers mobilized by external stakeholders demonstrating how professional judgment achieves influence through relational configurations rather than intrinsic authority. In Dutch psychiatric clinics, nurses exercised professional judgment to mediate and sometimes challenge the violence prediction algorithmâs conclusions. Yet their concerns remained subordinated because they conflicted with dominant policy and utility registers supported by administrators and national authorities. This contrast reveals the contingent nature of professional authority: identical forms of expertise-based challenge succeeded in one context but failed in another, depending on how professional discretion connected (or failed to connect) with other registers within the broader evaluation ecology.
These findings contribute to several bodies of literature. First, they extend recent STS scholarship on AI evaluation practices (Jaton, 2021; Luitse & Hansen, 2024; Engdahl, 2024; AUTHOR 2025) by demonstrating that evaluation cannot be understood solely through technical metrics or benchmarking practices but must account for how multiple registers compete for authority. Second, our framework advances work on valuation practices (Heuts & Mol, 2013; Felt, 2017, 2025) by showing how ML systems are valued across incommensurable registers rather than through singular logics. Third, we contribute to critical analyses of algorithmic governance (Ranchordas, 2021; Vonk & RanchordĂĄs 2025; Amoore, 2020; Halpern & Mitchell, 2023) by revealing how macro-political imperatives, including efficiency, perpetual testing, opacity, materialize through micro-political negotiations among professionals, administrators, and citizens. Finally, by foregrounding professional discretion as analytical category, we extend Suchmanâs (1987, 2002) work on situated action and located accountability into the domain of ML evaluation, demonstrating how embodied expertise remains crucial even within increasingly automated systems.
7 Concluding Remarks: The Politics of Evaluation Ecologies
This article introduces âevaluation ecologiesâ as a conceptual framework for understanding machine learning assessment in public sector contexts. Through analyses of ML implementations in Danish educational institutions and Dutch psychiatric clinics, we demonstrate how evaluation practices extend far beyond technical metrics to encompass complex negotiations of power, expertise, and accountability.
Our cases challenge narratives about the dominance of macro-political imperatives in technological governance. The Danish case, where legal and ethical concerns superseded technical achievements, and the Dutch case, where practical utility gained primacy despite professional reservations, reveal efficiency not as totalizing logic but as one register among many, generating friction and negotiation across multiple dimensions. ML experimentation in public institutions unfolds through intricate power relations that are fundamentally relational and contingent. By conceptualizing evaluation as ecologies rather than linear processes, we move beyond viewing ML assessments as purely technical or instrumental. This ecological perspective shows how evaluation practices are constituted through heterogeneous registers, each carrying distinct stakes, values, and tensions. Certain registers gain prominence and legitimacy while others face marginalization, reflecting broader power structures while creating spaces for contestation and resistance.
The politics of ML evaluation operates across both macro and micro dimensions. At the macro level, evaluation practices align with or resist broader political imperatives and institutional logics. At the micro level, they manifest through everyday negotiations, professional judgments, and situated decision-making. This dual perspective reveals how power operates both through alignment with dominant evaluation registers and through local acts of contestation. An ecological approach thus enriches theoretical understanding of ML evaluation by moving beyond instrumental concerns with âsuccessful implementationâ toward critical analysis of how evaluation practices themselves constitute and reproduce power relations in technological governance. Rather than prescribing pathways for âbetterâ ML deployment, our framework offers analytical tools for understanding how evaluation regimes emerge through contested negotiations across multiple social worlds. Future research might develop this perspective by examining how evaluation ecologies transform over time and differ across institutional and national contexts. Furthermore, and perhaps more pressing, it might open up to new research questions about how evaluation practices are conducted within experimental technologies and unfolding political situations where the object of evaluation is uncertain and / or contested (AI as well as emerging concerns with digital sovereignty).
Notes
Although our concept of âevaluation ecologiesâ advances a distinct analytical framework, it builds upon Marres and Starkâs (2020) influential work on âecologies of testing,â which similarly emphasizes the relational and multidimensional nature of technological assessment.
References
Amoore, L. (2020). Cloud ethics: Algorithms and the attributes of ourselves and others. Duke University Press.
Amoore, L. (2023). Machine learning political orders. Review of International Studies, 49(1), 20â36.
Campolo, A., & Schwerzmann, K. (2023). From rules to examples: Machine learningâs type of authority. Big Data & Society, 10(2), 20539517231188725.
Clarke, A. E., & Star, S. L. (2008). The social worlds framework: a theory/methods package. In E. J. Hackett, O. Amsterdamska, M. Lynch, & J. Wajcman (Eds.), The Handbook of Science and Technology Studies (pp. 113â137). MIT Press.
Carboni, C., Wehrens, R., van der Veen, R., & de Bont, A. (2025a). Doubt or punish: On algorithmic pre-emption in acute psychiatry. AI & SOCIETY, 40(3), 1375â1387.
Carboni, C., Wehrens, R., De Bont, A., & Van der Veen, R. (2025b). From attention to attunement: Data-driven efficiency and embodied care in the intensive care unit. Science, Technology & Human Values. Epub ahead of print October 17, 2025. https://doi.org/10.1177/01622439251384695
DâIgnazio, C., & Klein, L. (2018). Data Feminism. MIT Press.
Dahler-Larsen, P. (2011). The evaluation society. Stanford university press.
Dencik, L., & Kaun, A. (2020). Datafication and the welfare state. Global Perspectives, 1(1), 12912.
Dencik, L., Redden, J., Hintz, A., & Warne, H. (2019). The âgolden viewâ: Data-driven governance in the scoring society. Internet Policy Review, 8(2), 1â24.
Engdahl, I. (2024). Agreements âin the wildâ: Standards and alignment in machine learning benchmark dataset construction. Big Data & Society, 11(2), 20539517241242457.
Evans, T. (2016). Professional discretion in welfare services: Beyond street-level bureaucracy. Routledge.
Felt, U. (2017). Under the shadow of time: Where indicators and academic values meet. Engaging Science Technology and Society, 3, 53â63.
Felt, U. (2025). Environmental intelligence?! on the consequences of sidelining the materiality of AI. Harvard Data Science Review, 7(4).
Hacking, I. (1990). The taming of chance. Cambridge university press.
Halpern, O., & Mitchell, R. (2023). The smartness mandate. MIT press.
Heuts, F., & Mol, A. (2013). What is a good tomato? A case of valuing in practice. Valuation Studies, 1(2), 125â146.
Hoeyer, K. (2023). Data paradoxes: The politics of intensified data sourcing in contemporary healthcare. MIT press.
Hogan, R. L. (2007). The historical development of program evaluation: Exploring past and present. Online Journal for Workforce Education and Development, 2(4), 5.
Hood, C. (1991). A Public Management for all Seasons? Public Administration, 69(1), 3â19.
Jaton, F. (2021). The constitution of algorithms: Ground-truthing, programming, formulating. MIT press.
KPMG (2020). Inventarisatie AI-toepassingen in de gezondheid en zorg in Nederland: Onderzoek naar de stand van zaken in 2020. Retrieved March 24, 2025, from https://www.rijksoverheid.nl/documenten/rapporten/2020/10/05/inventarisatie-ai-toepassingen-in-gezondheid-en-zorg-in-nederland
Lipsky, M. (2010). Street-level bureaucracy: Dilemmas of the individual in public service. Russell sage foundation.
Luitse, D. M. R., & Hansen, A. S. (2024). The politics of machine-learning evaluation: from lab to industry. AoIR selected papers of internet research.
Marres, N., & Stark, D. (2020). Put to the test: For a new sociology of testing. The British Journal of Sociology, 71(3), 423â443.
Møllerhøj, J. (2015a). Gymnasieelev anmeldte CPR-hul i skolesystem til Datatilsynet - nu für han en advarsel af skolen. Version2, September 30. https://www.version2.dk/artikel/gymnasieelev-anmeldte-cpr-hul-i-skolesystem-til-datatilsynet-nu-faar-han-en-advarsel-af-skolen
Møllerhøj, J. (2015b). LÌs advokatfirmaet Bech-Bruuns privacy-kritik af gymnasie-it. Version2, September 23. https://www.version2.dk/artikel/laes-advokatfirmaet-bech-bruuns-privacy-kritik-af-gymnasie-it
Møllerhøj, J. (2015c). It-system til varsel af elevfrafald blev øjeblikkeligt standset af gymnasierne. Version2, August 17. https://www.version2.dk/artikel/it-system-til-varsel-af-elevfrafald-blev-oejeblikkeligt-standset-af-gymnasierne
Mølsted, H. (2014). Big Data-vÌrktøj rykker ind i gymnasierne og fortÌller hvem der dropper ud. Version2, June 2. https://www.version2.dk/artikel/big-data-vaerktoej-rykker-ind-i-gymnasierne-og-fortaeller-hvem-der-dropper-ud
Neyland, D. (2016). Bearing account-able witness to the ethical algorithmic system. Science Technology & Human Values, 41(1), 50â76.
Raji, I. D., Bender, E. M., Paullada, A., Denton, E., & Hanna, A. (2021). AI and the everything in the whole wide world benchmark. arXiv preprint arXiv:2111.15366.
Ranchordas, S. (2021). Empathy in the digital administrative state. Duke Law Journal, 71, 1341.
Ratner, H. F., & Elmholdt, K. (2023). Algorithmic constructions of risk: Anticipating uncertain futures in child protection services. Big Data & Society, 10(2), 20539517231186120.
Ratner, H. F., & Schrøder, I. (2024). Ethical plateaus in Danish child protection services: The rise and demise of algorithmic models. Science & Technology Studies, 37(3), 44â61.
Ratner, H. F., & Thylstrup, N. B. (2025). Citizensâ data afterlives: Practices of dataset inclusion in machine learning for public welfare. AI & SOCIETY, 40(3), 1183â1193.
Ravn, L., N'Diaye, B., Mackinnon, K., Thylstrup, N. B., & Muravyov, D. (2026). Governing by dismantling: tech oligarchy and the stifling of public data infrastructure. Science as Culture, 1â16.
Rella, L., et al. (2025). Hybrid materialities, power, and expertise in the era of general-purpose technologies. Distinktion: Journal of Social Theory, 26(1), 138â157.
Sara, N. B., Halland, R., Igel, C., & Alstrup, S. (2015). High-school dropout prediction using machine learning: a danish large-scale study. In ESANN 2015, 319â324.
Schjøtt, A., & Blanke, T. (2026). Making machine learning good enoughâstudying the political endeavour of finding ârightâmetrics and thresholds. Big Data & Society, 13(2). https://doi.org/10.1177/20539517261429198
Smith, G. M., Altenor, A., Altenor, R. J., Mack, D., Thomas, J., Laucius, J., Roth, C., & Schaufenbil, R. (2023). Effects of ending the use of seclusion and mechanical restraint in the Pennsylvania State Hospital System, 2011â2020. Psychiatric Services, 74(2), 173â181.
Star, S. L., & Griesemer, J. R. (1989). Institutional Ecology, âTranslationsâ and Boundary Objects: Amateurs and Professionals in Berkeleyâs Museum of Vertebrate Zoology, 1907-39. Social Studies of Science, 19(3), 387â420.
Steinert, T., Noorthoorn, E. O., & Mulder, C. L. (2014). The use of coercive interventions in mental health care in Germany and the Netherlands. A comparison of the developments in two neighboring countries. Frontiers in Public Health, 2, 141.
Suchman, L. A. (1987). Plans and situated actions: The problem of human-machine communication. Cambridge University Press.
Suchman, L. (2002). Located accountabilities in technology production. Scandinavian Journal of Information Systems, 14(2), 91â105.
Suchman, L., Trigg, R., & Blomberg, J. (2002). Working artefacts: ethnomethods of the prototype. The British Journal of Sociology, 53(2), 163â179.
Thylstrup, N. B., Hansen, K. B., Flyverbom, M., & Amoore, L. (2022). Politics of data reuse in machine learning systems: Theorizing reuse entanglements. Big Data & Society, 9(2). https://doi.org/10.1177/20539517221139785
Thylstrup, N. B., & Hartley, J. M. (2025). âArgh! the world doesn't fit the model!â: small acts of worldmaking in data annotation for news media. Media Theory, 9(2), 105â132.
Vonk, G., & RanchordĂĄs, S. (2025). Welfare state dystopia and the response of the law: An agenda for the 21st century A brief editorial introduction. European Journal of Social Security, 27(2), 77â81.
Wenzelburger, G., KĂśnig, P. D., Felfeli, J., & Achtziger, A. (2024). Algorithms in the public sector. Why context matters. Public Administration, 102(1), 40â60.
Acknowledgements
This work would not have been possible without the two foundationsâ commitment to advancing interdisciplinary scholarship on socio-technical systems. We are indebted to our research participants in the Danish educational institutions and Dutch psychiatric clinics who generously shared their experiences and insights. We extend our thanks to the editors of this special issue for their dedication and vision, to the participants of the preceding workshop for fostering a rich space for critical reflection on the politics of evaluation, and to the reviewers whose incisive comments and suggestions substantially strengthened the manuscript.
Funding
Open access funding provided by Copenhagen University. This research has received funding from the European Research Council (ERC) Grant no. 101078386 and VELUX Foundations (Jubilee grant), which has enabled critical research on digital infrastructures and algorithmic governance.
Author information
Authors and Affiliations
Corresponding author
Ethics declarations
Competing Interests
The authors have no competing interests to declare that are relevant to the content of this article.
Additional information
Publisherâs Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the articleâs Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the articleâs Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
About this article
Cite this article
Thylstrup, N.B., Ratner, H.F. & Carboni, C. When Algorithms Meet Institutions: Evaluation Ecologies in Public Sector Machine Learning. Digit. Soc. 5, 41 (2026). https://doi.org/10.1007/s44206-026-00282-2
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s44206-026-00282-2
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content â general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached â you'll always get the same 5 for this article.