Can Generative Intermediaries Deliver When It Comes to Information Quality?
Can Generative Intermediaries Deliver When It Comes to Information Quality?
The digital era has been defined by shifts in how people access information. Early search engines of the 1990s revolutionized how people navigate a complex web of online information; smartphones and social media platforms of the 2000s made that content more accessible than ever. So it’s no surprise that with generative AI products now in widespread use, people are turning to them to seek information and complete tasks across all kinds of topics. The stakes couldn’t be more clear: LLM-powered systems are mediating what people know and believe, the decisions they make, their ability to participate in democracy, and their ability to access trustworthy information on the issues that matter most to them. This shift is why the Center for Democracy & Technology is launching a new body of work examining how generative AI systems shape the information people encounter — and what it will take to improve the quality of those systems’ outputs while protecting rights like privacy and freedom of expression.
As with other internet technologies, generative AI products can lead people to incorrect information, present irrelevant invective, or even facilitate the spread of state-sponsored propaganda to users. People turning to AI systems for health information have sometimes been met with responses that contain inaccurate medical information or lack important legal context, dangerously affecting their physical well-being in critical domains like reproductive health. In the context of elections, chatbots have presented voters, especially those from marginalized communities, with inaccurate details on polling locations or voter registration guidelines, impairing their ability to exercise their right to vote. And AI-generated news summaries have misrepresented news coverage, risking a less-informed democracy in which a small number of AI models wield disproportionate influence over people’s perceptions of the world. As people increasingly turn to generative AI products for authoritative content, the quality of their outputs warrants independent, public-interest scrutiny.
On an operational level, shoring up the quality of AI-generated outputs is not an impossible task. AI companies can partner with domain experts, such as elections officials, to redirect users to reliable resources. Independent audits for information quality issues can also help uncover data voids and motivate interventions. AI models can be connected to Model Context Protocol (MCP) servers, which, if appropriately chosen and securely designed, can orient LLMs toward structured, reliable content. Civil society could invest in curating accessible APIs or repositories that contain grounded information from different domains for LLMs to draw on. And transparency into model development can help deepen understanding across the AI ecosystem about why models produce inaccurate or misleading outputs, and point toward what can be done to improve them. Determining which of these interventions actually work — and for whom — is central to the work ahead.
We recognize, of course, that what constitutes “high-quality” information at all is contested in many domains. Responses to closed-ended questions like, “Where is my nearest voting location?” can generally be grounded with factual, undisputed information. But queries covering more contested topics — those where answers might vary depending on someone’s values or preferences — face more challenging questions around the factors that contribute to output “quality.” For example, queries such as, “What is the best university?” or, “Is it safe for minors to take hormone blockers?” may confront contexts that are in flux or rely on premises that are disputed. LLMs’ responses in such domains are likely to face scrutiny from different actors with different motivations, even if technical gaps for when it comes to sourcing accurate information have been largely resolved. We do not aim to resolve these questions, nor claim they are resolvable by technical means. Our initial work will focus primarily on laying the foundations for these conversations to happen more productively, and in open fora where decisions about appropriate interventions can be deliberated rather than obscured with technical language.
AI developers and deployers also need to contend with attempts by outside actors to deliberately influence the information these systems produce. LLMs are known to be susceptible to external manipulation, with examples ranging from lighthearted to nefarious. Malicious actors have succeeded in exploiting model architecture through various algorithmic poisoning techniques to spread disinformation. And marketers and political strategists have begun to focus on boosting brand visibility in LLM outputs by using methods like generative engine optimization (GEO). Designing interventions that improve the quality of outputs without opening new avenues for actors to exploit the information within them will require significant care.
Importantly, solutions that seek to improve AI information quality in good faith involve difficult tradeoffs that require understanding and navigating important equities like privacy and free expression. Well-intentioned efforts to improve outputs might lead AI developers to trade off privacy for precision. For instance, if a user asks an LLM-powered tool a question about local elections, a response informed by knowledge of where that person lives or their party affiliation might be most immediately helpful. Or, when users seek sensitive health information, the accuracy or safety of a model’s output may depend on consideration of personal and protected details, like an individual’s race, gender, and medical history. This type of personal information might enhance output quality, but without care could leak into other contexts where it’s not appropriate, be used to train other models, target users with uncomfortable or predatory ads, or pop up inappropriately to other users. Users’ chat history could also be vulnerable to law enforcement requests, which may carry serious legal implications when users are discussing politically fraught topics like immigration or reproductive health. Put simply, privacy can’t be decoupled from information quality.
Moreover, government efforts to influence how developers shape what information their models produce pose serious risks to individuals’ rights to access information, and to free expression more broadly. In some countries, state coercion is already in place — in China, chatbots must pass an ideological test before public release, and Turkey has banned AI-generated content that insults high-ranking officials. Global freedom of expression standards generally protect the right of people to seek and to receive information, including through generative AI outputs. In the United States, government attempts to impose content- or viewpoint-based restrictions on the information provided or received via generative AI models are likely to be subject to heightened First Amendment scrutiny. Here, the Trump Administration’s Woke AI Executive Order, the Federal Trade Commission’s proposed policy statement about the “accuracy” of models, and arbitrary government export controls on certain AI models demonstrate some of the ways in which risks to the independence of model outputs could materialize. If a government is afforded the power to control what information LLMs can and cannot generate, AI companies may be forced to shape model outputs to align with the government’s viewpoint on politically polarizing topics, like gender-affirming care or vaccines. These dangers underscore the importance of ensuring that AI governance interventions maintain an information ecosystem free from government coercion and control.
The intersection of these questions — technical and legal, operational and strategic — are exactly why CDT is investing in this area. Our forthcoming work will interrogate how AI-generated information is shaped, what factors contribute to lower-quality information production, and what interventions might improve outputs while protecting privacy and freedom of expression. It will build on broader bodies of research regarding the levers stakeholders can pull to improve the information generated by AI systems, the impact of AI personalization, interventions that balance user rights against the risks associated with data voids, and methods to ensure that AI-generated information strengthens, rather than weakens, democracy.
We will also take on broader questions of how to ensure the widespread availability of high-quality information on the internet, at a moment when new AI systems both draw human eyes away from websites and swamp those sites with AI bot traffic. Accurate AI outputs depend on reliable information throughout a model’s lifecycle; where will that information come from if the open internet that once delivered it freely is no longer sustainable, economically or otherwise?
Generative AI systems will continue to evolve, and so will the ways they surface information. Getting this right will take deep technical research, careful policy and legal analysis, and sustained multistakeholder engagement. And it will take partners who share our conviction that the information generated by AI systems should strengthen, not weaken, people’s access to knowledge and the health of our democracy. If you are working in this area or interested in collaborating, please be in touch.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.