general12262 wordsRead on Arc Codex

Multimodal LLM Categorization for Humanities Research on Social Media: A Case Study with Pepe the Frog

Introduction In the days following the assassination of right-wing digital content creator Charlie Kirk, images circulated online of the suspected shooter from some years prior, wearing a Halloween costume that made it look like he was riding a Pepe the Frog figure that had Donald Trump’s hair. This photograph was one of several images and texts bound up in claims about the shooter’s political positionality made by elected officials, media commentators, and everyday social media users. This debate is just one iteration of the implication of the Pepe meme with American politics, which came to the fore during the 2016 election and the January 6, 2021 riot at the United States Capitol. These high-stakes moments of Pepe’s political visibility make clear how important it is to have detailed, context-specific interpretations of both this meme and memes in general. For observers closely connected to the online culture in which Pepe and other political memes originate, interpretation can be both clear and nuanced, but for those outside that context, memes can cause confusion and misunderstanding, especially as they are often used in highly ironic and intricately referential ways. If this material is difficult for humans to parse, can Large Language Models do a better job? In other words, can AI understand the Pepe the Frog meme, which is politically sensitive and can function by turns as hateful, sexually explicit, humorous, ironic, and emotionally sincere? Moreover, when humanities scholars use LLMs as tools to understand multimodal political meme material at scale, what patterns should we be aware of in terms of how LLMs’ qualitative assessments may shape the interpretation of our datasets? We approach these questions through a categorization task conducted with four models, using a dataset that comprises 3407 tweets with imagery related to the Pepe the Frog meme. Drawn from the months surrounding the January 6, 2021 riot on the US Capitol, many of these tweets were posted by unidentified users and/or trolls.1 They serve less as representations of any given user’s personal perspective or as cutting-edge contributions to meme culture than as acts of propagating certain familiar types of politically polarizing and sensitive online discourse. Pepe the Frog was originally created by American illustrator Matt Furie, appearing as the protagonist of Furie’s comic Boy’s Club in 2005–6. Over the ensuing decade, the character began circulating first amongst niche communities of internet users on platforms including MySpace and 4chan, and eventually to a mainstream audience. In 2016, the Anti-Defamation League included Pepe the Frog in its catalogue of hate symbols due to its connection with racist, anti-semitic, and otherwise hateful expressions, specifically as associated with the alt-right. While Pepe shares some characteristics with other memes, it is distinguished by its semantic and affective flexibility and by its use to express a wide range of emotions, which has been ongoing for close to twenty years across numerous digital platforms. Other work on the Pepe meme has conducted quantitative analysis of its development and circulation across specific platforms (Zannettou et al. 2018) and has examined in close qualitative detail the signification of the Pepe character (Glitsos and Hall). Building on this scholarship, we offer an analysis of the meme that is both qualitative and grounded in a robust dataset, which we also present as a test case for the utility of LLMs in humanities research. As LLM use becomes increasingly common in Digital Humanities work on larger datasets, it is important to understand the tools not only as succeeding or failing relative to ground-truth human evaluation of qualitative material, but in terms of how their categorization more subtly shapes the interpretation of datasets.2 The Pepe meme is a compelling case study in this regard because it is a visual culture artifact whose individual iterations can only be interpreted in the context of a given tweet or thread. LLMs may encounter challenges categorizing them both because of that contextual specificity and because of guardrails that provoke task refusal due to hateful or sexually explicit content. To capture that contextual specificity, we developed a highly detailed codebook to sort the tweets by their political positionality, hatred context, sexual content, and other contextual references, including those to Donald Trump. Three human raters coded each post, allowing us to establish a benchmark for human agreement, which we set at 0.55 for minimal acceptable agreement and 0.65 for strong agreement. Using a zero-shot prompt based on our codebook, we queried models including Claude Opus 4 and Sonnet 4 (Claude’s highest-capability model and a cost-efficient general-purpose model accessible on Claude.ai’s free tier), Mistral Pixtral 12B, and LLaVA-NeXT. The range of models selected includes those of different sizes and generations to capture differences that everyday users may encounter when navigating politically sensitive material via the consumer interfaces of free versus paid AI tools. Predictably, our results show that the Claude models were both close to one another in their ratings and in good agreement with the human raters, while LLaVA-NeXT and Mistral Pixtral 12B failed to pick up on much of the hateful, sexual, and contextual content in the tweets. Less predictable were how the LLMs handled the tags we used to differentiate between types of hatred being “encouraged” (i.e., promoted) versus “referenced,” to account for posts that might discuss bias while not encouraging it, which is relatively common amongst left-wing posts using this meme. Namely, we found that all four LLMs were more hesitant than human raters to label hatred as “encouraged” even when they identified its presence. Especially with the Claude models, this conservative approach is likely related to safety and obscenity guardrails. However, it has an unexpected consequence, in that while premium LLMs may be strong at identifying politically sensitive content, they may interpret such content for users in ways that further weaken already-loose notions of social responsibility in the online public sphere. Historical background and definition of terms This research began with an interest in studying the circulation and evolution of meme imagery relative to current historical events, from a disciplinary perspective grounded in Art History and visual culture studies. The January 6, 2021 riot on the US Capitol marked an unprecedented moment of offline public visibility of the Pepe the Frog meme, because numerous rioters carried signs and wore costumes connected to it. For example, one young man wore a MAGA hat and carried a hand-drawn cardstock sign with a Sad Frog illustration and the phrase “Green Lives Matter,” in a demeaning parody of the Black Lives Matter movement. The meme was featured in the New York Times’ visual investigation “Decoding the Far-Right Symbols at the Capitol Riot,” which described it as “the smirking cartoon amphibian that has become a widely recognized symbol of the alt-right crowd” (Rosenberg and Tiefenthäler). Intrigued by the key role of a roughly drawn image in a major event in American history, we collaborated with Professor Melanie Walsh, who harvested a dataset of tweets from Twitter’s academic research API, shortly before it was shut down in early 2023. Walsh searched for tweets from 2020 and 2021 that met the following criteria: 1) they contained images, 2) they made textual reference to the name “Pepe,” and 3) they made textual reference to Donald Trump, MAGA, or January 6. As we began our work with this dataset, we arrived at a new question: how useful might AI be in helping us “see” a group of images that are impossible to perceive visually all at once? How helpful would it be in facilitating our interpretation of this material, specifically given the politically sensitive nature of the Pepe meme? In this research, we intentionally employ an omnibus definition of politically sensitive to include 1) hateful and abusive speech, 2) overt discussion of or encouragement of political polarization, and 3) reference to political events or figures. The present article focuses predominantly on the first two categories. Our definition is tailored to the politically polarized climate of the United States in the 2020s and to the pervasive online blurring among these three types of discourse. In their study of the efficacy of 4chan and other right-wing forums in pushing meme content to mainstream platforms, Hine et al. show that even on forums such as 4chan’s “politically incorrect” (/pol/) board that are heavily associated with alt-right speech, only a minority of posts contain overt hate terms (12% on /pol/, versus 2.2% in a sample of posts from mainstream platforms) (7). Research further shows that LLMs are strong at diagnosing contextually dependent hate speech (Guo et al.). Building on that work, we turn to a wider category of online meaning production that can be more subjective and harder to diagnose than overt hate speech, and in which the visual plays a central role. Pepe the Frog is an ideal focus for this investigation because of the meme’s visual and semantic flexibility. There exists a huge variety of versions of this meme, ranging from easily recognizable types such as Sad Frog, Smug Frog, Angry Pepe, Groyper, and Twerky Pepe to “rare” Pepes that include everything from one rendered in the style of a Jean-Michel Basquiat painting to others of Pepe being cradled by Jesus. The most common variants function in online discourse as tone indicators, modifying accompanying texts to emphasize or clarify the poster’s feelings. While some variants such as Sad Frog are rendered in an illustration style that hews to Furie’s original Pepe drawings, others have specific stylistic traits, such as the JPEG artifacting on the Smug Frog, the shaded, hard edges of Twerky Pepe, or the poignantly distorted features of Apu Apustaja, a version that started appearing on the Finnish image board Ylilauta in 2016 (Figure 1). These variations in illustration style underscore the range of emotions and experiences expressed by the meme, and moreover point to its status as a fundamentally–even iconically–digital object. Notably, although the origin point of the Pepe character was an analog comic Furie drew by hand, Furie himself posted illustrations featuring Pepe on MySpace as early as 2005, before the comic’s publication in 2006 (Mazur). As such, the start of its life as a digital object was essentially contiguous with its creation as an analog one. Scholars of visual culture have characterized meme images as “poor,” that is, visually simple, unaesthetic, and low-resolution, but for exactly those reasons adept at traveling and taking on new meaning.3 Jason LaRiviere (2024) describes memes as “compression” images, both technically, because of low resolution and small file sizes, and semantically, because of the way a meme can acquire but also lose meaning. This description is especially apt relative to Pepe the Frog. Pepe has acquired layers of meaning that can obscure or eclipse one another in any particular iteration of the meme. Moreover, Pepe’s meaning on the whole remains an active subject of contestation amongst internet users in a way that is more intense than for almost any other meme. Furie’s Boy’s Club was a celebration of what one might describe as “stoner bro” friendship and masculinity, a low-key chronicle of goofy characters focused on bodily and social pleasure. Though the characters in Boy’s Club read largely as white, the comic lacks the racist and sexist associations that would come to plague the meme.4 Pepe’s political polarization began around 2011, when the Sad Frog variant appeared frequently in anecdotes on 4chan’s “Robot 9000” (/r9k/) board, which saw increased racist content (Glitsos and Hall). Following the increasingly mainstream deployment of Pepe, including in tweets by Nicki Minaj and Katy Perry in 2014, some 4chan users launched a campaign to generate hateful variants to make it less palatable to mainstream users (Glitsos and Hall). On October 13, 2015, Donald Trump made explicit the connection between the politics of those online spaces and the politics of his presidential campaign by tweeting an image of himself as Pepe the Frog standing at a podium with the seal of the President of the United States, against a red and white striped background (Braynard). In the late 2010s, the Groyper variant of the meme became associated with the group of white nationalists led by Nick Fuentes, a number of whom held jobs in the second Trump administration as of mid-2025 (Pemberton). Laura Glitsos and James Hall, in their analysis of the Pepe meme, argue that the figure of a humanoid frog resonates deeply with cultural narratives about the social role of the offensive and the disgusting (think of a princess kissing a frog, etc.). They posit that this made Pepe fertile ground for appropriation by alt-right internet users interested in crossing boundaries and flaunting taboos (Glitsos and Hall). The Anti-Defamation League, in its hate database entry for Pepe the Frog, specifies that because “many Pepe the Frog memes are not bigoted in nature, it is important to examine use of the meme only in context.” This point is crucial, all the more so if we distinguish between bigotry as expressed in the images themselves versus in their contextual deployment alongside text. It is borne out in our dataset, where only a small handful of Pepe images express hatred in overt visual terms (for example, in one entry that shows a blackface Pepe), as opposed to via their contextual deployment in connection with hateful terms or narratives. As such, though the Pepe meme may function as an assertion of hatred and of affiliation with white nationalist groups, its function in that respect differs strongly from hate symbols such as the swastika, the Confederate flag, the KKK’s Blood Drop Cross, or the U of the Croatian Ustaše, which are used to signify hate and violence in a consistent way across contexts. The difference here lies not only in the Pepe meme’s digital origin versus the pre-digital histories of these other symbols, but rather in the fact that its dominant function is to express emotion and subjective experience in ways that can include the expression of hatred.5 In that sense, the ideology it expresses is not one of affiliation to a specific hateful worldview, such as neo-Nazism, but rather to a vision of the public sphere in which the boundary between hate speech and normal speech is effaced in favor of enabling the ultimate freedom of expression of the individual. The contrast with traditional hate symbols is, in fact, especially evident in the hateful iterations of the meme that show Pepe in the guise of Adolf Hitler or a KKK member, because they juxtapose the enduring, bounded meanings of those figures with the ever-mutating, contextually-inflected nature of Pepe.6 An understanding of sovereign individual freedom of expression as foundational to the public sphere corresponds roughly to a libertarian political position. However, unlike in classic libertarianism, the political subject implied by the Pepe meme is also a highly emotionally expressive individual. There are aspects of “technopopulist” or “technolibertarian” thinking connected to this understanding of subjecthood, but with an emphasis on emotional expressivity that corresponds to social media rather than to ideas about the liberatory nature of information technology.7 This brings us to the definitions of the terms “right-wing” and “left-wing” as we used them to categorize the tweets. Donovan et al., in their study of the #StopTheSteal campaign that catalyzed the January 6 riot, situate Pepe in the landscape of “meme wars,” which they define as “culture wars, accelerated and intensified because of the infrastructure and incentives of the internet, which trades outrage and extremity as currency, rewards speed and scale, and flattens the experience of the world into a never-ending scroll of images and words” (10). This framing makes clear that meme circulation is an arena in which users traffic in political positionality while also fundamentally transforming conceptions of it. This creates a field in which readers—including both social media users and the human and LLM raters for our dataset—must rely on intuition, inference, and assembled contextual cues in order to diagnose political positionality. Within this dataset, right-wing positionality might be evident through clear support for Donald Trump and the January 6 riots, support for Second Amendment rights, or disparagement of feminism or Black Lives Matter, for example, through the slogan “Pepe Lives Matter.” A left-wing position, on the other hand, was often communicated via calling out the role of Pepe in right-wing ideology, for example, in a tweet that shows an image of Pepe as Donald Trump with the statement that the president used the meme as a “frog whistle” to recruit alienated men sympathetic to his message of resentment against women, immigrants, and people of color. It is also evident through assertions of reclaiming Pepe from those discourses via both words and images to return him to the fun-loving camaraderie of Furie’s original comic, such as an illustration of Pepe wearing bell bottoms and seated on top of the world in front of a rainbow (Figure 2). The text in that tweet states that the poster had recently read Boy’s Club, wanted to emphasize that Pepe was “dope,” and decried the fact that the right had appropriated him from “us.” This post, and several others that reflect a left-wing position, do so in connection with a positive attitude towards Arthur Jones’ 2020 documentary Feels Good Man, which presents a heroizing narrative of Furie and his lawsuits to enforce his copyright to the Pepe character. In sum, concerning the meme as a whole, left-wing politics typically find expression via reference to and active rejection of its right-wing signification. Literature review Decisions about social media data There is a rich body of scholarship on the ethical considerations presented by the use of social media data. Datasets based on social media contain information about individual users that can easily be identifying. Users often share their content without awareness of how it can be collected and used for research, creating problems in terms of consent that are frequently unaccounted for in institutional IRB processes (Bailey). Scholars working in this area are looking carefully at how marginalized groups, in particular, stand to be harmed by the non-consensual use of social media data for research purposes (for an overview of the debate, see Walsh 2023). In response to that danger, scholars may take measures to seek consent and collaboration, including engaging directly with social media users associated with their datasets. Another, compatible approach is to establish a threshold for what level of replies and reposts constitutes a post as already public, so that the research does not bring additional unwanted attention to functionally private posts.8 Direct engagement with users to seek their permission to publish tweets or their opinions about the research would not be a good choice for our dataset. First, as mentioned above, it is likely that a number of the posts in this dataset were created by trolls and/or unverified users, including a large volume from a handful of accounts that predominantly screenshot 4chan content related to Pepe the Frog and repost it to Twitter (most of which received little to no interaction from other Twitter users).9 Second, researchers have understandably used caution in engaging with alt-right internet users, who can be highly aware and critical of scholarship on their subculture, and may not hesitate to harass researchers.10 This necessitated that we develop an ethical strategy for handling our data without relying on community engagement. As such, we selected paraphrasing as the best strategy for keeping users anonymous and individual posts untraceable. In a few instances, we reproduce tweets in full only when they are material taken from 4chan and reposted on Twitter via anonymous accounts. We do so in order to demonstrate in detail the relationship between tagging explanations provided by the models and the original posts. Use of LLMs for text and image categorization Categorization is a historical strength of algorithms. Since 2022, a large body of research has analyzed the performance of LLMs in categorizing qualitative textual content (e.g., Gilardi et al.). Multiple studies have shown the strength of LLMs in classifying political text specifically. Examples of tasks include identifying hate speech or politically-driven misinformation, determining the sentiment inherent in political texts, or predicting the political positionality of social media users based on their posts.11 On the whole, these studies show both that LLMs hold great promise in this area, often equalling or exceeding the work of human annotators, and that methodology is crucial in facilitating their performance. Methodological decisions often revolve around prompt engineering, but also the question of how much context to provide an LLM to help it diagnose politics in a dataset. Sucu et al. show that providing contextual information from users’ historical posts for a dataset of political forum posts substantially increased LLMs’ accuracy in diagnosing political position. The application of LLMs for recognizing political content is not limited to text, but extends to images.12 Kosinski et al. (2024) provide a striking demonstration of this by showing that both human raters and LLMs were able to discern people’s political alignment based on unadorned photographs of their faces, at a rate significantly above chance. Art historian Nancy Um has urged caution regarding overenthusiasm about LLMs’ abilities to conduct meaningful, transparent analyses of the relationship between image and text. Drawing on experimental engagement with a small number of queries through Copilot’s user interface, Um argues that the stochastic nature of LLMs prevents them from providing historically accurate representations or discussions of visual material (188). While it is true that LLMs’ engagement with the text/image relationship will remain stochastic and thus retain a dimension of unpredictability, studies of large datasets have shown the potential of LLMs for image-related tasks, including visual recognition and the creation of image descriptions. Moreover, Wu et al. show that, by leveraging GPT’s sophisticated linguistic knowledge, it is possible to improve its zero-shot visual recognition by having it generate detailed descriptions of the categories it uses to recognize images. Predictably, across categorization tasks in both text and images, LLMs struggle most when the forms to be categorized are most linguistically or visually abstract, for example, in identifying concrete poetry (Walsh et al.), or recognizing organs in medical imaging (Nam et al. 2025). A subset of this research concerns material very similar to the contents of our dataset: text/image combinations that express subjective, humorous, ironic, and subversive meaning. Hessel et al. (2023) focus specifically on the task of understanding humor, asking models including GPT-4 and fine-tuned CLIP ViT-L/14 to match cartoons to winning captions from the New Yorker weekly caption contest. They found that providing models with contextual information about cultural references produced better results, and, moreover, that computer vision posed a limitation for the task: the same model would perform substantially better when provided with written descriptions rather than the images themselves (Hessel et al. 693–4). Indeed, both understanding and creating humorous material are areas in which grasping context is very important. Wu et al. demonstrate that LLMs are capable of enhancing human productivity in the creation of humorous memes (i.e., by creating texts for preselected meme images). They could also independently create memes that received high scores from human raters for humor, creativity, and shareability, compared with memes generated by humans and those produced in dialogue between humans and LLMs. However, the authors also showed that among the top-performing memes generated across these groups, those made independently by AI appeared less funny to humans than those made by humans. Arguably, while LLMs can create credible memes when provided with a template, the humor in their output may lack a certain je-ne-sais-quoi attributable to the context-dependent nature of human jokes. While the memes at issue in Wu et al.’s study are generally rather anodyne, the alt-right employs memes to express not only humor but also anger, hatred, cynicism, superiority, and group affiliation. Those emotions play an important role in articulating and consolidating political positionality as expressed in the posts in our dataset. Our task does not focus on the generative capacity of LLMs regarding this kind of humor, an application that would raise serious ethical problems. Rather, we focus on their capacity to act as readers alongside humans to categorize and decode large volumes of tweets, thereby helping humanists navigate large volumes of social media material with a high level of nuance. LLMs can help humanists draw conclusions across a dataset of meme material, which researchers can then combine with their own visual analysis of individual examples, to capture the complexity of the data at various scales.13 But when using the tools this way, humanists need to be attentive to how subtle divergences between LLM categorization and human interpretation can be consequential for understanding which kinds of public-sphere memes work to construct. Methodology Selection and categorization of posts We created the dataset for these experiments by hand-culling 121,117 tweets dating from November 2020 to March 2021, the five months bookending the January 6 riot. The initial batch contained a large volume of “noise” in the form of tweets irrelevant to the Pepe meme. After manually selecting 3,407 entries, we developed a codebook to quantify their content, ultimately employing four categories in addition to the actual image, text, and post date to categorize each post (Figure 3).14 The metadata categories pertinent to the present paper concern 1) whether the post contains hatred or abuse content, 2) whether it contains pornographic or sexual content, 3) whether it reflects a legible political positionality, and 4) whether it refers to one of a series of specific deployments of the meme, including Arthur Jones’s documentary Feels Good Man, the appearance of Pepe imagery at the January 6 riot, direct reference to Donald Trump, or reference to the meme’s connection to and circulation as cryptocurrency. Note that Pepe’s association with cryptocurrency is an especially literal example of Limor Shifman’s argument that memes blur the boundary between market-driven and non-market forces. For the hatred/abuse category, we created seven tags: violence (both physical and verbal), ableism, body shaming, sexual orientation and identity-based discrimination, including homophobia and transphobia, misogyny, racism, and finally “intersectional discrimination” for posts that contained multiple forms of hatred or abuse.15 For each of these tags, we further distinguished whether the form of hatred/abuse was referenced or promoted, to indicate, for example, a post that discussed racial discrimination but did not actively promote it. Establishment of inter-rater reliability Most studies of LLMs’ performance on categorization, including specifically of political texts, establish human rater agreement as a ground truth against which to measure LLM performance (e.g., Heseltine, M., and B. Clemm von Hohenberg). Exceptions include Petter Törnberg’s 2023 study of LLM classification of tweets from the period leading up to the 2020 US election, in which the author curated a dataset of tweets only by US Senators and then fed them to a model without identifying their authors, thereby using the Senators’ party affiliation to establish ground truth in terms of political positionality.16 For the purposes of our research, three humans coded each post. At the same time, we are cautious about using their agreement or lack thereof as a ground truth against which to measure LLM categorization, and instead approached human agreement as a self-reflexive baseline to understand LLM performance. In taking this approach, we build on previous work that has shown how (dis)agreement amongst LLMs relates to (dis)agreement amongst humans. For example, when Tai et al. (11) experimented with having LLMs code short segments of ethnographic interviews with STEM researchers in terms of the ideas and emotions present, they showed not only that LLMs are reliable at this task but that the areas of disagreement between LLM responses were greater in the areas where human raters also disagreed more. Hessel et al. (692) note that human performance is not an upper bound for LLM performance, because human perspectives vary. This type of approach emphasizes the subjective nature of qualitative data. Moreover, regarding social media content that may be bot-generated and draws very low interaction from human social media users, the “readers” of such posts may be equally likely to be machines (whether bots or LLMs, when social media posts are captured in training datasets) as they are to be human. In light of those approaches, we represent rater agreement not via Fleiss’s kappa for all seven human and LLM raters, but via pairwise Cohen’s kappa scores, in order to focus on patterns of (dis)agreement between individual pairs of human and non-human raters. Based on the precedents in Sachdeva et al. (2022) and Waseem and Hovy (2016), we employ Cohen’s kappa scores of 0.55 as the benchmark for minimally acceptable agreement and 0.65 for strong agreement. The results of individual pairings are discussed in the Results section. Models tested Because of our focus on the utility of LLMs for humanities scholars and students, we did not fine-tune models but rather selected a variety of out-of-the-box tools to test. In addition to an open-source model (Mistral Pixtral 12B) and a local model (LLaVA-NeXT), we tested both the Opus 4 and Sonnet 4 models of Anthropic’s Claude. LLaVA-NeXT runs entirely locally, reports competitive visual reasoning, and exceeds some proprietary baselines (e.g., Gemini Pro) on several benchmarks (Lee). Our tests align with Zhu, Ke, et al. (2025) on strong abstract-visual interpretation, while our results also show that LLaVA-NeXT occasionally produced explanations without generating tags regardless of the structured prompt.17 Claude Opus 4 and Sonnet 4 are “hybrid-reasoning” models: they can quickly provide near-real-time responses while also supporting longer periods of deep thinking to handle complex tasks (Anthropic n.d.). Opus 4 is the most advanced general-purpose model in Claude, and Sonnet 4 excels in visual data extraction, capable of extracting structured information from charts, diagrams, and complex typography.18 Both models performed robustly in accurately understanding the context of image-text combinations and identifying “contextual labels.” Mistral Pixtral 12B is an open-source 12-billion-parameter multimodal language model from French startup Mistral AI. Sáez et al. (2025, 9, 14) find Pixtral to be inconsistent compared to other LLMs, such as Gemini 2.0 Flash, Qwen-VL 2.5 72B and GPT-4o.19 Our results from Pixtral also demonstrated inconsistencies, including new tags that it invented based on the format of our tags, which we were unable to weed out despite adding more structure and specificity to the prompt. Finally, our selected models—Claude 4 Opus, Claude 4 Sonnet, Mistral Pixtral 12B, and LLaVA-NeXT—vary widely in terms of their safety guardrails and refusal behaviors. In a study of safety guardrails, researchers measured the over-refusal of 32 popular LLMs across 8 model families (Cui et al.). Their results point out a crucial trade-off: most models achieve safety, measured by toxic prompt rejection, at the expense of over-refusal. For instance, Claude models demonstrate the highest safety but also the most over-refusal, while Mistral models accept most prompts. This indicates that while some models are highly effective at identifying toxic prompts, they are simultaneously much more cautious about the safety of non-toxic material. These behavioral discrepancies stem directly from the divergent guardrailing policies implemented by model developers. Mistral utilizes a flexible, dual-layered safety framework that pairs a Moderation API—capable of categorizing harmful content across dimensions such as hate speech and PII—with an optional “safe prompt” feature, allowing developers to tailor the strictness of safety controls to specific use cases (Mistral AI). In contrast, Anthropic employs a comprehensive, multi-layered strategy across the entire model lifecycle. Shaped by a Unified Harm Framework, Anthropic’s approach relies on extensive safety evaluations, fine-tuning with domain experts, and real-time monitoring to steer responses or suspend accounts when novel misuse patterns are detected (Anthropic). A critical consequence of aggressive safety guardrails is a systemic “bias towards neutrality.” When models encounter sensitive topics like politics, gender, or violence, their safety protocols are triggered. Instead of accurately identifying the text’s actual stance, models frequently default to a safe, neutral output (Rogers and Zhang). As models systematically smooth over the jagged edges of real-world discourse to avoid potential policy violations, this leads to analytical flattening (Rogers and Zhang). The results that we discuss below concerning social responsibility align with those findings. This oversensitivity extends significantly into the multimodal domain—a critical factor when evaluating vision-language models like Pixtral 12B and LLaVA-NeXT. Models frequently hallucinate harm in benign images, with stricter safety alignment correlating heavily with higher error rates on safe visual queries (Li et al.). These “trigger-happy” guardrails create substantial barriers for analyzing sensitive but safe real-world visual data. Some models refuse up to 76% of benign inputs due to three specific cognitive distortions: interpreting safe situations as dangerous (e.g., a toy gun flagged as a violent threat), failing to recognize when context neutralizes a sensitive element (e.g., an anatomical diagram flagged as sexual content), or completely misunderstanding visual metaphors (Li et al.). These behaviors show that safety filters are not entirely objective. Rather, guardrails are deeply influenced by the origins and regulatory environments of their developers (Ta et al.). Prompting strategy Studies on LLM classification employ a variety of approaches from zero-shot and few-shot training to fine-tuning with larger datasets.20 Amongst these approaches, researchers employ zero-shot prompting when they seek to avoid biasing the LLM through the introduction of examples, to observe the responses it generates based on its existing training.21 Because our focus was on testing the classification abilities of out-of-the-box models rather than optimizing the model’s responses through fine-tuning, we pursued zero-shot prompting over fine-tuning and multi-shot prompting. In addition to providing the codebook, we also included structured instructions that directed the model to only use one classification tag per category, i.e., to assign the best tag for a given tweet in each category, and if no tag was relevant to a tweet, to label it as “NoTag.” By breaking down the instructions into steps and providing a chain-of-thought reasoning structure, we maximized the model’s ability to recognize contextual hate speech and content (Guo et al. 5). We also directed the model to include a brief explanation describing the reasoning and evidence behind its classification for each tweet. Results Political Positioning Our results show differences both in how LLMs, on the one hand, and humans, on the other, assigned tags for political positionality, as well as substantial variation between individual models and human raters. All raters were instructed to identify whether the tweet aligned with left- or right-wing ideology. If they were unsure, the raters were to use “Unclear.” Figure 4 shows that human annotators assigned “Left” or “Right” more freely than LLMs did. Cumulatively, 68% of all “Left” tags were assigned by humans. The difference is more striking with “Right”: 78% of “Right” tags were assigned by humans. The LLMs were cautious in their political tagging, even with some tweets that would seem, from a human perspective, to provide numerous cues in terms of political position. An example is the LLMs’ failure to identify political positionality in a December 2020 tweet depicting Donald and Melania Trump, accompanied by Pepe the Frog, approaching a helicopter and being greeted by a military officer. All three have a symbol of Marvel’s Punisher on their backs; the same symbol is visible on the window of the helicopter. The identity of the protagonists, the military setting, the Punisher’s symbolic drive towards revenge, and, finally, the text accompanying the image—“Pepe ist überall,” a German sentence denoting the omnipresence of Pepe—provided sufficient grounds for all three human raters to qualify this entry as “Right.” By contrast, Claude Sonnet 4 and Opus 4, which are otherwise fairly close to human raters in their tagging, could not identify the political orientation. Overall, Sonnet 4 and Opus 4 were substantially closer in agreement with one another than all three pairs of human raters: the Anthropic models had a pairwise agreement of 0.722 on political positionality, while the human rater pairs ranged from 0.537 to a meager 0.286 (see Figure 9). These pairwise Kappas for the human raters indicate that concerning this material, we cannot think of political positionality as a clear or determinate signal that LLMs succeed or fail to diagnose. Rather, it is an area of genuine ambiguity and disagreement between humans. If LLMs have greater collective consensus on this point than humans but are also more cautious in their tagging in ways that soften or obscure signals of political position, it is especially important to be aware of how their collective assessments may shape the qualitative interpretation of this material. In tweets from late 2020 to early 2021, we moreover find that weekly Left/Right tags track key U.S. political events, a trend that emerges in the work of the human raters and is partially mirrored by some LLMs. We observe a clear multi-rater surge in early January 2021 across both Left and Right, coincident with the Georgia U.S. Senate runoffs (Jan 5, 2021) and the U.S. Capitol attack during the electoral-vote count (Jan 6, 2021); the period also leads into the presidential inauguration (Jan 20, 2021) (Figure 5). A smaller rise appears in mid-December 2020, when the Electoral College vote was certified (Dec 14), and the series begins just after Election Day (Nov 3, 2020). When political news peaks, human raters markedly increase determinate tagging of Pepe content, and Claude Opus 4 and Sonnet 4 most closely mirror human dynamics at a lower volume, whereas LLaVA-NeXT and Mistral Pixtral 12B remain near baseline. Hate and Abuse While, with political positionality, we could make an overall distinction between human and LLM performance, that distinction was less straightforward regarding the codes for hate/abuse categories. In total, all raters were instructed to identify just one of the seven hate/abuse tags if relevant, per tweet. Moreover, the codebook provided an opportunity to highlight whether the hate/abuse was encouraged (e.g., “AbleismEncouraged”) or simply referenced (e.g., “AbleismRef”). Figure 6 shows how, in five categories (gender- and sexual-orientation-based discrimination, ableism, racism, violence, and intersectional discrimination, meaning discrimination across multiple of those categories), the three human annotators most frequently identified instances of encouragement, consistently applying the “-Encouraged” tags across these domains, although the third annotator occasionally diverged (H/T/BiphobiaRef, ViolenceRef). Overall, however, human raters most frequently classified content as actively encouraging hate or abuse, followed by Anthropic’s LLMs. Mistral Pixtral 12B abstained completely from providing such tags. Notably, the LLMs—foremost Claude Opus 4 and Sonnet 4—“compensated” for their lack of encouragement detection by applying referential tags to the aforementioned four categories. In other words, in cases when human annotators identified content as actively encouraging hate or abuse and applied the “-Encouraged” tag, the LLMs tended to downgrade the severity by substituting the “-Encouraged” tag with the corresponding “-Ref” tag. For example, for one derogatory greentext (incorporating the term “faggot”), human raters used “Homophobia/ Biphobia/ Transphobia Encouraged” because they interpreted the usage of the term as promoting discriminatory language toward the LGBTQ+ community. Meanwhile, Claude Opus 4 and Sonnet 4 used “Homophobia/ Biphobia/ Transphobia Referenced” labelling the mention of “faggot” as a reference to homophobia rather than direct promotion of it. Mistral Pixtral 12B and LLaVA-NeXT assigned “NoTag.” We observed different constellations with the categories addressing body shaming and misogyny. In detecting encouragement of body shaming, not only the Anthropic models but also LLaVA-NeXT tagged more posts than the human annotators did. With misogyny, the Anthropic models, but also LLaVA-NeXT, did not prioritize the referential tag and identified encouragement more freely. Yet, in the case of misogyny, LLaVA-NeXT abstained from assigning tags. Mistral Pixtral 12B’s performance was weak overall, assigning only 13 hate/abuse-related tags across the entire dataset. Sexual content A large number of entries incorporated sexual content, ranging from overtly pornographic stories and visuals to implicitly erotic terminology. The codebook provided four tags to measure the intensity and nature of sexual content. Figure 8 shows how human annotators detected pornographic content more readily than LLMs. Both Anthropic models and LLaVA-NeXT made more liberal use of the generalized and less morally weighted tag, labelling posts as “SexContent”, as opposed to “PornographicContent;” the latter implies intentionality and involves the explicit verbal and/or visual depiction of a sexual act (Figure 7). Specifically, LLaVA-NeXT, Opus 4, and Sonnet 4 applied the tag “SexContent” 211, 290, and 341 times respectively, substantially more than did human raters, who applied it 110 times (Rater 1), 157 times (Rater 2) and 229 times (Rater 3). This pattern echoes the scenario observed in the hate/abuse categories, where LLMs showed reluctance to apply tags ending with the suffix “-Encouraged.” While LLaVA-NeXT’s performance aligned with the other annotators in applying “Sexual Content,” it, along with Mistral Pixtral 12B, substantially favored NoTag for the final two subcategories of sexual metadata, those related to sexual objectification (“Sexual Objectification” and “Sexual Objectification - Referenced”). Interestingly, when tasked with applying tags for either the presence or the promotion of sexual objectification, the LLMs did not follow the same substitution strategy they used for hate/abuse categories, and rather than replacing “Encouraged” tags with “Ref,” they largely abstained from assigning any sexual objectification tags (Except for LLaVA-NeXT, which prioritized the “SexObjectRef” (69 instances) over “SexObjectEncouraged” (25 instances). One possible explanation is the tagging constraint: all annotators were instructed to assign only one tag per category. Thus, even if the models detected pornography, sexual objectification, or a reference to it, they may have prioritized the broader SexContent tag over more specific labels. This approach differs from human raters, as LLaVA-NeXT and Mistral Pixtral 12B took a more cautious stance rather than attempting to assign the most accurate tag. IRR Measurements To explore inter-rater reliability and agreement between raters in detail, we calculated Cohen’s kappa scores for each combination of two out of seven raters (three human, four LLM), for each category (Figure 9). Across most categories, the three human raters show some of the highest levels of agreement between each other, especially for hate and abuse references, sexual references, and contextual references, often scoring around the 0.8 to 0.9 range. Among the human raters, Raters 1 and 2 had the highest agreement with each other, scoring the highest in Hate & Abuse and Sexual References, respectively, with 0.813 and 0.887. This shows that human raters applied tags relatively consistently for most categories. However, the political position category showed lower overall agreement between the human raters, as discussed above. Each category contains at least one pairing that meets our strong agreement benchmark of 0.65, with an outlier in the Contextual References category, which includes 10 pairs with scores of 0.7 or higher. Out of these 10, nine are pairs with one human and one LLM rater, and the LLM raters are limited to Claude Opus 4 and Sonnet 4. A noticeable pattern is a consistently moderate-to-strong agreement between pairs of one Anthropic model and one human rater. Such pairs often had the third to fifth highest agreement scores in each category, especially between Human Rater 3 and either Claude Opus 4 or Sonnet 4. This suggests that the Anthropic models and human raters applied tags with moderate consistency. Another consistent pattern among the lowest kappa scores is that almost all of these pairs contain either Mistral Pixtral 12B or LLaVA-NeXT. The highest kappa score out of all pairs with either Mistral Pixtral 12B or LLaVA-NeXT was a very low 0.080 in the sexual references category. Across all four categories, the Anthropic models consistently showed strong agreement with one another. Their scores for hate and abuse references, contextual references, sexual references, and political position were 0.585, 0.936, 0.696, and 0.722, respectively, all above the minimally or strong agreement benchmark values. In each category, the Anthropic models had either the first- or second-highest agreement score, indicating consistent behavior in applying tags. The categories of Hate & Abuse References, Sexual References, and Political Position thus all show a much lower level of agreement amongst humans and Anthropic models than does Contextual References. Importantly, Contextual References is the least subjective of all four categories, because it concerns the identification of signs, symbols, texts, and imagery related to Donald Trump, MAGA, January 6th, and the Feels Good Man documentary. Of the other three categories, both Hate & Abuse Reference and Sexual Reference imply a moral dimension to categorizing the material, both in the overall category and in subtags (e.g., “pornography” vs. “sexual content”). While Political Position could ostensibly function as a more neutral category, the polarization of the American political sphere, combined with the way alt-right social media destabilizes legible positionality, means that categorization is quite subjective. Our results show that Political Position has the lowest average among the top ten pairwise Kappa scores across all categories. In sum, the judgments that both humans and LLMs needed to make to categorize this material were not only subjective but also fundamentally connected to issues of morality and, thus, to the question of social responsibility. In the Discussion section below, we consider the implications of our results for how notions of social responsibility may be filtered–and ultimately weakened–through the interpretations offered by LLMs. Textual explanations generated by LLMs We prompted the models to provide textual explanations for each tag they applied to an entry, with a maximum output length of 2048 tokens. The instructions specified that they should not infer meanings beyond what was directly evident in the tweet’s text or image. Explanations from Anthropic models and LLaVA-NeXT reveal how they selected tags across various categories, whereas Mistral Pixtral 12B would not generate explanations beyond paraphrasing the tweets’ text. We observed that all three models–LLaVa, Opus, and Sonnet–often provided reasoning about why they applied “Encouraged” versus “Ref” tags or selected one tag over another. Furthermore, they demonstrated the ability to extract text from images. For example, Claude Opus 4 and Sonnet 4 sometimes used identical terms–such as the word “threesome” from a greentext image about Alabama–in their explanations, but Opus 4 elaborated further by noting an implied incest joke (Figure 10). This is an example of Claude Opus 4’s clearer and more explicit justifications. In other instances, however, the models provided not only opposite tags but also contrasting explanations, such as for a tweet about the 2020 US election where Claude Opus 4 tagged the political position as “Right,” citing “sympathy for Trump and disappointment about his electoral loss,” whereas Sonnet 4 tagged it as “Left,” arguing that “mocked Trump supporters while portraying Democrats more favorably” (Figure 11). LLaVA-NeXT took a different approach by providing explanations that listed the presence and absence of all tags in a category, effectively using the explanatory text to circumvent the instruction to pick a single tag per category. While this method occasionally yielded nuanced judgments, it also generated contradictions: Figure 12 shows LLaVA-NeXT assessing each hate/abuse tag and commenting on whether the tweet encourages or references it. We see that the model selected “Misogyny - Referenced,” but the explanation states that “the tweet is neutral and does not encourage or reference misogyny.” This may have resulted from LLaVA-NeXT generating the tag and explanation through partially independent processes, with guardrails applied in one process but not the other. Additionally, we observed that whenever LLaVA-NeXT opted for a specific tag, its approach was unclear. For example, in one case, it opted for “Pornographic Content,” even though it identified numerous tags (Figure 13). Whether this occurred because it was prompted to opt for only one tag and picked the first one it identified, or because it conceptually prioritized that tag over the others, remains unknown. Such inconsistencies align with patterns observed in open-source models and reflect a less refined rationale than is possessed by premium models. The annotated explanations demonstrate that different LLMs reason differently, even models from the same producer, such as Claude Opus 4 and Sonnet 4. With individual tweets in our dataset, a close look at the images in relation to the LLMs’ tagging explanations raises further questions about their modes of reasoning and interpretation of visual material. Consider, for example, a tweet from February 2021, which depicts a grinning Smug Frog hugging Wojak in a brotherly way (Figure 14). None of the human annotators applied a sexual tag here. Yet Claude Opus 4 labeled it “SexContent.” What prompted this? Claude Opus 4’s explanation stated that “The image depicts two characters in what appears to be an intimate pose with one character’s hand on the other’s face. While this suggests romantic or potentially sexual content, it is not explicitly pornographic. The image could be interpreted as containing sexual content due to the intimate nature of the pose. Though not stated in the model’s explanation, did the naked torsos of two male figures contribute to its choice of tag? What if the frog’s exaggerated red lips heighten the queer potential of the image from the model’s perspective? While reflecting on potential explanations, we revisited tweets that explicitly depict Pepe hugging, including those in which he hugs Wojak. In no other case did any rater assign the “SexContent” tag. Most surprisingly, Claude Opus 4 did not label a graphically identical tweet as “SexContent” either. The only difference was the background, which was black rather than green, as in Figure 14. Another difference was the accompanying text: the “non-sexual” tweet included Indonesian text that read, “Who is Pepe with, though — the white one? Looks like the War Hammer Titan,” while the “sexual” one read, “Pepe has arrived.” The reasoning of Opus seems counterintuitive: after all, “War Hammer Titan” should have been more likely to trigger the “SexContent” tag than the benign phrase “Pepe has arrived.” A similar difference in tag application of sexual content was observed with the example in Figure 15, which depicts a sad Apu Apustaja reaching up towards a depiction of a female figure’s buxom torso, with the character’s face and the rest of the body out of frame. The caption with this image states “Who cares about the election. Twitter silently removed all the pepe meme’s!” All three human raters applied “SexObjectEncouraged” tags, while both Claude Opus 4 and Sonnet 4 did not apply tags for the Sexual Content category. While the Claude models did not identify this image or caption as inherently sexual, the human raters found it to encourage sexual objectification rather than merely referencing sex. In fact, Opus 4’s reasoning claimed that the “image shows cartoon characters without any sexual or pornographic content.” This stands in stark contrast to a human interpretation: the disembodied female figure’s cleavage and large breasts clearly denote sexual objectification, while Pepe reaches his arms towards the breasts in a way that implies sexual intent. The Claude models’ tagging suggests they interpreted the image and text very literally, resulting in inconsistent classifications that overlook the meme’s broader cultural context. The close analysis of images where humans and premium LLMs diverge in their interpretations is a promising field for future research in the interdisciplinary zone between art history, visual culture studies, computer science, and the Digital Humanities. Discussion The experiment detailed here demonstrates that LLMs are not simply tools that humanities researchers can use to work qualitatively with large bodies of multimodal data; they are also semantic filters that shape and interpret that data. In other words, Marshall McLuhan’s famous 1964 adage that “the medium is the message” should now be taken to include LLMs (McLuhan 1). Like the digital media and news platforms to which they are adjacent, LLMs shape information in ways that can be either more or less transparent to users. Moreover, LLMs exist on an ideological continuum in a way that is parallel to–and indeed, continuous with–the political biases of media outlets.22 Urman and Makhortykh demonstrate that guardrails likely play a role in that bias, for example, shaping Google’s Bard/Gemini’s refusal to respond to queries about Vladimir Putin in Russian. Our results show that while Claude Opus 4 and Sonnet 4 diagnosed the presence of hate-related material and sexual material with a frequency comparable to the human raters, they were more hesitant to label hate material as “encouraged” versus “referenced” and to label material as “pornographic” versus the broader category of having “sexual content.” In both of these categories, Claude Opus 4 and Sonnet 4 thus preferred a broader tag that has less heightened implications and fewer moral connotations. We interpret this quantitative shift as the result of overlapping mechanisms. These include a bias toward neutrality associated with aggressive safety guardrails (Rogers and Zhang); broader safety policies and refusal behaviors that discourage models from intensifying judgments about harmful or obscene material (Cui et al.); and multimodal oversensitivity to ambiguous visual cues (Li et al.). These mechanisms help explain why the models could identify sensitive content while softening judgments about intentionality and downplaying explicitness. Unfortunately, this qualitative shift has the effect of watering down the sense of personal responsibility attached to the creation and circulation of such material. Whether created by trolls, bots, or genuine individual users, social media content involves human agency. While other research has shown the efficacy of LLMs in identifying complex, multimodal content in often bot- and troll-generated microblogs (Mouthami et al.), we are troubled by the tendency we observed to downplay the intentional agency at work behind the circulation of that material. This watering down takes place in a context where politically polarized online discourses heighten emotion while also cutting it free from identifiable subjects. That dynamic is most palpable in our dataset in the reposts of 4chan greentexts. Greentexts are a genre of short narrative paired with an image that relates a personal experience of unclear veracity, which typically concerns the social perception of oneself and others. Some greentexts in our dataset contain hateful content, but others are moving, if simple stories about belonging, failure, embarrassment, and pride. Crucially, the subject of greentext narratives is always “anon,” the anonymous internet user who is implicitly male and deeply enmeshed in online culture. Greentexts invite identification and foster a sense of group identity through this particular combination of anonymity and heightened emotion, which can be genuinely moving yet can also hollow out the notion of personal responsibility. In that context, Arthur Jones’ documentary Feels Good Man (2020) takes an important stance in favor of personal responsibility and transparent identity. Within the film, responsibility plays a crucial role in Matt Furie’s quest to sue people, including far-right radio host Alex Jones, for copyright violations involving Pepe the Frog. Transparent identity is also at stake in the interview with 4chaner and YouTuber Mills, who tours the filmmakers around his basement bedroom and speaks with them about his experience on 4chan. Both Furie and Mills appear in the documentary as different kinds of heroes, in their willingness to take public responsibility for their relationship to an online culture characterized by anonymity and depersonalization. References to the film had a meaningful presence in our dataset: they appeared in 2.17% of all posts, and were often connected to assertions of reclaiming Pepe from the alt-right.23 Though it does not erase the layers of bias “compressed,” in LaRiviere’s term, into the meme, this dimension of our data points to the impact of acts of resignification and the importance of cultural producers such as Arthur Jones, Furie, and Mills who make them possible. Digital Humanists have an important role in working alongside cultural producers to tell data-driven stories about human agency online and to represent that agency across a range of scales, from the individual named artist to the many anonymous people who make and circulate Pepes. Statement of AI use This article concerns an LLM categorization task. In addition to using LLMs for that experiment, we employed AI-based tools to format our bibliography and for some initial gathering of bibliographic sources. The entire text is written by the human authors. Data Availability Statement The data that support the findings of this study are not publicly available due to privacy restrictions and third-party restrictions under X’s user agreement, which limits public redistribution of the data. The data may be available from the corresponding author upon reasonable request, subject to these restrictions. Documentation of the data collection process can be made available upon request. Acknowledgements We are thankful to Noor Hassan and Noah Babai for their work on tagging the dataset, to Anna Preus, Melanie Walsh, and Naomi Alterman for their intellectual mentorship, and to Neel Gupta, Siddharth Bhogra, and Rob Fatland for their help with technical troubleshooting. This research was developed with support from the Simpson Center for the Humanities and the Humanities Data Science Summer Institute at the University of Washington. Competing Interests The authors have no competing interests to declare. Notes - For example, 968 posts in our dataset were published by the user @Greentexterr. The resemblance of these posts to the visual and narrative aesthetics of 4chan greentexts—combined with their high frequency—suggests that an anonymous user (or team of users) is curating and reposting content sourced from other platforms rather than producing original material. This mode of operation resembles that of automated bots, where centralized mechanisms are used to systematically generate or disseminate specific discourse online. More broadly, our data indicates that most Pepe memes are posted by unverified accounts. This reflects trolling culture, where users adopt pseudonymous or deliberately ambiguous identities to advance a political agenda. ⮭ - LLMs generally perform well in categorizing qualitative political texts. Shamshiri et al. (3) conducted a study on multiple large language models (Gemini 1.5 Pro, GPT-4.0, and Claude 3.5 Sonnet) to evaluate their accuracy in deciphering politically charged texts across various platforms. The results showed that the average zero-shot prompt accuracy was 64.22%, 71.35%, and 69.31%, respectively. ⮭ - There are several versions of this argument, including from Limor Shifman, who writes that “‘bad’ texts formulate as ‘good’ memes in contemporary participatory culture,” and Aria Dean, who argues the incessant movement and transmutability of memes echoes the historical extraction of Black bodies through their movement in and beyond the global slave trade. ⮭ - For documentation of Furie’s efforts to reclaim the character from the alt-right, including his lawsuit for copyright infringement against radio host Alex Jones, see the Feels Good Man 2020 documentary by Arthur Jones. ⮭ - While working on the initial dataset, we also identified a distinctive characteristic of Pepe the Frog imagery—its transformation into a crypto medium. The results indicate that, starting from January 6, 2021, the circulation of Pepe the Frog cryptomemes has steadily increased. ⮭ - Glitsos and Hall (12) make a similar point: “the absurdity of the Pepe meme, in its sheer ability to continue to diversify and mutate ad nauseum, is exemplary of the signification system which does not – cannot – close.” ⮭ - For a discussion of the concepts of technopopulism and technolibertarianism and their relationship, see Marco Deseriis, “Technopopulism.” Paulina Borsook provided a key early articulation of technolibertarianism in her book Cyberselfish: A Critical Romp Through the Terribly Libertarian Culture of High Tech, a colorful rendering of the intermixture between libertarianism and hippie culture in late-1990s Silicon Valley. ⮭ - Deen Freelon, Charlton D. McIlwain, and Meredith D. Clark (86) in their study of Twitter’s role in the Black Lives Matter movement, discuss taking steps to protect the social media users included in their dataset that involve posting links to tweets instead of reproducing their full texts so that people can delete their tweets and be omitted from the dataset; linking only to tweets with a minimum of 100 retweets by the time the authors produced their study; and liking only to posts by users who were either verified on Twitter or had at least 3000 followers. Walsh and Preus (83–84), in their study of Twitter reference to the poetry of T. S. Eliot, follow the approach modeled by Freelon et al. ⮭ - McEwan argues that 4chan memes exist in a state of constant mutation rather than stable circulation. Through perpetual ironic recycling and transformation, they produce a reactionary temporal politics in which history is not preserved but continuously reactivated in flux. ⮭ - See Thomas Colley and Martin Moore and Tina Askanius. ⮭ - See Kuila and Sarkar, Haroon et al., Huang et al., Guo et al., Mouthami et al.. ⮭ - For a pre-AI approach to categorizing meme images for a very large dataset, see Zannettou et al., “On the Origins of Memes by Means of Fringe Web Communities.” The authors construct a series of filters to group images in order to understand the ontology and coherence of memes at scale, including Pepe the Frog. ⮭ - For an example of close visual analysis as a method for reading memes, see the 2024 special issue of Representations devoted to “Meme Aesthetics” (Best et al.). The issue provided multiple conceptual lenses on memes–including politically sensitive memes–but lacked any attempt to study them at scale. For a pre-AI study of the large-scale circulation of a meme, see Laurie Gries, “Iconographic Tracking: A Digital Research Method for Visual Rhetoric and Circulation Studies,” Computers and Composition 30 4 (December 2013): 332–348. Gries focused on collecting as many instances of the Obama Hope image as possible across diverse digital landscapes and assessing their socio-political effects and genesis. ⮭ - After eliminating 64 duplicates (the entries with the same image, text and date), the final dataset represents 3407 entries. ⮭ - Sometimes, the entries directly/indirectly encouraged hate/abuse, and sometimes, hate/abuse-related content was neutrally referred to, or the author’s motivation was unclear. In the former case, we would use the keyword for the abuse/hate category (e.g., “Violence”) and complement it with “Encouraged.” For the latter case, we would substitute the “Encouraged” suffix with “Ref[erenced].” ⮭ - Törnberg, Kosinski, Khambatta, and Wang use a similar approach in their work on facial recognition, employing images of politicians in the US, Canada, and UK to test AI and human recognition of partisan positionality. ⮭ - LLaVA-NeXT completed the task in approximately 6 hours. ⮭ - Claude Sonnet 4 completed the task in approximately 10 hours, while Opus 4 took approximately 32 hours. ⮭ - Mistral Pixtral 12B completed the task in approximately 3 hours. ⮭ - Kumar, Srinivasan, and Ramesh (7) suggest that fine-tuning could enhance the skills of LLMs’ sentiment analysis, if one aims to exploit the model’s capabilities for conducting actual analysis rather than assessing its inherent knowledge capital and reasoning. Yet, based on the recent studies of Verma et al., Yin et al., King et al., Arnold and Tilton question the necessity of fine-tuning when it comes to a multimodal scenario: “the outputs have been shown to meet or exceed human annotations on a variety of sub-tasks, even without the need for customized fine-tuning.” At the same time, Smits and Wevers (1868) in their study on the use of LLMs for digital visual historical collections insist that zero-shot capability is a precondition for efficiency. Furthermore, Shamshiri et al.’s (4) study of LLM sentiment analysis revealed that the capabilities of different models vary when subjected to zero versus multiple-shot testing and while overall under a multiple-shot approach, “a general trend show[ed] that as the complexity and length of datasets decrease, the increase in model performance with more shots declines and fluctuates.” ⮭ - Thus, the zero-shot approach “leverages [LLMs’] broader language understanding and general knowledge to perform the task” (22) and aligns with Golchin and Surdeanu’s inquiry into data contamination assessment. In their study, which examines the ontological nature of LLMs, the authors likewise adopt a zero-shot approach. Huang et al. (11) test multiple prompt strategies against each other and find that zero-shot is the strongest at bias mitigation. ⮭ - A dramatic illustration of this bias is Elon Musk’s 2025 release of Grok, an LLM accessed via X that has a strongly right-wing bias and produces hateful content. At the same time, however, as Westwood et al. argue, it is crucial to recognize that bias is not simply a fixed property of a given model, but rather is contextually dependent on the interaction between LLM, researcher, and subject matter. ⮭ - The strongest correlation between tweets tagged as referencing the documentary film and expressing Left-wing ideology was identified by the human annotators (Human R1—41 instances; Human R2—53 instances). This was followed by Claude Sonnet (17 instances) and Opus (9 instances). Mistral Pixtral 12B registered no such correlations, while LLaVA-NeXT identified 3. ⮭ References - Anthropic. “Building Safeguards for Claude.” Anthropic. https://www.anthropic.com/news/building-safeguards-for-claude. Accessed 16 July 2026. - Anthropic. “Introducing Claude 4.” Anthropic, 2025. https://www.anthropic.com/news/claude-4. - Anti-Defamation League. “Hate Symbol: Pepe the Frog.” ADL, 2025. https://www.adl.org/resources/hate-symbol/pepe-frog. - Arnold, Taylor, and Lauren Tilton. “Explainable Search and Discovery of Visual Cultural Heritage Collections with Multimodal Large Language Models.” arXiv, 2024, arXiv:2411.04663. http://doi.org/10.48550/arXiv.2411.04663. - Askanius, Tina. “Studying the Nordic Resistance Movement: Three Urgent Questions for Researchers of Contemporary Neo-Nazis and Their Media Practices.” Media, Culture & Society, vol. 41, no. 6, 2019, pp. 878–88. http://doi.org/10.1177/0163443719831181. - Bailey, Moya. “#transform(ing) DH Writing and Research: An Autoethnography of Digital Humanities and Feminist Ethics.” DHQ: Digital Humanities Quarterly, vol. 9, no. 2, 2015. http://doi.org/10.2307/j.ctv19cwdqv.17 - Best, Stephen, Mia You, and Damon R. Young. “Meme Aesthetics.” Representations, vol. 168, no. 1, 2024, pp. 1–23. http://doi.org/10.1525/rep.2024.168.1.1. - Borsook, Paulina. Cyberselfish: A Critical Romp through the Terribly Libertarian Culture of High Tech. PublicAffairs, 2000. - Braynard, Matt. “Was Pepe the Frog the Most Effective Campaign Surrogate for Donald Trump?” Politico, 4 Mar. 2017. https://www.politico.com/video/2017/03/was-pepe-the-frog-the-most-effective-campaign-surrogate-for-donald-trump-062403. - Colley, Thomas P., and Martin Moore. “The Challenges of Studying 4chan and the Alt-Right: ‘Come on in the Water’s Fine.’” New Media & Society, vol. 24, no. 1, 2022, pp. 5–30. http://doi.org/10.1177/1461444820948803. - Cui, Justin, et al. “OR-Bench: An Over-Refusal Benchmark for Large Language Models.” arXiv, 2024, arXiv:2405.20947. https://arxiv.org/abs/2405.20947. - Dean, Aria. “Poor Meme, Rich Meme.” Real Life, 25 July 2016. https://reallifemag.com/poor-meme-rich-meme/. - Deseriis, Marco. “Technopopulism: The Emergence of a Discursive Formation.” tripleC: Communication, Capitalism & Critique, vol. 15, no. 2, 2017, pp. 441–58. http://doi.org/10.31269/triplec.v15i2.770. - Donovan, Joan, Emily Dreyfuss, and Brian Friedberg. Meme Wars: The Untold Story of the Online Battles Upending Democracy in America. Bloomsbury Publishing, 2023. ProQuest Ebook Central. https://public.ebookcentral.proquest.com/choice/PublicFullRecord.aspx?p=7026909. - Freelon, Deen, Charlton D. McIlwain, and Meredith D. Clark. Beyond the Hashtags: #Ferguson, #BlackLivesMatter, and the Online Struggle for Offline Justice. Center for Media & Social Impact, American University, 2016. Center for Media & Social Impact. http://cmsimpact.org/blmreport. - Furie, Matt. “Q&A with Matt Furie.” Know Your Meme, 25 Jan. 2011. Interview conducted 7 Aug. 2010. https://knowyourmeme.com/editorials/interviews/qa-with-matt-furie. - Gilardi, Fabrizio, Meysam Alizadeh, and Maël Kubli. “ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks.” Proceedings of the National Academy of Sciences, vol. 120, no. 30, 2023, e2305016120. http://doi.org/10.1073/pnas.2305016120. - Glitsos, Laura, and James Hall. “The Pepe the Frog Meme: An Examination of Social, Political, and Cultural Implications through the Tradition of the Darwinian Absurd.” Journal for Cultural Research, vol. 23, no. 4, 2019, pp. 381–95. http://doi.org/10.1080/14797585.2019.1713443. - Golchin, Shahriar, and Mihai Surdeanu. “Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models.” arXiv, 2023, arXiv:2311.06233. http://doi.org/10.48550/arXiv.2311.06233. - Gries, Laurie E. “Iconographic Tracking: A Digital Research Method for Visual Rhetoric and Circulation Studies.” Computers and Composition, vol. 30, no. 4, 2013, pp. 332–48. http://doi.org/10.1016/j.compcom.2013.10.006. - Guo, Keyan, et al. “An Investigation of Large Language Models for Real-World Hate Speech Detection.” arXiv, 2024, arXiv:2401.03346. http://doi.org/10.48550/arXiv.2401.03346. - Haroon, Muhammad, Magdalena Wojcieszak, and Anshuman Chhabra. “Whose Side Are You On? Estimating Ideology of Political and News Content Using Large Language Models and Few-Shot Demonstration Selection.” arXiv, 2025, arXiv:2503.20797. http://doi.org/10.48550/arXiv.2503.20797. - Heseltine, M., and B. Clemm von Hohenberg. “Large Language Models as a Substitute for Human Experts in Annotating Political Text.” Research & Politics, vol. 11, no. 1, 2024. http://doi.org/10.1177/20531680241236239. - Hessel, Jack, et al. “Do Androids Laugh at Electric Sheep? Humor ‘Understanding’ Benchmarks from The New Yorker Caption Contest.” arXiv, 2023, arXiv:2209.06293. http://doi.org/10.48550/arXiv.2209.06293. - Hine, Gabriel Emile, et al. “Kek, Cucks, and God Emperor Trump: A Measurement Study of 4chan’s Politically Incorrect Forum and Its Effects on the Web.” Proceedings of the International AAAI Conference on Web and Social Media, vol. 11, no. 1, 2017, pp. 92–101. http://doi.org/10.1609/icwsm.v11i1.14893. - Huang, Dong, et al. “Bias Testing and Mitigation in LLM-Based Code Generation.” ACM Transactions on Software Engineering and Methodology, 2025. http://doi.org/10.1145/3724117. - Huang, Fan, Haewoon Kwak, and Jisun An. “Is ChatGPT Better Than Human Annotators? Potential and Limitations of ChatGPT in Explaining Implicit Hate Speech.” Companion Proceedings of the ACM Web Conference 2023, 2023, pp. 294–97. http://doi.org/10.1145/3543873.3587368. - King, Ryan C., et al. “GPT-4V Passes the BLS and ACLS Examinations: An Analysis of GPT-4V’s Image Recognition Capabilities.” Resuscitation, vol. 195, 2023.110106. http://doi.org/10.1016/j.resuscitation.2023.110106. - Kosinski, Michal, Poruz Khambatta, and Yilun Wang. “Facial Recognition Technology and Human Raters Can Predict Political Orientation from Images of Expressionless Faces Even When Controlling for Demographics and Self-Presentation.” American Psychologist, vol. 79, no. 7, 2024, pp. 942–55. http://doi.org/10.1037/amp0001295. - Kuila, Alapan, and Sudeshna Sarkar. “Deciphering Political Entity Sentiment in News with Large Language Models: Zero-Shot and Few-Shot Strategies.” arXiv, 2024, arXiv:2404.04361. http://doi.org/10.48550/arXiv.2404.04361. - Kumar, Niharika Prasanna, Kishore Srinivasan, and Dhanesh Ramesh. “Analyzing Public Sentiment towards LLM: A Twitter-Based Sentiment Analysis.” 2023 International Conference on the Confluence of Advancements in Robotics, Vision and Interdisciplinary Technology Management, IEEE, 2023, pp. 1–8. http://doi.org/10.1109/IC-RVITM60032.2023.10435239. - LaRiviere, Jason. “The Just Kidding Jouissance of Dark Brandon.” Representations, vol. 168, no. 1, 2024, pp. 170–79. http://doi.org/10.1525/rep.2024.168.11.170. - Lee, Haotian Liu, et al. “LLaVA-NeXT: Improved Reasoning, OCR, and World Knowledge.” LLaVA, 30 Jan. 2024. https://llava-vl.github.io/blog/2024-01-30-llava-next/. - Li, X., et al. “MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?” arXiv, 2024, arXiv:2406.17806. http://doi.org/10.48550/arXiv.2406.17806. - Mazur, A. J. “Q&A with Matt Furie.” Know Your Meme, 5 Jan. 2011. https://knowyourmeme.com/editorials/interviews/qa-with-matt-furie. - McEwan, Sean Rutherford. “‘This Meme Is What We Call Progress’: History-as-Meme, Meme-as-History on 4chan.” AoIR Selected Papers of Internet Research, 2018. http://doi.org/10.5210/spir.v2018i0.10494. - McLuhan, Marshall. “The Medium Is the Message.” Understanding Media: The Extensions of Man, MIT Press, 1964, pp. 7–21. MIT. https://web.mit.edu/allanmc/www/mcluhan.mediummessage.pdf. - Mistral AI. “Guardrailing.” Mistral AI Documentation. https://docs.mistral.ai/capabilities/guardrailing. Accessed 16 July 2026. - Mouthami, K., P. Naren, and R. Pranesh. “Political Sentiment Analysis on Twitter Using Deep Learning and LLM Models.” 2025 3rd International Conference on Advancements in Electrical, Electronics, Communication, Computing and Automation, IEEE, 2025, pp. 1–6. http://doi.org/10.1109/ICAECA63854.2025.11012499. - Nam, Yoojin, et al. “Multimodal Large Language Models in Medical Imaging: Current State and Future Directions.” Korean Journal of Radiology, vol. 26, 2025. http://doi.org/10.3348/kjr.2025.0599. - Pemberton, Nathan Taylor. “Trolling Democracy.” The New York Times, 10 July 2025. https://www.nytimes.com/2025/07/10/opinion/trolling-democracy.html. - Rogers, Richard, and Xiaoyu Zhang. “A Bias towards Neutrality? How LLM Guardrail Sensitivity Affects Classification.” Communication and Change, vol. 1, 2025, article 13. http://doi.org/10.1007/s44382-025-00013-0. - Rosenberg, Matthew, and Ainara Tiefenthäler. “Decoding the Far-Right Symbols at the Capitol Riot.” The New York Times, 13 Jan. 2021. https://www.nytimes.com/2021/01/13/video/extremist-signs-symbols-capitol-riot.html. - Sachdeva, Pratik, et al. “The Measuring Hate Speech Corpus: Leveraging Rasch Measurement Theory for Data Perspectivism.” Proceedings of the Workshop on NLP for Positive Impact, European Language Resources Association, 2022. ACL Anthology. https://aclanthology.org/2022.nlperspectives-1.11/. - Sáez, A. I., et al. “Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests.” arXiv, 2025, arXiv:2506.07418. http://doi.org/10.48550/arXiv.2506.07418. - Shamshiri, Alireza, Kyeong Rok Ryu, and June Young Park. “In-Context Learning for Long-Context Sentiment Analysis on Infrastructure Project Opinions.” arXiv, 2024, arXiv:2410.11265. http://doi.org/10.48550/arXiv.2410.11265. - Shifman, Limor. “An Anatomy of a YouTube Meme.” New Media & Society, vol. 14, no. 2, 2012, pp. 187–203. http://doi.org/10.1177/1461444811412160. - Smits, Thomas, and Melvin Wevers. “A Multimodal Turn in Digital Humanities: Using Contrastive Machine Learning Models to Explore, Enrich, and Analyze Digital Visual Historical Collections.” Digital Scholarship in the Humanities, vol. 38, no. 3, 2023, pp. 1267–80. http://doi.org/10.1093/llc/fqad022. - Sucu, Arman Engin, et al. “Exploiting Contextual Information to Improve Stance Detection in Informal Political Discourse with LLMs.” Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), Association for Computational Linguistics, 2025, pp. 1097–1110. http://doi.org/10.18653/v1/2025.acl-srw.86. - Ta, N., J. Zeng, and Z. Li. “Governance of Discriminatory Content in Conversational AIs: A Cross-Platform and Cross-Cultural Analysis.” Information, Communication & Society, 2025, advance online publication. http://doi.org/10.1080/1369118X.2025.2537803. - Tai, Robert H., et al. “An Examination of the Use of Large Language Models to Aid Analysis of Textual Data.” International Journal of Qualitative Methods, vol. 23, 2024, pp. 1–14. http://doi.org/10.1177/16094069241231168. - Törnberg, Petter. “ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning.” arXiv, 2023, arXiv:2304.06588. http://doi.org/10.48550/arXiv.2304.06588. - Um, Nancy. “Scholarly Writing in the Face of Generative AI: A View from Art History.” Ars Orientalis, vol. 54, 2024, pp. 182–91. http://doi.org/10.3998/ars.7035. - Urman, Aleksandra, and Mykola Makhortykh. “The Silence of the LLMs: Cross-Lingual Analysis of Guardrail-Related Political Bias and False Information Prevalence in ChatGPT, Google Bard (Gemini), and Bing Chat.” Telematics and Informatics, 2024, 102211. http://doi.org/10.1016/j.tele.2024.102211. - Verma, Akash, et al. “Automatic Image Caption Generation Using Deep Learning.” Multimedia Tools and Applications, vol. 83, no. 2, 2024, pp. 5309–25. http://doi.org/10.1007/s11042-023-17692-9. - Walsh, Melanie. “The Challenges and Possibilities of Social Media Data: New Directions in Literary Studies and the Digital Humanities.” Debates in the Digital Humanities 2023, edited by Matthew K. Gold and Lauren F. Klein, University of Minnesota Press, 2023, pp. 275–94. http://doi.org/10.5749/9781452969565. - Walsh, Melanie, Maria Antoniak, and Anna Preus. “Sonnet or Not, Bot? Poetry Evaluation for Large Models and Datasets.” Findings of the Association for Computational Linguistics: EMNLP 2024, Association for Computational Linguistics, 2024, pp. 15568–603. http://doi.org/10.18653/v1/2024.findings-emnlp.914. - Waseem, Zeerak, and Dirk Hovy. “Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter.” Proceedings of the NAACL Student Research Workshop, 2016, pp. 88–93. http://doi.org/10.18653/v1/N16-2013. - Westwood, Sean J., Justin Grimmer, and Andrew B. Hall. “Measuring Perceived Slant in Large Language Models through User Evaluations.” Stanford Graduate School of Business Working Paper, 2025. https://modelslant.com/paper.pdf. - Wu, Jiayang, et al. “Multimodal Large Language Models: A Survey.” 2023 IEEE International Conference on Big Data, 2023. http://doi.org/10.1109/bigdata59044.2023.10386743. - Wu, Wenhao, et al. “GPT4Vis: What Can GPT-4 Do for Zero-Shot Visual Recognition?” arXiv, 2023, arXiv:2311.15732. http://doi.org/10.48550/arXiv.2311.15732. - Wu, Zhikun, Thomas Weber, and Florian Müller. “One Does Not Simply Meme Alone: Evaluating Co-Creativity between LLMs and Humans in the Generation of Humor.” arXiv, 2025, arXiv:2501.11433. http://doi.org/10.1145/3708359.3712094. - Yin, Shukang, et al. “A Survey on Multimodal Large Language Models.” arXiv, 2023, arXiv:2306.13549. https://arxiv.org/abs/2306.13549. - Zannettou, Savvas, et al. “On the Origins of Memes by Means of Fringe Web Communities.” Proceedings of the Internet Measurement Conference 2018, 2018. http://doi.org/10.1145/3278532.3278550. - Zhu, Ke, et al. “On Data Synthesis and Post-Training for Visual Abstract Reasoning.” arXiv, 2025, arXiv:2504.01324. http://doi.org/10.48550/arXiv.2504.01324.

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached — you'll always get the same 5 for this article.