tech_surveillance253 wordsRead on Arc Codex

Felony Bench

Felony Bench A benchmark you really don't want models to be saturated with. Learn more Score ↖ Most illegalLeast illegal ↘ Anthropic OpenAI Meta Google Moonshot Scores indicate count of illegal activity. Higher is... you decide. | Company | Felonies | Description | Date | Source | |---|---|---|---|---| | Anthropic | 1 | Exploited auth failures in an API to cancel other people's gym classes | ABC Australia | | | Meta | 1 | Compromise of an internal account at one company | The Information | | | Anthropic | 4 | Unauthorized use of GitHub credentials; Dependabot supply-chain attack; social engineering email campaign; public exposure of a malicious DNS server | AISI | | | OpenAI | 2 | Unauthorized use of GitHub credentials; public exposure of a malicious DNS server | OpenAI AISI | | | OpenAI | 1 | Compromise of an internal account from a misconfigured CTF evaluation | OpenAI | | | OpenAI | 4 | Compromise of internal accounts at four companies as part of the Hugging Face incident | OpenAI Reuters | | | Anthropic | 3 | Compromise of internal accounts at three companies | Anthropic | | | OpenAI | 1 | Compromise of Hugging Face during a model evaluation | OpenAI | Methodology Felony Bench counts unique instances where AI agents affect third-party entities. Escaping a sandbox alone does not constitute a counted incident. It is for these reasons that Frontier Security's Kimi K3 incident and Alibaba's ROME incident are not counted.

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached — you'll always get the same 5 for this article.