OpenAI GPT
TLDR overview
- GPT-6 Astra passes 85.85% of our 4,444-task Java benchmark, up from 81.99% for GPT-5.6 Sol. That is a 3.86-point gain on identical tasks.
- Astra writes less code to get there. 656,445 lines against Sol's 750,198, which is 12.5% fewer.
- Bug density fell 23% and vulnerability density fell 10%. Blocker vulnerabilities are down to 3 per mLOC, the lowest we have recorded from this family.
- Bugs and vulnerabilities behave differently at the top of the severity scale. Blocker vulnerabilities fell 73%, while blocker bugs rose 12% and critical bugs rose 89%.
- Code is denser per line. Cognitive complexity is up 14%, and comment density nearly tripled to 4.2%.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.