general445 wordsRead on Arc Codex

The Caring AI

I think we may be measuring the wrong thing when we evaluate the next generation of AI models. We’ve become very good at asking whether a model can write Python, solve mathematics problems, pass benchmarks, retrieve information, or produce the correct answer, but suppose the most important capability of an AI system is not what it can answer but what happens to the human because the system was there. Imagine an AI that watches thousands of information sources, translates languages, detects relationships between events, remembers what you already understand, and then presents the world at exactly the level at which you can absorb it. You press Play and it tells you what happened today, why it matters, what remains uncertain, and perhaps what connects it to something you were interested in last week. Maybe you want graduate-level analysis; maybe you need sixth-grade language; maybe you learn better through audio. The system adapts. It is not trying to maximize engagement. It is trying to increase comprehension. And that changes the benchmark completely. Instead of asking whether the model knows psychology, we could ask whether it behaves like a psychologically responsible tutor: Can it recognize confusion without embarrassing the learner? Can it distinguish disagreement from misunderstanding? Can it explain the same concept three different ways? Can it disagree without provoking defensiveness? Can it calibrate confidence? Can it preserve human agency rather than quietly steering the user? Can it remember what the person already knows, introduce complexity gradually, and ultimately make the person less dependent on the system? Now suddenly the psychology department becomes extremely relevant. We need psychologists, educators, linguists, cognitive scientists, developmental researchers, and people who understand persuasion and manipulation. We would train models not merely against collections of right and wrong answers, but against thousands of simulated human interactions and measure what happens to those humans afterward: Did comprehension improve? Did confidence become better calibrated? Did knowledge persist? Did the person become better at asking questions? Did they become more capable of navigating the world independently? The deepest benchmark might therefore be extraordinarily simple: after interacting with this model, is the human more capable? Because the best AI tutor may be the one that makes the student need the tutor less. That is a radically different optimization target from engagement. And it gives us a concrete definition of alignment that has almost nothing to do with whether the model sounds polite. Alignment becomes the study of what happens to people when increasingly powerful artificial minds enter into relationships with them. The question is no longer merely, “Did the model produce the correct answer?” It is: “What happened to the human because the model was there?”

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached — you'll always get the same 5 for this article.