How Far Can AI Overlays Really Go?
How Far Can AI Overlays Really Go?
If you're not careful about the technology you bring into the firm today, you may find yourself living with the consequences for a very long time.
The swivel chair has been defeated for advisors, or so Anthropic says with the release of Claude for Financial Advisors. In some ways, they’re right: the company has built an amazing overlay to the existing stack. An advisor can connect their data and have Claude handle back-office busywork, which, for some firms, can occupy up to 80% of their time.
That’s a big win for anyone, but for everyone with a fiduciary duty to clients, it’s crucial to understand the juncture we now find ourselves at. Understanding this moment will determine which firms thrive and which ones get obliterated by regulatory crackdowns and client lawsuits.
Why “Almost Always Right” Isn’t Good Enough
Artificial intelligence models, at least the type everyone has been using for the last few years, i.e., large language models, are probabilistic systems. They are undeniably the most sophisticated probabilistic systems ever developed. That means they’re really good at guessing the answer to questions, and this generally holds true across many domains. But as it turns out, their educated guesses are based on patterns in their training data. For some questions, LLMs achieve very high accuracy because the model weights are decisive. But that’s generally not true for financial use cases.
Want to run a Monte Carlo analysis? Calculate maximum drawdown? Complete portfolio optimization or calculate rolling correlations? Better not use an LLM. Even the smartest LLMs have great difficulty addressing these questions. They’re statistical text networks, or auto-correct on steroids. They’re not built to perform mathematical calculations, and handing them a pile of skill files doesn’t truly address the gap. Hallucination rates have been shown in multiple studies to be in the mid-to-high 80s for these types of questions.
Imagine if someone approached you with a proposition: they would save you 80% of your working hours by doing it for you at low cost, but one time out of every hundred, the recommendations they produced for your clients would be wrong. And sometimes incredibly wrong. As a fiduciary, you would immediately balk at that offer. Now imagine that instead of being wrong one time in a hundred, it was wrong as many as 88 times in a hundred. That would be completely untenable, and the only reason anyone is pretending it’s not is the incredible amount of hype in the AI space right now.
It’s worth mentioning that the two most prominent AI labs are gunning for multi-trillion-dollar IPOs in the coming months, and arguably, according to Michael Burry, have been using the threat of civilizational collapse as a marketing tactic. They’re not profitable today, and open weight models are very close to parity. If you thought banks were too big to fail in 2008, imagine one of the biggest U.S. AI labs collapsing overnight. They’re so deeply entrenched everywhere right now that we essentially have to keep the hype train going because their spend is massive, while the average American is highly unlikely to want to foot that bill, particularly when the technology is supposed to be replacing them.
The Limits of the Overlay for an Advisor
On top of the advisor’s concerns about hallucinations, there are also other risks. Historically, new tech layered onto old systems inherits all the limitations and flaws of the foundation it’s built on. This constrains what AI can accomplish and raises questions about how much value an overlay can actually add.
Additionally, advisors don’t want to become prompt engineers. Open-ended prompting creates a “blank page” problem for professionals who want technology to simplify their work, not require them to develop a new technical skill.
Perhaps most importantly, hallucination rates limit what an AI overlay can do. Back office productivity is great, but being a one-stop shop is a much higher bar. If advisors can’t trust the numbers and outputs they see, there’s a ceiling on how deeply the technology can be used within core wealth management workflows.
Where AI Actually Fits in the Tech Stack
Maybe you’re wondering how fiduciaries should use AI. Surely there is a way. Luckily, the answer is yes, but maybe not exactly how you might expect.
It turns out, no matter how smart a probabilistic system gets, even if it can write top-tier quant-level code on the fly, it’s always going to hallucinate. Probabilistic calculations are a feature, not a bug. This means fiduciaries cannot use AI for investment workflows, and that’s okay. Most of the benefit of LLMs is in orchestration, interpreting intent, qualitative summarization and adjacent areas. To keep agents honest, they need a rigorous use-case and process ontology, deterministic tools, a consistent governance framework and contextual skills to bring it all together.
There is no panacea that can safely handle everything for an advisor and their clients. But there are amazing tools that cover large swathes of fiduciary workflows. If you’re modernizing your stack in the agentic era, the greatest existential risk is failing to understand the technology and whether it can provide answers a fiduciary can trust to inform decisions.
For the first time in the history of technology, a demo alone cannot address whether a tool can do what an advisor needs. Hallucinations appear plausible and are delivered with conviction. The advisor’s role in the agentic era is vetting the right stack against these AI realities. It’s a consequential decision, given how deeply these systems could become embedded in a firm’s operations. And if you’re not careful about the technology you bring into the firm today, you may find yourself living with the consequences for a very long time.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.