What Security Leaders Think About Frontier AI Models: Firsthand of Mythos
TL;DR
Frontier AI models are raising the ceiling for skilled attackers and lowering the bar for everyone else, and the defenders closest to this work say the advantage window is wider than most expect. The model is not the differentiator; the harness and the expertise behind it are. Whether you set security strategy or live in the SDLC every day, the calculus around risk, tooling, and resilience is changing faster than most programs are built to handle.
Frontier AI models like Mythos are no longer a research curiosity. They are writing code, finding vulnerabilities, and reshaping how security work gets done, and the people closest to that work are starting to share what it actually looks like in practice.
Adrian Peters (Managing Director & CISO, Vista Equity Partners) recently sat down with Vinnie Liu (CEO & Co-founder, Bishop Fox) and Jason Lish (SVP & CISO, Cisco) to discuss what frontier AI models like Mythos actually mean for security programs, budgets, and the people responsible for defending them. Their collective read: the model alone is not the differentiator; the harness built around it and the expertise driving it determine whether an organization finds real signal or not. Defenders who over index on any single control category will find gaps where they least expect them.
Here is their conversation:
Is This an Incremental Step or a Genuine Inflection Point?
JL: What I think these models demonstrate is that they can autonomously perform complex security work that previously required highly specialized expertise and a significant amount of time. The main difference is the speed and scale at which these can operate, as well as the ability to combine and chain vulnerabilities together and then write proof of concept for exploitation. It didn't really make attackers omnipotent, and it didn't make existing controls obsolete. But it did make time-intensive, expertise-dependent work dramatically faster. The classes of attack have not changed; the throughput has.
VL: It is a point of no return because now these capabilities exist and you have to reimagine your architecture, your program, your defensive methodologies. It's not like we're coming up with entirely new classes of attack. It's about existing attacks done faster, done in parallel. And for the experts that know how to build the harnesses and take advantage of it, it raises the ceiling of capability dramatically. On the flip side, it lowers the bar. People that previously wouldn't have been able to perpetrate some of these attacks now have some modest capability at their disposal, but it's undirected in a lot of ways. From an offensive perspective, I'm thinking about it through two lenses: a few skilled individuals who are now empowered with nation-state velocity, and then somebody who's really not skilled in this at all but now has some capability they may or may not be able to take advantage of.
What Are You Seeing These Models Do Well, and Where Do They Fall Short?
JL: They can explore a code base or binary for hours. We see them spin up agents and subagents to go and pull documentation to learn about a product in a matter of minutes. They can reverse engineer binaries, go look for secrets, generate hypotheses, test them, discard failures, and then continue without getting tired or losing interest. I see them increasingly effective at connecting weaknesses that would individually appear less significant. AI can identify how two or three issues combine into a critical attack path. That matters because most enterprises still prioritize vulnerabilities individually.
They still produce false positives, and depending on how the harness is structured, they still make incorrect assumptions, exaggerate impact, and sometimes claim success when they haven't actually succeeded. They also lack organizational context. A model may identify technically elegant vulnerabilities without understanding whether the asset is actually exposed, whether the affected code is running, whether another control blocks exploitation, or whether applying a fix could interrupt a critical business process.
VL: The entire calculus and math that we perform around risk is going to have to change because most enterprise organizations ignore low-risk issues. At Bishop Fox, we love low-risk issues because those are the things that we always use as stepping stones to get the more complex attack chains. You combine some of those together and they result in unauthorized access or data that we can pull down. And something we see a lot in dynamic testing: the model tends to stop once it finds a low-risk issue, treating the task as complete when the critical finding is often one or two steps further down the same path.
How Does the Existing Security Stack Change, and if Budget Forces a Tradeoff, What Goes First?
JL: I wouldn't use the most expensive frontier model to rediscover every issue that a traditional scanner can find cheaply. Traditional tools provide broad continuous coverage, smaller models help with triage and routine remediation, and frontier models should really be focused on high-value code, complex logic, exploitability, attack path chaining, and difficult-to-assess systems. If you're forced to cut something to cover token costs, I wouldn't begin by cutting SAST, DAST, or traditional pen testing. I'd first eliminate duplicate scanning, low-fidelity tools, and the large amount of manual labor we spend moving findings between systems or reviewing issues that have no realistic path to exploitation. The days of periodic pen testing, getting a PDF report, loading that into a risk register, assigning a ticket… those are the things we really need to focus on eliminating. And the right unit of measurement here is not cost per token. It's cost per independently validated and remediated risk.
VL: It's just not the model, it's how you're able to instrument it and use it. Even newer open-weight models are performing on par with frontier models in some cases, exceeding performance in certain benchmarks. Organizations that treat Mythos access as the goal and build nothing around it will get less out of it than teams running lighter models through a well-designed harness with real offensive expertise behind it.
How Long Do You Think Defenders Will Be at a Disadvantage?
VL: I used to think attackers have an advantage for the next 12 months or so. But I actually think it's +24 months that attackers will have the advantage. The ability for LLMs to be within environments to read your IT support wiki or read your architectural diagrams that are there to help your users…they can process all that in minutes, and then immediately know where to go and run all the routes in parallel. I think we're looking towards a future of a smash-and-grab style of attack, where it doesn't matter how fast your detection and response is because they're in and out so quickly that you have to rethink your defensive philosophy and architecture. And people are already building implants that don't even have a payload yet; they just call back to an inference engine and generate a different payload every time. Truly polymorphic. Before Mythos, the most commonly exploited vectors were credential theft, identity attacks, and phishing, not code-level vulnerabilities. Those have not gone away and are accelerating alongside everything else.
JL: Patching alone won't scale. Organizations are asking: can we temporarily shield the vulnerable system, can we segment it, restrict its communications, rehydrate from a clean image? The whole concept of resilience over the next couple of years is going to be critically important because in some cases vendors won't even have patches available when vulnerabilities get identified. The reaction simply isn't 'we need another AI security tool.' The larger realization is operational resilience and having those different defenses in place.
Watch the Full Session
See how Bishop Fox puts frontier models to work in your environment through our AI-Powered Application Penetration Testing service.
Watch the full recording to hear Vinnie Liu, Jason Lish, and Adrian Peters cover talent strategy, the harness build-vs-buy question, how PE firms are thinking about portfolio-wide adoption, and what "Mythos-ready" actually means as an operational posture.
Watch: Beyond the Hype: What Mythos Actually Means for Security Teams
Subscribe to our blog
Be first to learn about latest tools, advisories, and findings.
Thank You! You have been subscribed.
Recommended Posts
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.