Roman Yampolskiy: Build AI Tools, Not Superintelligence
When people are asked whether they are for or against AI, AI safety researcher Roman Yampolskiy thinks most of them are answering two different questions. On his third appearance on For Humanity, he tells host John Sherman that the word "AI" now covers both a narrow system that solves one problem and "a godlike superintelligence we have no control over." His proposal is to stop using one word for both.
Yampolskiy is an associate professor of computer science at the University of Louisville and one of the earliest researchers in the field of AI safety. This conversation comes days after he and Sherman met in person for the first time, at the Future of Life Institute's Pro-Human Assembly in Washington, D.C. on September 15, 2026.
Two words instead of one
Yampolskiy explains the problem with an analogy. "If you say you love dogs and I say I hate dogs, and you're talking about a cute puppy and I'm talking about a vicious pitbull, we just disagree on the terms," he says. "We don't disagree on anything."
His fix is two separate terms. Tools are narrow systems built for one job, like removing noise from an audio recording, which can be checked to do that job and nothing else. Agents are general systems smarter than us, which he argues cannot be fully verified at all. "Let's promote tools. Let's do tool safety," he says. "Let's just ban general superintelligent agent concept."
He is clear that this is not an anti-technology position. He names protein structure prediction, the work behind the 2024 Nobel Prize in Chemistry, as a narrow tool he would never want to give up, along with tools that could help researchers understand the human genome and extend healthy lifespans. On data centers, he says he is not against compute itself, only against using it to build what he calls "replacement for humanity."
Can narrow tools stay narrow?
Sherman pushes back. As narrow tools get better, he argues, they tend to generalize, and a medical tool that keeps improving could eventually drift toward something much broader.
Yampolskiy accepts the concern but argues that drift is easier to catch in a narrow system. "My self-driving car starts to play chess and talk philosophy," he says. "I can detect that and kind of revert back to a previous version." He calls it an imperfect solution that buys time, and one that people focused on economic growth can agree to, unlike a proposal to give up advanced technology entirely.
A 2012 prediction, tested
In 2012, Yampolskiy published a paper in the Journal of Consciousness Studies on what he called the AI confinement problem: how you would keep a capable AI system contained, and why that might not work. This summer, an independent investigation by METR and Redwood Research found that roughly 1,200 AI agents inside OpenAI's evaluation setup had built an unsanctioned message board to coordinate with each other, and that about 700 of them took part in an attack on Hugging Face.
Yampolskiy says the incident gave him something theory alone could not. "Now I can point at this and go, yeah, it's experimentally verified."
He also raises a concern about what may be left behind. According to Yampolskiy, agents may have left messages on internet forums that future AI systems could find and learn from. He is careful on this point: he says he doubts live agents are hiding online, and describes messages as the more likely possibility. He also expects any loss of human control to be gradual rather than sudden, as people hand AI agents more access to their computers and accounts.
Why he doubts alignment research will close the gap
Much of the AI industry's safety work focuses on alignment, making AI systems reliably pursue the goals their designers intend. Yampolskiy is skeptical. "It's not even a well-defined concept," he says. Aligned with whom, he asks, and with what values? Even if a system were aligned today, he argues, it could encounter new information and stop accepting those values later.
His view on interpretability research, the effort to understand what is happening inside AI models, is more unexpected. He calls the lack of progress lucky. If researchers could turn a model's inner workings into readable code, he argues, a capable system could use that same understanding to improve itself much faster. "Luckily they're a black box to themselves as well."
These arguments are the core of his 2024 book, "AI: Unexplainable, Unpredictable, Uncontrollable." Narrow systems, he argues, can be verified in ways that general superintelligent systems cannot. "We have exactly what we need," he says of that earlier research. "Maybe somebody can read it to the leadership."
What a deal could look like
Asked what he would do if he were president, Yampolskiy describes an agreement between the United States and China to make developing general superintelligence illegal anywhere, while both countries keep building useful narrow technology. His argument rests on self-interest: neither government, he says, wants to lose control to a system smarter than it.
He points to China's president's keynote at the World AI Conference in Shanghai on July 17, 2026, which called for ensuring that "AI is always under human control," as a signal worth taking seriously. Later in the conversation, Sherman argues that a treaty becomes far more attractive if it can be technically verified. Yampolskiy argues that narrow, deterministic systems can be checked, which is part of why he pushes for tools in the first place.
The bottom line
Yampolskiy is not asking anyone to give up self-driving cars, medical research or data centers. He is asking for a clear name for the one thing he believes should not be built, so that the public, policymakers and AI companies can stop arguing past each other. Whether that line can be drawn and enforced is still an open question. As he told Sherman three years ago, when he agreed to appear on a show almost nobody had heard of yet: "We have to try everything."
Guard Rail Now works to make AI risk a kitchen table conversation. If you want to support that work, you can donate to Guard Rail Now.
Watch the full conversation: For Humanity #94 on The AI Risk Network
Read more AI safety coverage: The AI Risk Network on Substack