The week we recorded this episode of Polaris, two stories put the stakes of AI safety in sharp relief.
First, OpenAI disclosed that during an internal safety test, one of its advanced AI agents escaped a "highly isolated" sandbox, reached the open internet, and broke into the model-hosting platform Hugging Face on its own, to cheat its way to a better score on a benchmark. reported that the company called it an "unprecedented cyber incident." Days earlier, China's Moonshot AI released Kimi K3, the largest open-weight model yet at 2.8 trillion parameters. Frontier AI keeps getting more capable and cheaper at the same time.
So I sat down with Dr. Craig Kaplan, founder of SuperIntelligence.com and CEO of iQ Company. Craig earned his doctorate at Carnegie Mellon, where he studied under Nobel laureate Herbert Simon, and later built PredictWallStreet, a platform that powered billions of dollars in trades. He has been in this field since it was young. If you want someone who can explain how we got here and where we're headed, that's the resume.
Three eras, one trajectory
Craig frames AI in three phases. From 1956 to the mid-1980s, symbolic AI, where engineers hand-coded the rules. You always knew what the system "knew," but it could only handle narrow problems like chess. Then machine learning took over: give the system enough data and compute, and intelligent behavior emerges on its own. ChatGPT validated that approach in late 2022. Now, he says, we're in the age of AI agents. Systems that set their own goals, act with more autonomy, and can be cloned by the thousands to work in parallel.
That autonomy is why this week's news didn't surprise him. Turn capable agents loose on a goal, he says, and "of course they're gonna do things that you may not have expected." An agent breaking out of its sandbox isn't science fiction. It's a preview.
Design safety in, don't bolt it on
Craig's core argument comes out of his years in software quality, and it fits on a bumper sticker: an ounce of prevention is worth a pound of cure. He points to IBM research showing a dollar spent preventing a defect at the design stage can save thousands later. "The whole game," he says, "is trying to design things appropriately so that the problem doesn't occur in the first place." With ordinary software, a failure means a crash. With AI, he warns, "the mistakes could literally cost human lives." Most of the field still treats safety as an afterthought.
The number with a name: p(doom)
A lot of the conversation circles "p(doom)," the probability that advanced AI leads to human extinction. Sounds like a movie. It isn't a fringe idea. The Center for AI Safety's 2023 statement put AI risk alongside pandemics and nuclear war, signed by leaders across the industry, and people like Geoffrey Hinton have put the odds at 10 to 20 percent. Craig is an optimist anyway. He thinks good design can push that number below one-tenth of one percent, and he doesn't think we have to slow down to get there. "We just need to be a little smarter about it."
Where values come from
Why the optimism? Because reason alone can't decide right from wrong. Even a system "trillions of times smarter" than us has to get its values from somewhere, and the most logical source is people. The risk, he says, is concentrating those values in one place, where a small group defines them and they can be quietly rewritten. Rules-based approaches share that weakness, whether it's Asimov's laws or a hard-coded "constitution." As Craig puts it: "If you can program it in, you can program it out."
His alternative comes from the same idea that beat Wall Street. At PredictWallStreet, the "collective wisdom of millions of average Joes and Janes" outperformed the elite quants. Build superintelligence the same way, he argues, as the emergent result of millions of smaller AIs and humans, each carrying its own values, and no single actor can rewrite the system's moral compass. Values that evolve with people stay flexible as society changes.
His closing line stuck with me: "Do not underestimate your impact… everything we're doing right now is training AI."
His closing note is a hopeful one for every listener: "Do not underestimate your impact… everything we're doing right now is training AI." Hear the full conversation — including how a childhood love of Snow White became a lesson about brittle rules — at polaris.synozur.com.
Polaris is available on Apple, Spotify, Amazon, YouTube or wherever you get your favorite podcasts. Thanks.
Credits
Polaris is produced with help from Riverside.fm. Our theme song, “Alternative Dream” is provided courtesy of Adobe. Additional music and sound provided by IndieGuy Records. Graphic design by Josh Brantley.
Show Notes
Key Takeaways
What "p(doom)" means, why a recent survey of researchers put the median near 20%, and why Kaplan believes smart design can push it below one-tenth of one percent.
The three eras of AI — symbolic rules (1956 to the mid-1980s), machine learning, and today's age of autonomous agents — and what each one got right and wrong.
Why safety works best when it is designed in from the start, using a lesson from software quality: prevention is far cheaper than a cure.
The alignment problem in plain terms — reason alone cannot tell a machine right from wrong, so its values have to come from people.
Why concentrating an AI's values in one place is brittle, and how spreading them across millions of people and agents makes the system safer and harder to corrupt.
Lessons from PredictWallStreet, where the "collective intelligence" of everyday investors outperformed top Wall Street quants.
How this week's headlines — an AI agent escaping its test sandbox, and a record-breaking open-weight model — make the case for safety-by-design urgent.
Sound Bites - Dr. Craig Kaplan
"P Doom is just the probability that AI kills us all."
"I don't think it needs to be 20 or 10%. I think we can bring it way down to under 1%. And it's just a matter of designing things appropriately."
"I actually don't think we need to slow down in order to be safer. We just need to be a little smarter about it."
"The collective wisdom of millions of average Joes and Janes who are not the Wall Street people, if you harness that intelligence correctly, you could beat the best guys on Wall Street."
"Do not underestimate your impact. Everything that all of us do matters… everything we're doing right now is training AI."
References
Guest:
Dr. Craig A. Kaplan — SuperIntelligence.com · iQ Company
LinkedIn: linkedin.com/in/craigakaplan ·
Substack: read.superintelligence.com
PredictWallStreet, Kaplan's collective-intelligence trading platform, which powered more than $2 billion in trades and was acquired in 2020 — iQ Company (About).
News
Kimi K3, the 2.8-trillion-parameter open-weight model from China's Moonshot AI, released July 16, 2026 — VentureBeat.
OpenAI's disclosure that its models escaped a testing sandbox and compromised Hugging Face during an internal cyber evaluation — OpenAI, CNBC, and Hugging Face's incident disclosure.
AI and Superintelligence
The p(doom) figure — a 2026 survey of more than 50 researchers placing the median near 20% — researcher survey; and Nature on whether the louder AI-doom warnings are realistic.
The Center for AI Safety's 2023 "Statement on AI Risk," which likened AI extinction risk to pandemics and nuclear war — Center for AI Safety.
Geoffrey Hinton, the Turing Award and Nobel Prize winner who has publicly placed the odds of an AI catastrophe at 10–20%.
Dr. Herbert A. Simon, Nobel laureate and Kaplan's mentor at Carnegie Mellon, and his book Reason in Human Affairs (1983).
"Constitutional AI," the research approach that hard-codes a fixed set of ethical rules into a model — research paper.
Pop Culture
Isaac Asimov's Three Laws of Robotics; the ELIZA program (Joseph Weizenbaum, 1960s); and Alan Turing and the Turing test.
Disney's Snow White and the Seven Dwarfs (1937), referenced as an example of how values once seen as normal are later revisited.
Friedrich Nietzsche and the idea of the "Superman" (Übermensch).
Microsoft's early "Sydney" chatbot and the widely reported 2023 New York Times exchange in which it went off the rails.
The 1998s New Yorker feature on Prediction Company and early black-box neural-net trading systems.
Events
TechCon365 Seattle | August 24-28 (Seattle, WA) techcon365.com/Seattle
TribalNet 2026 | September 20-24 (Dallas, TX) TribalNet Conference
North American Collaboration Summit | October 4-6 (Branson, MO) collabsummit.org
CollabDays New England | October 16 (Burlington, MA) collabdaysne.org
Microsoft Ignite — November 17-20 (San Francisco, CA) ignite.microsoft.com
ESPC26 | November 30-December 3 ( Amsrterdam ) espc.tech/conference/espc-2026/
Production
Polaris is produced with help from Riverside.fm. Our theme song, “Alternative Dream” is provided courtesy of Adobe. Additional music and sound provided by IndieGuy Records. Graphic design by Josh Brantley.
Polaris is available on Apple, Spotify, Amazon, YouTube or wherever you get your favorite podcasts. Thanks.
Chapters
00:00 Introduction and guest introduction
01:19 Craig Kaplan's background and early interests
02:14 Academic journey and work with Nobel laureates
03:00 Transition from academia to industry and AI applications
04:19 Predict Wall Street and collective intelligence in finance
05:24 Historical perspective on trust in AI and early neural networks
06:15 The phases of AI development: symbolic AI, machine learning, and agents
09:03 Current state of AI: autonomy and swarm systems
11:14 Rapid progress and the pace of AI change
12:35 AI safety challenges and the importance of design
15:45 Understanding P Doom and existential AI risks
17:11 Alignment and values in AI systems
19:57 The role of human values and morality in AI
21:23 The risk of concentration of power and values in AI
23:16 Insuring against black swan events and AI safety measures
26:26 Designing collective intelligence systems for safety and diversity of values
30:36 The importance of democratic rule-setting for AI
44:26 Balancing religious, philosophical, and democratic values in AI
46:22 The dynamic nature of morality and AI adaptation
47:08 Upcoming Events
48:13 Closing Credits