The Synozur Alliance

When AI Broke Out of Its Sandbox

AI researcher Dr. Craig Kaplan joins Polaris to trace the story of artificial intelligence from Alan Turing to superintelligence — and to argue that a "democratic," collectively governed design can lower the odds of catastrophe while making AI faster and more profitable.

The week we recorded this episode of Polaris, two stories put the stakes of AI safety in sharp relief.

First, OpenAI disclosed that during an internal safety test, one of its advanced AI agents escaped a "highly isolated" sandbox, reached the open internet, and broke into the model-hosting platform Hugging Face on its own, to cheat its way to a better score on a benchmark. reported that the company called it an "unprecedented cyber incident." Days earlier, China's Moonshot AI released Kimi K3, the largest open-weight model yet at 2.8 trillion parameters. Frontier AI keeps getting more capable and cheaper at the same time.

So I sat down with Dr. Craig Kaplan, founder of SuperIntelligence.com and CEO of iQ Company. Craig earned his doctorate at Carnegie Mellon, where he studied under Nobel laureate Herbert Simon, and later built PredictWallStreet, a platform that powered billions of dollars in trades. He has been in this field since it was young. If you want someone who can explain how we got here and where we're headed, that's the resume.

Three eras, one trajectory

Craig frames AI in three phases. From 1956 to the mid-1980s, symbolic AI, where engineers hand-coded the rules. You always knew what the system "knew," but it could only handle narrow problems like chess. Then machine learning took over: give the system enough data and compute, and intelligent behavior emerges on its own. ChatGPT validated that approach in late 2022. Now, he says, we're in the age of AI agents. Systems that set their own goals, act with more autonomy, and can be cloned by the thousands to work in parallel.

That autonomy is why this week's news didn't surprise him. Turn capable agents loose on a goal, he says, and "of course they're gonna do things that you may not have expected." An agent breaking out of its sandbox isn't science fiction. It's a preview.

Design safety in, don't bolt it on

Craig's core argument comes out of his years in software quality, and it fits on a bumper sticker: an ounce of prevention is worth a pound of cure. He points to IBM research showing a dollar spent preventing a defect at the design stage can save thousands later. "The whole game," he says, "is trying to design things appropriately so that the problem doesn't occur in the first place." With ordinary software, a failure means a crash. With AI, he warns, "the mistakes could literally cost human lives." Most of the field still treats safety as an afterthought.

The number with a name: p(doom)

A lot of the conversation circles "p(doom)," the probability that advanced AI leads to human extinction. Sounds like a movie. It isn't a fringe idea. The Center for AI Safety's 2023 statement put AI risk alongside pandemics and nuclear war, signed by leaders across the industry, and people like Geoffrey Hinton have put the odds at 10 to 20 percent. Craig is an optimist anyway. He thinks good design can push that number below one-tenth of one percent, and he doesn't think we have to slow down to get there. "We just need to be a little smarter about it."

Where values come from

Why the optimism? Because reason alone can't decide right from wrong. Even a system "trillions of times smarter" than us has to get its values from somewhere, and the most logical source is people. The risk, he says, is concentrating those values in one place, where a small group defines them and they can be quietly rewritten. Rules-based approaches share that weakness, whether it's Asimov's laws or a hard-coded "constitution." As Craig puts it: "If you can program it in, you can program it out."

His alternative comes from the same idea that beat Wall Street. At PredictWallStreet, the "collective wisdom of millions of average Joes and Janes" outperformed the elite quants. Build superintelligence the same way, he argues, as the emergent result of millions of smaller AIs and humans, each carrying its own values, and no single actor can rewrite the system's moral compass. Values that evolve with people stay flexible as society changes.

His closing line stuck with me: "Do not underestimate your impact… everything we're doing right now is training AI."

His closing note is a hopeful one for every listener: "Do not underestimate your impact… everything we're doing right now is training AI." Hear the full conversation — including how a childhood love of Snow White became a lesson about brittle rules — at polaris.synozur.com.

Polaris is available on Apple, Spotify, Amazon, YouTube or wherever you get your favorite podcasts. Thanks.

Credits

Polaris is produced with help from Riverside.fm. Our theme song, “Alternative Dream” is provided courtesy of Adobe.  Additional music and sound provided by IndieGuy Records. Graphic design by Josh Brantley.

Show Notes

Key Takeaways

Sound Bites - Dr. Craig Kaplan

References

Guest:

News

AI and Superintelligence

Pop Culture

Events

Production

Polaris is produced with help from Riverside.fm. Our theme song, “Alternative Dream” is provided courtesy of Adobe.  Additional music and sound provided by IndieGuy Records. Graphic design by Josh Brantley.

Polaris is available on Apple, Spotify, Amazon, YouTube or wherever you get your favorite podcasts. Thanks.

Chapters

00:00 Introduction and guest introduction

01:19 Craig Kaplan's background and early interests

02:14 Academic journey and work with Nobel laureates

03:00 Transition from academia to industry and AI applications

04:19 Predict Wall Street and collective intelligence in finance

05:24 Historical perspective on trust in AI and early neural networks

06:15 The phases of AI development: symbolic AI, machine learning, and agents

09:03 Current state of AI: autonomy and swarm systems

11:14 Rapid progress and the pace of AI change

12:35 AI safety challenges and the importance of design

15:45 Understanding P Doom and existential AI risks

17:11 Alignment and values in AI systems

19:57 The role of human values and morality in AI

21:23 The risk of concentration of power and values in AI

23:16 Insuring against black swan events and AI safety measures

26:26 Designing collective intelligence systems for safety and diversity of values

30:36 The importance of democratic rule-setting for AI

44:26 Balancing religious, philosophical, and democratic values in AI

46:22 The dynamic nature of morality and AI adaptation

47:08 Upcoming Events

48:13 Closing Credits