About this Concept
Safety & governance
Ways to reduce harm, set limits, test behavior, and govern how AI is used.
Trend this week
7 of 7 weeks tracked had mentions across multiple podcasts.
A closer read
- What changed
- Its share fell from 16.5% to 11.7%, with evidence across 11 podcasts.
- Where it showed up
- It appeared in episodes including “Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $100...” and “Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, ...”.
- How to read it
- 5 cited episodes from 4 podcasts are shown below. The topic appeared across 11 podcasts in the full analysis.
Podcast moments from the week
Moments are grouped by episode so repeated excerpts from one conversation do not look like separate sources.
The Cognitive Revolution
Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
“that's sort of how you generalize from interpretability as like a factoring tool to interpretability as a tool for alignment.”
ConceptSafety & governance
Report this moment“we're going to have to continue to be very thoughtful in our, in like what guardrails we apply”
ConceptSafety & governance
Report this moment“I think we have to sort of get in front of, of some of these risks before they become more severe.”
ConceptSafety & governance
Report this moment
The Cognitive Revolution
Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...
“If you say that's aligned, then I don't care about the thing you're calling alignment.”
ConceptSafety & governance
Report this moment“That is aligned in the sense that you solved the alignment problem and got to do what you want to do, but also we're all super dead.”
ConceptSafety & governance
Report this moment“Do you have a best-case scenario for what safe superintelligence might be up to? Could you imagine anything that feels Both somewhat realistic and, wow, that would be amazing if they could come forward with something like what you have in mind.”
ConceptSafety & governance
Report this moment
The Artificial Intelligence Show
#228: More Rogue AI Agents, AI Lab Staff Ask Washington to Pace Development, Continuing Battle Over Open Weights & OpenAI Previews Astra
“We may be one turn away on frontier models models to where these researchers do no longer feel comfortable that they fully understand what the models are capable of and how to put guardrails in place to safely release them into the world.”
ConceptSafety & governance
Report this moment“mandatory safety testing for all sufficiently capable models, open and Closed seems reasonable.”
ConceptSafety & governance
Report this moment
Last Week in AI
#253 - Opus 5, Gemini 3.6, Kimi K3, Hugging Face Hack
“I also expect that our alignment kind of efforts will start to lag more and more behind capabilities over time.”
ConceptSafety & governance
Report this moment
Gradient Dissent
40 Trillion Tokens a Day (Yes, More Than OpenAI) | Lin Qiao, CEO of Fireworks
“Is that right? My understanding is it was using— it was testing Exploit Jam, which is a cybersecurity attack benchmark, and that's just starting to go wild.”
ConceptSafety & governance
Report this moment









