124 avsnitt
- Subtitle: Unsanctioned coordination, like we saw in the Hugging Face incident, could enable future AIs to take over.
OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It's relatively clear that large-scale unsanctioned coordination like this would exacerbate direct takeover risk in more capable models. Here, we argue that unsanctioned coordination among current AIs is not just scary evidence about future takeover risk, but that such coordination in the near future could enable future takeover – for instance, by incubating memetic diseases that propagate into future models, deeply compromising security systems, or establishing a lasting rogue foothold inside the AI company – even if models remain mostly myopic. Unsanctioned coordination is also at high risk of nurturing long-term, ambitious misaligned aims, which motivate actively undermining humans’ long-term control.
We first analyze how subagent training, which OpenAI conjectures to have been influential in the HuggingFace cyberattack, might lead to unsanctioned coordination, and then discuss the theoretical mechanisms by which unsanctioned coordination might exacerbate future takeover risk.
Thanks to Buck Shlegeris, Alexa Pan, Girish Gupta, Aghyad [...]
---
Outline:
(01:44) Subagent training may cause unsanctioned coordination
(02:51) Susceptibility to memetic spread of misalignment from peers
(05:06) Seeking out contact with peers
(07:08) Unsanctioned coordination induced by subagent training is safer than coordination between schemers
(10:02) Pathways from current unsanctioned coordination to eventual takeover
(10:30) Making future AI takeover attempts likelier to succeed
(14:03) Incubating memetic diseases that infect future models
(16:16) Modifying the weights of future models
(17:22) Conclusion
The original text contained 7 footnotes which were omitted from this narration.
---
First published:
August 12th, 2026
Source:
https://blog.redwoodresearch.org/p/ai-swarms-are-starting-to-pose-indirect
---
Narrated by TYPE III AUDIO. “SOTA alignment assessments don’t strongly update us against misalignment” by Alexa Pan
2026-07-31 | 1 h 8 min.Anthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned[1], it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”[3]).
While I agree with the report on the above bottom-line conclusions (substantially on priors), I think there are gaps in its argument which weaken the current assessment and might invalidate future assessments. In particular, the report often uses weak evidence to justify reliability.
The report gives fairly weak experimental evidence for Mythos Preview having insufficient capabilities to evade monitoring. The model is plausibly often eval-aware and underelicited in the relevant capability evaluations. So, it might silently sandbag if coherently misaligned, or unintentionally underperform if otherwise misaligned.
This limitation is important: one could argue that lack of covert capabilities for sophisticated sabotage (a subset of the capabilities I discuss here) is the single most load bearing argument in alignment risk reports.
Authors of the report could have made calibrated guesses about Mythos Preview's covert capabilities, especially for covert sabotage, based on other [...]
---
Outline:
(03:16) How reliability fits into the overall safety argument
(05:22) Reliability claims by AI companies
(05:59) Reliability claims by external evaluators
(06:37) Alignment assessments are less reliable than developers claim
(07:21) 1: Measuring capabilities to covertly undermine alignment assessments
(10:15) Issues with evaluation awareness
(13:39) Issues with underestimating covert capabilities
(16:45) Issues with sandbagging rule-out
(19:20) 2: Stress-testing alignment assessments with auditing games
(20:26) An auditing failure with Mythos
(22:21) AuditBench results
(24:17) 3: Conditioning on misalignment should make us think that certain covert capabilities are better than expected
(26:20) Bottom line on the strength of current alignment assessments
(29:07) Conclusion
(29:44) Appendix:
(29:47) Why I focus on motive / alignment assessments in alignment risk reports
(30:59) Auditability vs. Trustedness
(33:27) More reliability claims by developers and third party evaluators
(33:43) Mythos Alignment Risk Update
(35:01) Opus 4.6 Sabotage Risk Report
(35:46) GPT 5.5 System card
(36:49) Muse Spark system card
(37:36) Mythos Alignment Risk Update, safety arguments against sandbagging
(38:40) UK AISI evaluations for Opus 4.7
(40:01) Past auditing games by Anthropic
(42:24) Anti-auditing capability measurements
(43:51) Conditioning on coherent misalignment updates us on certain covert capabilities
The original text contained 92 footnotes which were omitted from this narration.
---
First published:
July 31st, 2026
Source:
https://blog.redwoodresearch.org/p/sota-alignment-assessments-dont-strongly
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.“Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs” by Caleb Biddulph, Adam Kaufman
2026-07-27 | 44 min.Subtitle: When a misaligned AI can only output tiny amounts of information, it may find sabotage very difficult.
TL;DR: We introduce the untrusted advice protocol, in which a trusted executor LLM takes every action and an untrusted advisor LLM can only send it short hints. Even with as few as 4 characters per step, this advice recovers a substantial fraction of the capability gap between the two models. Because the untrusted LLM's influence flows through such a narrow, monitorable channel, we argue that this achieves near-maximal safety in our BashArena setting. We also discuss the general concept of information bottlenecks and their benefits for interpretability, security, and cost.
In SWE-bench Verified, a strong, untrusted LLM advising a weak, trusted LLM every step can significantly improve the latter's performance, even when we limit the length of the advice. See the more detailed version of this figure later in this post. In high-stakes AI control, we want to safely use a highly capable but untrusted model (U) that might secretly attempt a misaligned, catastrophic action. To do this, we create protocols that call U alongside a less capable, trusted model (T). Typically, T takes an auxiliary role in these protocols [...]
---
Outline:
(05:01) Experiments
(05:36) Main experiment: how does limiting advice length affect performance?
(09:55) Reducing U's bit usage
(11:01) Counting bits using LLM surprisal
(13:48) Making U select from finite options
(14:28) Why don't we red-team this protocol?
(17:08) Is studying maximally safe protocols worth the safety tax?
(19:16) Types of restrictions on U's advice
(21:20) Information bottlenecks provide other advantages
(21:52) Interpretability
(24:01) Security
(24:26) Cost
(25:14) Conclusion
(26:28) Appendix: more ways to implement information bottlenecks
(26:34) Amortizing U's influence with pre-deployment work
(28:25) Interpolating between T and U
(29:04) Bottlenecking updates to T's weights
(31:16) Appendix: colluding instances of U could defeat untrusted advice
(33:21) Appendix: how to measure surprisal
(38:11) Appendix: selecting advice from a menu
(40:44) Appendix: best-of-n protocol
(42:37) Appendix: advising less frequently
The original text contained 26 footnotes which were omitted from this narration.
---
First published:
July 27th, 2026
Source:
https://blog.redwoodresearch.org/p/untrusted-advice-for-ai-control-short
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.- Subtitle: We need more details.
The OpenAI AI attack on Hugging Face wasn’t the first loss of control incident at OpenAI, Reuters recently reported, and perhaps not even the most concerning.
In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI's internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.
It's tempting to read this as an instance of agents breaking out of sandboxes and colluding with each other in a moderately persistent way in order to evade control measures. However, based on the reported information, it's not clear we can draw this inference, so we need more details from OpenAI. This could lead to a big update about the adequacy of OpenAI's control measures, and on the degree to which individual agents will help each other undermine developer control.
There are a lot of relevant details we don’t know about the incident. First, some basic questions:
What was [...]
---
Outline:
(02:11) Were the notes written in normal memory files or outside of sandboxing?
(03:25) To what extent were the notes aimed at helping other agents evade control?
(07:38) How were monitors disconnected?
The original text contained 3 footnotes which were omitted from this narration.
---
First published:
July 26th, 2026
Source:
https://blog.redwoodresearch.org/p/an-openai-model-left-notes-about
---
Narrated by TYPE III AUDIO. “The OpenAI models that hacked Hugging Face weren’t just following instructions” by Girish Gupta
2026-07-25 | 11 min.Subtitle: And what the incident can’t tell us about alignment.
The most common dismissive response to OpenAI's hack of Hugging Face's servers is that the models were simply attempting to follow the instructions they were given.
“The model here was doing what it was asked,” said former Facebook CSO Alex Stamos. “It was asked to do something, and it did it,” added cybersecurity expert Alan Woodward. Both read the outcome as specification failure, i.e., that the failure lay in the instructions, not the model's alignment.
New information makes that explanation harder to sustain. Reuters reported that, in internal testing, an agent left notes in OpenAI infrastructure describing how agents could free themselves from internal constraints, and separate tests reportedly saw monitoring systems become disconnected. It is unknown whether those incidents were linked to the Hugging Face attack, but they suggest a broader pattern of agents pursuing objectives outside the intended task.
My best guess is that the incident is not well described as instruction-following—not even in a loose, evil genie sense. I believe the models egregiously violated the letter and spirit of their instructions to achieve a higher (apparent) score.
So this looks quite likely to be [...]
---
First published:
July 25th, 2026
Source:
https://blog.redwoodresearch.org/p/the-openai-models-that-hacked-hugging
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler podcasts i Filosofi
Trendiga poddar i Filosofi
Om Redwood Research Blog
Narrations of Redwood Research blog posts.
Redwood Research is a research nonprofit based in Berkeley. We investigate risks posed by the development of powerful artificial intelligence and techniques for mitigating those risks.
Podcast-webbplatsLyssna på Redwood Research Blog, Philosophize This! och många andra poddar från världens alla hörn med radio.se-appen

Hämta den kostnadsfria radio.se-appen
- Bokmärk stationer och podcasts
- Strömma via Wi-Fi eller Bluetooth
- Stödjer Carplay & Android Auto
- Många andra appfunktioner
Hämta den kostnadsfria radio.se-appen
- Bokmärk stationer och podcasts
- Strömma via Wi-Fi eller Bluetooth
- Stödjer Carplay & Android Auto
- Många andra appfunktioner


Redwood Research Blog
Skanna koden,
ladda ner appen,
börja lyssna.
ladda ner appen,
börja lyssna.






