Fables - Pt II
Tea leaf reading and a tentative case for optimism.
Not two weeks ago, I published a post trying to extract the moral lesson – or fable, if you will – from the Claude Fable ban. At the time, Fable had been indefinitely retracted from the public market in order to comply with the Trump Administration’s export control directive to suspend access by foreign nationals. I rounded out my previous blog by concluding that the situation was evolving, and the ending to this fable was therefore as-of-yet unwritten:
If there is a cautionary tale at the heart of this saga, it remains unfinished. In the bad case, policymakers conclude that Fable is a singular menace, but fail to discern the broader trend of AI capabilities that are running out of our control. They continue to play a game of whack-a-mole that identifies specific models as threats based on politics, personal grudges or random contingent facts of history like who-happened-to-call-who. The better ending – and one that extricates us from the genre altogether – is that this tale proves a vivid example that finally makes this urgent, and increasingly untenable, situation clear.
In the weeks since, there have been two important developments: OpenAI announced that its most recent model, GPT-5.6, will be subject to staggered deployment at the request of the US government, and public access to Fable has been restored.
GPT-5.6 rollout
The first of these developments appears to rule out the worst possible ending, that the administration considered Fable a “singular menace”, or was simply targeting Anthropic out of spite. It doesn’t entirely foreclose the possibility that individuals in the government harbour some specific saltiness towards Anthropic, since the mechanism by which Fable was restricted – an export control directive with no apparent timeline or clear justification – is far more heavy-handed than the voluntary regime into which OpenAI has now entered. But at the very least, it is now clear that the administration has internalised a more general principle than “Claude is bad and scary”.
The staggered rollout of GPT-5.6 is also not that surprising; it looks a bit like a sneak preview of the regime we should expect under Trump’s June Executive Order “Promoting advanced AI innovation and security”, which directs various bodies to design a voluntary scheme whereby AI companies “provide the Federal Government with access to covered frontier models…for a period of up to 30 days before they plan to release such models to other trusted partners”. This directive is not in effect yet, since the order mandates the creation of such a scheme within 60 days, but we should likely expect rollouts broadly of this nature in the future. Still, the GPT-5.6 rollout will require more government intervention than we might have expected – the admin will be approving access to 5.6 on a customer-by-customer basis during the preview period, which is not mandated under the EO. So the update here is that the admin is Getting Serious, and looks set to proceed with more Seriousness than it planned to not a month ago.
Does this mean we’re on track for the good ending? I was somewhat surprised to see many safety-aligned people react very negatively to the GPT-5.6 announcement. Zvi Mowshowitz calls staggered deployment a “maximally terrible policy”, while Andrew Curran points out that safety advocates should pause before taking a victory lap, since this new regime will not slow down development, but will only cause the gap between internally-deployed and publicly-available capabilities to “steadily widen”.
I can certainly see the downsides of this policy, while remaining confused about how it is “maximally bad”, namely:
It makes existing capabilities less publicly auditable, which is bad both for public awareness of AI risk, and because external safety organisations will now have a harder time conducting research at the frontier.
It could bring us closer to a scenario that Daniel Kokotajlo describes in a manifesto against training AGI in secret (I helped write up an article based on Daniel’s original piece here). In this scenario, the number of people within an AI company with access to frontier capabilities dwindles as the process of automated AI R&D gets underway due to strict information siloing, borne out of fear of regulatory scrutiny or public backlash. This results in a tiny group responsible for both aligning AGI and deciding how to use the power it grants them, assuming they manage to keep it under control at all. The staggered deployment regime only gives the government visibility into models that a company plans to publicly release, so one could imagine this incentivising increasingly furtive development of very powerful systems to avoid triggering any government involvement at all.
Even absent existential risk, the government arbitrating who can and cannot access the most powerful AI models is obviously bad from a concentration-of-power perspective.
It reveals that the government still lacks the Situational Awareness not to endlessly chase red herrings; they are still conceiving of AI models as things that are exclusively dangerous in the hands of the public, foreign actors, cybercriminals etc. This is despite the AI safety community being at pains to point out that internal deployment of powerful capabilities poses just as much risk, if not more.
Still – and maybe I am just stupid or grasping at straws – I can’t help but feel ever-so-slightly heartened by this development. As I repeatedly said in my recap of the last year’s AI safety events, I am resting a lot of my hope in some sort of Lightbulb Moment among the US government or national security establishment that alerts them to the direness of the situation. I can imagine this wakeup being profound enough that it obviates many of the bad policy decisions that came before it. And it’s hard to argue that a staggered deployment regime wouldn’t make a wakeup of this nature somewhat more likely. This alone makes it clear to me that the current policy isn’t “maximally terrible”.
Information on how the GPT-5.6 rollout will actually work is pretty thin, but we know per reporting from Axios that Commerce Secretary Howard Lutnick “[wants] to be sure all relevant parts of the government have tested and approved the model”. We can reasonably guess that “relevant parts of the government” will include the NSA, CISA, CAISI, the OTSP, and perhaps others. I don’t have a particularly comprehensive mapping of the US government in my head – but that’s a lot of Serious People with well-trained muscle memory for responding to Serious Threats with eyes on (and hands on) one of the world’s most capable AI models. The Director of the CIA has apparently made a “rare public appearance” to compare frontier AI to “digital nuclear weapons”. Things are happening. To be fair, we don’t know what capabilities these agencies will be most focused on evaluating, and I could certainly imagine that they will over-index on misuse – and particularly cybersecurity – risks at the expense of all the weirder loss-of-control threat models that are the bread and butter of safety organisations working out of Berkeley and San Francisco. We still have a way to go on educating the policy establishment about misalignment risks. But come on guys, this is something! We should be calibrating our optimism-metres based not just on what the government is currently doing, but also on how it might act once it has substantially more information. And if there’s one thing we can now be sure of, it is that information is starting to flow.
Though Zvi calls the GPT5.6 rollout announcement “maximally terrible”, he later lays out a position that I largely agree with:
If we are wise, we will use this as an opportunity to Pick Up The Phone. We have sent a costly signal that we see real issues with these models and are willing to make real sacrifices in the name of security. We should try to use that to get things in return, or at least lay the foundation for joint action.
Most of all, we are out of the ‘people don’t do things’ phase of the game, where basically all meaningful actions were outside the Federal overton window. That’s done. We should expect a lot more actions, many of them similarly ill-executed, at least at first, in ways that we did not anticipate.
I like the angle that this could provide an opportunity for international coordination; the US government is loudly and publicly signalling a willingness to take its foot off the gas pedal. This signal could be even louder and clearer, of course, and maybe it will be in the future. It is also true that the “people just don’t do things” era is over. Much as I often encourage AI-sceptics to envision how they would have felt two years ago in the presence of AI capabilities that exist today, I think the pro-regulation crowd should consider just how much the Overton Window has shifted in that time. We’re no longer in a position where the government is asleep at the wheel while companies YOLO their way to superintelligence with less regulatory oversight than Louisiana-based flower arrangers. Things are not as bad as they were.
Fable returns
The other significant development of the last two weeks has been the (mostly) ceremonious return of Fable to the public market. People are mostly happy about this, while simultaneously being pissed off that Anthropic appears to have cut Fable’s capabilities off at the knees by downgrading users to Opus 4.8 for anything that looks even tentatively cyber or bio-related.
The reasoning behind the retraction – and re-release – of Fable still seems fairly opaque, and the entire process appears to have been quite chaotic. It appears as if Anthropic has spent a good part of the last few weeks demonstrating to the US government that their initial reaction had been the result of a misunderstanding. Per their own blog, they tested a series of models, including several less capable Claudes, and found that all of them could identify the same software vulnerability that had been the genesis of the original Amazon-mediated freakout. They developed a new classifier that blocks the technique discovered by Amazon in over 99% of cases, but will also come at the cost of blocking some benign cyber requests, “out of an abundance of caution”. The US government offered a vaguer narrative events, with Secretary Lutnick stating in a letter reported on by the New York Times that Anthropic had “taken steps in close coordination with the US government to address the risks posed by the model” – though it appears from Anthropic’s account of events that there weren’t specific threats posed by Fable, and that the attendant mitigations have not really mitigated much of anything at all.
As part of the negotiations, Anthropic has also agreed to set up a collaboration with HackerOne that will let security researchers submit vulnerabilities, and (re)committed to incident reporting in line with the June EO mentioned earlier. So we have some better-than-nothing safety infrastructure that has come from this debacle. If safety-aligned people have negatively updated on this series of events, it appears to be because the government doesn’t seem to have a clue what they’re doing, and their risk assessment appears widely uncalibrated. I agree that this is the case – I just don’t think this revelation ought to outweigh all the Situational Awareness benefits it could deliver in the future. I don’t think we were ever going to see a path to sensible regulation that appeared downstream of sensible reasoning every step of the way. One could object to my tentative optimism here by pointing out that what is flowing between AI companies and the government is more noise than signal, but I have a little more faith in humanity than that. We shouldn’t rule out policymakers picking out the signal from this cacophony at some point.
I’ve been yapping on so much here that I’d forgotten about the purpose of this mini-blog series: to extract the “fable” from current events. The fable-adjacent frame I keep returning to here is forests for trees – which I am aware is more of a proverb than a fable, but such is the organic process of writing. So, to rather clumsily round out this metaphor, we’ve excluded the terrible ending where the US government fixates on Fable’s particular capabilities, but we still risk an almost-as-bad ending where it plays jailbreak-wackamole (albeit with models from multiple developers) up until the singularity, while overlooking the threat of an internal intelligence explosion. Their purview has widened beyond a single tree to include several, but there’s still a lot of forest out there.


