What happened to superalignment after OpenAI disbanded its team

OpenAI's Superalignment team launched in 2023 with a four-year deadline and a 20% compute pledge, then dissolved in 2024 amid resignations and claims the compute never arrived. The research outlived the team. Here is where it went, and whether it is progress or spin.

By Zain

Published

A complex network of cables in a data center with a monitor in the foreground
Photo: panumas nikhomkhai / Pexels

What superalignment was supposed to be

On July 5, 2023, OpenAI announced a team with an unusually concrete promise. It said superintelligent AI could arrive within the decade, and that humanity had no reliable way to steer it. To close that gap the company gave itself four years and committed 20 percent of the compute it had secured to that point. Ilya Sutskever, OpenAI's chief scientist, and Jan Leike, its alignment lead, would run the effort together.

The framing mattered. OpenAI was not claiming to have solved anything. It argued that today's main method, reinforcement learning from human feedback, would break down once models outpaced the people grading them, because a human cannot reliably judge work they do not understand. The proposed fix was to build a roughly human-level "automated alignment researcher" and use it to help supervise systems beyond human level. The compute pledge was the part that made outsiders believe the company meant it.

Why the team fell apart

Less than a year later the structure collapsed. Sutskever and Leike both resigned in May 2024, and OpenAI confirmed the Superalignment team would no longer exist as a standalone unit. Leike was blunt about why. In a public thread he wrote that at OpenAI "safety culture and processes have taken a backseat to shiny products," that his group had been "sailing against the wind," and that "sometimes we were struggling for compute."

That last point cut closest to the founding promise. Reporting by Fortune, citing six people familiar with the team, found the 20 percent compute commitment was effectively never honored and that requests were repeatedly denied. The pledge that had signaled seriousness turned out to be the thing that quietly failed. Sutskever's exit was bound up in the November 2023 board fight that briefly ousted Sam Altman, but Leike's complaint was narrower and harder to wave away: the resources never matched the rhetoric.

“The pledge that had signaled seriousness turned out to be the thing that quietly failed.”

Detailed view of a magnifying glass examining a small circuit board with soldering tools nearby
Interpretability research tries to read what a model represents internally rather than trusting its outputs. Photo: https://kaboompics.com/ / Pexels

Where the work actually went

Alignment research at OpenAI did not stop, but it stopped being a single team with its own budget line. The company folded the work into other groups and said co-founder John Schulman would oversee alignment rather than run a dedicated team. In practice the specialists scattered.

They scattered productively, if you are an optimist. Leike joined Anthropic within weeks to continue similar research. Sutskever launched Safe Superintelligence Inc. in June 2024, a lab whose entire pitch is that it will build one thing, safe superintelligence, and ship no products until it does; by 2025 reports valued it at roughly 32 billion dollars with nothing on the market. Anthropic, meanwhile, had already made interpretability and its Responsible Scaling Policy central to its identity. OpenAI has since pointed to new internal oversight structures, but it has not reconstituted a single team with a public compute commitment of the original kind. So the honest answer to "where did superalignment go" is uncomfortable: the branded team is gone, and the agenda now lives across at least three organizations with different incentives.

The techniques people are really trying

Beneath the personnel drama, the technical menu is real and specific. Scalable oversight is the umbrella: methods that let weaker supervisors, including humans, check stronger systems. The main candidates are debate, where two models argue so a judge can catch the weaker case, and recursive critique, where models critique other models' critiques. OpenAI's most concrete contribution was weak-to-strong generalization, published in December 2023. It tested whether a GPT-2-level model could supervise GPT-4, and found it partly could: the strong model recovered close to GPT-3.5-level performance even when its teacher was far weaker.

Mechanistic interpretability is the other major bet, and it is largely Anthropic's. In May 2024 the company reported extracting millions of interpretable "features" from Claude 3 Sonnet using sparse autoencoders, and showed that turning a feature up or down changed the model's behavior. It was the first detailed look inside a production model of that scale. The appeal is obvious: if you can read what a model represents, you do not have to take its outputs on trust.

Progress or marketing?

The fair read is that these are early results, and their authors mostly say so. Weak-to-strong generalization recovers only part of the gap between weak and strong supervision, and researchers studying recursive oversight keep finding success rates well below 100 percent once the capability gap is large. No number of critique layers fully closes it. The old worry about reward hacking, where a model optimizes the proxy it is graded on rather than the thing you wanted, has not gone away either.

Interpretability carries the same caution in its own fine print. Anthropic noted the features it found are "a small subset" of what the model has learned, that cataloging all of them could cost more than training the model, and that "understanding the representations the model uses doesn't tell us how it uses them." That is a long way from certifying a system as safe.

So what should a reader conclude? Superalignment as a crisp, well-funded project with a deadline did not survive contact with commercial priorities. The questions it raised are legitimate and are still being worked on, now in more places and with less transparency about resources. Treat the specific results, weak-to-strong transfer, feature extraction, debate, as genuine progress on narrow pieces of the problem. Treat any claim that the whole problem is nearly solved as marketing, at least until a lab shows its work and its compute.

Frequently asked questions

Why did OpenAI's Superalignment team shut down?

It was disbanded in May 2024 after co-leads Jan Leike and Ilya Sutskever resigned. Leike said safety had taken "a backseat to shiny products," and Fortune reported that the promised 20% of compute was never delivered to the team.

What is weak-to-strong generalization?

An OpenAI method published in December 2023 that tests whether a weaker model can supervise a stronger one. A GPT-2-level model was used to elicit close to GPT-3.5-level performance from GPT-4, a partial result rather than a full solution.

What happened to Jan Leike and Ilya Sutskever?

Leike joined Anthropic in May 2024 to continue alignment research. Sutskever co-founded Safe Superintelligence Inc. in June 2024, a lab reported at roughly a 32 billion dollar valuation by 2025 with no product on the market.

What is mechanistic interpretability?

Research that reverse-engineers a model's internals. In May 2024 Anthropic used sparse autoencoders to extract millions of interpretable features from Claude 3 Sonnet, though it stressed these were only a small subset of what the model had learned.

Sources

  1. 2

    Weak-to-strong generalization, OpenAI (December 13, 2023)