AI safety jobs and salaries: what Anthropic, OpenAI and the UK AI Security Institute actually post
Posted pay for alignment, interpretability and safeguards roles at Anthropic and OpenAI runs from about $190,000 to $850,000 in base salary, as of October 9, 2026. Here are the roles, the bands and what they leave out.
By Yash Malviya
Published

People searching for AI safety jobs and salaries tend to find aggregator estimates built from self-reports. The better evidence is on the labs' own job boards, where pay transparency laws push companies to post ranges. We pulled the live postings from Anthropic's and OpenAI's job boards on October 9, 2026 and read the UK AI Security Institute's careers page. The numbers below are posted base salary ranges, not what any person is paid, and they will change.
The top bands: research roles
At Anthropic, the research roles with a safety focus post the highest ranges. Research Engineer or Scientist, Alignment lists $350,000 to $500,000. Research Engineer, Interpretability lists $315,000 to $560,000. Research Scientist, Interpretability lists $350,000 to $850,000, the widest in this set, which suggests the top of the band is reserved for senior hires. Research Scientist or Engineer, Biological Safety lists $300,000 to $405,000. All are in San Francisco, and the postings carry USD ranges labeled as annual salary.
OpenAI's board shows a similar shape. Researcher, Alignment lists $295,000 to $500,000, and so do the Alignment CoT Monitorability and Alignment Interpretability roles. Researcher, Interpretability, Researcher, Safety Oversight, Researcher, Multimodal Safety and Researcher, Recursive Self-Improvement Safety each list $380,000 to $500,000. Every one of these OpenAI postings says it offers equity, which is where total compensation can pull well above the base.
The middle: engineering, red teams and data
Below the research scientists sit the engineers and analysts who build safety systems. Anthropic lists ML or Research Engineer, Safeguards at $350,000 to $500,000, Red Team Engineer, Safeguards at $320,000 to $405,000, and Data Scientist, Safeguards at $285,000 to $380,000. A Staff Plus Software Engineer for Safeguards Infrastructure in London lists £325,000 to £395,000.
OpenAI lists Software Engineer, AI Safety at $207,000 to $385,000, Data Scientist, Preparedness at $345,000 to $385,000, Data Scientist, Safety at $265,000 to $380,000, and Red Team Specialist, Cyber at $198,000 to $320,000. The overlap between the two labs is large for comparable jobs, which is what you would expect in a market where candidates hold offers from both.

The lower bands: policy, operations and enforcement
Many of the safety jobs are not research. Anthropic lists Safeguards Enforcement Analyst, Cyber Harm at $285,000 to $330,000, and Head of Policy Design, Societal Harms at $330,000 to $395,000. OpenAI lists a Safety Transparency Editor at $220,000 to $245,000, Model Policy Manager, Agentic Safety at $207,000 to $335,000, and User Safety and Risk Operations analyst roles at $189,000 to $280,000. They sit below the top research ranges.
The practical message for a career changer is that the label safety covers three very different families: research, engineering, and policy or operations. The first has the highest pay and the steepest requirements. The latter two may suit people without a machine learning research record, though the postings should be read role by role.
What the UK AI Security Institute offers
Government pay works differently. The UK AI Security Institute's careers page says technical job advertisements show a base salary and a technical talent allowance, and that the allowance counts as part of total pay. It states that the institute contributes 28.97 percent of base salary to pensions, and lists benefits including 25 days of annual leave and eight public holidays. On October 9, 2026 the page said there were no open roles, so we cannot report a current band. Our earlier explainer on the UK and US institutes covers what each tests. The tradeoff for candidates is clear: lower headline cash than the labs, plus public-sector pension and a mandate to evaluate models independently.
What these numbers leave out
Posted ranges are not offers. Companies post ranges that cover several seniority levels, and where in the range a person lands depends on level and negotiation. Equity is the biggest unknown. OpenAI's postings note equity is offered but do not value it, and the Anthropic ranges we read are labeled annual salary and do not value any equity. Aggregators that publish total compensation estimates are using other data and often disagree with each other, so we have not used them.
Two other limits apply. Postings change weekly, and the Anthropic postings we read carry update dates between August and early October 2026. And the boards list only open roles. A lab can hire into a safety team through an internal move or an expression of interest without a posting.
How to read this if you are applying
Start with the family of job you want. If it is research, expect a track record in machine learning, interpretability or evaluations. If it is policy or operations, look for roles such as policy design or enforcement analyst. Our explainers on what AI safety evaluations test and miss and on superalignment give the vocabulary the postings assume.
A note on method
We pulled every posting from each board's public API on October 9, 2026 and filtered titles for safety, alignment, interpretability, red team, preparedness, safeguards and trust. We read the pay range stated in each posting and report it as written, converting nothing. Currency is as posted, so the London Anthropic role is in pounds. We did not average ranges, because a midpoint of a band that spans several seniority levels does not describe any real offer, and we did not include roles whose postings listed no range.
Our take
The posted numbers support the claim that frontier labs pay a premium for alignment and interpretability research, with base salaries reaching the mid six figures and, at the top of one Anthropic band, $850,000. They also show that a safety job is not one job. We have no on-the-record quote for this piece, because the evidence is the posted ranges themselves. The fair conclusion is narrower than the hype: well paid, concentrated in research, and dependent on equity we cannot see.
Frequently asked questions
How much do AI safety jobs pay?
At the labs we checked on October 9, 2026, posted base salaries for research roles range from about $295,000 to $850,000, depending on the lab and role. Enforcement, policy and operations roles post lower, around $189,000 to $395,000.
How much does Anthropic pay an alignment researcher?
Anthropic's posting for Research Engineer or Scientist, Alignment lists $350,000 to $500,000 in annual salary. Its Research Scientist, Interpretability role lists $350,000 to $850,000.
How much does OpenAI pay safety researchers?
OpenAI's postings list $295,000 to $500,000 for Researcher, Alignment and $380,000 to $500,000 for roles such as Safety Oversight and Interpretability. All of these postings say they offer equity.
Does the UK AI Security Institute pay as much as the labs?
We could not compare directly because its careers page listed no open roles on October 9, 2026. It pays a base salary plus a technical talent allowance and contributes 28.97 percent of base to pensions.
Are these posted salaries what people actually earn?
No. They are ranges covering multiple levels, and equity and bonuses are not quantified in the postings. Individual pay depends on level and negotiation.
Do you need a PhD to get an AI safety job?
The postings we read do not answer that for all roles. Research positions emphasize ML or interpretability experience, while many policy, enforcement and operations roles are separate career paths.
Sources
What each one is, and whose it is.
- 1
Anthropic job board, public postings API, Anthropic (October 9, 2026)
OtherThe vendor’s own - 2
OpenAI job board, public postings API with compensation, OpenAI (October 9, 2026)
OtherThe vendor’s own - 3
Careers at the AI Security Institute, UK AI Security Institute (October 9, 2026)
Documentation