Reflection AI Beam open-weight model: a $20 billion lab finally names its first model
Reflection AI announced Beam on October 5, a 501 billion parameter open-weight model, with weights promised later this month. The scores are self-reported and trail Qwen's best. The release has to prove far more than that.
Published

Reflection AI Beam is the first model from a lab that has raised billions and shipped nothing. On October 5, 2026, Reflection announced Beam, a 501 billion parameter open-weight model, and said the weights will arrive under an Apache 2.0 license "later this month." Until they do, everything below is the company grading its own homework.
The announcement matters because Reflection is the best-funded American bet on a specific idea: that Western companies and governments want an open model they can inspect and run themselves, and that Chinese labs should not own that category. Beam is the first test of whether the bet produces anything usable.
What Reflection actually announced
According to its own Beam post, the model is a sparse mixture-of-experts design with 501 billion total parameters and 23 billion active per token, aimed at coding, reasoning and agentic work. It was pretrained on 23.8 trillion tokens drawn from the web and licensed proprietary data, with a context window of 256,000 tokens. Reinforcement learning used more than 100 million rollouts on 10,500 Nvidia GB300 GPUs over four weeks.
The post lists four benchmark scores: 77.2 on SWE Bench Pro v2-Hard, 80.1 on Terminal Bench v2.1, 97.8 on AIME 2026 and 90.5 on GPQA Diamond. It calls Beam competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks, with three to four times less inference compute than GLM 5.2 for comparable scores.
Note what is missing. As of October 6, we could not find a model card, a technical report or downloadable weights. Those are promised for later in the month. The benchmark numbers are vendor-run, and no outside lab has reproduced them. A claim of "approaching" Qwen's top model also concedes the point: Beam is, by Reflection's own wording, behind.
“DeepSeek and Qwen and all these models are our wake-up call because if we don't do anything about it, then effectively, the global standard of intelligence will be built by someone else.”
How a $20 billion lab got here
Reflection raised $2 billion at an $8 billion valuation in October 2025, with Nvidia leading through an $800 million investment, according to Sherwood News. At the time CEO Misha Laskin told TechCrunch that DeepSeek and Qwen were "a wake-up call." The Turing Post profile reported that the company was raising another $2 billion at roughly a $20 billion valuation in March 2026, and noted it had no frontier model out and no published research papers.
The first release target was early 2026. It slipped. Beam arrives roughly nine months after that window, which is the kind of delay that matters when Chinese labs ship new generations every few months.
Compute is not the constraint. Mobile World Live, citing reporting from CNBC and Bloomberg, says Reflection agreed in June 2026 to pay SpaceX $150 million a month from July 1 through 2029 for GB300 capacity at the Colossus 2 site in Memphis, about $6.3 billion if the full term runs. Either side can exit on 90 days' notice after the first three months. We did not open the underlying contract, so treat those terms as reported. That is a large fixed cost for a company with no revenue we could find disclosed.

What the release has to prove
A lab at this valuation does not get credit for matching last quarter's best open model. Four things decide whether Beam matters.
- The weights exist and are open. Apache 2.0 is the most permissive mainstream license. If the release comes with usage carve-outs or arrives late, the positioning collapses.
- The numbers reproduce. Independent evaluators need to confirm the SWE Bench and Terminal Bench scores without Reflection's own scaffolding. Agentic benchmarks are especially sensitive to harness choices.
- The efficiency claim holds in deployment. Three to four times less inference compute than GLM 5.2 is the commercial pitch to enterprises. It is also the easiest claim to test, since anyone with GPUs can measure it.
- The safety testing is shown. The company's site says the models are rigorously tested for safety. A 501 billion parameter model with strong coding and agentic scores can be fine-tuned by anyone, so we want to see the evaluations, not the adjective. Our piece on an open Chinese model that trails the US frontier by four months shows why cyber capability in downloadable weights is a live question.
The open-weight gap
The competitive picture is the real story. A July 2026 Pure AI column framed the question as whether the US will produce a leading independent open-weight company, pointing to DeepSeek and Moonshot releasing increasingly capable models while Western options such as Meta's Llama family and France's Mistral lag. Reflection's own comparison set for Beam is GLM and Qwen, both Chinese. No American model is the yardstick, which tells you where the frontier of open weights currently sits.
Reflection says it will make money from large enterprises and governments building sovereign AI systems, per TechCrunch's 2025 reporting. That is a different product from a benchmark leaderboard. It needs support contracts, deployment tooling and trust, and the company's site does list air-gapped and on-premises deployment. Whether buyers choose a new, unproven American vendor over a cheaper, stronger Chinese model is a procurement question that benchmarks will not settle.
What we could not verify
We could not open the Axios report that prompted this coverage, which returned an access error, so we rely on Reflection's own post for the announcement. The $20 billion valuation comes from a single newsletter profile, not a filing. We found no on-record quote from Reflection executives tied to the Beam launch. Nothing here confirms Genesis Mission or other federal involvement, and we are not implying any.
Our take
Announcing a model with weights still weeks away is a press cycle, not a release. Beam may well be good: the efficiency claim, if it survives outside testing, is the most interesting part. But a $20 billion valuation and a $6.3 billion compute commitment demand more than a benchmark table that concedes second place to Qwen. Judge it when the weights, the model card and independent evaluations land, and not before.
Frequently asked questions
What is the Reflection AI Beam open-weight model?
Beam is a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active, announced by Reflection AI on October 5, 2026. It targets coding, reasoning and agentic work. Reflection says weights will be released under Apache 2.0 later in October.
Are Beam's weights available to download yet?
Not as of October 6, 2026. Reflection's post says weights, a technical report, a model card and developer materials are coming later this month. Early access sign-up is open.
How does Beam compare with Qwen and GLM?
Reflection says Beam is competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks. These are the company's own benchmark results, and no independent reproduction exists yet.
How much has Reflection AI raised and who backs it?
It raised $2 billion at an $8 billion valuation in October 2025, led by Nvidia with $800 million. A Turing Post profile reported a further raise of about $2 billion at roughly $20 billion in March 2026; we have not seen a filing confirming that.
What is the Reflection and SpaceX Colossus 2 deal?
Reportedly, Reflection pays SpaceX $150 million a month from July 1, 2026 through 2029 for Nvidia GB300 capacity at Colossus 2 in Memphis, about $6.3 billion in total. Either party can terminate on 90 days' notice after three months. The terms come from press reports.
What does Beam need to prove?
That the Apache 2.0 weights really ship, that independent evaluators reproduce the benchmark scores, that the claimed 3 to 4 times lower inference compute versus GLM 5.2 holds in deployment, and that safety testing is published rather than asserted.
Sources
What each one is, and whose it is.
- 1
Introducing Beam, Reflection AI (October 4, 2026)
Vendor announcement - 2
Reflection AI homepage, Reflection AI (October 4, 2026)
Vendor announcement - 3
SpaceX inks computing deal with Reflection AI, Mobile World Live (June 22, 2026)
Press reportIndependent of the vendor - 4
Inside Reflection AI: The $20B Open-Model Startup That Has Yet to Ship, Turing Post (February 28, 2026)
Press reportIndependent of the vendor - 5
Reflection raises $2B to be America's open frontier AI lab, challenging DeepSeek, TechCrunch (October 8, 2025)
Press reportIndependent of the vendor - 6
Nvidia backs Reflection AI in $2 billion fundraising round, Sherwood News (October 8, 2025)
Press reportIndependent of the vendor - 7
America is looking for an open weight AI champion, Pure AI (July 28, 2026)
Press reportIndependent of the vendor