TL;DR
Reflection AI unveiled Beam on October 5: a 501-billion-parameter open-weight model (23 billion active per token) that the company says matches China's GLM-5.2 on advanced reasoning while using 3-4x less inference compute. The catch: the numbers are company-reported, the weights aren't public yet, and an Apache 2.0 release is promised later this month.

What Beam is
Reflection AI is a two-year-old Brooklyn startup founded in March 2024 by former Google DeepMind researchers Misha Laskin (CEO) and Ioannis Antonoglou (president and CTO, an AlphaGo co-creator). Beam is its first frontier model and the first American open-weight challenger aimed squarely at the Chinese labs that have dominated open-model downloads this year.
Beam is a text-only mixture-of-experts model: 501 billion total parameters with only 23 billion active per token, pretrained on 23.8 trillion tokens with a 1-million-token context window. Reflection says it trained the model end-to-end from scratch and will release the weights under the Apache 2.0 license.
The company is well capitalized for the fight: TechCrunch reports it has raised roughly $4.7 billion per PitchBook, backed by Nvidia, Sequoia, and Lightspeed, and it signed more than $7 billion in compute deals this summer with SpaceX and Nebius to lock up Nvidia GB300 capacity through 2029.
The efficiency claim
Reflection's headline isn't that Beam beats China's best outright — it doesn't claim that. The claim is that Beam matches Z.ai's GLM-5.2 (744 billion total parameters, 40 billion active) on advanced reasoning benchmarks at roughly one-quarter to one-third the inference compute, and beats today's leading Western open models while doing it.
The company-published scores: Terminal-Bench v2.1 at 80.1 vs. GLM-5.2's 81.0, SWE-Bench Pro v1 at 65.5 vs. 62.1, GPQA Diamond at 90.5 vs. 91.2, and Humanity's Last Exam (no tools) at 36.2 vs. 40.5. Reflection credits the gap to high-compute reinforcement learning, including training with a length penalty that rewards reaching correct answers in fewer tokens.
For anyone paying inference bills, this is the number that matters more than the leaderboard position: every point of reasoning quality held while burning a quarter of the tokens is margin for whoever runs the model.
The fine print
Every score above is company-reported and has not been independently verified — TechCrunch notes this explicitly, and the weights are not yet public, so nobody outside Reflection's early-access group can check. Artificial Analysis was given access and says early indicators suggest Beam will be one of the most token-efficient open models it has seen, but no independent evaluation has been published.
Reflection's own tables also show the honest part: newer Chinese models — GLM-5.3 and Moonshot's Kimi K3 — are ahead of Beam on most tests where both report scores, and the company acknowledges Kimi K3 remains ahead on raw capability. Beam's play is efficiency, not the crown.
And it isn't downloadable yet. The model is still in final red-teaming; early access runs through a waitlist, with the open weights, technical report, and model card promised later this month.
What to do with this news
Don't plan production around an unreleased model. But if you're evaluating open models for coding or agentic work — especially where the model must run on infrastructure you control — build your evaluation harness now against a current Qwen or DeepSeek release so you can drop Beam in and test it the day the weights ship.
The bigger trend is the pitch itself. As CEO Misha Laskin told Semafor, organizations that want sovereign AI systems they can inspect, adapt, and operate themselves “don't really have very good options today.” Beam is aimed at enterprises and governments that want exactly that, without relying on Chinese labs. The race is shifting from raw benchmark scores to scores per token — and that's the race worth watching.