Reflection's Beam is a 501B open-weight model. With open weights, the operating duty moves to you
Reflection announced Beam, a 501B-parameter mixture-of-experts model for coding and agents under Apache 2.0. Weights arrive later this month, after final red-teaming. Plan for what you inherit.
Listen to this article · 5 min
AI-generated narration of the full article.

On October 5, Reflection introduced Beam, its first open-weight model, built for coding, reasoning and agentic work. If its claims hold, enterprises get another capable model they can run on their own infrastructure. Before anyone downloads it, it is worth being clear about what open weights actually transfer: not just control, but the operating duties a model API used to hide.
What Reflection announced, and what it has not shipped yet
- Architecture. A sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token. Context of 256K tokens during reinforcement learning, extended to 1 million through midtraining.
- Training. Pretraining on 23.8 trillion tokens on 6,144 NVIDIA GB300 GPUs in under four weeks; a reinforcement learning run with more than 100 million rollouts on about 10,500 GB300 GPUs over four weeks.
- License. Apache 2.0, a permissive license that allows commercial use.
- Status. Beam is undergoing final red-teaming and evaluations. Reflection says it will release the weights, technical report, model card and developer artifacts later this month. Today there is a waitlist for early access on Reflection’s platform.
- Benchmarks. Reflection reports scores such as 77.2% on SWE Bench Pro v2-Hard and 90.5% on GPQA Diamond. These are self-reported until independent evaluations appear.
So the honest status today is an announcement with a license and a date, not a downloadable model. That is useful: it gives platform teams a few weeks to prepare.
What moves to you when you run the weights
- Serving capacity and uptime
- Security patches and upgrades
- Safety testing for your use
- Evaluation on your tasks
- Data boundary and retention
- Cost per accepted output
Yours either way
Capacity is the first surprise. “23 billion active” describes compute per token, not memory. All 501 billion parameters must be loaded to serve the model: about 500 GB of weights at 8 bits per parameter, about a terabyte at 16 bits, before the memory needed for long contexts. That is a multi-GPU server at minimum, and production traffic with failover means several. Plan capacity from the total parameter count, not the active one.
Upgrades become your release. With an API, the provider ships fixes and new versions on its schedule. With weights, nothing changes until you change it, which is good for stability and bad if a safety issue is found later. Someone has to watch for new versions, re-run your evaluations and roll out the update, as with any model change.
Safety testing is partly yours. Reflection is finishing its red-teaming before release. It cannot test your agent’s tools, data and permissions. A coding agent that can run commands needs your own adversarial testing in your environment.
The data boundary is the payoff. Prompts, code and outputs never leave your infrastructure. For regulated U.S. sectors (health, defense contractors, financial services) and for proprietary codebases, that can justify the work. It is the same custody logic we described for self-hosted coding agents.
Where Beam could fit
Not every task needs a frontier generalist, and not every model needs to be the biggest. In our framework for a model portfolio, an open-weight model like Beam would be a candidate for the roles where custody or cost at volume matter more than having the absolute best answer: internal coding assistance on sensitive repositories, or agent steps that run thousands of times a day. Only your own evaluation set can decide whether it earns that role.
What to do now
- Prepare an evaluation set from your real coding and agent tasks before the weights arrive.
- Size capacity from total parameters, including failover and long-context memory.
- Read the model card and technical report when they are published, before any pilot.
- Verify the weights you download against the publisher’s checksums, and store them as a controlled artifact.
- Plan your own red-teaming for the tools and permissions your agents will have.
- Compare cost per accepted output against your current API, not cost per token.
The bottom line
Beam may add a strong, permissively licensed option to enterprise model portfolios, but it is not shipping weights yet, and its benchmarks are its own. Use the weeks before release to decide whether you want the duties that come with running a model, because with open weights, they are yours.