← All insights

News analysisLens: United States3 min read

Reflection's Beam is a 501B open-weight model. With open weights, the operating duty moves to you

Reflection announced Beam, a 501B-parameter mixture-of-experts model for coding and agents under Apache 2.0. Weights arrive later this month, after final red-teaming. Plan for what you inherit.

Listen to this article · 5 min

AI-generated narration of the full article.

The blades of a jet turbine seen from the front, in La Madre duotone, beside the words You run the model
Photo: rawpixel (CC0)

On October 5, Reflection introduced Beam, its first open-weight model, built for coding, reasoning and agentic work. If its claims hold, enterprises get another capable model they can run on their own infrastructure. Before anyone downloads it, it is worth being clear about what open weights actually transfer: not just control, but the operating duties a model API used to hide.

What Reflection announced, and what it has not shipped yet

  • Architecture. A sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token. Context of 256K tokens during reinforcement learning, extended to 1 million through midtraining.
  • Training. Pretraining on 23.8 trillion tokens on 6,144 NVIDIA GB300 GPUs in under four weeks; a reinforcement learning run with more than 100 million rollouts on about 10,500 GB300 GPUs over four weeks.
  • License. Apache 2.0, a permissive license that allows commercial use.
  • Status. Beam is undergoing final red-teaming and evaluations. Reflection says it will release the weights, technical report, model card and developer artifacts later this month. Today there is a waitlist for early access on Reflection’s platform.
  • Benchmarks. Reflection reports scores such as 77.2% on SWE Bench Pro v2-Hard and 90.5% on GPQA Diamond. These are self-reported until independent evaluations appear.

So the honest status today is an announcement with a license and a date, not a downloadable model. That is useful: it gives platform teams a few weeks to prepare.

What moves to you when you run the weights

Who owns each duty: model API vs open weights you run01Servingcapacity anduptime02Securitypatches andupgrades03Safetytesting foryour use04Evaluation onyour tasks05Data boundaryand retention06Cost peracceptedoutputYours either way
  1. Serving capacity and uptime
  2. Security patches and upgrades
  3. Safety testing for your use
  4. Evaluation on your tasks
  5. Data boundary and retention
  6. Cost per accepted output

Yours either way

Open weights do not add new duties so much as remove the provider who was quietly doing half of them.

Capacity is the first surprise. “23 billion active” describes compute per token, not memory. All 501 billion parameters must be loaded to serve the model: about 500 GB of weights at 8 bits per parameter, about a terabyte at 16 bits, before the memory needed for long contexts. That is a multi-GPU server at minimum, and production traffic with failover means several. Plan capacity from the total parameter count, not the active one.

Upgrades become your release. With an API, the provider ships fixes and new versions on its schedule. With weights, nothing changes until you change it, which is good for stability and bad if a safety issue is found later. Someone has to watch for new versions, re-run your evaluations and roll out the update, as with any model change.

Safety testing is partly yours. Reflection is finishing its red-teaming before release. It cannot test your agent’s tools, data and permissions. A coding agent that can run commands needs your own adversarial testing in your environment.

The data boundary is the payoff. Prompts, code and outputs never leave your infrastructure. For regulated U.S. sectors (health, defense contractors, financial services) and for proprietary codebases, that can justify the work. It is the same custody logic we described for self-hosted coding agents.

Where Beam could fit

Not every task needs a frontier generalist, and not every model needs to be the biggest. In our framework for a model portfolio, an open-weight model like Beam would be a candidate for the roles where custody or cost at volume matter more than having the absolute best answer: internal coding assistance on sensitive repositories, or agent steps that run thousands of times a day. Only your own evaluation set can decide whether it earns that role.

What to do now

  1. Prepare an evaluation set from your real coding and agent tasks before the weights arrive.
  2. Size capacity from total parameters, including failover and long-context memory.
  3. Read the model card and technical report when they are published, before any pilot.
  4. Verify the weights you download against the publisher’s checksums, and store them as a controlled artifact.
  5. Plan your own red-teaming for the tools and permissions your agents will have.
  6. Compare cost per accepted output against your current API, not cost per token.

The bottom line

Beam may add a strong, permissively licensed option to enterprise model portfolios, but it is not shipping weights yet, and its benchmarks are its own. Use the weeks before release to decide whether you want the duties that come with running a model, because with open weights, they are yours.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case