In short
Reflection has announced a model trained from scratch in the US and promises to release its weights under Apache 2.0. For those choosing a model for coding and agents, the key question isn’t just the benchmarks: will it be possible to independently verify the claimed results?
American developers are getting another major open model, but for now the announcement looks more like a statement saying “we’re from the US and we’re doing everything ourselves” than a bid for leadership. Reflection says Beam scored 80.9 on SWE-bench Verified, but next to the leading models, that doesn’t look particularly convincing.
Architecturally, Beam is an MoE model: 501 billion total parameters, with 23 billion active. It is aimed at coding, agents, and scientific tasks. The company promises to release the full weights under Apache 2.0 later this month.
There is one practical benefit here: if the weights actually appear, it will be possible to run the model yourself and test it on your own tasks without being tied to someone else’s API. For those who value an open license and an origin trained from scratch in the US, this expands the options. But the origin itself says nothing about whether Beam will handle a particular project.
The weak points are already obvious. In reviews, Beam is placed roughly at the level of GLM 5.2, while some consider it weaker than DeepSeek V4 Flash on certain benchmarks. Reflection claims high inference efficiency, but there has been neither independent verification nor the promised technical report so far. So the figures in the announcement are best viewed as a reason to pay attention, not as proof that the model is better than the rest.
Would you choose a model with open weights and verifiable origins if it loses to the benchmark leaders on your tasks?
Source: Latent.Space