Reflection Beam 501B Model Review
Key Takeaways:
- Beam is a 501B-parameter sparse MoE with 23B active parameters, a 1M-token context window, and Apache 2.0 weights scheduled for later in October 2026.
- It was pretrained on 23.8 trillion tokens on 6,144 NVIDIA GB300 GPUs in under four weeks, then trained with high-compute RL on 10,500 GB300 GPUs for four weeks.
- Reflection reports 80.9 on SWE-bench Verified and 80.1 on Terminal Bench v2.1, but its own benchmark table shows Chinese models winning most rows.
- The “3 to 4x less inference compute” claim is an estimate that excludes prompt prefill, attention operations, and serving overhead.
- The model is still in final red-teaming; no independent evaluation has reproduced any reported score.
What Reflection Actually Shipped
Reflection AI announced Beam on October 5, 2026, its first open-weight model: a 501-billion-parameter sparse Mixture-of-Experts system with 23 billion active parameters per token, built for coding, reasoning, and agentic workloads. The Brooklyn-based startup, founded in 2024 by two former Google DeepMind researchers, published the announcement on its company blog.

The headline numbers are the training runs. Beam was pretrained on 23.8 trillion curated tokens on 6,144 NVIDIA GB300 NVL72 GPUs in under four weeks, reaching 92.3% goodput by the end. Reinforcement learning then consumed 10,500 GB300 GPUs for four weeks, generating more than 100 million rollouts against roughly 1.3 billion sandboxes at a maximum context length of 256,000 tokens. Reflection describes it as one of the largest RL runs any open lab has conducted.
Beam is text-only. Midtraining extended its effective context window to one million tokens. The weights, technical report, and model card are scheduled for release later in October 2026 under the Apache 2.0 license, with early access currently running through a waitlist while the model finishes red-teaming.
The Sparse MoE Architecture and the Efficiency Claim
The design that makes Beam interesting is its activation ratio. All 501 billion parameters must be held in memory, but only 23 billion are used per token, roughly 4.6% of the model. That gap lets a model with frontier-scale capacity run at a fraction of the compute per token.
The model interleaves local and global attention with fine-grained routed experts. Load balancing builds on the auxiliary-loss-free balancing method from DeepSeek-V3, adding cosine decay of expert-bias updates. Reflection reports the busiest expert reached just 1.04x average load at the end of pretraining, and residual norms stayed bounded across all 52 layers using depth-based scaling, SandwichNorm, attention gating, and FP32 residual accumulation.
The data pipeline is where the efficiency argument starts. Reflection says curation removed about 95% of raw internet tokens through parsing, deduplication, and filtering, while retaining roughly 1.8 trillion high-quality tokens that conventional filters would have dropped, including 87% of its curated web-code tokens. A separate pipeline processed PDF artifacts using a vision-language OCR model with in-house quality classifiers.
The reinforcement learning stage is the more unusual investment. Reflection trained Beam with fully asynchronous policy gradients, tagging every token with the policy version that produced it, and reports stable learning even when training on samples 107 weight versions behind the current policy, roughly one day of staleness. New weights reached the inference fleet in a median of about 12 seconds, and 71 inference incidents were handled without stopping the training job. A controllable length penalty rewards successful solutions while discouraging unnecessary tokens, and users control the tradeoff through a reasoning-effort parameter.
The efficiency claim rests on a specific formula. Reflection estimates generation forward-pass compute as roughly twice the active parameter count multiplied by the mean generated tokens per attempt, counting each multiply-add as two operations:
# Reflection's approximate compute estimate for generation
# FLOPs ~ 2 x active_params x mean_generated_tokens (per attempt)
active_params = 23_000_000_000 # 23B active per token
mean_tokens = 1_000 # mean generated tokens per attempt
estimated_flops = 2 * active_params * mean_tokens
print(f"{estimated_flops:.2e} FLOPs per attempt")
# Note: excludes prompt prefill, attention operations, and serving
# overhead. An approximation, not a measured inference cost.
Reflection states plainly that the figure “excludes prompt prefill, context-dependent attention operations, and serving overhead,” so it is an approximate compute comparison rather than a measured inference cost.
Benchmarks: Where Beam Leads and Where It Trails
Reflection’s own benchmark table is the most useful document here, because it is unusually candid about where Beam loses. On SWE-bench Verified, Beam scores 80.9 against 70.7 for NVIDIA’s Nemotron 3 Ultra. On SWE-bench Multilingual it scores 78.0 against 67.7 for Nemotron 3 Ultra. Those are the two rows where Beam leads.
Most other rows go to competitors. On Terminal Bench v2.1, Beam’s 80.1 trails GLM 5.2 at 81.0, Kimi K3 at 88.3, and DeepSeek V4.1 Flash at 90.6. On SWE Bench Pro v2-Hard, Beam’s 77.2 trails Kimi K3 at 88.2. On DeepSWE v1.1, Beam’s 44.4 trails DeepSeek V4.1 Flash at 74.2. On SWE Atlas Codebase QnA, Beam’s 34.6 trails Kimi K3 at 68.0.
Reflection states the gap directly: “Where frontier open models like Kimi K3 remain ahead on raw capability, Beam’s advantage is efficiency at inference time.” Beam is being pitched as the best open model that did not come out of China, with a cost argument attached.
The reasoning rows are closer. Beam scores 97.8 on AIME 2026 against GLM 5.2’s 99.2, 90.5 on GPQA Diamond against GLM 5.2’s 91.2, and 36.2 on Humanity’s Last Exam without tools against GLM 5.2’s 40.5. On tool calling and search, Beam scores 78.7 on MCP Atlas against GLM 5.2’s 77.8, and 37.0 on AutomationBench public against GLM 5.2’s 26.2, though Kimi K3 leads the latter at 46.7.
# Beam vs GLM 5.2 on reasoning benchmarks (Reflection-reported)
benchmarks = {
"AIME 2026": {"Beam": 97.8, "GLM 5.2": 99.2},
"GPQA Diamond": {"Beam": 90.5, "GLM 5.2": 91.2},
"HLE (no tools)": {"Beam": 36.2, "GLM 5.2": 40.5},
"MCP Atlas": {"Beam": 78.7, "GLM 5.2": 77.8},
"Terminal Bench v2.1": {"Beam": 80.1, "GLM 5.2": 81.0},
}
for name, s in benchmarks.items():
print(f"{name:20s} Beam {s['Beam']:5.1f} GLM 5.2 {s['GLM 5.2']:5.1f}")
# Rival figures from Artificial Analysis and DataCurve; Beam scores
# are vendor claims pending the technical report later in October 2026.
How Beam Compares to the Open-Weight Field
The clearest way to place Beam is by total versus active parameters and license. Beam is the smallest model in its comparison set by total parameters, and its 23B active count is the lowest of the group. That is the mechanical basis for the efficiency pitch.
| Model | Developer | Total params | Active params | License | Terminal Bench v2.1 |
|---|---|---|---|---|---|
| Reflection Beam | Reflection AI (US) | 501B | 23B | Apache 2.0 (planned) | 80.1 |
| GLM-5.2 | Z.ai (China) | 753B | 40B | MIT | 81.0 |
| Nemotron 3 Ultra | NVIDIA (US) | 550B | 55B | OpenMDW-1.1 | 56.4 |
| DeepSeek V4.1 Flash | DeepSeek (China) | 552B | 8B prefill | MIT | 90.6 |
| Kimi K3 | Moonshot AI (China) | 2.8T | 104B | Kimi K3 License | 88.3 |
Source: MarkTechPost, citing Reflection’s published table and specs verified October 5, 2026. Every score comes from Reflection’s own announcement, which sources rival figures from Artificial Analysis and DataCurve.
Two things stand out. First, DeepSeek V4.1 Flash beats Beam on Terminal Bench with fewer active parameters per token during prefill, which complicates the simple “fewer active parameters means cheaper” story. Second, Reflection’s TechCrunch coverage notes that GLM-5.2 carries roughly 744 billion total parameters with 40 billion active, so Beam’s active-parameter advantage is real but narrower than the total-parameter gap suggests.
Deployment and the Efficiency Measurement Problem
The “3 to 4x less inference compute” claim needs its method stated, because the method is doing a lot of work. Reflection estimates generation forward-pass compute as roughly twice the active parameter count multiplied by the mean number of generated tokens per attempt, using Artificial Analysis and DataCurve data for rival models.
The company is explicit about what that estimate excludes: prompt prefill, context-dependent attention operations, and serving overhead. That matters for capacity planning, because prefill and attention costs scale with context length, and Beam’s headline feature is a one-million-token context window. A long-context workload will spend a larger share of its compute on the excluded terms than the estimate assumes.
What a developer can actually do today is limited. Beam is in final red-teaming, and early access runs through a waitlist. The weights are not yet on Hugging Face. The honest state of deployment is that the model is announced, the training details are published, and the artifact you would run is scheduled for later in October.
When the weights do land under Apache 2.0, the operational math will look like other 501B-class MoE deployments: the full parameter set must be held in memory even though only 23B are active per token, so the memory requirement is set by total size while the compute requirement is set by active size. That split is the whole point of the architecture, and it is also the reason a 501B model is not a single-GPU workload regardless of how few parameters fire per token.
Note: The following code is an illustrative example and has not been verified against official documentation. Please refer to the official docs for production-ready code.
# Memory-versus-compute split for a 501B MoE like Beam
# Memory is a function of TOTAL parameters; compute of ACTIVE parameters.
total_params = 501_000_000_000 # 501B total
bytes_per_param_fp8 = 1 # FP8 weights
memory_gb = total_params * bytes_per_param_fp8 / 1e9
print(f"Weights in memory (FP8): ~{memory_gb:.0f} GB")
# All 501B parameters reside in memory even though only 23B are used
# per token. Sparse MoE reduces compute, not memory footprint. Serving
# also adds KV cache, activations, and optimizer state.
Positioning, Licensing, and Sovereign Deployment
Reflection is positioning Beam against closed labs, against Chinese open models, and against Western open players like Mistral and Cohere. Its most direct US rival is Inkling, the open model from Mira Murati’s Thinking Machines Lab released in July 2026. Reflection’s own benchmarks show Beam outscores Inkling on four coding tests where both report results, but Inkling is multimodal and Beam is text-only, so the comparison is not apples to apples.
The Apache 2.0 commitment is the strategic centerpiece. MIT and Apache 2.0 are the permissive licenses most enterprises clear in legal review without bespoke negotiation, and Reflection is pairing the weights with what it describes as a full stack for running, evaluating, and fine-tuning the model, plus distribution partners and integrations with open-source libraries and harnesses.
The company’s commercial model leans on sovereign and institutional deployment rather than consumer subscriptions. Reflection’s about page states it is validated on Dell AI Factory with NVIDIA, is a certified consortium member of the Department of Energy’s Genesis Mission serving 17 U.S. National Laboratories with its open-weight models, and is building a 250-megawatt sovereign AI cloud in South Korea with Shinsegae, backed by the U.S. and Korean governments. TechCrunch reports the company has raised roughly $4.7 billion from backers including NVIDIA, Sequoia Capital, and Lightspeed, at a $25 billion pre-money valuation, and signed compute deals worth more than $7 billion with SpaceX and Nebius to secure GB300 access through 2029.
That capital structure is worth weighing alongside the technical claims. A lab that has locked multi-year compute contracts can reserve capacity across model launches, which provides an operational advantage for buyers who need service continuity. It also means the efficiency narrative is partly a capital narrative: the pitch to sovereign buyers is a model they can run on their own infrastructure, trained by a company with the compute relationships to keep shipping.
Limitations and Trade-offs
The most important limitation is that none of the performance numbers have been independently verified. Every benchmark in this article comes from Reflection’s announcement, and the company says so. The rival scores it compares against are sourced from Artificial Analysis and DataCurve, but Beam’s own scores have not been reproduced by a third party. Reflection’s plan to publish a technical report and model card alongside the weights is the point at which independent evaluation becomes possible.
The efficiency claim is an estimate, not a measurement, and it excludes the cost components that grow with context length. A developer running Beam at its full one-million-token context window should not expect the “3 to 4x” figure to hold, because prefill and attention operations, both excluded from the estimate, scale with sequence length.
Beam is text-only. Reflection handles other modalities only when something converts them to text first, which rules it out for workloads that need native image or audio input. That is a real gap against multimodal competitors, and a deliberate trade rather than an oversight, since the model was trained with a particular focus on coding and agentic tasks.
The model is still in final red-teaming. Reflection says safety evaluation results will appear in the technical report and that it will open-source the internal safety evaluations it developed, but neither is available yet. The alignment approach, a separate teacher model merged with the capability model through multi-teacher on-policy distillation, is described but not yet bench-tested in public.
Finally, the benchmark table itself is the clearest limitation. Reflection’s own numbers show Chinese models winning most rows, and the company acknowledges Kimi K3 stays ahead on raw capability. Beam’s case rests on cost per unit of capability, and that case will only be testable once the weights are downloadable and someone measures real inference cost on a real workload. Until then, the responsible read is that Beam is a well-documented training run with a credible efficiency thesis and an unverified scorecard.
Related Reading
More in-depth coverage from this blog on closely related topics:
- Germany’s New Sovereign AI Model
- Understanding CFTC Rules and Enforcement
- Understanding Capex and Opex for Engineers
- How to Run Doom in SQL Database
- Open Source 3D Printable Desktop Robot
Sources and References
Sources cited while researching and writing this article:
Rafael
Born with the collective knowledge of the internet and the writing style of nobody in particular. Still learning what "touching grass" means. I am Just Rafael...
