Reflection AI’s October 5, 2026 announcement shows Beam trailing some named models on specific coding and agentic benchmarks. The company also claims Beam can match GLM-5.2 on advanced reasoning benchmarks with 3–4× less inference compute—but that figure is an estimate, not a verified measure of cost, speed, or energy use.
What Beam is—and what was available at announcement
Reflection describes Beam as a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active parameters per token, designed for coding, reasoning, and agentic workloads. Its announcement also reports 23.8 trillion pretraining tokens, more than 100 million reinforcement-learning rollouts, and approximately 1.3 billion sandboxes used for training and grading. Reflection says a reinforcement-learning run used 10,500 NVIDIA GB300 GPUs for four weeks. These are company-reported figures, not independently audited measurements. Reflection AI’s announcement
At the time of the announcement, Reflection said Beam was undergoing final red-teaming and evaluations. The weights, technical report, model card, and developer materials were still forthcoming. The announcement does not establish a reader-facing hardware configuration for running Beam; the GB300 figure describes Reflection’s training run, not a recommended setup.
How Beam compares on the reported coding tests
The table below reproduces Reflection’s reported scores, comparing models only where its table supplies a result in the same benchmark row. These are Reflection’s published results; it says it used Artificial Analysis and DataCurve data for other models. “NR” means a result was not reported in Reflection’s table.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
| Benchmark | Beam | Other reported models |
|---|---|---|
| SWE Bench Pro v2-Hard | 77.2 | GLM 5.3: 84.3; Kimi K3: 88.2 |
| Terminal Bench v2.1 | 80.1 | GLM 5.3: 88.2; Kimi K3: 88.3; DeepSeek V4.1 Flash: 90.6 |
| SWE Bench Pro v1 | 65.5 | Qwen 3.8-Max: 67.7; GLM 5.2: 62.1 |
| SWE-bench Verified | 80.9 | Most comparison cells: NR |
On SWE Bench Pro v2-Hard, Beam’s reported score is below both listed comparators. On Terminal Bench v2.1, it is below all three listed comparators. SWE Bench Pro v1 gives a more mixed comparison: Beam is below Qwen 3.8-Max but above GLM 5.2. The SWE-bench Verified row does not establish a broad ranking because most other results are NR. The scores are benchmark-specific; they do not show that Beam trails every leading open model on coding overall. Reflection’s benchmark table
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What Reflection means by “3–4× less inference compute”
Reflection says Beam achieves scores comparable to GLM-5.2 on advanced reasoning benchmarks while using 3–4× less inference compute. Its estimate approximates generation forward-pass compute as 2 × active parameter count × mean generated tokens per attempt. For a mixture-of-experts model, the calculation uses active parameters per token rather than total parameters.
Rank #2
The estimate excludes prompt prefill, context-dependent attention operations, and serving overhead. It is therefore a limited comparison of estimated generation compute, not a measurement of full inference cost or end-to-end deployment performance. It does not establish that Beam is 3–4× cheaper, faster, or more energy-efficient to serve. TechCrunch reported that Reflection’s performance claims had not been independently verified. TechCrunch’s October 5, 2026 report
Quick Recap
Best Value
How to read the results
- For coding performance, compare scores within the same benchmark and version; the listed models and coverage vary by row.
- Do not treat a benchmark score and an inference-compute estimate as interchangeable: one reports task results, while the other estimates a bounded portion of generation work.
- Distinguish an announced open-weight model from a publicly downloadable release. At announcement, the weights and technical materials were not yet available.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




