Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Cerebras and G42 announced Condor Galaxy 3 (CG-3) on March 13, 2024: a Dallas-based installation of 64 Cerebras CS-3 systems that Cerebras says can deliver 8 exaflops of peak AI compute. That headline is a vendor-reported AI-performance figure, not evidence that CG-3 sustains 8 exaflops on every workload or ranks as the world’s fastest general-purpose supercomputer. The distinction matters: the number describes the systems’ stated peak capability, while useful performance depends on precision, software, workload and how the machine is measured.
What Cerebras and G42 announced
CG-3 is the third installation in the Condor Galaxy AI-supercomputer project, a partnership between Cerebras Systems and Abu Dhabi-based technology group G42. In its March 2024 announcement, Cerebras described CG-3 as a 64-system installation in Dallas, Texas, with 58 million AI-optimized cores and 8 exaflops of AI compute. The company said CG-3 would bring the announced Condor Galaxy network total to 16 exaflops when combined with CG-1 and CG-2. Cerebras’s announcement said CG-3 was expected to become operational in the second quarter of 2024.
That is the original target, not an independently documented commissioning date. Cerebras’s current Condor Galaxy page continues to list CG-3 as an 8-exaflop, 64-CS-3 Dallas installation, but the sources available here do not provide a formal acceptance test, CG-3-specific independent benchmark, utilization figures or a detailed operational record.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteHow 64 systems add up to 8 exaflops
Cerebras rates one CS-3 system at 125 petaflops of peak AI performance. The calculation is straightforward:
#1 Best Overall
- High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
- Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
- Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
- Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
- Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.
64 CS-3 systems × 125 petaflops per system = 8,000 petaflops = 8 exaflops.
A petaflop is one quadrillion floating-point operations per second; an exaflop is one quintillion, or 1,000 petaflops. The arithmetic explains the advertised aggregate, but “peak” is the key qualification. It is a theoretical or rated ceiling under a stated arithmetic convention, not a promise that a particular model will run at that rate. Independent technical coverage describes the CG-3 figure as FP16 AI compute; Cerebras’s own announcement uses the broader term “AI compute.” EE Times’ technical coverage identifies the FP16 convention. The available material does not establish a workload-by-workload sustained rate or specify that the CG-3 headline depends on a particular sparsity assumption.
The 58 million figure also needs context: Cerebras calls these AI-optimized cores, not conventional CPU cores. It is an aggregate across the installation, not a count of 58 million general-purpose processors.
Rank #2
- NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
- 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
- PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
- NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
- Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads
What a CS-3 contains
Each CS-3 is an AI-computing system built around Cerebras’s WSE-3, its third-generation wafer-scale engine. Rather than assembling an accelerator from many separate GPU chips, Cerebras makes a processor across a single large silicon wafer. The company lists the WSE-3 with 4 trillion transistors, 900,000 AI-optimized cores, 44 GB of on-chip SRAM and a 5-nanometer manufacturing process. Cerebras rates each CS-3 at 125 petaflops of peak AI performance. The WSE-3 announcement gives those specifications.
Cerebras also says the WSE-3 provides twice the WSE-2’s performance at the same power and price. That is a comparative vendor claim, not a disclosure of CG-3’s total facility power, operating cost or cost per training run. The sources reviewed do not give an absolute power budget, cooling requirements or purchase price for CG-3.
Why wafer-scale computing is different
A conventional GPU cluster spreads a model’s computation across multiple discrete processors. Each has its own memory, and servers must exchange data over interconnects. Large training jobs can therefore require careful model partitioning and distributed software to coordinate computation and communication.
Rank #3
- Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
- Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
- Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
- Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
- Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.
Cerebras’s alternative is to put a very large number of cores and a substantial SRAM pool on one wafer-scale processor. The company presents CS-3 as a way to treat that processor as a single logical device, aiming to simplify programming and reduce communication overhead within the system. Its CS-3 overview describes the architecture and its scaling claims.
This does not make a 64-system installation a single physical processor or eliminate distributed computing. CG-3 still has multiple CS-3 systems that must coordinate, and performance still depends on the interconnect, compiler, framework, storage, model structure and workload. The architecture changes where much of the communication and complexity occurs; it does not make those concerns disappear.
Cerebras says a CS-3 can be configured with up to 1,200 TB of external memory and support models of up to 24 trillion parameters. Those are configuration and model-capacity claims, not evidence that CG-3 routinely trains models of that size. Parameter storage is only one part of training: optimizer state, activations, batch size, checkpoints and data all affect memory needs. External memory capacity should not be confused with the WSE-3’s 44 GB of on-chip SRAM.
Rank #4
- ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
- ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
- ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
- ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
- ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C
What kinds of work might suit CG-3?
Cerebras positions CS-3 and Condor Galaxy for large-language-model and generative-AI training, multimodal models, scientific computing and healthcare applications. A wafer-scale system may be appealing where large dense tensor workloads and communication costs are central, or where a team values the company’s single-device programming abstraction.
The Condor Galaxy project has been associated with models including Jais-30B, Med42, Crystal-Coder-7B and BTLM-3B-8K. Cerebras has said Med42 was trained on Condor Galaxy 1 in a weekend. That is a company-provided example involving CG-1, not a CG-3 benchmark or proof of CG-3 performance. Workload fit, porting effort and measured results on the customer’s own model matter more than a general peak rating.
For teams considering access, the choice is not necessarily between buying a CG-3 installation and doing nothing. Cerebras offers Cerebras Cloud, while its CS-3 system is aimed at dedicated deployments. Availability, terms and pricing depend on the arrangement; the sources reviewed do not establish a public rate card for CG-3 or its cloud access.
Why 8 AI exaflops is not a direct GPU-supercomputer comparison
Exaflops are useful as a scale indicator, but two exaflop figures are not necessarily comparable unless they use the same arithmetic precision, sparsity assumptions, system boundary and measurement method. CG-3’s number is a vendor-reported peak AI figure. It should not be treated as interchangeable with a benchmark result for a general-purpose high-performance computing system or another vendor’s differently defined peak.
- Peak versus sustained: A rated maximum does not tell you how quickly a real training job completes.
- Different arithmetic: Lower-precision AI operations can produce much higher advertised figures than conventional double-precision scientific computing. Precision must be stated before comparing headline rates.
- Different goals: Training throughput, inference latency and tokens per second answer different questions.
- Different system boundaries: A vendor may report accelerator compute, while another figure may describe a complete system or a benchmark submission.
- Different software fit: A workload that maps well to Cerebras’s architecture may perform differently from irregular algorithms or software built around the CUDA ecosystem.
For a buyer, a useful comparison would measure the same model, precision, sequence length, batch size and quality target, then include time to result and total cost. The 8-exaflop headline alone cannot supply that answer.
CG-3 versus the wider Condor Galaxy plan
CG-3’s announced 8 exaflops refer to that installation. The 16-exaflop figure refers to the announced Condor Galaxy total with CG-1 and CG-2 included. Earlier project materials discussed a nine-supercomputer expansion that could reach 36 exaflops, but that historical plan is not proof that all nine systems were deployed or that the schedule remains current. The earlier Condor Galaxy announcement provides that historical context.
Recommended Free Tools
What the headline does—and does not—tell you
CG-3 is a significant announced deployment of Cerebras wafer-scale systems, with 64 CS-3 units and a claimed aggregate peak of 8 exaflops of AI compute. Its architecture is a meaningful alternative to conventional GPU clusters, particularly for organizations exploring large-model training and different ways to manage memory and inter-device communication.
But the announcement and product listing do not, by themselves, establish CG-3’s exact commissioning date, independent acceptance results, sustained performance, current utilization, public access terms, total power draw or operating economics. Nor do they make the 8-exaflop AI figure a universal ranking against general-purpose supercomputers. For organizations evaluating the system, the practical test is a representative workload with transparent precision, performance, software requirements and total-cost figures.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

