Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

xAI is expanding a family of AI supercomputer projects in the Memphis area to train and run its Grok models. The effort began with Colossus, announced with 100,000 NVIDIA Hopper GPUs, and now includes the much larger Colossus 2 project. NVIDIA has said Colossus 2 is expected to house more than half a million NVIDIA GPUs. Those figures describe different systems and stages of development—not one independently verified count of active chips.

The distinction matters: announcements, planned capacity, installed accelerators and performance-normalized “H100 equivalents” are not interchangeable. Power, networking, cooling and utilization will determine how much useful compute xAI can actually get from the hardware.

What xAI is building

Colossus is xAI’s AI training cluster in Memphis, Tennessee. Its stated purpose is to develop the company’s Grok models. NVIDIA described the initial system as containing 100,000 Hopper-generation GPUs, and later said xAI planned to expand Colossus to a combined 200,000 Hopper GPUs. NVIDIA’s Colossus announcement also detailed the networking equipment used to connect the cluster.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Colossus 2 is a larger, related expansion project in the Memphis area. NVIDIA said it is expected to contain more than half a million NVIDIA GPUs, with public descriptions associating the project with newer Blackwell systems. That is an announced expectation, not proof that every planned accelerator has been delivered, installed or placed into production.

#1 Best Overall
ASUS Ascent GX10 Personal AI Supercomputer, NVIDIA GB10 Grace Blackwell Superchip, 128GB LPDDR5x Unified Memory, 2TB NVMe SSD, DGX OS, Wi-Fi 7, 10GbE, AI Workstation for Local LLM and RAG
  • [Personal AI Supercomputer]: Built for AI developers, researchers, data scientists, startup labs, and university labs, the ASUS Ascent GX10 is designed for local AI development, model testing, inferencing, RAG workflows, and agentic AI experimentation beyond a standard mini PC.
  • [NVIDIA GB10 Grace Blackwell Superchip]: Powered by the NVIDIA GB10 Grace Blackwell Superchip with Blackwell GPU architecture and a 20-core Arm CPU, GX10 delivers up to 1 PetaFLOP of FP4 AI performance for generative AI prototyping and local model workflows.
  • [128GB Unified Memory for Large AI Workloads]: 128GB LPDDR5x unified memory helps support demanding AI development and testing scenarios, including workflows for large language models, multimodal AI, local inference, fine-tuning experiments, and model evaluation.
  • [2TB NVMe Storage for AI Projects]: The 2TB M.2 2242 NVMe SSD provides high-speed local storage for AI model libraries, datasets, Docker containers, checkpoints, development environments, and RAG or vector database workflows.
  • [DGX OS and Advanced Connectivity]: DGX OS and the NVIDIA AI software stack help streamline CUDA, PyTorch, TensorFlow, TensorRT, NVIDIA NIM, and AI Blueprint workflows, while Wi-Fi 7, 10GbE, USB-C, HDMI, and NVIDIA ConnectX-7 support modern lab and desktop deployments.

It is therefore more accurate to describe xAI’s effort as a growing set of related facilities and clusters—Colossus I and II—than as one giant machine with a single confirmed chip count. The projects can include different hardware generations and workloads.

What the headline GPU numbers mean

Figure What it refers to Status and caveat
100,000 Hopper GPUs The initial Colossus cluster NVIDIA’s 2024 description; not a current total for all xAI infrastructure.
200,000 Hopper GPUs Planned expansion of Colossus NVIDIA described xAI’s plan to double the system. A plan does not establish that all units were operating simultaneously.
More than 500,000 NVIDIA GPUs Colossus 2 NVIDIA’s announced expectation for the facility; not an independently audited operational count.
More than 1 million H100 GPU equivalents Colossus I and II infrastructure at the end of 2025 xAI’s January 2026 company statement. “Equivalent” is a performance-normalized measure, not necessarily a count of physical H100 chips.
2 gigawatts Reported target for the broader planned data-center development Associated with a Mississippi project in AP reporting; a project target, not a statement of current GPU power consumption or completed capacity.

These numbers should not be added together. They describe different facilities, dates, units and degrees of certainty. xAI’s January 2026 Series E announcement reported more than one million H100 GPU equivalents across Colossus I and II at the end of 2025. That measure cannot be directly compared with NVIDIA’s projected physical-GPU count for Colossus 2.

GPU, superchip, server and equivalent are different units

A physical GPU is an accelerator device. A GPU equivalent converts compute capacity into the performance of a chosen reference GPU. A superchip is a package combining processor components; a server or node can contain multiple accelerators, and a rack contains multiple servers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

NVIDIA’s GB200 Grace Blackwell Superchip combines a Grace CPU with Blackwell GPU components. So a report counting GB200 superchips, systems or nodes is not automatically counting the same thing as a report counting standalone GPU devices. The exact final hardware mix and operational count for Colossus 2 have not been established by the announcements cited here. NVIDIA’s GB200 product description explains the package architecture.

Which NVIDIA hardware is involved?

The original Colossus announcement centered on NVIDIA H100 accelerators from the Hopper generation. Public descriptions of subsequent expansion have referred to H200-class hardware as well. For Colossus 2, announcements and reporting point to Blackwell-generation systems, including GB200 and GB300 configurations. The final installed mix should not be treated as settled unless xAI or NVIDIA provides a dated, specific confirmation.

These facilities depend on more than accelerator chips. NVIDIA’s Blackwell platform integrates CPUs, GPUs, high-speed interconnects, networking and software. The point is to make many processors work together as a system, not simply to accumulate a large number of cards.

Why networking, power and cooling matter

Distributed AI training requires accelerators to exchange data rapidly. A cluster with abundant GPUs can still waste expensive capacity if communication between them, data delivery or job scheduling becomes a bottleneck. NVIDIA said Colossus uses Spectrum-X Ethernet networking, Spectrum switches, BlueField-3 SuperNICs and Remote Direct Memory Access networking. These components help move data between machines so accelerators can stay productive.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

At this scale, the facility’s electrical and thermal systems are equally important. A large deployment requires utility connections and substations, power conversion, backup arrangements, high-capacity cooling, storage and network infrastructure, as well as construction and permits. The reported 2-gigawatt target is a measure of planned facility-scale power capacity, not a claim that GPUs alone consume that amount. A data center’s total demand includes cooling, networking, storage and other equipment.

xAI says the original Colossus was built in 122 days. That is a company-described milestone for the first system; it does not establish that later, larger facilities can be completed on the same timetable. xAI’s Colossus page describes the project.

Rank #2
NVIDIA Tesla L4 24GB PCIe Graphics ACELLERATOR HH/HL 75W GPU 900-2G193-0000-000
  • 24GB Video Memory
  • Fourth Generation Tensor Cores
  • HALF HEIGHT BRACKET ONLY

Why xAI wants this much compute

Large accelerator clusters can support model pretraining, reinforcement learning, synthetic-data generation, evaluation and other development work. More capacity can let a company run larger or more experiments, or complete them faster. xAI has positioned Colossus and Grok as part of a broader consumer and enterprise product strategy, while Grok is also integrated with X.

Compute is only one input into model capability. Results also depend on data quality, model architecture, algorithms, software efficiency, networking, electricity and how consistently the hardware is used. A larger cluster does not by itself guarantee that Grok will outperform competing models.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Training and inference are also distinct jobs. Training creates or updates models; inference runs a model to answer users. xAI has indicated that some inference is handled through cloud providers rather than solely by its training supercluster. There is no public basis here for assuming that all Colossus capacity is devoted to consumer queries or offered for outside rental. Colossus itself is infrastructure, not a public GPU-hosting service; people access Grok through xAI’s products and integrations.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

The investment behind the buildout

xAI announced a $20 billion Series E financing round in January 2026 and named NVIDIA and Cisco Investments among strategic investors. The company described computing infrastructure as central to its ambitions for Grok. Separately, the Associated Press reported a planned Mississippi data-center project valued at roughly $20 billion and linked to a broader 2-gigawatt target. These are distinct figures and announcements: the fundraising amount is not a disclosed complete construction budget, and reported project costs should not be read as audited final spending.

The total cost of a project like this includes more than accelerators: complete rack systems, buildings, electrical equipment, cooling, networking, storage, maintenance, staffing and financing all matter. Public GPU list prices cannot yield a reliable total project bill because actual configurations, discounts, leases, financing and infrastructure costs are not fully disclosed. The capital raise confirms substantial funding, but it does not reveal the full financing structure or economics of every facility.

What the scale does—and does not—tell us

  • It signals an unusually compute-intensive strategy. xAI is investing in infrastructure that could support rapid model development and growing product demand.
  • It increases dependence on power and construction. Grid access, cooling, permitting, hardware deliveries and facility readiness can constrain deployment even when capital and chips are available.
  • It does not prove a model-quality lead. Useful compute depends on software, data, workload design and utilization—not headline chip totals alone.
  • It does not establish an exact live fleet size. Company and vendor announcements use different units and describe plans, facilities or equivalents.
  • It does not establish “world’s fastest” status. That claim would require a defined benchmark and date; AI training capacity is not the same as a traditional high-performance-computing ranking.

There are also operational risks. Hardware can sit idle if data pipelines, scheduling or networking cannot feed it effectively. A power connection or cooling system may lag behind accelerator deliveries. And newer chip generations can change the economics of older equipment, even though previous-generation systems can remain useful for many workloads. NVIDIA’s announced Vera Rubin platform is a vendor roadmap, not evidence that xAI has adopted it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why NVIDIA is central

The partnership is about a hardware and software stack, not just a chip order: accelerators, CPUs, interconnects, networking, DPUs and CUDA-related software all contribute. NVIDIA has described Colossus 2 as expected to use more than half a million of its GPUs, and xAI has publicly described large-scale use of its systems. NVIDIA’s ecosystem maturity and integrated offerings help explain the choice, but the available evidence does not show that NVIDIA is best for every workload or that competing accelerators could not be used in other systems.

For the broader AI market, a buildout of this size adds demand for advanced accelerators, high-speed networking, electricity and data-center capacity. It also illustrates the growing advantage—and cost—of controlling compute infrastructure. Whether that investment translates into better or more widely used products depends on what xAI can train, how efficiently it operates the systems and whether demand justifies the continuing expense.

Sources and dates

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.