Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

NVIDIA’s cuLitho is a GPU-accelerated software library for computational lithography—not a chipmaking machine or a consumer GPU. On March 18, 2024, NVIDIA said TSMC and Synopsys were taking cuLitho-integrated workflows into production. The significance is that a foundry and a major EDA supplier were incorporating NVIDIA acceleration into real manufacturing software and processes. It does not mean every TSMC fab or chip now uses cuLitho, nor does it promise cheaper chips or higher yields.

What cuLitho does

Before a chip design can be printed onto silicon, its intended patterns must be translated into photomask patterns that account for the behavior of light and materials during exposure. At advanced dimensions, the printed result can differ from the ideal layout because of diffraction, optical proximity effects and process variation. Computational lithography uses physical models and intensive computation to adjust mask patterns so the wafer is more likely to receive the intended shapes.

A simplified path is: chip layout → lithography modeling and correction → mask synthesis → photomask → wafer exposure → inspection and process correction. cuLitho accelerates the computational part of that chain. It is a CUDA-X library of GPU-optimized algorithms and tools for workloads such as optical proximity correction (OPC), inverse lithography technology (ILT), geometric operations, optimization and distributed computing. NVIDIA describes the library on its cuLitho developer page.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

OPC modifies mask shapes to compensate for predictable printing distortions. ILT takes a more computationally intensive approach, calculating mask shapes that are likely to produce the desired wafer pattern. Neither replaces the physical lithography equipment, process models, mask writing, metrology or engineering sign-off.

#1 Best Overall
NVD RTX PRO 6000 Blackwell Professional Workstation Edition Graphics Card for AI, Design, Simulation, Engineering - 96GB DDR7 ECC Memory - 4th Gen RT/5th Gen Tensor Core GPU - OEM Packaging
  • PLEASE NOTE: Exporting an NVIDIA RTX Pro 6000 GPU outside the US requires strict adherence to the U.S. Export Administration Regulations (EAR) and issuance of an export license from the Bureau of Industry and Security (BIS). Compliance and Know Your Customer (KYC) screening may be required as a condition of order acceptance. [NVIDIA Blackwell Streaming Multiprocessor] The new SM features increased processing throughput, and new neural shaders that integrate neural networks inside of programmable shaders | DLSS 4: Multi Frame Generation ensures ultra-smooth frame pacing for lifelike simulations.
  • [Double-Flow-Through Design] The RTX PRO 6000 Blackwell features a double-flow-through cooling design, optimizing efficiency and airflow to sustain peak performance under 600W power loads. | [5th Gen Tensor Cores] Deliver up to 3X the performance of the previous generation and support for FP4 precision for faster AI model processing times with reduced memory usage, enabling local fine-tuning of LLMs and generative AI | [4th Gen Ray Tracing Cores] Double the ray-triangle intersection rate of the previous generation to create photoreal, physically accurate scenes and immersive 3D designs with RTX Mega Geometry, which enables up to 100X more ray-traced triangles.
  • [PCIe Gen 5] Support for PCIe Gen 5 provides double the bandwidth of PCIe Gen 4, improving data-transfer speeds from CPU memory and unlocking faster performance for data-intensive tasks like AI, data science, and 3D modeling. | [GDDR7 Memory] With 96 GB of GPU memory and 1.8 TB ps bandwidth, it can tackle massive 3D and AI projects, fine-tune AI models locally, explore large-scale VR environments, and drive larger multi-app workflows.
  • [DisplayPort 2.1] Achieve unparalleled visual clarity and performance, driving high resolution displays at up to 8K at 240 Hz and 16K at 60 Hz. Increased bandwidth enables seamless multi-monitor setups while HDR and higher color depth support ensures superior color accuracy for precision work, such as video editing, 3D design, and live broadcasting.
  • [Universal MIG] Divide a single RTX PRO 6000 Blackwell into multiple isolated instances, each with dedicated resources, allowing for concurrent execution of multiple workloads, optimized GPU utilization, and secure isolation of different applications or users. [WARRANTY] 3 YR Manufacturer's Warranty. Bulk OEM Packaging. Retail Packaging is NOT included.

What TSMC and Synopsys are contributing

TSMC is the manufacturing partner. Its role is to integrate GPU-accelerated computation into foundry lithography workflows, where process data, validation and manufacturing requirements matter. In its March 2024 announcement, NVIDIA said TSMC and NVIDIA had integrated cuLitho with software, manufacturing processes and systems and were moving into production.

Synopsys supplies the specialized EDA application layer. Its Proteus mask-synthesis software was described as running with the cuLitho library. That matters because an acceleration library becomes more useful to manufacturers when it is integrated into established mask-synthesis software, rather than standing alone as a research or GPU-programming project. Synopsys outlined the collaboration in its 2024 announcement.

“Production” is meaningful, but the public announcements do not identify the TSMC fabs, process nodes, customer chips or product families involved. They also do not establish that every TSMC advanced node uses cuLitho. The evidence supports production integration, not universal deployment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why GPUs may help

Many computational-lithography operations involve applying calculations across large numbers of pattern regions or data points. GPUs can perform many such operations concurrently, while CUDA supplies the programming environment for mapping workloads to NVIDIA hardware. NVIDIA says the industry spends tens of billions of CPU hours annually on computational lithography, illustrating the scale of the problem in its 2024 release.

Faster computation can increase throughput, shorten iterations and make more demanding methods—including curvilinear mask flows—more practical. A Manhattan-style mask uses predominantly horizontal and vertical edges; a curvilinear mask uses curves and more complex shapes. Curves can improve pattern fidelity in some approaches, but they create heavier computation and data-processing demands.

Rank #2
NVIDIA RTX PRO 4000 Blackwell Graphics Card - 24GB GDDR7 ECC Memory, PCIe 5.0 x16, 4X DisplayPort 2.1b, Single Slot Full Height AI Workstation GPU, Retail Packaging
  • Professional GPU with Blackwell Architecture
  • Blackwell Architecture
  • 24GB GDDR7 with PCIe 5.0 & Ray Tracing
  • AI Workstation

These gains are conditional. GPU acceleration does not remove the need for accurate physical models, validated process data, compatible EDA software, mask-writing equipment, inspection, process control or production sign-off. Nor does less energy per computational job automatically mean lower total fab energy if the fab uses the extra capacity to run more jobs.

How to read the speedup figures

The figures publicized by NVIDIA, TSMC and Synopsys refer to different workloads, baselines or outcomes. They should not be treated as interchangeable guarantees.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Figure What it refers to Important qualification
Up to 40× NVIDIA’s broad cuLitho acceleration claim for computational lithography, including inverse-lithography workloads. A platform claim; the workload, baseline and system configuration matter.
45× A TSMC/NVIDIA result for a curvilinear workflow announced in 2024. A reported shared-workflow benchmark, not a universal fab result.
Nearly 60× A separate TSMC/NVIDIA result for a Manhattan-style workflow. Different workflow and comparison; not directly interchangeable with 45×.
350 H100 systems versus 40,000 CPU systems An illustrative infrastructure comparison in NVIDIA’s 2024 announcement. Not a general purchasing recommendation or proof of equivalent performance for every workload.
15× Synopsys’ reported OPC speedup for an H100-optimized Proteus implementation integrated with cuLitho. A Synopsys-specific test; see its 2025 announcement.
20%–50% A TSMC/NVIDIA 2026 claim about improved cost effectiveness or cycle time versus CPU-based computational lithography. A different metric from raw workload acceleration; it is not a contradiction of the earlier multipliers.

NVIDIA’s original 2023 cuLitho announcement also described 500 DGX H100 systems doing work associated with 40,000 CPU systems and said work that could take roughly two weeks might be completed overnight. Those were company-provided illustrative claims, not a promise of current results for every mask or production environment.

A workload running 45 times faster does not make a chip 45 times cheaper. Lithography computation is one portion of a much larger chain that includes design, masks, equipment, materials, wafer processing, inspection and packaging. Other bottlenecks can dominate, and public figures do not disclose enough detail to calculate a fab’s total cost savings.

What has changed since the 2024 announcement

In March 2025, Synopsys reported a 15× OPC speedup in its H100-optimized Proteus implementation with cuLitho and said NVIDIA Blackwell would accelerate computational lithography further. That is a distinct Synopsys-reported result, not a replacement for or confirmation of NVIDIA’s 45× and nearly 60× figures.

Rank #3
PNY NVIDIA RTX A6000
  • NVIDIA Ampere Architecture-based CUDA Cores - Double-speed processing for single-precision floating point (FP32) operations and improved power efficiency provide significant performance improvements for graphics and simulation workflows, such as complex 3D computer-aided design (CAD) and computer-aided engineering (CAE), on the desktop.
  • Second-Generation RT Cores - With up to 2X the throughput over the previous generation and the ability to concurrently run ray tracing with either shading or denoising capabilities, second-generation RT Cores deliver massive speedups for workloads like photorealistic rendering of movie content, architectural design evaluations, and virtual prototyping of product designs. This technology also speeds up the rendering of ray-traced motion blur for faster results with greater visual accuracy.
  • Third-Generation Tensor Cores - New Tensor Float 32 (TF32) precision provides up to 5X the training throughput over the previous generation to accelerate AI and data science model training without requiring any code changes. Hardware support for structural sparsity doubles the throughput for inferencing. Tensor Cores also bring AI to graphics with capabilities like DLSS, AI denoising, and enhanced editing for select applications.
  • Third-Generation NVIDIA NVLink - Increased GPU-to-GPU interconnect bandwidth provides a single scalable memory to accelerate graphics and compute workloads and tackle larger datasets.
  • 48 Gigabytes (GB) of GPU Memory - Ultra-fast GDDR6 memory, scalable up to 96 GB with NVLink, gives data scientists, engineers, and creative professionals the large memory necessary to work with massive datasets and workloads like data science and simulation.

On May 31, 2026, NVIDIA said TSMC was using cuLitho and other CUDA-X libraries and AI models across a wider set of fab workloads, including lithography, transistor and process simulation, process control and fab-operation optimization. The companies cited a 20%–50% improvement in cost effectiveness or cycle time for cuLitho compared with CPU-based computational lithography. This later statement broadens the story from a lithography workflow to GPU- and AI-assisted fab operations; it does not make the percentage directly comparable with benchmark speedups. See the 2026 NVIDIA–TSMC announcement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Where ASML fits—and what NVIDIA is becoming

ASML was part of NVIDIA’s 2023 cuLitho ecosystem announcement. NVIDIA said ASML was working with it on GPU support and intended to integrate GPU support into computational-lithography software products, with high-NA EUV among the relevant developments. ASML is important because it supplies lithography equipment and is connected to the broader computational-lithography ecosystem, but the central production-integration news here is TSMC and Synopsys.

This is an expansion of NVIDIA’s role in semiconductor infrastructure, but it does not mean NVIDIA is replacing EDA companies. NVIDIA provides accelerated-computing hardware and foundational libraries; Synopsys supplies specialized mask-synthesis software; TSMC brings manufacturing processes and fab integration. NVIDIA’s 2025 industry announcement also discussed work with companies including Cadence, KLA, Siemens and Synopsys. These firms’ wider products can complement or compete across different parts of EDA and manufacturing, but the cuLitho announcement does not establish that they are equivalent substitutes for Proteus in this specific workflow.

What this does—and does not—mean for chip production

  • It can improve computational throughput. Completing computational-lithography jobs faster may allow more iterations or shorten a step on the path from design to wafer learning.
  • It may make complex methods more practical. Additional compute capacity can help with demanding algorithms and curvilinear patterns.
  • It does not prove better yields. Yield depends on the complete design and manufacturing process; faster computation alone is not a yield guarantee.
  • It does not prove lower chip prices or faster shipments. No public announcement quantifies an effect on finished-chip costs, production volumes or delivery dates.
  • It does not prove every NVIDIA chip uses cuLitho. NVIDIA has said the work could support future advanced architectures, including Blackwell, but the disclosures do not establish that every Blackwell chip or other NVIDIA product was manufactured with it.
  • It is not a normal consumer software purchase. The public materials do not provide a standalone cuLitho price or a self-service route for independent developers. The evidence points to enterprise collaborations and integration with EDA and fab workflows.

For a semiconductor company, the practical decision is about the complete stack: GPU servers and interconnects, data and cooling capacity, compatible EDA software, process models, validation, support and integration with sensitive manufacturing systems. GPU clusters can require substantial capital and specialist staff, and a speedup on one OPC or ILT task may not carry over to another. Existing CPU systems may remain appropriate where workloads are not GPU-optimized or migration risk outweighs the expected benefit.

cuLitho is therefore best understood as an infrastructure layer for industrial semiconductor computing. TSMC’s production integration and Synopsys’ Proteus integration are stronger evidence of practical relevance than a general endorsement, while the undisclosed deployment scope and company-reported benchmark methodology put clear limits on what outsiders can conclude.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.