Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Huawei’s Da Vinci was an AI-processor architecture, not a single chip or accelerator card. It underpinned the company’s Ascend processors, which Huawei packaged into Atlas products for edge inference, data-center workloads and larger AI systems. The Hot Chips 31 topic belongs to 2019; its significance is best understood as an early view of Huawei’s effort to build an integrated AI hardware and software stack.

What Da Vinci, Ascend and Atlas mean

The names describe different layers of Huawei’s AI-computing portfolio:

  • Da Vinci was Huawei’s AI-processor architecture, described by the company as a “3D Cube” design.
  • Ascend was the family of AI processors built around that architecture.
  • Atlas was the product and infrastructure platform using Ascend processors, spanning modules, cards, edge systems and clusters.
  • CANN and MindSpore were part of the software and programming environment used to develop for Ascend.

Huawei said it launched Da Vinci in 2018 and described Ascend processors as using the Da Vinci 3D Cube architecture. Ascend was one part of a broader processor portfolio that also included Kunpeng, Kirin and Honghu. Huawei’s Atlas launch announcement and its September 2019 computing strategy announcement establish that positioning.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The distinction matters: Da Vinci was not another name for Ascend 310 or Ascend 910, and it was not a conventional general-purpose GPU. It was the architectural foundation used across products with different capabilities and intended roles.

#1 Best Overall
Sale
HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
  • NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
  • 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
  • PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
  • NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
  • Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads

Why Huawei discussed it at Hot Chips

Hot Chips is an architecture-focused conference, making Da Vinci relevant as more than a product launch: it represented Huawei’s approach to organizing AI computation, memory and control around neural-network workloads. The surviving AnandTech result identifies the historical item as “Hot Chips 31 Live Blogs: Huawei Da Vinci Architecture,” but the old tag URL now redirects to the forums, so the original live-blog text is not readily accessible there. AnandTech’s Da Vinci tag provides that archival trail.

That access limitation also sets a boundary on what can responsibly be attributed to the conference session. Huawei’s later product announcements help explain how it positioned Da Vinci and Ascend, but they should not be mistaken for a transcript of what was shown at Hot Chips. Publicly available material supports discussion of the architecture’s broad components and product strategy; it does not establish every microarchitectural detail or exact implementation dimension presented at the event.

What was inside an Ascend processor

A technical chapter in Ascend AI Processor Architecture and Programming describes an Ascend SoC as combining a Control CPU, AI Core, AI CPU, cache and buffer hierarchy, and Digital Vision Preprocessing (DVPP), alongside Da Vinci-based AI computation. The book describes Da Vinci coverage across the compute unit, memory system, control, instruction-set design and convolution acceleration. The architecture chapter is a useful technical reference, though it does not substitute for an unavailable Hot Chips slide deck.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
MX3 M.2 AI Accelerator
  • High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
  • Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
  • Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
  • Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
  • Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.

Control CPU and AI CPU

The Control CPU handles general orchestration around accelerator work. The AI CPU provides a processor element for control, preprocessing, postprocessing or other operations that do not map efficiently to the main matrix engine. These roles help explain why an AI SoC is not just a block of tensor arithmetic: tasks still need coordination and operations outside the primary accelerator.

AI Core and the Da Vinci compute engine

The AI Core is the main high-throughput engine associated with AI computation, including matrix- and tensor-oriented operations. Neural networks frequently rely on matrix multiplication and convolution, so specialized parallel hardware can accelerate those patterns. The gain depends on how well the model’s operations map to the engine and how efficiently operands are supplied.

Memory, buffers and data movement

Moving tensors can be as important as calculating with them. A hierarchy of caches and buffers lets hardware reuse data close to the compute engine rather than repeatedly fetching it from external memory. If a workload cannot keep data local or move it efficiently, nominal arithmetic throughput may not translate into application speed.

Rank #3
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
  • ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
  • ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
  • ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
  • ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
  • ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C

DVPP for vision pipelines

Huawei’s developer documentation describes DVPP functions including color-space conversion, normalization and cropping. Offloading such image and video preparation can reduce CPU work in computer-vision pipelines; it does not make every pipeline faster automatically, because the result also depends on supported formats, transfers and the rest of the workload. Huawei’s DVPP introduction outlines the documented functions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What Huawei meant by “3D Cube”

“3D Cube” is Huawei’s terminology for its matrix/tensor computation approach, not a claim that the processor is a literal three-dimensional geometric device. A useful way to understand the idea is that neural-network computation combines dimensions such as input values, weights and outputs. Processing blocks of these values in parallel can increase throughput, while local buffers help reuse data and limit costly off-chip movement.

The available sources do not establish a precise cube size, pipeline width or instruction encoding for the Hot Chips implementation. Those details should not be inferred from the name alone. The defensible point is that Huawei positioned Da Vinci around specialized AI computation, memory organization and control—not just a headline count of operations.

Rank #4

Early Ascend and Atlas products

Huawei’s 2019 portfolio illustrates how a shared architecture could appear in products for different deployment scales. The specifications below are Huawei’s launch claims, not independent benchmark results; figures such as TOPS also need precision and workload context to be comparable.

Product Intended role or form Huawei-stated detail
Ascend 310 Lower-power, inference-oriented processor associated with embedded and edge deployments. Specific performance and power figures are not stated in the cited Atlas launch page.
Ascend 910 Higher-performance processor positioned for AI training. Specific performance figures are not stated in the cited Atlas launch page.
Atlas 200 Accelerator module for terminal devices such as cameras, robots and drones. Huawei presented it as a module rather than a general-purpose consumer card.
Atlas 200 DK Developer kit for building Ascend applications. Huawei claimed applications could be developed for device, edge and cloud deployment without code modification; actual portability depends on operators, frameworks, compiler behavior and target-specific tuning.
Atlas 300 Accelerator card. Huawei reported 64 TOPS INT8, 32 GB of memory, 67 W power consumption and support for up to 64-channel real-time HD video analytics.
Atlas 500 Edge AI appliance. Huawei reported 16 TOPS INT8, less than 1 kWh per day of power consumption and an operating range of −40°C to +70°C; these are vendor-stated specifications, not a comparative efficiency test.
Atlas 900 Large-scale AI training cluster. Huawei said it combined thousands of Ascend processors and trained ResNet-50 in 59.8 seconds, which it described as ten seconds faster than the previous record in its September 2019 announcement.

Huawei’s April 2019 announcement described Atlas accelerator modules, cards, development kits, edge stations and appliances for applications including smart cities, carriers, finance, internet services and electric power. Its September announcement placed the Atlas 900 in scientific and enterprise contexts such as astronomy, weather forecasting, autonomous driving and oil exploration. These are examples of Huawei’s intended markets, not evidence that every listed system used identical hardware or delivered identical performance. Atlas launch details and Atlas 900 details are from Huawei.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why software and memory decide real-world results

For an engineer comparing AI processors, peak operations per second are only a starting point. The practical questions are whether a model’s operators are supported, how much conversion or graph rewriting is needed, how data moves through memory, and whether tools make errors and bottlenecks diagnosable.

Best Value
ASRock Radeon AI PRO R9700 Creator 32GB Professional Graphics Card, 2920 MHz Boost Clock, GDDR6, AMD RDNA 4, AI-Accelerators, DisplayPort 2.1a, PCIe 5.0, Blower Cooler
  • Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
  • Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
  • Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
  • Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
  • Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.
  • Operator coverage: Unsupported or poorly optimized operators may require custom work or execute on a less suitable processor element.
  • Model conversion: Huawei documentation describes converting models from frameworks such as Caffe and TensorFlow into formats supported by Ascend.
  • Data layout and memory: Input-format requirements and buffer management affect whether computation can stay on the fast path.
  • Optimization level: Graph-level compiler optimizations can help broadly, while custom operators and hand-tuned kernels may be necessary for unusual workloads.
  • Portability: A common architecture across product tiers does not guarantee that every model moves without changes or retuning.
  • Tooling and maintenance: Compilers, profilers, debugging tools and vendor-specific programming layers affect engineering cost as well as speed.

Huawei’s documented development process describes offline model generation and framework conversion, including fixed input-format requirements for some Da Vinci-related paths. The same announcement that promoted the Atlas 200 DK’s cross-scenario workflow claimed zero code modification; that should be read as Huawei’s product claim, not a universal guarantee of portability.

How to interpret the Atlas 900 result

Huawei’s 59.8-second ResNet-50 result is a historical vendor-reported cluster benchmark from September 18, 2019. It illustrates the intended move from a processor architecture to a distributed system, but it is not a universal measure of AI performance. A fair comparison needs the benchmark rules and configuration, including model and dataset, batch and precision, software stack, power and whether data input or preprocessing is included. The available Huawei announcement reports the time and its claimed improvement over a prior record; it does not independently establish broad superiority across workloads.

Da Vinci and GPU-centric approaches

Da Vinci is better compared with GPU-based systems by category than by a single winner/loser claim. Specialized tensor hardware can be efficient on supported neural-network operations, while a general-purpose GPU ecosystem may offer broader familiarity and software support for developers. Actual differences depend on the hardware generation, model, precision, software stack and deployment conditions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Throughput: INT8 TOPS cannot be compared directly with FP16, BF16 or FP8 figures, nor with results that assume sparsity, unless the measurement conditions match.
  • Flexibility: Specialized hardware may excel on supported operations and be less convenient for uncommon operators.
  • End-to-end performance: Memory bandwidth, preprocessing, batch size and data movement can dominate the result.
  • Software cost: Moving from CUDA-oriented workflows to Ascend can require framework conversion, operator validation and team familiarity with a different toolchain.
  • Deployment fit: Edge inference, data-center inference and large-scale training have different constraints and should not be collapsed into one performance comparison.

Huawei’s “full-scenario” framing described a strategy across deployment contexts, not identical silicon, performance or software behavior in every device.

What the surviving public record does not establish

The contemporary AnandTech tag entry is identifiable, but its former page is not readily accessible at the cited address. The sources that remain support the broad architecture and product story, but not a complete reconstruction of the Hot Chips talk.

  • The exact Da Vinci core dimensions, microarchitectural widths and instruction encoding shown at Hot Chips are not established here.
  • Independent performance-per-watt testing for the listed early products is not established by Huawei’s launch claims.
  • Complete operator-coverage data for the original Ascend products is not provided by the cited sources.
  • Current physical-hardware availability cannot be inferred from 2019 product announcements.

Huawei Cloud currently presents Ascend-based AI compute and related development services, but access is subject to service, region and account conditions. Its claims about cost effectiveness and migration remain vendor claims; the cloud service is a different route to experimentation from buying a physical accelerator. Huawei Cloud’s Ascend AI Cloud Service page describes its current offering.

Quick Recap

Bestseller No. 2
MX3 M.2 AI Accelerator
MX3 M.2 AI Accelerator
Software and Documentation can be accessed at the MemryX developer website
$169.00
Bestseller No. 3
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
✅Scalable, enabling simultaneous processing of multi-streams & multi-models; ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
$219.99
Bestseller No. 4
Tesla L40S 48GB AI HPC Graphics Accelerator
Tesla L40S 48GB AI HPC Graphics Accelerator
48GB AI graphics accelerator
$5,999.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.