What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Les Kohn’s “multiple big chips” forecast is an argument about wide-operational-design-domain Level 4 autonomy—not a claim that every L4 vehicle must use a fixed number of processors. In a July 2023 EE Times interview, Ambarella’s then-CTO argued that the combination of richer sensor fusion, expanding AI workloads, safety redundancy and vehicle power limits will eventually make several high-performance automotive processors more practical than one monolithic device.
What Kohn’s L4 claim actually means
Level 4 automated driving means the vehicle can perform the driving task within a defined operational design domain (ODD). It does not mean unrestricted autonomy on every road, in every weather condition and without geographic or operational limits.
Kohn’s statement is narrower still: he was discussing wide-ODD L4, where the system must handle a comparatively broad range of roads, environments, traffic patterns and edge cases. A narrowly mapped L4 shuttle operating in favorable conditions may have very different computing requirements from a vehicle expected to work across a large region and a wide variety of situations.
The interview, published on July 5, 2023, presents Kohn’s strategic view of Ambarella’s automotive roadmap. It does not establish a universal industry requirement, specify how many chips a vehicle needs or provide independent benchmarks proving that Ambarella’s architecture is superior to competing approaches.
#1 Best Overall
- ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
- ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
- ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
- ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
- ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C
His underlying argument is that autonomous-driving compute does not scale only with the number of neural-network operations. It also scales with the amount of sensor data that must be moved, the need to run independent safety paths, the demand for future software headroom and the energy limits of an electric vehicle.
Why L4 workloads keep expanding
An advanced vehicle may combine many cameras with radar and other sensors. The system must then perform several jobs:
- detect and classify road users, signs, lanes and obstacles;
- estimate the vehicle’s position and the surrounding scene;
- fuse camera, radar and other observations;
- predict how nearby road users may behave;
- plan a safe trajectory; and
- monitor the main system and provide a fallback when something fails.
AI is also moving beyond isolated camera perception. Kohn described growing interest in using neural networks for deeper multi-sensor fusion and additional parts of the autonomy stack. More capable models can improve the system’s interpretation of complex scenes, but they generally bring greater demands for computation, memory and data movement.
Free tools Windows power users keep installed
One-click scans. No signup required.
Designers must also reserve capacity for difficult peak scenarios and future software updates. A processor sized only for today’s average workload may leave too little margin when models become larger, sensor configurations change or a safety monitor must run concurrently.
Sensor-level processing versus a domain controller
In a sensor-level design, each camera or other sensor has a relatively fixed processing allocation. That can be efficient when workloads are predictable, but it creates a mismatch between available compute and actual demand. A difficult scene may overwhelm one sensor processor while compute assigned to another sensor sits unused.
A centralized automotive domain controller can pool more of that capacity. It can also combine richer sensor data before independent processing has discarded information that might be useful to another algorithm.
| Architecture | Potential benefit | Engineering cost |
|---|---|---|
| Sensor-level AI | Local processing can reduce the amount of data sent onward and isolate workloads. | Fixed allocations may be inefficient, and early preprocessing can limit later fusion. |
| Single domain controller | Shared compute and centralized fusion can support more flexible workload balancing. | Bandwidth, memory, thermal, latency and safety concerns become concentrated in one system. |
| Multi-chip domain controller | Workloads can be partitioned across processors, with room for monitoring or independent paths. | Inter-chip communication, synchronization, software orchestration and system-level safety analysis become harder. |
Centralization is therefore not a free solution. It moves the design challenge toward high-bandwidth interconnects, memory architecture, deterministic scheduling, thermal management and fault containment.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →What Ambarella’s CV3-AD brings together
The interview describes Ambarella’s CV3-AD family as automotive domain-controller hardware for perception, multi-sensor fusion and path planning in L2+ through L4 applications. According to the article, the platform can process data from up to 20 image streams.
Rank #2
- High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
- Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
- Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
- Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
- Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.
Its heterogeneous processing architecture includes:
- a neural vector processing (NVP) engine for AI workloads;
- a general vector processor (GVP), described as especially suitable for radar processing;
- an image signal processor (ISP);
- stereo-processing engines;
- optical-flow engines; and
- video encoder engines.
This is not simply a general-purpose GPU replacement. The design combines different engines for different parts of the automotive pipeline. The intended advantage is that a workload can use a suitable block rather than forcing every task through the same large, power-hungry accelerator.
Why data movement can matter more than peak AI throughput
Kohn presented Ambarella’s NVP as an accelerator built around a data-flow programming model. Instead of treating a neural network as a conventional sequence of low-level instructions, the system represents operations such as convolution and matrix multiplication as a graph showing how data moves between operators.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsIn the described design, communication between operators can occur through on-chip memory rather than repeatedly going to external DRAM. That matters because moving data can consume substantial energy and bandwidth, even when the arithmetic itself is relatively inexpensive.
Kohn claimed this approach could be more than 10 times as efficient as a GPU-style approach for some data-movement patterns. That is an attributed executive and company claim, not a neutral benchmark reported by the interview. “Efficiency” can also mean different things: energy per inference, memory traffic, latency, silicon utilization or total accelerator power. Those metrics should not be treated as interchangeable.
Why raw-data fusion increases the burden
Independent sensor processing produces separate interpretations: a camera reports objects and lanes, while radar produces its own measurements. Centralized raw-data fusion can compare the observations earlier and preserve relationships that may be difficult to reconstruct after each sensor has compressed its information into a local result.
That potential benefit comes with a cost. The controller must ingest, store, synchronize and process a much larger volume of data. It needs sufficient memory capacity and bandwidth, predictable latency and a software architecture that can coordinate the different sensor streams.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Transformer-based networks are part of this discussion. Kohn said transformer support was becoming an important customer requirement, particularly for deep fusion across sensors, and said CV3-AD supported transformers. However, hardware support does not mean that every transformer architecture, sequence length or deployment configuration will run equally efficiently. It also does not by itself demonstrate production validation in a safety-critical vehicle.
The case for more than one large processor
Kohn’s roadmap argument combines four pressures.
- More compute: perception, fusion, prediction and planning are becoming more demanding.
- More redundancy: high-assurance systems need independent monitoring and fallback paths.
- More flexibility: different processors can handle different workloads and vehicle tiers.
- More headroom: production systems must accommodate peak scenarios and software evolution.
Several large processors can also distribute work and, potentially, heat across a system. Multiple devices may be easier to scale across L2+, L3 and wide-ODD L4 products than one enormous design. In some cases, using several moderately large dies can reduce the manufacturing and product-segmentation risks associated with one exceptionally large die.
Those are architectural implications, not benefits that the interview quantified. Multiple processors can introduce their own problems: duplicated memory, inter-chip traffic, synchronization overhead, additional board and power-delivery complexity, and a more difficult system-level safety case.
“Multiple chips” could mean several identical accelerators, heterogeneous processors, separate autonomy and safety computers, distributed domain controllers or a multi-die package. Kohn’s interview does not specify one of these implementations.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Redundancy is not automatically safety
Kohn argued that increasingly complex L3 and L4 systems need redundancy because both classical algorithms and deep-learning systems can make mistakes. He discussed using a classical checker alongside a learned system and suggested that, ultimately, two independent deep-learning implementations may be needed.
The important word is independent. Two copies that share the same model, data, software defect or sensor failure may share the same failure mode. Genuine diversity requires careful treatment of implementation, training, inputs, fault containment, diagnostics and common-cause failures.
Kohn’s view is that independent implementations could provide diversity for high-assurance operation. It should not be paraphrased as a claim that two neural networks automatically deliver ASIL-D compliance or replace a complete safety case. Functional safety also involves requirements, architectural analysis, verification, validation, diagnostic coverage and applicable automotive standards.
Sparse computation: less work, with conditions
Ambarella’s approach, as characterized by Kohn, emphasizes random sparsity: any weight may become zero, rather than only weights in a prescribed channel or position. Some other approaches use structured pruning, such as removing entire channels, or fixed patterns that retain a limited number of nonzero values within a group.
When more than half the weights are zero, the described hardware can avoid processing the zero values. In principle, this reduces arithmetic and memory movement. But nominal sparsity does not guarantee a real-world speedup. The compiler, model structure, memory layout and hardware utilization all matter.
Rank #4
- ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
- ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
- ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
- ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
- ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C
Sparsity can also reduce accuracy, especially in rare cases that matter greatly to an autonomous vehicle. The interview says Ambarella’s toolchain gradually sparsifies networks and retrains after each step to limit the loss. That is a company description, not an independently demonstrated result across production workloads.
Precision: why 4-bit is not the whole story
The NVP supports 16-bit, 8-bit and 4-bit precision, according to the interview. Lower precision can reduce storage, bandwidth and arithmetic cost, but neural networks do not respond identically at every layer.
Weights are generally easier to compress below 8 bits than activations. Some layers may work entirely in 4-bit arithmetic, while others may require 16-bit activations to preserve accuracy. A mixed-precision design can therefore be more practical than forcing the whole network to use one format.
Quantization may sometimes be calibrated using representative data without full retraining. Pushing closer to the hardware’s performance limits, however, may require quantization-aware retraining and renewed validation. In an automotive system, the question is not merely whether a compressed model runs faster; it is whether it maintains acceptable behavior across ordinary and rare safety-relevant conditions.
Why the GVP matters
Kohn described the GVP as particularly suitable for radar algorithms. Workloads with relatively little convolution or matrix multiplication can reportedly run on the GVP at similar speed to the NVP while using less power because the GVP is a smaller silicon block.
That comparison is another attributed architectural claim rather than a published independent benchmark. Its practical value depends on the actual radar algorithms, input rates, memory behavior and system integration.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Why not specialize every workload?
Highly specialized accelerators can be extremely efficient when workloads are stable and well understood. The risk is that an automotive chip remains in service for many years while neural-network architectures, sensor configurations and software requirements change.
More programmable hardware can adapt more easily, though it may sacrifice some peak efficiency. More specialization can improve efficiency for today’s model while making tomorrow’s model awkward or expensive to support. New hardware and software configurations may also require additional safety validation.
Best Value
- DEEPX DX-M1M NPU: Powered by the DEEPX DX-M1M neural processing unit, purpose-built for efficient on-device AI inference workloads.
- COMPACT M.2 2242 FORM FACTOR: Fits the standard M.2 2242 slot, making it easy to integrate into embedded systems, edge devices, and compact computing platforms.
- EDGE AI ACCELERATION: Designed to accelerate deep learning inference at the edge, enabling real-time AI applications without relying on cloud connectivity.
- RADXA AICORE MODULE: The Radxa AICore DX-M1M delivers a plug-and-play AI compute solution ideal for robotics, smart cameras, and industrial automation.
- WARRANTY AND ORIGIN: Backed by a 1-year manufacturer warranty and crafted with quality components for reliable long-term performance in demanding environments.
Kohn’s 2023 assessment was that further specialization was premature because AI workloads were still changing rapidly. That is a time-specific strategic judgment, not a permanent rule against dedicated engines.
The RISC-V question
Kohn said Ambarella had considered RISC-V but identified obstacles in matching the performance of high-end Arm processors, meeting automotive functional-safety requirements and winning customer acceptance in a conservative market.
Ambarella had internal processor designs based on OpenRISC, which predates RISC-V, and Kohn suggested those designs could potentially be adapted. He described a broader goal of using a common architecture for the main processor and other on-chip components.
An open instruction-set architecture can offer flexibility and reduce dependence on a particular licensing model, but openness alone does not solve performance, safety certification, toolchain maturity, software compatibility or automaker confidence. Kohn’s comments should therefore be read as his assessment of the adoption barriers at the time, not as a standards-based conclusion that RISC-V is unsuitable for automotive use.
What the roadmap implies
The roadmap described in the interview separates product needs by vehicle capability:
- smaller, more cost-effective processors for L2 and L2+ systems;
- larger and faster devices as workloads rise; and
- multiple large chips for wide-ODD L4.
This is the direct context for the headline. It is Ambarella’s 2023 roadmap direction and Kohn’s forecast, not a confirmed production configuration or an industry-wide consensus.
What remains unproven
The interview does not provide the system-level data needed to decide whether a multi-chip architecture is better for a particular vehicle. It does not disclose:
Recommended Free Tools
- required or delivered TOPS;
- actual chip or vehicle-level power consumption;
- thermal-design-power figures;
- memory capacity and bandwidth;
- inter-chip bandwidth or end-to-end latency;
- cost, packaging or reliability data;
- independent comparisons with competing platforms; or
- a completed safety case demonstrating the proposed redundancy.
It is also important to separate several meanings of “efficiency.” Arithmetic efficiency, memory-bandwidth efficiency, energy per inference, accelerator power, total compute-system power and software productivity are different measurements. A lower-power accelerator may not reduce vehicle-level energy use if additional processors, memory, networking and cooling offset the saving.
Bottom line
Kohn’s “multiple big chips” prediction is best understood as a strategic hardware forecast: wide-ODD L4 autonomy will need to balance growing AI workloads with raw-data fusion, safety diversity, software evolution and electric-vehicle power limits. Several processors may offer a more scalable path than one enormous chip, but they also create communication, thermal, software and safety challenges. The interview makes a compelling architectural case; it does not prove that every L4 vehicle will require the same multi-chip design.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

