Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

A peer-reviewed study found that eight pedestrian-detection systems used in autonomous-driving research missed children more often than adults and darker-skinned pedestrians more often than light-skinned pedestrians. The measured gaps were 20.14 and 7.52 percentage points, respectively. The researchers tested image-based detectors—not named commercial robotaxis or complete vehicle systems—so the findings identify a serious weakness in technology relevant to self-driving cars, not the error rates of every vehicle on the road.

What the researchers tested

The paper, “Bias Behind the Wheel: Fairness Testing of Autonomous Driving Systems,” was published online on November 2, 2024, in ACM Transactions on Software Engineering and Methodology. The team evaluated eight pedestrian detectors used in autonomous-driving research: YOLOX, RetinaNet, Faster R-CNN, Cascade R-CNN, ALFNet, CSP, MGAN and PRNet. King’s College London’s publication record identifies the peer-reviewed paper and its publication details.

The analysis covered 8,311 real-world images, with 16,070 gender labels, 20,115 age labels and 3,513 skin-tone labels. Researchers compared how often the detectors failed to identify pedestrians across demographic groups. The final paper reports 8,311 images; some earlier coverage used 8,111, while the 2023 preprint also reported a slightly different age disparity. The figures below are from the final publication. The full paper describes the methods and results.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A detector is only one part of a vehicle’s perception and safety system. It identifies a person in an image or video frame. Tracking, sensor fusion, movement prediction, path planning, braking and fallback behavior are separate stages. The study evaluated pedestrian detection, not that entire chain.

#1 Best Overall
LK COKOINO Arduino Robot Car Kit - 4WD Smart Robot Car Chassis with Motors, Wheels and Battery Case for Arduino R3/R4/Leonardo/Raspberry Pi 5/4B/3B+/3B/2B/1B+
  • This is a newly designed 4-wheel car frame that can be used with other devices to realize function of tracing, obstacle avoidance, distance testing, autonomous driving, wireless remote control, etc.
  • The smart robot car chassis has plenty of fixed mounting holes and room for expansion to add various sensors, actuators and controllers (such as Arduino, Raspberry Pi, Micro bit).
  • 4WD Robot Car Kit maximum load 1KG; size of robot car chassis: 10*6*2.5 inches; wheel diameter: 2.56 inches
  • 4 pcs TT Robot Gear Motor; Operating voltage: 3V~12VDC (recommended operating voltage of about 6 to 8V) Wires Length: 0.8 inch 24 AWG; Maximum torque: 800gf cm min (3V) ; No-load speed: 1:48 (3V)
  • The DIY car kit will be easy to assemble according to the instructions we provide.It also comes with a battery case that can hold two 18650 batteries (batteries not included)

How large were the reported gaps?

Comparison Result in the tested systems
Children versus adults Children had a 20.14-percentage-point higher miss rate.
Darker-skinned versus light-skinned pedestrians Darker-skinned pedestrians had a 7.52-percentage-point higher miss rate.
Male versus female pedestrians The average difference was about 1.1 percentage points.

These are percentage-point differences in miss rates, not statements that a group was a corresponding percentage less likely to be detected. For example, a 20.14-point gap means the measured miss rate for children exceeded the adult miss rate by 20.14 points; it does not mean the relative chance of detection fell by 20.14 percent. The preprint’s age-gap figure was 19.67 points; the peer-reviewed version reports 20.14. The earlier version is available on arXiv.

What “bias” means here

In this study, bias means unequal measured performance: the detectors missed members of some groups more often than others. It does not mean the software had an intention or attitude toward those people. Uneven training examples, image quality, detector design and the conditions in a scene can all contribute to a performance gap.

Rank #2
HIWONDER Robot Car with ChatGPT Large AI Models, 3D Depth Camera Ackermann Chassis ROS2-HUMBLE Lidar SLAM Mapping Navigation Autonomous Driving, MentorPi A1 Standard Kit with Raspberry Pi5 8GB
  • For Raspberry Pi 5 & ROS2 Robot Car. MentorPi A1 smart AI robot car is powered by Raspberry Pi 5, compatible with ROS2, and programmed in Python, making it an ideal platform for AI robot development.
  • High-Performance Hardware. Equipped with Ackerman chassis, closed-loop encoder motors, TOF lidar, depth camera, AI voice interaction box, and other advanced components to ensure optimal performance and efficiency.
  • Advanced AI Capabilities. Supports SLAM mapping, path planning, multi-robot coordination, vision recognition, target tracking, and more, covering a wide range of AI applications.
  • Autonomous Driving with Deep Learning. Utilizes YOLO model training to enable road sign and traffic light recognition, along with other autonomous driving features, helping users explore and develop autonomous driving technologies.
  • Empowered by Large AI Model, Human-Robot Interaction Redefined. MentorPi AI robot car deploys multimodal models with ChatGPT at its core, integrating 3D vision and Al voice interaction box. This synergy enhances its perception, reasoning, and actuation capabilities, enabling advanced embodied AI applications and delivering natural, context-aware human-robot interaction.

The study grouped people by skin tone. Skin tone is a continuum, and those analytical categories are not interchangeable with race, ethnicity or nationality. The results should not be stretched into a claim about every racial or ethnic group.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why children may be harder for detectors to identify

The measured disparity is the study’s finding; the reasons for it are not all established by the result alone. Children are generally shorter than adults and may occupy fewer pixels in a camera image, particularly at a distance. Their bodies may also be partly hidden by adults, vehicles or street furniture. Differences in proportions, movement and the amount or variety of child examples in training data are other plausible factors. These possibilities help explain what should be investigated, but do not by themselves prove why each detector missed a particular child.

Rank #3
HIWONDER Robot Car with ChatGPT Large AI Models, 3D Depth Camera Ackermann Chassis ROS2-HUMBLE Lidar SLAM Mapping Navigation Autonomous Driving, MentorPi A1 Advanced Kit with Raspberry Pi5 8GB
  • For Raspberry Pi 5 & ROS2 Robot Car. MentorPi A1 smart AI robot car is powered by Raspberry Pi 5, compatible with ROS2, and programmed in Python, making it an ideal platform for AI robot development.
  • High-Performance Hardware. Equipped with Ackerman chassis, closed-loop encoder motors, TOF lidar, depth camera, AI voice interaction box, and other advanced components to ensure optimal performance and efficiency.
  • Advanced AI Capabilities. Supports SLAM mapping, path planning, multi-robot coordination, vision recognition, target tracking, and more, covering a wide range of AI applications.
  • Autonomous Driving with Deep Learning. Utilizes YOLO model training to enable road sign and traffic light recognition, along with other autonomous driving features, helping users explore and develop autonomous driving technologies.
  • Empowered by Large AI Model, Human-Robot Interaction Redefined. MentorPi AI robot car deploys multimodal models with ChatGPT at its core, integrating 3D vision and Al voice interaction box. This synergy enhances its perception, reasoning, and actuation capabilities, enabling advanced embodied AI applications and delivering natural, context-aware human-robot interaction.

The broad label “child” can also conceal variation: a toddler, a school-age child and a teenager differ in size, proportions and movement. Evaluations that report only an overall child category may miss distinct failure patterns within it.

Why lighting and contrast matter

The researchers found that the disparity for darker-skinned pedestrians increased in low-brightness or low-contrast scenarios. King’s College London’s explanation of the study also points to imbalanced training data, with major pedestrian datasets containing more light-skinned than dark-skinned people.

Rank #4
HUPILAN LDW Target Board Compatible with Benz,ADAS Camera Calibration Tool
  • 1.Fit For: LDW ADAS calibration tool compatible with Benz,-Please confirm whether your car model match before purchasing
  • 2.Without Stand: Please note that this product does not include a set of stand
  • 3.Size And Color:100% match in size and color of the original manufacturer calibration boards. This ensures accurate and reliable calibration results for your LDW system
  • 4.Material: Unlike soft paper alternatives, our calibration boards are tangible and hard aluminum alloy , providing a solid surface for precise calibration
  • 5.Easy To Use: LDW Pattern Board for precise static front camera aiming and ADAS calibration

This is not simply a matter of cameras being unable to see dark skin. Detection depends on the whole image: exposure and dynamic range, contrast with the background, clothing, shadows, glare, motion blur, occlusion, weather and image processing all matter. Night driving, backlighting, poorly lit roads, wet pavement and dark clothing are examples of conditions that merit careful testing; the study does not show that darkness alone explains every miss.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Does this prove commercial self-driving cars are unsafe?

No. The researchers tested eight research detectors on annotated image datasets, not proprietary perception systems from Waymo, Cruise, Tesla, Zoox or another named manufacturer. The study therefore does not establish that every production robotaxi or driver-assistance system has these same miss rates. A commercial vehicle may use different models, cameras, additional sensors and safety layers. The available results do not show how those specific systems perform for each group.

Best Value
BEV Vision Kit for reComputer GMSL Series
  • BEV Vision Kit for reComputer GMSL series

A missed detection is a potential safety concern, but it is not a collision measurement. A detector may miss a person, identify them late, draw an imprecise box or detect them correctly while a later subsystem fails. The paper measured subgroup differences in pedestrian detection; it did not calculate real-world collision rates or establish a particular increase in crash risk.

The findings remain relevant despite older accounts describing the work as not yet peer-reviewed: the paper later appeared in the ACM journal in 2024. They also fit a wider machine-learning fairness problem. The stakes are especially high in driving because a detection failure can affect decisions made by a moving vehicle, but this study does not show that autonomous vehicles alone have demographic performance gaps.

What would reduce the risk?

The study does not prove that any single remedy will eliminate the measured disparities. It does point toward safeguards that manufacturers, evaluators and regulators can require or investigate:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Broader training and evaluation data: Include varied skin tones, ages, body sizes and geographies, and check whether the examples are sufficiently numerous and representative.
  • Subgroup testing under difficult conditions: Report results separately across age and skin-tone groups, and test low light, glare, shadows, rain, fog, occlusion and low contrast rather than relying only on aggregate accuracy.
  • Multiple safety layers: Evaluate how detection interacts with tracking, prediction, planning and braking. Where appropriate, sensor fusion and conservative fallback behavior can add defenses, though the study did not test their effectiveness.
  • Independent audits and public reporting: Third-party assessments and clear subgroup performance reporting would make it easier to detect gaps that an overall score can hide.
  • Monitoring after deployment: Systems need ongoing incident review, mechanisms for reporting failures and a clear process to correct or withdraw software when evidence shows a safety problem.

The authors say they released code, data and results to support further work, according to the publication record. Reproducible evaluation can help other researchers test whether the disparities persist across datasets and newer systems.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.