Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
OpenAI’s Whisper speech-recognition model was found to generate fluent text that did not appear in the source audio, including invented medications, racial descriptions, violent statements and sexual content. The concern became especially serious because Whisper-based technology was being used in medical documentation workflows, where an apparently plausible mistake can become part of a patient’s permanent health record.
The findings came from a 2024 investigation and related research. They demonstrate a significant reliability and governance risk—not proof that every hospital record was wrong or that a specific patient was harmed.
The short version
- Whisper is an automatic speech-recognition model from OpenAI. It converts speech into text; it is not itself a diagnostic system.
- Nabla Copilot, a medical documentation product, used Whisper-based speech recognition to help draft clinical notes.
- Researchers and developers reported that Whisper sometimes produced words, phrases or entire sentences that were never spoken.
- Reported fabrications included nonexistent treatments or medications, racial commentary, violent statements, sexual content and unrelated phrases generated after the audio had ended.
- More than 30,000 clinicians across roughly 40 health systems were reported to have used the Whisper-based Nabla tool in 2024.
- The evidence establishes a dangerous capability and a deployment risk, not a universal failure rate or a documented list of patient injuries.
The original investigation was published in October 2024. The reported deployment figures and product behavior should therefore not be treated as verified measurements of the product’s status in 2026.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →What Whisper is—and what it is not
OpenAI Whisper is an automatic speech-recognition system designed to transcribe spoken audio and, in some cases, translate speech. It predicts the text that most likely corresponds to an audio recording. It does not have a human-like guarantee that every word in its output was actually spoken.
#1 Best Overall
- 【PCM Recording and Automatic Noise Reduction】:This digital voice recorder is equipped with advanced dual noise reduction microphones and supports 1536 kbps PCM HD audio recording, ensuring crystal-clear sound capture in any environment. Recorder device with automatic noise reduction and voice-activated recording, the recorder only picks up the sound when there’s speech, reducing background noise,Excellent sound quality can meet the needs of students, journalists, music lovers and more people
- 【136GB Memory and Long Battery Life】Voice Recorder with Playback with 8GB built-in storage and includes a complimentary 128GB TF card, this digital voice recorder can hold up to 9775 hours of recordings in MP3 format or WAV format;Recorder for lectures with a built-in 1100mAh rechargeable lithium battery, this voice recorder can continuously record for up to 68 hours on a single charge, making it perfect for back-to-back meetings, interviews, or extended classroom sessions
- 【One Click Record and Save】: Our voice recorder supports one click recording and saving functions. Even when the product is in a powered-off state, simply push up the side recording button to immediately enter recording mode, and push down the recording button to save the recording. This allows for capturing as much information as possible.Easily transfer your recordings to your computer using the USB-C connection, allowing for fast and secure file management
- 【Easy-to-Use】This portable voice recorder is designed with a simple, user-friendly interface featuring a large, easy-to-read LCD screen. The voice-activated recording (VOR) feature makes hands-free operation a breeze. With one-touch recording, users can start or stop recording instantly, even during busy moments. A-B repeat function and password protection ensure that important segments are easily accessible and secure
- 【Portable and Durable Design】Designed with portability in mind, this lightweight screen recorder fits comfortably in your pocket or bag, weighing only 97 grams. Its sleek and durable metal casing ensures longevity and protection from everyday wear and tear. Whether you’re traveling, in the office, or attending a lecture, this compact recorder is always ready to capture clear, high-quality audio
That distinction matters because “transcription,” “medical scribe” and “clinical record” describe different layers:
| Layer | Function | Primary risk |
|---|---|---|
| Speech recognition | Converts audio into text | Mishearing or inventing text |
| Medical scribe | Organizes, summarizes and formats a visit into a draft note | Adding, omitting or misinterpreting information |
| Clinical record | Stores information used by future clinicians and administrators | False information persisting and influencing later decisions |
Whisper itself was not autonomously diagnosing patients in the reported workflow. The issue was that its output could feed a documentation tool, which could then produce a polished note that looked more authoritative than the underlying audio justified. OpenAI introduced Whisper as a general speech-recognition model in its original model announcement.
How the hospital documentation workflow worked
The relevant chain can be simplified as:
Patient visit → audio capture → speech recognition → medical formatting or summarization → clinician review → electronic health record
Nabla developed a medical documentation tool using Whisper-based speech recognition and fine-tuned technology for medical language. Its purpose was to reduce the time clinicians spend taking notes and allow more attention during patient visits.
The Associated Press reported that more than 30,000 clinicians across approximately 40 health systems had used the tool. The reporting named the Mankato Clinic in Minnesota and Children’s Hospital Los Angeles among the users. It also reported an estimate of approximately seven million medical visits processed, attributed to the company and reporting rather than independently audited as a count of affected or erroneous records.
Rank #2
- 【Simple Operation】- switch on your voice recorder, one button for recording. press the "REC", start the recording, press "STOP", end the recording, press “PLAY”, listen what you just recorded, and then Press A-B, select your important section to repeat. Easy to playback with inner powerful speaker, support external sound speaker playback, let you enjoy superior recording quality.
- 【Clear Voice Record】- high quality recording with noise redution, you will get super clear recorded voice, the sensitive microphone help you to catch speaker's words in an interview, lectures, meetings.
- 【Voice Activated Recording】- automatic voice reduction function, it starts recording when sound is detected or turn to standby state, saving recording time and reduce power consumption.
- 【 Player Function】- this voice recorder can be used as an music player, you could enjoy the music after your tired study, meeting and so on. Also can function as a detachable data storage device.you can take along your favorite pictures and documents whenever you go.Simply cut-and-paste or drag-and -drop files to or from it via USB connection, the player will appear as a removeable drive in Windows.
- 【High quality and long time】 uses DSP noise reduction technology to filter out environmental noise, has high-quality recording, 【1536kbps】to restore the real scene. It can continuously record for more than 30 hours and play for 7 hours.
“Used by hospitals” should not be interpreted as “Whisper made autonomous medical decisions.” In the described workflow, it helped transcribe or draft documentation from conversations. The safety problem arose when unsupported text entered a note that clinicians, other providers, insurers or legal authorities might later treat as factual.
What kinds of details did Whisper make up?
Researchers, engineers and developers interviewed by the AP said they found hallucinated content—text unsupported by the recording—in several categories:
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall- Nonexistent medications or medical treatments
- Racial descriptions that were absent from the audio
- Violent statements or events
- Sexual content and descriptions of sexual acts
- Extra sentences generated after the spoken audio had ended
- Unrelated phrases such as “like and subscribe”
- Invented context that changed the apparent meaning of a conversation
These examples should not be read as the most common Whisper errors, nor as evidence that every Nabla note contained such material. They are significant because the consequences of a rare but highly plausible fabrication can be far more severe than an ordinary misspelling.
A garbled word may prompt a clinician to ask for clarification. A fluent sentence saying that a patient takes a particular medication, made a violent statement or disclosed a sexual act can pass a quick visual review—especially when the reviewer remembers the visit imperfectly or cannot access the original recording.
How often did hallucinations occur?
There is no single error rate that can be applied to every Whisper deployment. The reported investigations used different model versions, audio conditions, languages, recording lengths, datasets and definitions of “hallucination.” Their results must be kept separate.
Rank #3
- Clear PCM Recording: Adopts upgraded noise cancelling microphone with professional recording chip. Capture 1536Kbps premium quality sound. Voice recorder with playback function, which is well designed for the users to easily access. Customer Service includes real life phone call from a specialist to give instructions on this high-quality recording device. We ensure your satisfaction on this product.
- 128GB Digital Recorder, Computers Compatible: stores 9296hours of recording, or 40,000songs, up to 54 hours of continuous recording with full battery. Recording can be pre-set into mp3 128kbps,192kbps, or wav 1536kbps format. A wonderful voice recording device for lectures, meetings, and conversations.
- Voice Activated Recorder: This recorder device can set voice decibels at 6 different levels. Regardless the level of the volume, with correct voice decibel level, this recorder will catch talking voice only, reduce blank and whispering snippet.
- Powerful Feature: Multi-usage as a voice recorder, an USB flash drive, and a Mp3 Player. Newly developed 4-folder storage(A/B/C/D) for file management make your recording and other files more organized. Many other helpful features like password protection, A-B repeat, auto record, bookmark, ideal recorder for lectures, meetings, speeches, and interviews.
- Fast File Download: V618 can easily transfer files onto computers. A rechargeable voice recorder that can be quickly recharged, suit for students, teachers, seniors, businesspeople, writers, and bloggers
The AP described observations including:
- One machine-learning engineer who found hallucinations in roughly half of more than 100 hours of Whisper transcriptions he examined.
- Another developer who reported hallucinations in nearly all of 26,000 transcripts.
- A University of Michigan researcher who found hallucinations in eight of ten inspected public-meeting transcriptions before attempting to improve the system.
- A study by researchers associated with Cornell University and the University of Virginia that examined thousands of short audio samples and documented hallucinations, including harmful insertions.
These observations do not mean that half of hospital visits were corrupted, or that nearly every clinical transcript was unusable. Public meetings and other research samples may have very different acoustics and speaking patterns from clinical encounters. Conversely, aggregate accuracy can conceal serious failures in specific conditions such as silence, overlapping speech or medical terminology.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
The paper Careless Whisper: Speech-to-Text Hallucination Harms examines harmful hallucinations and disparities, including issues affecting speech-impaired users. Research has also investigated long-form transcription and hallucinations caused by non-speech audio, including long-form Whisper behavior and hallucinations induced by non-speech audio.
Why does a speech model invent words?
Whisper is a predictive model, not a recording verifier. When audio is unclear, it attempts to produce likely language rather than reliably stopping at “inaudible.” Several conditions can increase uncertainty:
- Silence or dead air
- Background noise, music, television or alarms
- Low-volume speech
- Accents and unfamiliar pronunciations
- Interruptions and overlapping speakers
- Speech impairments such as aphasia or dysarthria
- Long recordings processed in segments
Long-form transcription can introduce repetition, drift or invented continuation. A model may continue generating text after a speaker has stopped. Background sounds can also be interpreted as speech-like input. The result is not that the model is “lying”; the more accurate description is model-generated text unsupported by the source audio.
Medical context can make this output look even more credible. A plausible drug name, symptom or treatment may fit the surrounding note so naturally that a reader does not recognize it as an insertion.
Recommended Free Tools
Rank #4
- ENHANCED CONTEXT WITH MULTIMODAL INPUT: Capture audio, type notes, add images, and press to highlight key moments for richer context. During recording, instantly mark key moments with a single button press. Simultaneously enrich your audio by snapping photos of important documents or typing in ideas
- CHAT WITH YOUR RECORDINGS USING "ASK Plaud": Unlock deeper insights with this interactive AI. Ask questions, extract key points, draft emails, and get next-step suggestions—all grounded in your original audio for reliable, ready-to-use answers
- INTELLIGENT RECORDING WITH AI DIRECTIONAL AUDIO: Enjoy seamless, intelligent recording with Plaud Note Pro. Its AI automatically switches between call and meeting modes while recording, while directional audio and real-time spatial awareness minimize noise to capture voices with crystal clarity
- Everything Included: Includes Plaud Note Pro, magnetic case, magnetic ring, charging cable, and a free Starter Plan with 300 transcription minutes per month. Upgrade anytime in the Plaud app to Pro Plan (1,200 min/mo) or Unlimited Plan(Up to 24 hours of transcription per user per day)
- PREMIUM ULTRA-SLIM DESIGN WITH INSTANTVIEW DISPLAY: Meticulously designed, the AI Note Taker is just 0.12 inches thin and 1.06 oz —about the size of a credit card. Its sleek aluminum body with a textured wave finish features a vivid AMOLED display, letting you check battery and recording status at a glance, while it seamlessly works with Apple Find My to ensure you never misplace it
Why a hallucination is more dangerous in a medical record
The central risk chain is:
Audio ambiguity → false transcript → clinician-approved note → persistent medical record → downstream clinical decision.
A fabricated detail could:
- Misstate a patient’s symptoms or medical history
- Add a medication the patient never took
- Influence a diagnosis or treatment plan
- Mislead the next clinician who relies on the previous note
- Create racial, psychiatric, sexual or demographic bias
- Affect referrals, billing, disability claims or insurance disputes
- Damage a patient’s reputation or relationships
- Create consent, privacy, compliance and liability problems
Once a false statement enters a chart, it may be copied into later notes. That “copy-forward” effect can make an unverified AI output appear to have been independently confirmed over time.
The AP quoted experts describing the potential consequences of fabricated medical documentation as grave. But the available reporting does not establish a specific patient death, misdiagnosis, prescription error or lawsuit caused by a Whisper hallucination. A serious safety risk should not be overstated into a proven injury claim.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Human review is necessary—but not automatically sufficient
Nabla reportedly required clinicians to review and approve generated notes. That is an important safeguard, but “a human was in the loop” does not prove that errors were reliably detected.
Review becomes weaker when:
- The clinician is rushed or reviewing a long note
- The fabricated sentence is fluent and medically plausible
- The clinician remembers the conversation imperfectly
- The source audio is unavailable
- The error appears in a summary or templated section rather than an obvious transcript
- The clinician assumes the software is accurate
A meaningful review process needs time, training, clear responsibility and enough source material to verify disputed claims. Approval should not be treated as independent confirmation unless the reviewer actually had a practical way to compare the note with the encounter.
Best Value
- 【One Click Record and Save】This voice recorder features instant one-click recording and saving. Even when powered off, simply push up the side button to start recording and push down to save. Designed with ergonomic controls, this digital voice recorder ensures fast operation so you never miss important moments—perfect as a voice recorder with playback, mini recorder device, or portable recorder for interviews, lectures, and field work
- 【64GB Memory & High-Capacity Battery】Equipped with a built-in 64GB TF card, this recorder device stores up to 4,600 hours of recordings. Its 600mAh battery supports up to 48 hours of continuous use (MP3 at 32kbps). Ideal for students, journalists, and professionals, this tape recorder portable mini excels in lectures, meetings, interviews, and even for paranormal sound research
- 【PCM Recording & Automatic Noise Reduction】Capture audio in WAV format with up to 1536kbps PCM quality. Advanced noise reduction minimizes background sounds, delivering crystal-clear playback on headphones or professional gear. This makes it an excellent audio recorder, digital audio recorder, or sound recorder for music creation, interviews, and high-detail sound archiving
- 【Voice-Activated Recorder, Big Screen & Password Protection】The voice activated recorder automatically starts/stops when sound reaches your set level, helping save storage and battery. A large 1.44-inch screen offers easy navigation, while password protection safeguards your files—perfect for storing personal memos and important audio files when using it as a dictaphone voice recorder or recording device for professional use
- 【Multi-Function Recorder】This versatile digital recorder supports internal and external recording, file segmentation, scheduled recording, A-B loop playback, MP3 music, and bookmarking. Functions as a USB storage drive and MP3 player with quick transfer via USB cable. Great as a pocket recorder, lecture recorder, mini voice recorder, or recording devices for travel and daily use
The source-audio retention trade-off
The AP reported that Nabla deleted original audio recordings for data-safety reasons. Nabla’s CTO said clinicians were expected to quickly edit and approve the resulting notes.
Deleting audio can reduce the amount of sensitive information exposed in a breach and limit retention risks. But the recording is also the closest available ground truth for investigating a disputed note. If the audio disappears immediately, a hospital may be unable to determine whether a statement was spoken, misheard or invented.
A responsible policy must balance:
- Patient privacy and security
- Consent and legally required retention rules
- Auditing and incident investigation
- Correction of inaccurate records
- Access controls and encryption
- Retention limits proportionate to clinical need
Deleting recordings is not automatically privacy-preserving if it leaves no defensible audit trail. Hospitals should document what is retained, for how long, who can access it and how patients can challenge an inaccurate note.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsOpenAI’s high-risk-use warning
The AP reported that OpenAI warned Whisper should not be used in “high-risk domains,” including contexts where errors could have serious consequences. That warning is a central governance issue in the story, but it should not be turned into a blanket legal conclusion.
There are important distinctions between:
- Using the open-source Whisper model directly
- Using an API or hosted service
- Using a vendor’s medically fine-tuned product
- Deploying the system inside a hospital’s own environment
- Using AI to draft a note versus allowing it to generate orders, prescriptions, diagnoses or triage decisions
A vendor or hospital may add controls that are not present in the base model. Those controls still need independent validation. A general model warning does not answer whether a particular clinical workflow is safe, and medical fine-tuning does not guarantee that every output is grounded in the audio.
What hospitals should require before deployment
- Test representative clinical audio. Include different accents, languages, ages, speech impairments, microphones, room layouts and background-noise conditions.
- Measure hallucinations separately from word-error rate. A transcript can have a good average accuracy score while still inserting one dangerous medication or allegation.
- Require review before signing. AI-generated notes should remain drafts until an appropriately trained clinician verifies them.
- Make verification practical. Provide secure access to source audio or a defensible alternative audit trail, with documented retention and deletion rules.
- Pin model versions. Record which model and software produced each note, and require change notification and revalidation after updates.
- Block autonomous clinical action. Do not allow an unverified draft to place orders, prescribe medication, diagnose a condition or make triage decisions.
- Log and investigate incidents. Fabricated medications, diagnoses, allegations and demographic details should trigger defined reporting and correction procedures.
- Protect patient rights. Explain the use of encounter recordings and AI documentation, obtain consent where required and provide a clear route to amend inaccurate records.
- Assess vendors beyond accuracy claims. Review security, subcontractors, data retention, EHR integration, auditability, business-associate terms and clinician-review requirements.
Questions patients and clinicians can ask
- Is the AI-generated note reviewed before it becomes part of the official record?
- Can the reviewer compare the note with the original audio?
- How long is encounter audio retained, and who can access it?
- What happens when the vendor changes its model or software?
- Are hallucinations and corrections logged?
- Can the system generate or trigger medication orders, diagnoses or referrals?
- Which vendors and subcontractors handle the encounter audio?
- How can a patient request correction of an inaccurate AI-generated statement?
What remains unknown
The 2024 reporting leaves several questions open:
- Whether later Whisper or Nabla versions materially reduced hallucinations
- How often fabricated text entered finalized medical records
- Whether patients experienced measurable clinical harm
- How individual hospitals handled audio deletion, auditing and correction
- Whether comparable commercial medical scribes experience the same failure mode at similar rates
- What the products’ deployment scale and model architecture were after the 2024 investigation
Those unknowns are precisely why hospitals should not rely on a vendor’s overall transcription-accuracy number. They need workflow-specific evidence about unsupported insertions and whether staff can detect them before they propagate.
The broader lesson
The important distinction is not between an AI note and a human note. Humans also make documentation errors. The issue is whether the system produces errors that are fluent, plausible and difficult to trace back to the source.
A tool that saves clinicians time is not safe merely because a clinician can theoretically review its output. Safety depends on whether the workflow makes errors visible, whether source evidence remains available when appropriate, whether changes are auditable and whether false information can be corrected before it influences future care.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

