Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Phishing simulations and course-completion figures can show what employees did in a particular test or whether they finished a course. On their own, they do not show that learning lasted or that real-world risk fell. Research points instead to a harder problem: measuring behavior fairly across messages, people, and work contexts, then linking that behavior to meaningful security outcomes.
Why common awareness metrics leave important questions unanswered
Organizations often measure activity because it is easy to count. In its 2022 study of U.S. federal cybersecurity awareness programs, the National Institute of Standards and Technology (NIST) found that 84% used training-completion rates as a common effectiveness measure. The report also found that 85% performed phishing simulations, which respondents often described as a successful aspect of their programs. Those figures describe reported practices and perceptions, not proof that training reduced phishing risk. NISTIR 8420A
As an Amazon Associate I earn from qualifying purchases.
The same report illustrates the measurement gap: 44% of participants said they faced challenges determining program effectiveness, and 48% reported difficulty correlating security-incident data with behaviors their awareness programs targeted. More than half of surveyed programs used behavior-based measures such as clicks or phishing reports, but collecting those measures did not automatically establish a connection to downstream incidents.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchThree kinds of measures answer different questions
- Activity measures, such as course completion or the number of campaigns, show whether a program was delivered or completed. They do not establish what participants learned.
- Proximal behavior measures, such as clicking a lure, submitting credentials, or reporting a suspicious message, show what happened during a particular interaction. They are more directly tied to a targeted behavior, but depend on the test and its circumstances.
- Downstream outcomes, such as incident data, concern security consequences. Linking those outcomes to a specific training behavior can be difficult, as federal-program participants reported to NIST.
This is a practical way to distinguish the measures described in NISTIR 8420A, not a standardized NIST scoring framework. A useful evaluation asks which question each measure can answer, rather than treating all counts as interchangeable evidence of effectiveness.
#1 Best Overall
Why a click rate depends on the message and the workplace
A simulation result is partly a measure of the lure. In a 2024 NIST study at a university, Phish Scale difficulty ratings tracked observed click rates closely. That means two click rates are hard to compare fairly if one message was more difficult to recognize than the other. The finding is specific to that study setting; it does not establish a universal conversion between lure difficulty and click probability. NIST’s repeat-click study
The study sent participants eight messages over four weeks: four phishing messages and four controls. It also reported associations between repeat clicking and participant characteristics, including less time working online, checking email more often, a more internally oriented locus of control, and lower need for cognition. These are associations, not evidence that any one characteristic causes susceptibility. The researchers noted that the study took place soon after COVID-19 shutdowns of in-person classes, which may have influenced its results.
Rank #2
Work context matters too. NIST researchers examined approximately 70 stratified staff members at a U.S. government research institution, drawing on 4.5 years of embedded phishing-exercise data and focusing on the last three exercises with participant feedback. Their 2018 study found that workplace context shaped how participants interpreted email cues, including when the same cues appeared with different premise-and-context alignment. This small, specific workplace study offers insight into interpretation, not a population-wide estimate of susceptibility. NIST’s user-context study
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →What studies say about courses, games, and embedded exercises
The available findings do not identify one training format as universally best. They examine different interventions, participants, settings, and outcomes, so their results should be read within those boundaries.
| Study | Setting and scope | Reported finding |
|---|---|---|
| 2024 education-versus-game study | 115 participants; quantitative comparison of awareness strategies | Authors reported benefits from both traditional courses and interactive games and called for continuous reevaluation. Predictive models using demographic characteristics do not by themselves establish that those characteristics cause vulnerability. |
| 2025 healthcare field experiment | One large healthcare organization; randomized experiment over eight months, ten campaigns, and more than 19,500 employees | The abstract reports no significant relationship between recent annual training and simulation failure, very small absolute differences in failure rates across embedded-training content, and minimal time spent interacting with embedded material in the wild. |
The healthcare experiment’s results matter, but they describe one organization and the outcomes reported in its abstract. They do not establish that every course or training design is ineffective. Likewise, the 2024 study’s reported benefits do not prove that either format produces lasting changes in real-world behavior.
A separate 2024 paper on the human factor in phishing describes Spamley, a system intended to collect and share user behavior as people read emails with varied phishing features and attack strategies. It treats richer observation of user behavior as a research area; its available abstract does not establish a general effect size for training. The human-factor paper
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to evaluate an awareness program more carefully
Before interpreting a campaign result, define the behavior the program is meant to change and the evidence that would count as progress. These questions help keep an evaluation tied to its actual goal:
- What behavior is targeted? Distinguish recognizing a suspicious message from reporting it, avoiding a link, or withholding credentials. Pick a measure that corresponds to the behavior of interest.
- What exactly is being counted? Record whether the result concerns completion, clicks, credential submission, reports, or an incident-related behavior. Do not present one as a substitute for another.
- How difficult was the test? Use a consistent approach to describing lure difficulty. When comparing campaigns, account for differences in the messages rather than assuming every click rate reflects the same challenge.
- Who was tested, and in what work context? Consider role, work demands, email habits, and whether the message premise fits participants’ responsibilities. A result from one team or institution may not transfer directly to another.
- What intervention did participants actually experience? Separate a passive course from an interactive activity, immediate feedback, or an embedded exercise. For an embedded intervention, consider whether participants engaged with it.
- How long was behavior followed? A single simulation captures a moment. Repeated measurement and a stated follow-up period can help establish whether observed behavior persists, though they do not alone prove changes in organizational risk.
- Can the measure be connected to a meaningful outcome? If the program claims to reduce incidents, explain how incident data relate to the behavior being trained and acknowledge any limits in that connection.
- Could the evaluation support learning without unfairly labeling people? Treat an individual result as evidence about a specific interaction, not a permanent risk ranking. Provide safe ways to report messages and use results to improve the program.
This approach is a practical synthesis of issues raised across the studies, not a validated universal KPI or scoring system. NIST’s federal-program report also found that participants wanted additional guidance and government-specific benchmarking data, underscoring the difficulty of deciding what counts as success across organizations.
What the evidence can—and cannot—settle
The evidence challenges simple success claims based on simulation activity or course completion, but it does not support the opposite blanket claim that awareness training never works. NIST’s 2022 report describes programs that use simulations and several kinds of measures while reporting difficulty assessing effectiveness. Later studies show why results can vary with message difficulty, work context, intervention, and study design.
NIST’s report also notes that awareness programs may be perceived by employees as “boring, ‘check-the-box’ activity.” That characterization describes a possible workforce perception, not a finding that every program is viewed that way. NISTIR 8420A
A stronger evidence base would follow behavior over time, account for context and test difficulty, and make the path from training to security outcomes clearer. The current studies leave room for better evaluation rather than a single metric or format that can stand in for every organization.
Recommended Free Tools
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




