Recommended Free Tools
CitePulse is presented as a way to audit several different questions about how a website appears in AI-generated answers—and whether a browser-driven agent can use it. Its central lesson is that readability, accurate citations, citation frequency, competitive visibility, and task completion are separate outcomes. In a DEV Community case study dated September 24, 2026, CitePulse maintainer Lawrence reports three anonymized audits; the results illustrate distinct failure profiles, not general performance benchmarks.
What does CitePulse audit?
The case study describes CitePulse v1.7.0 as a local-first, open-source tool under the MIT license. It says the cited run used Ollama with a local llama3.1:8b model. Those project and execution details are claims in Lawrence’s article; the repository, license file, and implementation were not independently verified here.
As an Amazon Associate I earn from qualifying purchases.
The instrument organizes its assessment around five principles: a machine can read the site; cited claims are supported by their cited pages; the site appears in real prompts relative to competitors; an autonomous agent can complete a task; and the tool reports “not determined” when a value cannot be measured honestly. The article describes nine KPIs spanning crawl accessibility, schema, llms.txt, citation correctness, citation rate, share of voice, interaction readiness, and task completion.
These categories describe different stages of an answer-layer audit. A site may be accessible to a crawler but absent from tested answers. It may earn a citation that accurately supports a claim, yet appear infrequently. It may be visible in answers while a browser agent cannot complete a task. And an access barrier may prevent a fair task assessment altogether.
#1 Best Overall
How to read the measures
Crawl and machine-readability checks
Accessibility, schema, and llms.txt probes concern whether a system can inspect or interpret site material. Passing such checks does not show that an answer engine will retrieve or cite the site. Lawrence’s article also notes a specific crawl-probe limitation: a WAF challenge page can return HTTP 200, making a nominally successful response a poor signal of usable access.
Citation correctness versus citation rate
Citation correctness asks whether the cited page supports the statement attributed to it. Citation rate asks how often the target appeared as a citation among the tested answers. Correct citations do not imply frequent citations; a low citation rate does not, by itself, show that any citation made was inaccurate. If there are no citations to judge, correctness is not determined rather than scored as a failure.
Share of voice
Share of voice measures relative visibility across the tested prompt set. Raw and weighted share are separate reported measures, but the article’s reported figures should not be read as market-wide visibility or as a universal measure of brand presence. They describe the tested prompts and comparison set.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchRank #2
Interaction readiness and task completion
Interaction readiness concerns whether a site is prepared for browser actions; task completion concerns whether an agent finishes the tested task. Neither is a measure of how persuasive or accurate a model’s answer is. A site can have a high interaction-readiness figure and still produce a low task-completion result, especially when the task sample is small or site access imposes constraints.
What powers the answer measurements
The article says citation and share metrics come from a local model synthesizing live web-search results. Lawrence calls this “a proxy for AI-answer-engine behavior, not a live query to ChatGPT, Perplexity, Gemini, or Copilot.” The reported answer metrics therefore are not direct tests of those named services. They reflect the described local-model workflow and should not be presented as results from those commercial answer engines.
What the three audits reported
All figures below are reported outputs from Lawrence’s CitePulse case study on DEV Community, published September 24, 2026, for CitePulse v1.7.0 runs dated September 24, 2026. The targets are anonymized, and the article’s author discloses that he maintains CitePulse. These three cases do not establish population-level benchmarks.
| Target | KPIs measured | Citation correctness | Citation rate | Raw share of voice | Weighted share of voice | Interaction readiness | Task completion |
|---|---|---|---|---|---|---|---|
| A, AI search-monitoring SaaS | 9 of 9 | 100.0% (N=10) | 55.6% (N=18) | 91.3% (N=18) | 89.1% (N=18) | 74.3% (N=35) | 33.3% (N=3) |
| B, European staffing and recruitment firm | 6 of 9 | Not determined: no citations to judge | 0.0% (N=18) | 0.0% (N=18) | 91.7% (N=18) | 85.7% (N=7) | Not determined: sample below the floor |
| C, cooperative bank | 5 of 9 | 100.0% (N=5) | 33.3% (N=18) | 86.5% (N=18) | 91.2% (N=18) | Not determined: authentication gated the probes | Not determined: authentication gated the probes |
Target A: accurate citations, limited task completion
For Target A, Lawrence reports that all 10 judgeable citations were supported by the cited pages: 100.0% citation correctness (N=10). The site appeared as a citation in 55.6% of 18 answers. Its raw share of voice was 91.3% (N=18), while weighted share was 89.1% (N=18). The reported task-completion rate was 33.3% (N=3), alongside interaction readiness of 74.3% (N=35). These outputs show why a strong citation-quality result and high reported visibility need not mean an agent can finish the tested task.
Target B: readable, but not cited in the tested answers
The article describes Target B as crawl-accessible but not cited in its tested prompt set: citation rate was 0.0% (N=18), and raw share of voice was 0.0% (N=18). Citation correctness was not determined because there were no citations to assess. Weighted share of voice was nevertheless reported as 91.7% (N=18), and interaction readiness as 85.7% (N=7). Task completion was not determined because the sample fell below the floor. The contrast is a reminder not to collapse distinct metrics—or infer citation quality from the absence of citations.
Target C: visibility with authentication-gated tasks
For the cooperative-bank target, the case study reports 100.0% citation correctness (N=5), a 33.3% citation rate (N=18), raw share of voice of 86.5% (N=18), and weighted share of voice of 91.2% (N=18). The article says only 6 of 18 answers cited the bank and that coverage varied by query. In the tested set, the bank was not cited for the basic identity question, “What is the bank?” Authentication gated the interaction and task probes, so those outcomes were not determined. An access barrier is not evidence that the agent would succeed or fail if it could reach the relevant workflow.
Rank #4
Why the verdict should not be a single average
Averaging these KPIs into one score can hide the question a site owner actually needs answered. Crawl access does not substitute for citation frequency; citation frequency does not establish correctness; relative share does not establish successful browser interaction; and an inaccessible or undersampled probe cannot responsibly be turned into a speculative score.
Lawrence summarizes the approach this way: “The verdict band is never the average of nine numbers; it is the report’s statement of the weakest load-bearing principle.” That framing is useful if the report keeps the underlying measurements visible. In Target A, citation correctness and task completion point to different strengths. In Target B, crawl access coexists with no citations in the tested answers. In Target C, citation and share measures can be reported even while authentication prevents task measurement.
How to compare audits responsibly
A comparison is meaningful only when the measurement conditions are sufficiently alike. Before treating a score change as progress or regression, check the following:
Best Value
- Prompt set and scheme: compare the same queries and prompt-construction method; a different set can change retrieval and share results.
- Model and version: record the local model for each run. The article warns that historical runs using different local models may not form a like-for-like trend.
- Dates and live-search conditions: record run dates and search context, since the answer metrics use live search results.
- Access conditions: note crawler responses, WAF behavior, authentication requirements, and whether a nominal HTTP success actually exposes usable content.
- Metric definitions: keep correctness separate from rate, raw share separate from weighted share, and both separate from interaction readiness and task completion.
- Sample sizes and confidence floor: preserve each N and report when a sample is below the tool’s floor. A small sample can make an apparently precise percentage misleading.
- Undetermined values: retain “not determined” for blocked, absent, gated, or insufficient data rather than replacing it with zero or an estimate.
- Uncertainty: the article cautions that score changes without confidence intervals should not be treated as significant.
The case study says the three public-site targets were audited without prior arrangement and anonymized. That makes the examples useful as illustrations of how the measures can diverge, but it also means readers cannot inspect named targets or independently verify each target-specific condition from the article alone.
What the case study does—and does not—establish
The article reports that CitePulse runs locally and that no data leaves the machine; it also identifies the project as open-source and MIT-licensed. These are the maintainer’s claims, not independently verified operational guarantees or a review of the repository. The numbers likewise come from the maintainer-authored article, not an independent replication. Three anonymized audits cannot establish typical performance for websites, local models, or answer engines.
Read CitePulse’s results as a structured diagnostic of a particular test setup, not a universal ranking of how well a site performs in AI search. Its most useful contribution in this case study is the separation of failure modes—and the willingness to leave a result undetermined when the test cannot support a score.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




