October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk6 min

CitePulse: Auditing the Answer Layer

CitePulse’s three anonymized audits show why site readability, accurate citations, citation frequency, visibility, and agent task completion must be measured separately—and why its local-model results are only a proxy for commercial AI answer engines.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

CitePulse is presented as a way to audit several different questions about how a website appears in AI-generated answers—and whether a browser-driven agent can use it. Its central lesson is that readability, accurate citations, citation frequency, competitive visibility, and task completion are separate outcomes. In a DEV Community case study dated September 24, 2026, CitePulse maintainer Lawrence reports three anonymized audits; the results illustrate distinct failure profiles, not general performance benchmarks.

What does CitePulse audit?

The case study describes CitePulse v1.7.0 as a local-first, open-source tool under the MIT license. It says the cited run used Ollama with a local llama3.1:8b model. Those project and execution details are claims in Lawrence’s article; the repository, license file, and implementation were not independently verified here.

As an Amazon Associate I earn from qualifying purchases.

The instrument organizes its assessment around five principles: a machine can read the site; cited claims are supported by their cited pages; the site appears in real prompts relative to competitors; an autonomous agent can complete a task; and the tool reports “not determined” when a value cannot be measured honestly. The article describes nine KPIs spanning crawl accessibility, schema, llms.txt, citation correctness, citation rate, share of voice, interaction readiness, and task completion.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

These categories describe different stages of an answer-layer audit. A site may be accessible to a crawler but absent from tested answers. It may earn a citation that accurately supports a claim, yet appear infrequently. It may be visible in answers while a browser agent cannot complete a task. And an access barrier may prevent a fair task assessment altogether.

How to read the measures

Crawl and machine-readability checks

Accessibility, schema, and llms.txt probes concern whether a system can inspect or interpret site material. Passing such checks does not show that an answer engine will retrieve or cite the site. Lawrence’s article also notes a specific crawl-probe limitation: a WAF challenge page can return HTTP 200, making a nominally successful response a poor signal of usable access.

Citation correctness versus citation rate

Citation correctness asks whether the cited page supports the statement attributed to it. Citation rate asks how often the target appeared as a citation among the tested answers. Correct citations do not imply frequent citations; a low citation rate does not, by itself, show that any citation made was inaccurate. If there are no citations to judge, correctness is not determined rather than scored as a failure.

Share of voice

Share of voice measures relative visibility across the tested prompt set. Raw and weighted share are separate reported measures, but the article’s reported figures should not be read as market-wide visibility or as a universal measure of brand presence. They describe the tested prompts and comparison set.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Interaction readiness and task completion

Interaction readiness concerns whether a site is prepared for browser actions; task completion concerns whether an agent finishes the tested task. Neither is a measure of how persuasive or accurate a model’s answer is. A site can have a high interaction-readiness figure and still produce a low task-completion result, especially when the task sample is small or site access imposes constraints.

What powers the answer measurements

The article says citation and share metrics come from a local model synthesizing live web-search results. Lawrence calls this “a proxy for AI-answer-engine behavior, not a live query to ChatGPT, Perplexity, Gemini, or Copilot.” The reported answer metrics therefore are not direct tests of those named services. They reflect the described local-model workflow and should not be presented as results from those commercial answer engines.

What the three audits reported

All figures below are reported outputs from Lawrence’s CitePulse case study on DEV Community, published September 24, 2026, for CitePulse v1.7.0 runs dated September 24, 2026. The targets are anonymized, and the article’s author discloses that he maintains CitePulse. These three cases do not establish population-level benchmarks.

Target KPIs measured Citation correctness Citation rate Raw share of voice Weighted share of voice Interaction readiness Task completion
A, AI search-monitoring SaaS 9 of 9 100.0% (N=10) 55.6% (N=18) 91.3% (N=18) 89.1% (N=18) 74.3% (N=35) 33.3% (N=3)
B, European staffing and recruitment firm 6 of 9 Not determined: no citations to judge 0.0% (N=18) 0.0% (N=18) 91.7% (N=18) 85.7% (N=7) Not determined: sample below the floor
C, cooperative bank 5 of 9 100.0% (N=5) 33.3% (N=18) 86.5% (N=18) 91.2% (N=18) Not determined: authentication gated the probes Not determined: authentication gated the probes

Target A: accurate citations, limited task completion

For Target A, Lawrence reports that all 10 judgeable citations were supported by the cited pages: 100.0% citation correctness (N=10). The site appeared as a citation in 55.6% of 18 answers. Its raw share of voice was 91.3% (N=18), while weighted share was 89.1% (N=18). The reported task-completion rate was 33.3% (N=3), alongside interaction readiness of 74.3% (N=35). These outputs show why a strong citation-quality result and high reported visibility need not mean an agent can finish the tested task.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Target B: readable, but not cited in the tested answers

The article describes Target B as crawl-accessible but not cited in its tested prompt set: citation rate was 0.0% (N=18), and raw share of voice was 0.0% (N=18). Citation correctness was not determined because there were no citations to assess. Weighted share of voice was nevertheless reported as 91.7% (N=18), and interaction readiness as 85.7% (N=7). Task completion was not determined because the sample fell below the floor. The contrast is a reminder not to collapse distinct metrics—or infer citation quality from the absence of citations.

Target C: visibility with authentication-gated tasks

For the cooperative-bank target, the case study reports 100.0% citation correctness (N=5), a 33.3% citation rate (N=18), raw share of voice of 86.5% (N=18), and weighted share of voice of 91.2% (N=18). The article says only 6 of 18 answers cited the bank and that coverage varied by query. In the tested set, the bank was not cited for the basic identity question, “What is the bank?” Authentication gated the interaction and task probes, so those outcomes were not determined. An access barrier is not evidence that the agent would succeed or fail if it could reach the relevant workflow.

Why the verdict should not be a single average

Averaging these KPIs into one score can hide the question a site owner actually needs answered. Crawl access does not substitute for citation frequency; citation frequency does not establish correctness; relative share does not establish successful browser interaction; and an inaccessible or undersampled probe cannot responsibly be turned into a speculative score.

Lawrence summarizes the approach this way: “The verdict band is never the average of nine numbers; it is the report’s statement of the weakest load-bearing principle.” That framing is useful if the report keeps the underlying measurements visible. In Target A, citation correctness and task completion point to different strengths. In Target B, crawl access coexists with no citations in the tested answers. In Target C, citation and share measures can be reported even while authentication prevents task measurement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to compare audits responsibly

A comparison is meaningful only when the measurement conditions are sufficiently alike. Before treating a score change as progress or regression, check the following:

  • Prompt set and scheme: compare the same queries and prompt-construction method; a different set can change retrieval and share results.
  • Model and version: record the local model for each run. The article warns that historical runs using different local models may not form a like-for-like trend.
  • Dates and live-search conditions: record run dates and search context, since the answer metrics use live search results.
  • Access conditions: note crawler responses, WAF behavior, authentication requirements, and whether a nominal HTTP success actually exposes usable content.
  • Metric definitions: keep correctness separate from rate, raw share separate from weighted share, and both separate from interaction readiness and task completion.
  • Sample sizes and confidence floor: preserve each N and report when a sample is below the tool’s floor. A small sample can make an apparently precise percentage misleading.
  • Undetermined values: retain “not determined” for blocked, absent, gated, or insufficient data rather than replacing it with zero or an estimate.
  • Uncertainty: the article cautions that score changes without confidence intervals should not be treated as significant.

The case study says the three public-site targets were audited without prior arrangement and anonymized. That makes the examples useful as illustrations of how the measures can diverge, but it also means readers cannot inspect named targets or independently verify each target-specific condition from the article alone.

What the case study does—and does not—establish

The article reports that CitePulse runs locally and that no data leaves the machine; it also identifies the project as open-source and MIT-licensed. These are the maintainer’s claims, not independently verified operational guarantees or a review of the repository. The numbers likewise come from the maintainer-authored article, not an independent replication. Three anonymized audits cannot establish typical performance for websites, local models, or answer engines.

Read CitePulse’s results as a structured diagnostic of a particular test setup, not a universal ranking of how well a site performs in AI search. Its most useful contribution in this case study is the separation of failure modes—and the willingness to leave a result undetermined when the test cannot support a score.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.