DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
World desk10 min

RAG Citation Verification: Building Deterministic Byte-Span Validators in TypeScript

How to check that a quoted RAG passage really sits at the cited byte range: preserve source bytes, validate offsets, compare bytes under one encoding policy, and keep tolerant matches and semantic support as separate verdicts.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To verify a RAG citation deterministically, keep the original source bytes, record every citation as a byte range into those bytes, and check that the slice at that range equals the cited text encoded with the same policy. Do the check outside the model, in plain code, so the same inputs always give the same verdict.

The usual reason this fails in TypeScript is that JavaScript string indices count UTF-16 code units, while byte offsets count bytes in the encoded file. They agree only for ASCII. This article covers the data model, a working validator, how to get chunk offsets right when chunks overlap, and where a passing check stops proving anything.

What a byte-span citation is

SitePoint’s September 18, 2026 tutorial (by “SitePoint Team”) defines a byte span as a (start, end) range in the original source buffer. Its assertion carries sourceId, byteStart, byteEnd and citedText. The validator resolves the source, slices the asserted range, encodes citedText with the same encoding, and compares the two byte sequences. The tutorial uses the outcomes VERIFIED, PARTIAL_MATCH and UNGROUNDED. This is one tutorial’s model, not a standardized RAG protocol, but it is a sound way to get a literal provenance check.

The guarantee is narrow and that is its value. For a fixed source, a fixed range convention and a fixed encoding policy, bounds checks plus byte equality have exactly one answer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why string indices and byte offsets disagree

JavaScript strings are sequences of UTF-16 code units, so "…".length and slice() work in those units. UTF-8 uses one to four bytes per code point. A citation produced from string positions and checked against a Buffer will drift as soon as non-ASCII text appears earlier in the document.

Text Code points UTF-16 units (.length) UTF-8 bytes
A 1 1 1
é (U+00E9, precomposed) 1 1 2
é as e + U+0301 (decomposed) 2 2 3
日本語 3 3 9
😀 (U+1F600) 1 2 (surrogate pair) 4

The precomposed and decomposed rows matter separately: they look identical on screen but are different byte sequences. Unicode normalization is its own transformation, distinct from encoding.

Decide which bytes the offsets point to

Do this before writing any validator code. Retain the source at ingestion together with its identity, byte length, encoding and a stable content hash or version, so offsets cannot be checked against a replaced document.

Rank #2
TypeScript Programming Language - Software Engineer & Coder T-Shirt
  • TypeScript implements a superset of syntax for strictly typed development, facilitating deep static analysis and enhanced development environment integration. The compiler translates source into standard script formats, ensuring parity across any runtime.
  • TypeScript is ideal for front-end developers, full-stack engineers, and software architects who build large-scale web applications. It serves those looking to improve code excellence, reduce bugs through static checking, and maintain complex projects more.
  • Lightweight, Classic fit, Double-needle sleeve and bottom hem
  • Original file bytes. Strongest provenance, but only practical when the file is already text (plain text, Markdown, JSON, source code).
  • Canonical extracted-text bytes. For PDF, HTML or Office input, offsets into the original file are not offsets into the text you extracted. If the pipeline stores extracted text, treat that UTF-8 buffer as the source of record, hash it, and describe citations as offsets into it. Re-extracting with a different parser version creates a new source version.
  • Normalized text. If you apply NFC or any other normalization, either do it once before offsets are captured and store the result as the canonical bytes, or do not do it at all. Normalizing one side later silently changes byte identity.

The WHATWG Encoding Standard recommends UTF-8 for new protocols and formats and warns of security problems when a producer and consumer disagree about encoding, which is the same hazard here in miniature. If a file is not valid UTF-8, or was decoded and re-encoded before offsets were recorded, UTF-8 offsets may not match the original file’s byte positions. A leading byte-order mark is also part of the bytes: keep it in the buffer and count it, rather than letting a decoder strip it and shifting everything by three.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Ingestion: store bytes and reject malformed UTF-8 early

Node’s TextDecoder accepts fatal: true, which makes malformed input throw instead of being silently replaced with U+FFFD. Using it at ingestion means a corrupt document is caught once, rather than producing mysterious mismatches later. (Node’s Buffer.toString() does the replacing, so avoid it for integrity checks.)

import { createHash } from "node:crypto";

const encoder = new TextEncoder(); // Node's TextEncoder only supports UTF-8
const strictUtf8 = new TextDecoder("utf-8", { fatal: true });

export interface SourceRecord {
  id: string;
  version: string; // sha256 of the exact bytes offsets refer to
  bytes: Buffer;
}

export function ingest(id: string, raw: Buffer): SourceRecord {
  strictUtf8.decode(raw); // throws TypeError on malformed UTF-8
  const version = createHash("sha256").update(raw).digest("hex");
  return { id, version, bytes: raw };
}

Node’s documentation states plainly that “All instances of TextEncoder only support UTF-8 encoding.” That is convenient here, since it pins the citation side to the same encoding as the stored source.

Getting chunk offsets right

Retrieval returns chunks, and the model quotes from them, so a citation’s offsets must be translated into source positions. This is where most real-world breakage happens.

Contiguous chunks: accumulate byte lengths

If chunks are adjacent, non-overlapping and cover the document, each chunk’s start is the sum of the previous chunks’ encoded byte lengths. Add a consistency check that the final total equals the source length. The tutorial gives exactly this caveat: accumulation only holds for contiguous, non-overlapping chunks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Overlap, gaps or repeated text: capture boundaries from the splitter

With overlap, summing lengths overstates every later offset. Searching for chunk text with Buffer.indexOf from a maintained cursor works, but is ambiguous when identical text occurs more than once, and may land on the wrong occurrence. The safer design is for the splitter itself to emit its boundaries. A splitter that works directly on bytes makes offsets true by construction, as long as it never cuts inside a multi-byte character:

const isContinuation = (b: number) => (b & 0xc0) === 0x80;

export interface Span { byteStart: number; byteEnd: number }

// maxBytes must be at least 4 so every UTF-8 character can fit.
export function splitBytes(buf: Buffer, maxBytes: number, overlap: number): Span[] {
  const spans: Span[] = [];
  let start = 0;
  while (start < buf.length) {
    let end = Math.min(start + maxBytes, buf.length);
    while (end < buf.length && isContinuation(buf[end])) end--;
    spans.push({ byteStart: start, byteEnd: end });
    if (end >= buf.length) break;

    let next = end - overlap;
    while (next > start && isContinuation(buf[next])) next--;
    start = next > start ? next : end; // always make progress
  }
  return spans;
}

This cuts at character boundaries, not word or sentence boundaries. If your real splitter works on strings (most do), convert each boundary it reports from a UTF-16 index to a byte offset instead of guessing later:

// O(n) per call; fine at ingestion, precompute a table for very large documents.
// utf16Index must not fall between the two halves of a surrogate pair.
export const byteOffsetOf = (text: string, utf16Index: number) =>
  Buffer.byteLength(text.slice(0, utf16Index), "utf8");

A boundary that splits a surrogate pair leaves a lone surrogate, which encodes as the three-byte replacement character and corrupts the count. That is the usual cause of “citations break with emoji.”

The validator

The function below validates inputs first, then compares bytes, and only then, if you enable it, tries a nearby-window search that returns a different verdict. It returns structured reason codes rather than a single generic failure, so operators can tell a model that fabricated a quote from a pipeline bug.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
export type Verdict = "VERIFIED" | "PARTIAL_MATCH" | "UNGROUNDED" | "INVALID_INPUT";
export type Reason =
  | "EXACT_MATCH" | "NEARBY_MATCH"
  | "UNKNOWN_SOURCE" | "BYTES_DIFFER"
  | "STALE_SOURCE" | "BAD_OFFSETS" | "OUT_OF_BOUNDS"
  | "EMPTY_CITATION" | "MALFORMED_CITATION_TEXT";

export interface Assertion {
  sourceId: string;
  sourceVersion: string;
  byteStart: number; // inclusive
  byteEnd: number;   // exclusive: half-open [start, end)
  citedText: string;
}

export interface Result {
  verdict: Verdict;
  reason: Reason;
  matchedStart?: number;
  matchedEnd?: number;
}

const encoder = new TextEncoder();

export function verify(
  a: Assertion,
  sources: Map<string, SourceRecord>,
  opts: { tolerant: boolean; window: number } = { tolerant: false, window: 256 },
): Result {
  const src = sources.get(a.sourceId);
  if (!src) return { verdict: "UNGROUNDED", reason: "UNKNOWN_SOURCE" };
  if (src.version !== a.sourceVersion)
    return { verdict: "INVALID_INPUT", reason: "STALE_SOURCE" };

  const { byteStart: s, byteEnd: e } = a;
  if (!Number.isSafeInteger(s) || !Number.isSafeInteger(e) || s < 0 || s > e)
    return { verdict: "INVALID_INPUT", reason: "BAD_OFFSETS" };
  if (e > src.bytes.length)
    return { verdict: "INVALID_INPUT", reason: "OUT_OF_BOUNDS" };

  if (a.citedText.length === 0)
    return { verdict: "INVALID_INPUT", reason: "EMPTY_CITATION" };
  // Requires Node 20+ and an ES2024 lib setting for the type.
  if (!a.citedText.isWellFormed())
    return { verdict: "INVALID_INPUT", reason: "MALFORMED_CITATION_TEXT" };

  const want = encoder.encode(a.citedText);
  if (src.bytes.subarray(s, e).equals(want))
    return { verdict: "VERIFIED", reason: "EXACT_MATCH", matchedStart: s, matchedEnd: e };

  if (opts.tolerant) {
    const trimmed = encoder.encode(a.citedText.trim());
    if (trimmed.length > 0) {
      const lo = Math.max(0, s - opts.window);
      const hi = Math.min(src.bytes.length, e + opts.window);
      const at = src.bytes.subarray(lo, hi).indexOf(trimmed);
      if (at !== -1)
        return {
          verdict: "PARTIAL_MATCH",
          reason: "NEARBY_MATCH",
          matchedStart: lo + at,
          matchedEnd: lo + at + trimmed.length,
        };
    }
  }
  return { verdict: "UNGROUNDED", reason: "BYTES_DIFFER" };
}

Several choices in this sketch are policy, not law, and you should write yours down:

  • Range convention. It uses half-open [start, end). The tutorial’s example treats a zero-length slice as valid; the sketch rejects empty citations because an empty span cannot ground anything.
  • Unknown source. Treated as ungrounded because a model can invent an ID. If your IDs come from a registry you control, you may prefer an input error.
  • Trim only. The tolerant branch trims surrounding whitespace and searches nearby. The tutorial’s whitespace and punctuation rules are examples, not universal ones.
  • Do not swallow exceptions. If a bug or corrupt buffer throws, let it surface as an operational error rather than folding it into UNGROUNDED, which would blame the model for your data problem.

Note also that TextEncoder.encodeInto() reports read (UTF-16 code units consumed) and written (UTF-8 bytes produced). Only written is a byte length; mixing them up reintroduces the original bug.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Verdicts and what they mean

Verdict / reason Meaning Suggested handling
VERIFIED / EXACT_MATCH The cited bytes sit exactly at the asserted range of that source version. Render as a literal-quote citation.
PARTIAL_MATCH / NEARBY_MATCH The text occurs close by, after trimming. The submitted offsets were not right. Show as weaker evidence, optionally rewrite offsets to matchedStart/matchedEnd, and count separately in metrics.
UNGROUNDED The source is unknown or the bytes differ and no tolerant match was found. Block, annotate, or retry generation.
INVALID_INPUT The assertion is malformed, out of bounds, empty, or points at a replaced source version. Investigate: often a pipeline or schema bug, not a hallucination.

A window match means the bytes exist nearby, not that the model’s offsets were correct; keep it out of any “exact” statistic. Because it only searches near the asserted range, it also cannot be used as a substitute for a global search, which would happily “verify” a quote copied from the wrong passage.

What a passing check does not prove

An exact match proves the literal text exists at that location in that source version. It does not prove that the passage supports the claim it is attached to, that the right document was retrieved, that the source is authoritative or current, or that the answer interprets the passage correctly. Entailment, authority, freshness and citation completeness need separate evaluation, such as a natural-language-inference model or an LLM judge run after this deterministic gate. The byte check is the cheap, reproducible first layer: it removes fabricated and misplaced quotes so the expensive semantic checks only see real ones.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Design choices and trade-offs

Choice Stricter option Looser option Trade-off
Matching Exact bytes only Tolerant/nearby match Provenance strength versus recovery from formatting drift
Offset capture Taken from the splitter Reconstructed later with indexOf Reliable identity versus convenience and duplicate-text ambiguity
Offset target Original file bytes Canonical extracted text Fidelity to the input versus offsets that are easier to use in text workflows
Decoding fatal: true Replacement characters Fail-fast integrity versus continuing with possible byte/text disagreement
On failure Block the answer Annotate or retry User trust versus availability, latency and operational complexity

Placing it in the pipeline

The SitePoint tutorial runs validation as post-generation middleware in a LangChain sequence, after retrieval, prompting and generation. Its pipeline example uses placeholder retriever, prompt and validator declarations, so treat it as showing where the step goes, not as a ready integration. The pieces you still need:

  1. Structured citation output. Have the model emit citations in a schema (sourceId, sourceVersion, byteStart, byteEnd, citedText) and validate the schema before calling verify. Models are unreliable at counting bytes, so a practical pattern is to give them chunk IDs and have your code supply the offsets for chunk-relative quotes.
  2. A complete extractor for every citation format the model might produce, so nothing escapes verification by being unparsed.
  3. A failure policy: block, annotate or retry. Expose exact and partial outcomes to the UI so rendering can distinguish them.
  4. Privacy-aware logging. Log source IDs, versions, offsets and reason codes; avoid storing cited text unless you need to and are permitted to.
  5. Streaming behavior. Decide whether you validate after the full response or as each citation completes; a citation that fails after the user has already seen it needs a retraction path.

Performance claims

The tutorial describes a benchmark fixture of 1,000 citations across 50 documents totaling roughly 200 KB and says performance depends on hardware, document size and citation density. No reproducible results table or independent benchmark accompanies it, so there is no latency figure to quote. The operation itself is a bounds check, a slice and a byte comparison; profile it on your own documents and citation volume before making any service-level promise. The code in this article is a sketch to adapt, not a tested library.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.