Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchTo verify a RAG citation deterministically, keep the original source bytes, record every citation as a byte range into those bytes, and check that the slice at that range equals the cited text encoded with the same policy. Do the check outside the model, in plain code, so the same inputs always give the same verdict.
The usual reason this fails in TypeScript is that JavaScript string indices count UTF-16 code units, while byte offsets count bytes in the encoded file. They agree only for ASCII. This article covers the data model, a working validator, how to get chunk offsets right when chunks overlap, and where a passing check stops proving anything.
What a byte-span citation is
SitePoint’s September 18, 2026 tutorial (by “SitePoint Team”) defines a byte span as a (start, end) range in the original source buffer. Its assertion carries sourceId, byteStart, byteEnd and citedText. The validator resolves the source, slices the asserted range, encodes citedText with the same encoding, and compares the two byte sequences. The tutorial uses the outcomes VERIFIED, PARTIAL_MATCH and UNGROUNDED. This is one tutorial’s model, not a standardized RAG protocol, but it is a sound way to get a literal provenance check.
The guarantee is narrow and that is its value. For a fixed source, a fixed range convention and a fixed encoding policy, bounds checks plus byte equality have exactly one answer.
Recommended Free Tools
#1 Best Overall
Why string indices and byte offsets disagree
JavaScript strings are sequences of UTF-16 code units, so "…".length and slice() work in those units. UTF-8 uses one to four bytes per code point. A citation produced from string positions and checked against a Buffer will drift as soon as non-ASCII text appears earlier in the document.
| Text | Code points | UTF-16 units (.length) |
UTF-8 bytes |
|---|---|---|---|
A |
1 | 1 | 1 |
é (U+00E9, precomposed) |
1 | 1 | 2 |
é as e + U+0301 (decomposed) |
2 | 2 | 3 |
日本語 |
3 | 3 | 9 |
😀 (U+1F600) |
1 | 2 (surrogate pair) | 4 |
The precomposed and decomposed rows matter separately: they look identical on screen but are different byte sequences. Unicode normalization is its own transformation, distinct from encoding.
Decide which bytes the offsets point to
Do this before writing any validator code. Retain the source at ingestion together with its identity, byte length, encoding and a stable content hash or version, so offsets cannot be checked against a replaced document.
Rank #2
- TypeScript implements a superset of syntax for strictly typed development, facilitating deep static analysis and enhanced development environment integration. The compiler translates source into standard script formats, ensuring parity across any runtime.
- TypeScript is ideal for front-end developers, full-stack engineers, and software architects who build large-scale web applications. It serves those looking to improve code excellence, reduce bugs through static checking, and maintain complex projects more.
- Lightweight, Classic fit, Double-needle sleeve and bottom hem
- Original file bytes. Strongest provenance, but only practical when the file is already text (plain text, Markdown, JSON, source code).
- Canonical extracted-text bytes. For PDF, HTML or Office input, offsets into the original file are not offsets into the text you extracted. If the pipeline stores extracted text, treat that UTF-8 buffer as the source of record, hash it, and describe citations as offsets into it. Re-extracting with a different parser version creates a new source version.
- Normalized text. If you apply NFC or any other normalization, either do it once before offsets are captured and store the result as the canonical bytes, or do not do it at all. Normalizing one side later silently changes byte identity.
The WHATWG Encoding Standard recommends UTF-8 for new protocols and formats and warns of security problems when a producer and consumer disagree about encoding, which is the same hazard here in miniature. If a file is not valid UTF-8, or was decoded and re-encoded before offsets were recorded, UTF-8 offsets may not match the original file’s byte positions. A leading byte-order mark is also part of the bytes: keep it in the buffer and count it, rather than letting a decoder strip it and shifting everything by three.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Ingestion: store bytes and reject malformed UTF-8 early
Node’s TextDecoder accepts fatal: true, which makes malformed input throw instead of being silently replaced with U+FFFD. Using it at ingestion means a corrupt document is caught once, rather than producing mysterious mismatches later. (Node’s Buffer.toString() does the replacing, so avoid it for integrity checks.)
import { createHash } from "node:crypto";
const encoder = new TextEncoder(); // Node's TextEncoder only supports UTF-8
const strictUtf8 = new TextDecoder("utf-8", { fatal: true });
export interface SourceRecord {
id: string;
version: string; // sha256 of the exact bytes offsets refer to
bytes: Buffer;
}
export function ingest(id: string, raw: Buffer): SourceRecord {
strictUtf8.decode(raw); // throws TypeError on malformed UTF-8
const version = createHash("sha256").update(raw).digest("hex");
return { id, version, bytes: raw };
}
Node’s documentation states plainly that “All instances of TextEncoder only support UTF-8 encoding.” That is convenient here, since it pins the citation side to the same encoding as the stored source.
Getting chunk offsets right
Retrieval returns chunks, and the model quotes from them, so a citation’s offsets must be translated into source positions. This is where most real-world breakage happens.
Contiguous chunks: accumulate byte lengths
If chunks are adjacent, non-overlapping and cover the document, each chunk’s start is the sum of the previous chunks’ encoded byte lengths. Add a consistency check that the final total equals the source length. The tutorial gives exactly this caveat: accumulation only holds for contiguous, non-overlapping chunks.
Overlap, gaps or repeated text: capture boundaries from the splitter
With overlap, summing lengths overstates every later offset. Searching for chunk text with Buffer.indexOf from a maintained cursor works, but is ambiguous when identical text occurs more than once, and may land on the wrong occurrence. The safer design is for the splitter itself to emit its boundaries. A splitter that works directly on bytes makes offsets true by construction, as long as it never cuts inside a multi-byte character:
const isContinuation = (b: number) => (b & 0xc0) === 0x80;
export interface Span { byteStart: number; byteEnd: number }
// maxBytes must be at least 4 so every UTF-8 character can fit.
export function splitBytes(buf: Buffer, maxBytes: number, overlap: number): Span[] {
const spans: Span[] = [];
let start = 0;
while (start < buf.length) {
let end = Math.min(start + maxBytes, buf.length);
while (end < buf.length && isContinuation(buf[end])) end--;
spans.push({ byteStart: start, byteEnd: end });
if (end >= buf.length) break;
let next = end - overlap;
while (next > start && isContinuation(buf[next])) next--;
start = next > start ? next : end; // always make progress
}
return spans;
}
This cuts at character boundaries, not word or sentence boundaries. If your real splitter works on strings (most do), convert each boundary it reports from a UTF-16 index to a byte offset instead of guessing later:
// O(n) per call; fine at ingestion, precompute a table for very large documents.
// utf16Index must not fall between the two halves of a surrogate pair.
export const byteOffsetOf = (text: string, utf16Index: number) =>
Buffer.byteLength(text.slice(0, utf16Index), "utf8");
A boundary that splits a surrogate pair leaves a lone surrogate, which encodes as the three-byte replacement character and corrupts the count. That is the usual cause of “citations break with emoji.”
The validator
The function below validates inputs first, then compares bytes, and only then, if you enable it, tries a nearby-window search that returns a different verdict. It returns structured reason codes rather than a single generic failure, so operators can tell a model that fabricated a quote from a pipeline bug.
Best Value
export type Verdict = "VERIFIED" | "PARTIAL_MATCH" | "UNGROUNDED" | "INVALID_INPUT";
export type Reason =
| "EXACT_MATCH" | "NEARBY_MATCH"
| "UNKNOWN_SOURCE" | "BYTES_DIFFER"
| "STALE_SOURCE" | "BAD_OFFSETS" | "OUT_OF_BOUNDS"
| "EMPTY_CITATION" | "MALFORMED_CITATION_TEXT";
export interface Assertion {
sourceId: string;
sourceVersion: string;
byteStart: number; // inclusive
byteEnd: number; // exclusive: half-open [start, end)
citedText: string;
}
export interface Result {
verdict: Verdict;
reason: Reason;
matchedStart?: number;
matchedEnd?: number;
}
const encoder = new TextEncoder();
export function verify(
a: Assertion,
sources: Map<string, SourceRecord>,
opts: { tolerant: boolean; window: number } = { tolerant: false, window: 256 },
): Result {
const src = sources.get(a.sourceId);
if (!src) return { verdict: "UNGROUNDED", reason: "UNKNOWN_SOURCE" };
if (src.version !== a.sourceVersion)
return { verdict: "INVALID_INPUT", reason: "STALE_SOURCE" };
const { byteStart: s, byteEnd: e } = a;
if (!Number.isSafeInteger(s) || !Number.isSafeInteger(e) || s < 0 || s > e)
return { verdict: "INVALID_INPUT", reason: "BAD_OFFSETS" };
if (e > src.bytes.length)
return { verdict: "INVALID_INPUT", reason: "OUT_OF_BOUNDS" };
if (a.citedText.length === 0)
return { verdict: "INVALID_INPUT", reason: "EMPTY_CITATION" };
// Requires Node 20+ and an ES2024 lib setting for the type.
if (!a.citedText.isWellFormed())
return { verdict: "INVALID_INPUT", reason: "MALFORMED_CITATION_TEXT" };
const want = encoder.encode(a.citedText);
if (src.bytes.subarray(s, e).equals(want))
return { verdict: "VERIFIED", reason: "EXACT_MATCH", matchedStart: s, matchedEnd: e };
if (opts.tolerant) {
const trimmed = encoder.encode(a.citedText.trim());
if (trimmed.length > 0) {
const lo = Math.max(0, s - opts.window);
const hi = Math.min(src.bytes.length, e + opts.window);
const at = src.bytes.subarray(lo, hi).indexOf(trimmed);
if (at !== -1)
return {
verdict: "PARTIAL_MATCH",
reason: "NEARBY_MATCH",
matchedStart: lo + at,
matchedEnd: lo + at + trimmed.length,
};
}
}
return { verdict: "UNGROUNDED", reason: "BYTES_DIFFER" };
}
Several choices in this sketch are policy, not law, and you should write yours down:
- Range convention. It uses half-open
[start, end). The tutorial’s example treats a zero-length slice as valid; the sketch rejects empty citations because an empty span cannot ground anything. - Unknown source. Treated as ungrounded because a model can invent an ID. If your IDs come from a registry you control, you may prefer an input error.
- Trim only. The tolerant branch trims surrounding whitespace and searches nearby. The tutorial’s whitespace and punctuation rules are examples, not universal ones.
- Do not swallow exceptions. If a bug or corrupt buffer throws, let it surface as an operational error rather than folding it into
UNGROUNDED, which would blame the model for your data problem.
Note also that TextEncoder.encodeInto() reports read (UTF-16 code units consumed) and written (UTF-8 bytes produced). Only written is a byte length; mixing them up reintroduces the original bug.
Verdicts and what they mean
| Verdict / reason | Meaning | Suggested handling |
|---|---|---|
VERIFIED / EXACT_MATCH |
The cited bytes sit exactly at the asserted range of that source version. | Render as a literal-quote citation. |
PARTIAL_MATCH / NEARBY_MATCH |
The text occurs close by, after trimming. The submitted offsets were not right. | Show as weaker evidence, optionally rewrite offsets to matchedStart/matchedEnd, and count separately in metrics. |
UNGROUNDED |
The source is unknown or the bytes differ and no tolerant match was found. | Block, annotate, or retry generation. |
INVALID_INPUT |
The assertion is malformed, out of bounds, empty, or points at a replaced source version. | Investigate: often a pipeline or schema bug, not a hallucination. |
A window match means the bytes exist nearby, not that the model’s offsets were correct; keep it out of any “exact” statistic. Because it only searches near the asserted range, it also cannot be used as a substitute for a global search, which would happily “verify” a quote copied from the wrong passage.
What a passing check does not prove
An exact match proves the literal text exists at that location in that source version. It does not prove that the passage supports the claim it is attached to, that the right document was retrieved, that the source is authoritative or current, or that the answer interprets the passage correctly. Entailment, authority, freshness and citation completeness need separate evaluation, such as a natural-language-inference model or an LLM judge run after this deterministic gate. The byte check is the cheap, reproducible first layer: it removes fabricated and misplaced quotes so the expensive semantic checks only see real ones.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Design choices and trade-offs
| Choice | Stricter option | Looser option | Trade-off |
|---|---|---|---|
| Matching | Exact bytes only | Tolerant/nearby match | Provenance strength versus recovery from formatting drift |
| Offset capture | Taken from the splitter | Reconstructed later with indexOf |
Reliable identity versus convenience and duplicate-text ambiguity |
| Offset target | Original file bytes | Canonical extracted text | Fidelity to the input versus offsets that are easier to use in text workflows |
| Decoding | fatal: true |
Replacement characters | Fail-fast integrity versus continuing with possible byte/text disagreement |
| On failure | Block the answer | Annotate or retry | User trust versus availability, latency and operational complexity |
Placing it in the pipeline
The SitePoint tutorial runs validation as post-generation middleware in a LangChain sequence, after retrieval, prompting and generation. Its pipeline example uses placeholder retriever, prompt and validator declarations, so treat it as showing where the step goes, not as a ready integration. The pieces you still need:
- Structured citation output. Have the model emit citations in a schema (
sourceId,sourceVersion,byteStart,byteEnd,citedText) and validate the schema before callingverify. Models are unreliable at counting bytes, so a practical pattern is to give them chunk IDs and have your code supply the offsets for chunk-relative quotes. - A complete extractor for every citation format the model might produce, so nothing escapes verification by being unparsed.
- A failure policy: block, annotate or retry. Expose exact and partial outcomes to the UI so rendering can distinguish them.
- Privacy-aware logging. Log source IDs, versions, offsets and reason codes; avoid storing cited text unless you need to and are permitted to.
- Streaming behavior. Decide whether you validate after the full response or as each citation completes; a citation that fails after the user has already seen it needs a retraction path.
Performance claims
The tutorial describes a benchmark fixture of 1,000 citations across 50 documents totaling roughly 200 KB and says performance depends on hardware, document size and citation density. No reproducible results table or independent benchmark accompanies it, so there is no latency figure to quote. The operation itself is a bounds check, a slice and a byte comparison; profile it on your own documents and citation volume before making any service-level promise. The code in this article is a sketch to adapt, not a tested library.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




