Surgical subgraphs are a way to give a coding agent a small, relationship-aware view of a repository instead of whole files or isolated text chunks. In a September 2026 DEV Community article, author cos white reports that this approach, called LKIO, reduced average task tokens from 12,698 with full-file context to 545 on one 4,899-file codebase. Those are author-reported benchmark results, not independently verified evidence that other projects will see the same savings.
What “surgical subgraphs” retrieve
Ordinary text retrieval finds passages that resemble a query. A subgraph approach instead tries to model how code elements connect, then returns a bounded portion of that structure around symbols relevant to the task.
LKIO’s described pipeline parses source with Tree-sitter into symbols such as classes, methods, interfaces, and Vue single-file-component blocks. It records relationships including calls, imports, DTO field lineage, and REST route mappings. Starting from relevant “anchor symbols,” it traverses a bounded, cycle-safe graph and returns the resulting subgraph. The project describes its repository state as a copy-on-write in-memory snapshot and its agent integration as read-only MCP tools over stdio.
For example, a request to “trace this API call” could mean following a Vue form submission through an API client, REST route, Spring controller, service, DTO, and database table. That is a cross-file and cross-layer relationship problem: a matching text chunk may mention one part without showing the path connecting the rest.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minute#1 Best Overall
- Students build unmatched deductive-reasoning skills as they become crime-solving stars
- Most scenarios have more than one plausible outcome, allowing individuals or groups to broadly interpret evidence
- Includes interpretive handwriting, body language, fingerprinting, and many more activities
What the reported comparison found
Cos white compared naive full-file context, chunk retrieval with top-k=10, and LKIO’s subgraph retrieval on a stated 4,899-file business application with a Vue 3 frontend, Spring Boot microservices, and an enterprise dashboard. The September 2026 article reports these results:
| Measure | Full-file context | Chunk RAG (top-k=10) | LKIO subgraphs |
|---|---|---|---|
| Average task tokens | 12,698 | 4,266 | 545 |
| P95 tokens | 24,012 | 5,000 | 590 |
| Reported cost per 1,000 tasks | $38.09 | $12.80 | $1.64 |
| Cross-stack link recall | not stated by cos white | 0/12 | 12/12 |
| Hop precision | not stated by cos white | not stated by cos white | 72/72 hops with zero spurious hops |
The reported cost row uses the article’s September 2026 pricing assumption of $3.00 per million input tokens for Claude 3.5 Sonnet. It is not a current-price quote; model pricing can change, and the comparison’s cost figures depend on that stated rate.
Rank #2
The reported recall and hop figures are small-sample results, not general accuracy guarantees. The author gives a Wilson 95% confidence interval of 75.8%–100.0% for LKIO’s 12/12 cross-stack recall. That interval describes uncertainty around the reported sample; it does not replace independent testing on other repositories.
What the benchmark does—and does not—establish
The article describes the measurements as rigorous synthetic benchmarks on a real codebase, but also identifies them as self-reported and invites independent reproduction. That distinction matters: using a real repository does not by itself make a benchmark an independently verified test, nor does it show how a system performs across a sustained production workload.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
Cos white reports additional measurements, including calibration on 120 decision samples (expected calibration error of 0.1850 before temperature scaling and 0.0469 after, with a 0.0583 Brier score), governance tests (8/8 adversarial scenarios blocked and 32/32 benign changes passed), and a signal-to-noise change from about 3% to about 27%. These are also the author’s reported results, and the small, stated samples limit how broadly they can be applied.
The performance figures are tied to a specific laptop configuration: Intel Core Ultra 9 275HX, 32 GB DDR5, Windows 11, and Python 3.12.10. On that machine, the author reports a 1,000-file cold start of 10.61 seconds, 128.9 MB peak RSS, save-to-queryable time of 56.4 ms including a 50 ms filesystem debounce, symbol lookup around 2 microseconds, and depth-two impact analysis P50 of 0.121 ms. The article also reports 11.62 MB of additional memory over 1,000 update cycles while retaining 50 snapshots. These machine- and workload-specific numbers should not be assumed for other hardware or repositories.
The author says the project began in September 2026 and describes a two-week dogfooding effort involving one or two engineers as underway, with a field report expected later. The article therefore does not establish completed production validation. Cos white summarizes the distinction this way: “We hold the invariant `Implementation Complete ≠ Benchmark Validated ≠ Production Gate Passed`, and we will not dress benchmark scores up as production proof.”
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.When this retrieval design may be useful
A graph-based context strategy is most compelling when a task depends on tracing explicit relationships across a codebase—for example, from a UI event through a client and route to backend logic and data structures. It may be less useful when the needed context is a short, self-contained passage or when the repository’s relevant links cannot be parsed and modeled reliably. The reported benchmark does not establish which approach wins across different languages, architectures, repositories, or agent tasks.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
For teams evaluating it, useful checks include whether the parser recognizes their language and framework constructs, whether traversals stay bounded without omitting necessary links, and whether retrieved paths can be reviewed against known answers. Compare token use and task quality on the same representative tasks, and treat cost as a calculation based on the model’s current input-token rate—not a fixed property of the retrieval technique.
Source and reproducibility
The primary source is cos white, “Surgical Subgraphs: How We Cut Coding-Agent Token Costs by 95%,” DEV Community, September 29, 2026. The author says the full methodology, Wilson confidence intervals, and machine-readable results are available in the project’s report and explicitly invites independent reproduction. No independent reproduction is established in that article.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




