Recommended Free Tools
A coding agent can finish a task with every test passing and still leave the codebase worse off. A utility lands in the wrong module, a client gets imported into a domain layer, or a public interface gets bypassed on the way to a working feature. Tests check the behaviors they were written to check. They do not check whether the structure around those behaviors still holds.
One published answer is Archkeel, a deterministic architecture gate that runs alongside the test suite. It checks three things separately: whether the scan saw everything it claims to see, whether the code obeys a declared contract, and whether a change matched an expectation written before the implementation existed. The description below comes from the tool author’s own DEV Community article by Alex (2026). The figures and behaviors it reports have not been independently verified, and the article is clear about where the evidence stops.
As an Amazon Associate I earn from qualifying purchases.
Why a green suite does not certify structure
Test suites are good at answering one question: does the tested behavior still work? They are weak at answering a different question: is the program still shaped the way its designers intended? An agent optimizing for passing tests has no particular reason to respect module boundaries that no test enforces. The author describes three recurring patterns: agents placing utilities in modules where they do not belong, agents crossing public interfaces instead of using them, and agents importing client code into layers that should not know about it. Each of these can leave every test green.
The gap is not only about rule violations. A change can also weaken the analyzer’s ability to understand the program. If the static scan can no longer resolve some calls, the architecture checks silently see less, and a clean result means less than it appears to. Archkeel is built around that second concern as much as the first.
#1 Best Overall
What the gate checks
The target architecture contract
The architecture is described as a contract. It names the components in the system, the packages each component owns, the public names each component exposes, and the dependency rules between components. Every ordered pair of components receives an explicit decision: allowed or forbidden, with a written reason.
If a pair has no decision, it stays open, and validation stays red until someone resolves it. An unresolved pair is treated as an unfinished design question, not as permission. The architect remains responsible for the intended target architecture; the tool enforces the decisions that have been recorded.
Interview mode and auto mode
The packaged skill that helps write the contract works in two modes:
Rank #2
- Interview mode: the skill prepares recommendations from existing architecture documents and asks about conflicts and gaps before anything is decided.
- Auto mode: the skill makes the decisions itself and labels who decided each rule, so a reviewer can tell an inferred rule from one a person agreed to.
Three separate verdicts
Archkeel keeps three verdicts apart rather than folding them into one score. Each answers a different question, and each can fail on its own.
| Verdict | Question it answers | What a failure means |
|---|---|---|
observation_complete |
Did the scan see everything it claims to see? | Some code or calls were not resolved, so the other verdicts rest on partial evidence. |
declared_rules |
Does the code obey the contract? | A component crosses a forbidden relationship, owns a package it should not, or breaks a public-name rule. |
expectation_fulfilled |
Did the change match what was declared, without regressions? | The change did something the agent did not declare, or it made the evidence weaker. |
Why completeness is a verdict of its own
The most instructive case in the account is a change that breaks no rule at all. In the author’s fixture, a refactor replaces two statically resolved calls with a dictionary lookup. The test suite still passes, and no forbidden import or dependency cycle appears. But the analyzer now reports one unresolved call where it previously reported none.
A rule check alone would call that clean. Archkeel treats the weaker evidence as a regression and rejects the change unless the change was declared. The unresolved ratio is compared using integer cross-multiplication rather than rounded percentages, so a small change in the count cannot be hidden by rounding. Visibility loss is therefore reported as a concern in its own right, not allowed to pass as an apparently clean result.
Rank #3
Publication order: committing the expectation first
The gate also checks process, not just code state. The idea is that an agent states what it intends to change before it submits the implementation, so the expectation cannot be written to match the result afterward. The workflow the author describes is:
- The agent commits an expectation describing the intended architecture change.
- The implementation is submitted afterward, as a merge request on the host.
- Archkeel checks Git ancestry and the host’s merge request history to confirm the expectation was published before the implementation, and rejects an expectation written after the fact.
Exit codes are part of the contract for automation. A pass returns 0, a rejection returns 1, and 2 means the input could not be verified. Unknown evidence never becomes a pass; the author summarizes this as “Unknown never becomes green.”
Reported results and how far they reach
The article reports several figures. They describe one application and one self-check, and they are measured by the author rather than an independent benchmark.
Rank #4
| Reported figure | What it describes | Scope and qualification |
|---|---|---|
| 140 of 156 component-pair decisions matched (89.7%) | Agreement between the tool’s auto-mode decisions and the comparison set in one exercise | One service, measured once. The author says this is not a general accuracy estimate for auto mode. |
| 13 components | Size of the field-service application used in the account | One application; a reported example, not a benchmark. |
| 162 violations in the first report against the final target | Initial findings when the target architecture was checked | Includes 148 on the use-case-to-persistence-adapter dependency. |
| 630 of 3,303 calls unresolved (Archkeel itself); 998 of 4,318 (the field-service application) | Calls the analyzer could not resolve statically | Counted and reported, not guessed. Unresolved calls are a stated limit on what the verdicts can see. |
| 6 components, 30 component pairs, 46 rules in Archkeel’s self-check contract | The tool’s own architecture contract | The author reports planting a violation to prove that each enforcing rule actually fires. |
The field-service example was reported as running on Python 3.12 with FastAPI, async SQLAlchemy, PostgreSQL with PostGIS, Redis, Taskiq, and OR-Tools. These describe the environment in the account, not a requirement of the tool.
Known blind spots
The author is explicit about what the gate does not observe or establish. Treat these as part of the design, not as footnotes:
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →- Runtime behavior, data flow, and performance are not observed. The gate reasons about static structure only.
- Two competing implementations of the same idea are not detected unless a rule or regression exposes them.
- Private access through a package import, such as
import pkg; pkg._member, can slip through. - Decision reasons are checked for existence, not truth. The tool confirms that a reason was written, not that it is correct.
- Publication order is not proof of no private edits. Ancestry and merge request history show when the expectation was published; they do not show that nobody changed code privately beforehand.
- Determinism was tested narrowly. Reports were run repeatedly across two clones with varied paths, hash seeds, working directories, time zones, and locales, and produced byte-identical output on one machine and one Python build. Behavior across other operating systems or Python versions was not established.
- Host support is GitLab-only for merge request evidence. No GitHub adapter existed at the time of publication.
How it differs from snapshot tests and rule tools
The author contrasts the approach with snapshot architecture tests and with rule tools such as ArchUnit, import-linter, and dependency-cruiser. The differences are easiest to see along a few axes:
Best Value
- Comparison basis: snapshot checks compare the current code against a recorded picture; a baseline-to-candidate comparison asks whether a change made the structure or the evidence worse.
- Scope of checks: declared dependency rules are only part of the picture; the gate also asks whether the analyzer’s evidence became weaker.
- Process evidence: code-state checks say what the code looks like now; the publication-order check says whether the stated expectation came before the implementation.
- Output: separate verdicts and diagnostics, rather than a single aggregate score that hides which concern failed.
- Observed scope: static structure is checked; runtime behavior, data flow, and performance are outside what the tool sees.
- Host and portability: merge request evidence is GitLab-only, and cross-platform determinism has not been shown.
Where it fits, and what it does not replace
Archkeel is presented as a guardrail, not a substitute for the rest of engineering practice. It does not replace tests, human ownership of the architecture, runtime validation, or code review. Its value is narrower: it makes structural drift and lost analyzer visibility visible in the same place a passing test suite would otherwise hide them.
The project is described as MIT-licensed and distributed through GitHub and PyPI. A starting point given in the account is uvx archkeel --help. Packaging, commands, and supported hosts can change, so confirm the current repository and release notes before adopting it in a pipeline.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




