Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →CodeSmith’s central idea is that a model needs more than a prompt to complete engineering work: it needs a harness that governs how it acts, checks what it claims to have done, and gives the user visibility into interventions. A concrete example is its streaming filter for text that looks like a tool call but did not arrive through the API’s tool-call channel. The design described here is tied to CodeSmith v0.5.0, commit 3a74c82f, as examined by DogeKing in a 2026 DEV Community essay.
Why tool-call-shaped text is not a tool call
A model can print text that resembles an instruction to run a tool. That text alone does not mean the tool was invoked. In an API interaction, an actual invocation must come through the API’s tool-call channel; otherwise, a model—or a later part of an agent pipeline—could mistake a textual imitation for an action and proceed as though it had received real results.
DogeKing’s essay examines a filter in crates/agent-runtime/src/engine/streaming.rs that addresses this case. The filter watches for five opening markers: [TOOL_CALL], <codesmith:tool_call, <tool_call, <invoke , and <function_calls>, as well as their matching closing markers. It strips wrapper text from the stream and sends a notice to the interface.
The implementation’s state machine is designed to handle markers split across streaming chunks, rather than assuming a complete wrapper will arrive in one piece. When it removes one, the notice reproduced in the essay reads: “Stripped non-API tool-call wrapper from model output (use the API tool channel).” The point is not merely to suppress misleading text: the user is also told that the system intervened.
#1 Best Overall
What CodeSmith means by a harness
The CodeSmith README describes the project this way: “A model answers a question; an agent finishes a task. CodeSmith is the harness in between.” In this framing, the harness is the layer that guides and constrains a model as it works through a multi-step task. It is not the model itself, nor just a prompt; it is the surrounding rules and mechanisms that shape what the agent may do and how its work is managed.
In the v0.5.0 snapshot discussed by DogeKing, that design includes a written constitution, a nine-level authority hierarchy, three operating modes, OS-level sandboxing, a side-git snapshot for each turn, and optional concurrent sub-agents. Together, these illustrate several different jobs a harness can take on: establish priorities, set the degree of autonomy, limit execution, preserve a recoverable state, and divide work. They are features as described for that snapshot, not confirmation of current support across platforms or later releases.
Rank #2
Rules and authority
A constitution and authority hierarchy give the agent a way to resolve competing instructions. That matters in engineering work, where a task request, project constraints, and safety rules may point in different directions. The essay names these mechanisms but does not reproduce their full contents, so the useful takeaway is their role in the architecture: the model operates within an explicit ordering of rules rather than treating every instruction as equal.
Modes and autonomy
CodeSmith’s three named modes are Plan, Agent, and YOLO. The names signal that the project provides distinct operating modes, but the essay’s account does not establish detailed behavior for each mode. They should therefore be understood as evidence of configurable operation, not as a promise about exactly which approvals or actions each mode permits.
Sandboxing and snapshots
OS-level sandboxing is presented as a boundary around execution, while a side-git snapshot each turn is a way to preserve work state. These address different risks: a sandbox limits what processes can access or change, whereas a snapshot can help make changes inspectable or recoverable. The essay identifies both in the source snapshot but does not establish a universal sandbox platform matrix or quantify recovery guarantees.
Concurrent sub-agents
Optional concurrent sub-agents offer a way to split work among agents. Concurrency can expand the amount of work handled in parallel, but it does not by itself guarantee correctness; the harness still needs to define boundaries and manage the results. The essay names this as an available design element in the described snapshot, not as a benchmarked productivity claim.
Rank #4
What the project’s scale figures do—and do not—show
DogeKing describes CodeSmith as the successor to CodeWhale, formerly called deepseek-tui, and reports the following counts for the source snapshot discussed in the essay:
| Reported measure | Figure | Qualification |
|---|---|---|
| Rust workspace crates | 21 | Author-reported for the examined snapshot. |
| Rust source files | 548 | Author-reported for the examined snapshot. |
| Lines of code | 356,193 | Author says this was counted with find and wc, including comments and inline tests. |
| Test functions | 5,429 | Author-reported for the examined snapshot. |
The essay also names components such as agent-runtime, tui, agent / providers, execpolicy, index, mcp, hooks, and extensions. The counts convey the reported size of that codebase snapshot; they are not independently verified current project metrics. A file count or test-function count also cannot, on its own, establish software quality, reliability, model performance, or cost.
Best Value
What the essay establishes about “cheap brains”
The title’s “cheap brains” framing points to using lower-cost or open-source models with an agent harness, but the essay does not provide model prices, controlled cost comparisons, or benchmark results. It therefore supports an architectural argument, not a measured claim that a particular model is cheaper or performs as well as another.
The practical lesson is narrower and more useful: models can produce plausible-looking output that is not evidence of an action, so an agent system needs a reliable distinction between text and actual tool execution. CodeSmith’s filter is one example of a harness making that distinction, while its user notice makes the intervention visible. Rules, autonomy modes, sandboxing, snapshots, and delegated agents illustrate other layers that can help an agent stay within the task’s boundaries.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




