DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
World desk4 min

CodeSmith: A Harness for Cheap Brains

CodeSmith’s design shows how a coding-agent harness can distinguish real tool execution from text that only looks like a tool call.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

CodeSmith’s central idea is that a model needs more than a prompt to complete engineering work: it needs a harness that governs how it acts, checks what it claims to have done, and gives the user visibility into interventions. A concrete example is its streaming filter for text that looks like a tool call but did not arrive through the API’s tool-call channel. The design described here is tied to CodeSmith v0.5.0, commit 3a74c82f, as examined by DogeKing in a 2026 DEV Community essay.

Why tool-call-shaped text is not a tool call

A model can print text that resembles an instruction to run a tool. That text alone does not mean the tool was invoked. In an API interaction, an actual invocation must come through the API’s tool-call channel; otherwise, a model—or a later part of an agent pipeline—could mistake a textual imitation for an action and proceed as though it had received real results.

DogeKing’s essay examines a filter in crates/agent-runtime/src/engine/streaming.rs that addresses this case. The filter watches for five opening markers: [TOOL_CALL], <codesmith:tool_call, <tool_call, <invoke , and <function_calls>, as well as their matching closing markers. It strips wrapper text from the stream and sends a notice to the interface.

The implementation’s state machine is designed to handle markers split across streaming chunks, rather than assuming a complete wrapper will arrive in one piece. When it removes one, the notice reproduced in the essay reads: “Stripped non-API tool-call wrapper from model output (use the API tool channel).” The point is not merely to suppress misleading text: the user is also told that the system intervened.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What CodeSmith means by a harness

The CodeSmith README describes the project this way: “A model answers a question; an agent finishes a task. CodeSmith is the harness in between.” In this framing, the harness is the layer that guides and constrains a model as it works through a multi-step task. It is not the model itself, nor just a prompt; it is the surrounding rules and mechanisms that shape what the agent may do and how its work is managed.

In the v0.5.0 snapshot discussed by DogeKing, that design includes a written constitution, a nine-level authority hierarchy, three operating modes, OS-level sandboxing, a side-git snapshot for each turn, and optional concurrent sub-agents. Together, these illustrate several different jobs a harness can take on: establish priorities, set the degree of autonomy, limit execution, preserve a recoverable state, and divide work. They are features as described for that snapshot, not confirmation of current support across platforms or later releases.

Rules and authority

A constitution and authority hierarchy give the agent a way to resolve competing instructions. That matters in engineering work, where a task request, project constraints, and safety rules may point in different directions. The essay names these mechanisms but does not reproduce their full contents, so the useful takeaway is their role in the architecture: the model operates within an explicit ordering of rules rather than treating every instruction as equal.

Modes and autonomy

CodeSmith’s three named modes are Plan, Agent, and YOLO. The names signal that the project provides distinct operating modes, but the essay’s account does not establish detailed behavior for each mode. They should therefore be understood as evidence of configurable operation, not as a promise about exactly which approvals or actions each mode permits.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sandboxing and snapshots

OS-level sandboxing is presented as a boundary around execution, while a side-git snapshot each turn is a way to preserve work state. These address different risks: a sandbox limits what processes can access or change, whereas a snapshot can help make changes inspectable or recoverable. The essay identifies both in the source snapshot but does not establish a universal sandbox platform matrix or quantify recovery guarantees.

Concurrent sub-agents

Optional concurrent sub-agents offer a way to split work among agents. Concurrency can expand the amount of work handled in parallel, but it does not by itself guarantee correctness; the harness still needs to define boundaries and manage the results. The essay names this as an available design element in the described snapshot, not as a benchmarked productivity claim.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the project’s scale figures do—and do not—show

DogeKing describes CodeSmith as the successor to CodeWhale, formerly called deepseek-tui, and reports the following counts for the source snapshot discussed in the essay:

Reported measure Figure Qualification
Rust workspace crates 21 Author-reported for the examined snapshot.
Rust source files 548 Author-reported for the examined snapshot.
Lines of code 356,193 Author says this was counted with find and wc, including comments and inline tests.
Test functions 5,429 Author-reported for the examined snapshot.

The essay also names components such as agent-runtime, tui, agent / providers, execpolicy, index, mcp, hooks, and extensions. The counts convey the reported size of that codebase snapshot; they are not independently verified current project metrics. A file count or test-function count also cannot, on its own, establish software quality, reliability, model performance, or cost.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the essay establishes about “cheap brains”

The title’s “cheap brains” framing points to using lower-cost or open-source models with an agent harness, but the essay does not provide model prices, controlled cost comparisons, or benchmark results. It therefore supports an architectural argument, not a measured claim that a particular model is cheaper or performs as well as another.

The practical lesson is narrower and more useful: models can produce plausible-looking output that is not evidence of an action, so an agent system needs a reliable distinction between text and actual tool execution. CodeSmith’s filter is one example of a harness making that distinction, while its user notice makes the intervention visible. Rules, autonomy modes, sandboxing, snapshots, and delegated agents illustrate other layers that can help an agent stay within the task’s boundaries.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.