In four runs analyzing two short articles about logical fallacies, Miguel Diaz Kusztrich found that workflow cost depended not just on input reuse but also on how many terms, classifications, and explanatory tokens the system produced—and whether it repeated work. His results are a useful case study, not a general benchmark: the quality review was preliminary, the estimated costs were specific to his setup, and one comparison included a repeated-call incident.
What the workflow asked the application and models to do
Kusztrich’s AIDBDeveloper workflow assigned orchestration, storage, and deterministic operations to the application, reserving model calls for interpretation. It extracted sentences, split text into words, numbers, and punctuation, extracted multi-word terms, then performed syntactic, secondary, and free-form classifications. Token classifications ran in batches of five, with ten model instances processing different sentences in parallel. Later steps reused earlier information where possible to narrow what the model had to decide.
The reported setup used GPT 5.6 Sol with low reasoning effort for sentence extraction, GPT 5.4 mini for tokenization, and GPT 5.6 Terra at medium reasoning effort for term extraction and later classification. These are the models and settings in this account, not a recommendation for model selection today.
As Kusztrich put it, “The application should do everything it already knows how to do.” He added: “The model should be used for the uncertain parts.”
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
What changed across the four runs
The test processed two previously written short articles about logical fallacies twice each. TEXT 1’s first run used shorter system messages to reduce input tokens; its second used more explicit instructions. TEXT 2’s first run used essentially the improved configuration, while its second removed an instruction to end function calls with only a single full stop, permitting explanatory final messages. A repeated-function-call loop also occurred in one step of that second TEXT 2 run.
| Text and run | Configuration change or context | Reported observation |
|---|---|---|
| TEXT 1, trial 1 | Shorter system messages | Kusztrich reported cache misses in some steps, overly permissive term extraction, and excessive classifications. |
| TEXT 1, trial 2 | More explicit system messages | He reported better cache use and fewer extracted terms and classifications. |
| TEXT 2, trial 1 | Essentially the improved configuration | Used as the comparison point for TEXT 2’s second run. |
| TEXT 2, trial 2 | Allowed explanatory final output after function calls | Output volume rose in a classification step; a repeated-call issue also occurred. |
What the reported numbers show—and do not show
The following figures are Kusztrich’s reported results for these runs. Costs are theoretical estimates tied to his setup, not current API price quotations, independently reproduced measurements, or general benchmarks.
Rank #2
| Measure | Reported result | Scope and qualification |
|---|---|---|
| Tokenization | 1,650 tokens for TEXT 1; 1,762 for TEXT 2 | Kusztrich, 2026; unchanged between the two runs for each text. |
| Extracted terms | 1,114 to 431 | Kusztrich, 2026; TEXT 1 comparison after instructions became more explicit. |
| Classifications | 15,673 to 9,580 | Kusztrich, 2026; TEXT 1 comparison across the instruction change. |
| Workload scale | Approximately 3–8 million tokens and roughly 2,000–3,000 requests | Kusztrich, 2026; the reported scale for relevant trials. |
| Uncached-input cost | Almost 73% lower | Kusztrich, 2026; estimated TEXT 1 trial comparison. |
| Combined input-related cost | Approximately 18% lower | Kusztrich, 2026; combines uncached input, cached input, and cache writes in the TEXT 1 comparison. |
| Output cost | Almost 15% lower | Kusztrich, 2026; TEXT 1 comparison. Output tokens accounted for about 64% of total estimated cost in that comparison. |
| Total theoretical cost | $11.39 to $9.59, approximately 16% lower | Kusztrich, 2026; TEXT 1 comparison, specific to the author’s model mix and cost assumptions. |
| TEXT 2 theoretical cost | $11.67 to $14.97 | Kusztrich, 2026; the latter run permitted explanatory output and included a repeated-call issue, so the difference cannot be attributed to output wording alone. |
| Output in one classification step | Roughly 234,000 to 426,000 tokens | Kusztrich, 2026; across the TEXT 2 comparison. |
| Hypothetical model-price substitution | Roughly $42–65, or around 4.5 times the actual-model-mix estimate | Kusztrich, 2026; hypothetical GPT 6 Astra pricing applied to logged usage. The author cautions this does not establish that Astra would use the same tokens or produce identical results. |
The TEXT 1 results associate clearer instructions with fewer terms and classifications and better cache use. But these were four runs on two short texts, not a randomized experiment, and the comparison does not isolate one prompt change as the cause of every cost difference. The figures describe this workflow, not a promise that the same edits will yield the same savings elsewhere.
Why output and repeated calls matter as much as caching
Input caching can reduce the cost of sending stable context again, but it does not make unnecessary work valuable. In TEXT 2, allowing explanatory final output coincided with a large increase in output tokens in one classification step and a higher total estimate. Because the same run also had a repeated-call issue, the reported cost change reflects more than one difference.
A loop can also replay context that is cacheable while repeating a function call that should not have happened. As Kusztrich wrote, “You can cache an error very efficiently.” In automated workflows, call counts, outputs, retries, and step-level token usage therefore matter alongside cache percentages.
Quality remained an important limitation
The reported quality review was preliminary, not a formal benchmark. Sentence extraction was described as extremely consistent, and tokenization was identical across equivalent trials. Word-level syntactic classification still needed refinement, though Kusztrich considered it reasonably good.
- Multi-word terms: extraction remained weak; the initial TEXT 1 run over-extracted according to the author.
- Term classification: syntactic classification was poorer than word classification, and secondary classification of terms was described as clearly inadequate.
- Free-form word tags: appeared more promising, but remained subjective.
That mix of results matters: fewer outputs and lower estimated cost do not demonstrate better analysis. A cheaper step that misses useful terms, or classifies them poorly, may not be an improvement. Kusztrich said a larger follow-up effort was still planned; the four runs do not establish a general accuracy rate.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to apply the lessons to your own workflow
Treat the case study as a set of design questions to test against your own texts, models, and current API behavior—not as a recipe with guaranteed savings.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
- Move deterministic work out of model calls. Let application code handle orchestration, storage, token splitting, and other operations whose answer is already specified by the input and rules.
- Keep interpretive tasks narrow. Make each model step answer a bounded question, and pass forward reliable prior results rather than asking later steps to infer them again.
- Constrain output to what the next step consumes. Where your interface and API support it, suppress unused natural-language final prose in function-call workflows. Check that the constraint does not prevent required information or valid completion.
- Instrument each operation. Record configuration, start and end times, input and output, token usage, and the context used. Attribute uncached input, cached input, cache writes, output, and retries to individual steps so a total bill can be explained.
- Check invocation behavior, not only cache behavior. Inspect logs for duplicate calls, retries, and loops; set appropriate stopping conditions and investigate unexpected call counts.
- Evaluate quality and cost together. Check whether each step produces linguistically valid results, particularly for multi-word term extraction and classification. Compare models for reliability on the specific task instead of assuming that a cheaper or newer model is interchangeable.
- Prioritize steps that are both costly and weak. If a step generates a large volume while producing poor results, redesigning its role or the surrounding process may be more effective than further prompt tuning.
The model-price substitution Kusztrich calculated is a cost simulation on recorded token counts, not a head-to-head model test. More generally, the trial figures cannot tell another team which model, prompt, or cache strategy will perform best. The transferable lesson is to measure the work each step actually does, including its output and retries, and verify that result quality justifies it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




