Abstract
The adoption of large language model (LLM) agents has moved the economic unit of analysis from isolated generation to multi-step task execution. In an agentic workflow, a system may read files, plan, edit, run commands, inspect failures, retry, and request review before reaching a usable result. Under this execution model, evaluating an AI system by token price or single benchmark score is insufficient.
This paper proposes Cost per Completed Task (CPCT) as a system-level measurement framework for comparing AI model configurations in agentic workflows. CPCT aggregates input/output token cost, cache usage, tool and sandbox cost, wall-clock runtime, failed-attempt cost, and human-review effort into a single task-centred denominator: accepted work.
The paper (i) situates CPCT against Total Cost of Ownership (TCO) and Activity-Based Costing (ABC) and against emerging cost-aware agent metrics such as cost-of-pass; (ii) decomposes task cost into operational components with an operable rubric for the human and quality terms; (iii) resolves the failure-mode edge cases (total-failure division-by-zero, sunk cost of timeouts) that a naïve ratio leaves undefined; and (iv) specifies a repeatable protocol with a per-run data schema.
1. Introduction
Large language models are increasingly deployed as agentic systems that receive a goal, inspect an environment, invoke tools, take actions, observe results, and iterate until a task is completed or abandoned. This is true across software engineering (issue resolution, refactoring, CI triage) and, increasingly, across knowledge-work domains such as legal and forensic analysis.
This shift changes how cost and productivity should be measured. In stateless API usage, models are naturally compared by input and output token price. In agentic execution, the relevant economic unit is no longer the isolated prompt or token; it is the completed task.
In agentic workflows, the cheapest model is not necessarily the one with the cheapest token. The economically preferable configuration minimizes the quality-adjusted cost per accepted task — and much of that cost is operational overhead (tool-output verbosity, context re-reads, turns) that is independent of unit-token price.
1.1 Contributions
- It proposes Cost per Completed Task as a system-level metric, arguing that CPCT is a property of the {model + scaffold + prompt + tools + repository/data + tests + reviewer} system, not of a model in isolation.
- It decomposes task cost into token, cache, tool, runtime, retry, and human-review components, with an operable rubric for the review-cost and quality terms.
- It provides a repeatable protocol — a controlled-variable specification and a per-run data schema — for comparing configurations.
- It presents a real, measured cross-domain case study (forensic report generation, not code) whose central result — an 8.6× cost swing on an identical model and task driven purely by operational overhead — provides clean evidence that token price is a weak proxy for cost.
2. Background
2.1 From code completion to agents
Early LLM coding tools suggested completions; controlled studies of GitHub Copilot reported gains such as 55.8% faster completion on some standardized tasks (Peng et al., 2023). Later work complicated this: a randomized controlled trial by METR on experienced open-source developers on mature projects found early-2025 tooling increased task time by 19%, even as developers perceived a speed-up (Becker et al., 2025).
2.2 Benchmarks for real tasks
SWE-bench evaluates whether a produced patch resolves a real GitHub issue without breaking existing behavior, across 2,294 problems in 12 Python repositories (Jimenez et al., 2024). AutoCodeRover pairs LLMs with program-analysis search at low dollar cost per task; Agentless shows a simpler localize–repair–validate process can be competitive at low cost; OpenHands provides an open platform with sandboxed execution and evaluation.
2.3 Token consumption and cost-aware evaluation
Efficient Agents frames an efficiency–effectiveness trade-off and introduces cost-aware ideas such as cost-of-pass (Wang, N. et al., 2025). Work on token consumption in agentic coding reports that agentic tasks can consume far more tokens than chat or reasoning, that higher consumption does not imply higher accuracy.
3. Positioning CPCT Against Existing Cost Concepts
3.1 Relationship to TCO and Activity-Based Costing
CPCT is a domain-specific specialization of two established ideas. Total Cost of Ownership (TCO) holds that acquisition price is only part of a resource's cost. Activity-Based Costing (ABC) assigns indirect and overhead cost to the activities that consume them. CPCT applies both to agentic work: the "purchase price" (token price) is one component, and the framework attributes operational overhead — turns, cache, tools, runtime, retries, review — to the activity that consumes it (the task).
3.2 A ladder of cost-aware metrics
| Metric | Unit | Captures | Misses |
|---|---|---|---|
| Cost per token | 1M tokens | Nominal unit price | Everything operational |
| Cost per call | API request | Request-level price | Multi-turn loops, retries |
| Cost per generated patch | patch | Output produced | Whether it works |
| Cost per resolved issue | passing task | Test-passing resolution | Review, quality, maintainability |
| Cost-of-pass | passing task | Expected cost to a pass | Human review, quality grading |
| CPCT | accepted task | All operational + review cost | (time — see TCPCT) |
| Quality-adjusted CPCT | quality-weighted task | Above + reviewer quality | (rubric subjectivity) |
4. Research Problem and Questions
Central problem. How should teams evaluate the economic efficiency of AI model configurations in agentic workflows where the objective is not token generation but accepted task completion?
- RQ1. Does lower token price reliably imply lower end-to-end cost?
- RQ2. Which operational variables most influence total task cost beyond token price?
- RQ3. How can teams normalize comparison around completed tasks rather than model calls?
- RQ4. What governance decisions can a CPCT metric support?
5. The CPCT Framework
5.1 Task and success definitions
| Level | Criterion | Notes |
|---|---|---|
| L1 — Technical pass | Target output produced; automated checks pass | SWE-bench-style fail-to-pass / pass-to-pass |
| L2 — Review pass | A human expert accepts the output | Minimality, security, correctness, coverage |
| L3 — Production pass | Integrated / filed without observed regression | The economically meaningful outcome |
5.2 Per-run cost decomposition
Cᵢ = Cᵢ(tok) + Cᵢ(cache) + Cᵢ(tool) + Cᵢ(run) + Cᵢ(rev) + Cᵢ(retry)
where the terms are, respectively, direct token cost; billed cache read/write cost; external tool/sandbox/CI/hosted-execution cost; wall-clock or compute cost; human review cost; and failed-attempt/rework cost.
5.3 Basic CPCT and the failure-mode fix
CPCT = Σ min(Cᵢ, Cᵢ(max)) / Σ Sᵢ, together with the acceptance rate Σ Sᵢ / n
Always report CPCT with the acceptance rate. When Σ Sᵢ = 0, CPCT is reported as "undefined" — no accepted tasks — with total sunk cost stated, never as a finite number that hides total failure.
5.4 Quality-adjusted CPCT
QCPCT = Σ min(Cᵢ, Cᵢ(max)) / Σ (Sᵢ · Qᵢ), with Qᵢ ∈ [0,1]
| Qᵢ | Reviewer decision |
|---|---|
| 1.0 | Accepted without changes |
| 0.7 | Accepted after minor human-suggested refactor |
| 0.3 | Requires partial rewrite |
| 0.0 | Rejected |
5.5 Time-adjusted CPCT
TCPCT = Σ [min(Cᵢ, Cᵢ(max)) + Vₜ · Tᵢ] / Σ Sᵢ
Tᵢ is elapsed time; Vₜ is the value assigned to time. The case study confirms that time is a separable axis: parallelizing independent stages cut wall time ~40% at identical dollar cost.
6. Experimental Protocol
6.1 Minimal unit and repetition
The minimal unit is a tuple (model, scaffold, task, run). Because agentic execution is non-deterministic, at least three runs per task per configuration are recommended, reported with dispersion.
6.4 Per-run data schema
| Field | Definition |
|---|---|
| task_id / stage | Task or pipeline-stage identifier |
| model | Model name + version |
| agent_scaffold | Agent/tool harness |
| input_tokens / output_tokens | Generation and context cost |
| cache_read / cache_write | Context dependence |
| runtime_seconds | Operational and time cost |
| success_level | L1 / L2 / L3 / fail |
| review_minutes / review_score | C(rev) and Qᵢ |
| total_cost | Aggregated Cᵢ |
7. Measured Case Study: A Cross-Domain Forensic Pipeline
7.1 Domain and apparatus
The case study is deliberately outside software engineering, to test CPCT's generality. The domain is agentic forensic expert-report (laudo pericial) generation on the DraIara/OpenClaw harness: the agent reads an OCR'd evidence dossier, calls forensic tools via MCP, produces maps and crops, writes a sixteen-section report, and self-reviews. Two named models were used: Claude Opus 4.8 (premium, ≈US$15/M output) and Claude Sonnet 4.6 (budget, ≈5× cheaper per token).
7.2 The isolation experiment: same model, same task, 8.6× cost
| Configuration (identical model & task) | Cost (US$) | Turns | Input tokens |
|---|---|---|---|
| Sonnet, no I/O cap | 7.93 | 20 | 766,715 |
| Sonnet + I/O cap | 1.18 | 10 | 24 |
| Sonnet + I/O cap + tool-lean | 0.92 | 6 | 16 |
Cost fell 88% (an 8.6× swing) with no change of model. The driver was tool-output verbosity: nine IP-geolocation maps returned ~766,715 input tokens that were re-ingested into context. This is the paper's strongest single piece of evidence that token price is a weak proxy for cost.
7.3 Stage-decomposed pipeline: the irreducible cost floor
| Stage | Model | Cost (US$) | Turns | Runtime | Output tok | Cache read |
|---|---|---|---|---|---|---|
| extract | Sonnet 4.6 | 0.67 | 28 | 181 s | 9,344 | 792,255 |
| geo | Sonnet 4.6 | 1.81 | 9 | 203 s | 4,454 | 308,518 |
| integrity | Sonnet 4.6 | 0.45 | 15 | 135 s | 7,348 | 394,178 |
| osint | Sonnet 4.6 | 0.46 | 21 | 125 s | 7,043 | 374,040 |
| biometrics | Sonnet 4.6 | 0.21 | 5 | 145 s | 768 | 156,058 |
| synthesis | Opus 4.8 | 3.82 | 21 | 720 s | 50,398 | 2,238,068 |
| Total | hybrid | 7.42 | 99 | ~19 min | 79,355 | 4.26 M |
The single premium synthesis stage is 51.5% of total cost. It behaves as an irreducible floor: no amount of cheaper tokens on the other five stages can push total cost below it.
7.4 End-to-end scenarios and the time axis
| Scenario | Cost (US$) | Wall time |
|---|---|---|
| Original (all-Opus monolith) | ~9.83 | ~30 min |
| Decomposed naïve (all-Opus, estimated) | ~12–16 (est.) | ~30+ min |
| Optimized (Sonnet stages + Opus synthesis) | 6.29 | ~25 min |
| + Parallelized | 6.29 | ~15 min |
Moving 5/6 stages to the budget model cut cost only 36% (9.83 → 6.29), not ~90%. Parallelizing independent stages cut wall time ~40% (25 → 15 min) at identical dollar cost.
7.5 Findings mapped to the research questions
- RQ1 — No. A ~5× per-token discount on most of the pipeline produced a 36% end-to-end saving; and the same-model 8.6× swing shows token price is not even the main lever.
- RQ2 — Operational overhead dominates. The measured dominant drivers were tool-output verbosity fed back into context, the turns it induced, and cache/context re-reads — all above unit-token price.
- RQ3 — Decompose and account per stage. Per-stage cost accounting localizes the irreducible floor and the movable work.
- RQ4 — Route by stage, cap I/O, parallelize. Reserve the premium model for the accepted-work-producing stage; run cheaper stages on the budget model; cap tool output; parallelize independent stages for time without cost.
8. Discussion
8.1 The false economy of cheap tokens
Cheap tokens suit simple, high-volume tasks. For complex work, the measured floor result (§7.3) is decisive: the cost of the stage that produces accepted work sets a floor that per-token discounts elsewhere cannot break.
8.2 CPCT is a system property, not a model property
CPCT = f(model, scaffold, tools, data, prompt, tests, reviewer)
The isolation experiment proves this quantitatively: the same model spanned US$0.92–7.93 on one task purely from scaffold-level I/O discipline.
8.3 Operational overhead is the primary cost lever
The practical corollary is that context hygiene is a first-class cost control: capping tool-output verbosity, trimming re-ingested context, and reducing turns produced an 8.6× saving with no model change and no quality loss at L2.
9. Threats to Validity
Single-run measurements. Section 7 reports single runs on a few real cases; figures need repeated runs and dispersion before being treated as general point estimates.
Cross-domain transfer. Evidence is from forensic report generation; transfer to coding agents is argued (via the domain-independent isolation mechanism) but not directly measured here.
Model and price volatility. Models, prices, context windows, and caching policies change quickly. A CPCT figure is time-stamped and records model version, date, pricing snapshot, and scaffold configuration.
10. Practical Recommendations
- Measure cost per completed task, not per token.
- Audit tool-output verbosity and re-ingested context first — it can dominate cost more than model choice.
- Decompose pipelines and account cost per stage; find the irreducible floor.
- Route each stage to the cheapest tier that still passes L2; reserve the premium model for the accepted-work-producing stage.
- Parallelize independent stages for time without cost.
- Track failed attempts and include their capped cost.
- Log model version, scaffold, prompt version, data snapshot, and acceptance level.
- Use a published quality rubric; run ≥3 repetitions and report dispersion.
- Include standardized human-review time.
- Re-run evaluations periodically as models and prices change.
11. Conclusion
Agentic workflows change the economics of AI-assisted work. Token price is an incomplete and sometimes misleading proxy. This paper proposed Cost per Completed Task as a system-level measurement framework, resolved the failure-mode edge cases a naïve ratio leaves undefined, and specified a protocol and data schema.
A real, measured cross-domain case study grounded the framework: holding the model fixed, operational discipline alone moved the cost of one task by 8.6×; a premium synthesis stage formed a 51.5% irreducible cost floor that per-token discounts elsewhere could not break; and parallelism cut time ~40% at equal cost.
The broader conclusion is that AI configuration selection should be governed by completed-task economics: the true cost of AI is the cost of correct, accepted, and maintainable work — most of which is operational, not per-token.
Data Availability
The per-stage and per-configuration measurements in Section 7 are reported in full in this manuscript. An anonymized per-run dataset following the schema in Section 6.4, together with the figure-generation scripts, accompanies the archival deposit.
Conflict of Interest
The author declares no conflict of interest. The case study uses commercial models (Anthropic Claude Opus 4.8 and Sonnet 4.6) accessed via a third-party router; this is a factual description of the measurement environment and no sponsorship or commercial relationship influenced the results.
Suggested Citation
Filho, O. J. (2026). Cost per Completed Task: A Measurement Framework and Cross-Domain Case Study for Evaluating AI Model Configurations in Agentic Workflows. Zenodo. https://doi.org/10.5281/zenodo.21198419
References
- Bai, L., Huang, Z., Wang, X., Sun, J., Mihalcea, R., Brynjolfsson, E., Pentland, A., & Pei, J. (2026). How Do Coding Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks. Preprint.
- Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089.
- Chowdhury, N., et al. (2024). Introducing SWE-bench Verified. OpenAI.
- Dong, Y., et al. (2025). A Survey on Code Generation with LLM-based Agents. arXiv:2508.00083.
- Jimenez, C. E., Yang, J., Wettig, A., et al. (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv:2310.06770.
- Jin, H., Huang, L., Cai, H., et al. (2025). From LLMs to LLM-based Agents for Software Engineering. arXiv:2408.02479.
- Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv:2302.06590.
- Wang, N., et al. (2025). Efficient Agents: Building Effective Agents While Reducing Cost. arXiv:2508.02694.
- Wang, X., et al. (2025). OpenHands: An Open Platform for AI Software Developers as Generalist Agents. arXiv:2407.16741.
- Xia, C. S., Deng, Y., Dunn, S., & Zhang, L. (2024). Agentless: Demystifying LLM-based Software Engineering Agents. arXiv:2407.01489.
- Zhang, Y., Ruan, H., Fan, Z., & Roychoudhury, A. (2024). AutoCodeRover: Autonomous Program Improvement. arXiv:2404.05427.