Cost per Completed Task: A Measurement Framework and Cross-Domain Case Study for Evaluating AI Model Configurations in Agentic Workflows

Proposta do CPCT como métrica para avaliar o custo real de agentes de IA — não por token, mas por tarefa concluída e aceita

Publicado em 04/07/2026 • Tags: agentic workflows, large language models, AI agents, cost per completed task, software engineering economics, LLM evaluation, AI governance

Interessado neste tema? Agende uma conversa rápida sobre Cost per Completed Task: A Measurement Framework and Cross-Domain Case Study for Evaluating AI Model Configurations in Agentic Workflows.

Agendar 20 min

Abstract

The adoption of large language model (LLM) agents has moved the economic unit of analysis from isolated generation to multi-step task execution. In an agentic workflow, a system may read files, plan, edit, run commands, inspect failures, retry, and request review before reaching a usable result. Under this execution model, evaluating an AI system by token price or single benchmark score is insufficient.

This paper proposes Cost per Completed Task (CPCT) as a system-level measurement framework for comparing AI model configurations in agentic workflows. CPCT aggregates input/output token cost, cache usage, tool and sandbox cost, wall-clock runtime, failed-attempt cost, and human-review effort into a single task-centred denominator: accepted work.

The paper (i) situates CPCT against Total Cost of Ownership (TCO) and Activity-Based Costing (ABC) and against emerging cost-aware agent metrics such as cost-of-pass; (ii) decomposes task cost into operational components with an operable rubric for the human and quality terms; (iii) resolves the failure-mode edge cases (total-failure division-by-zero, sunk cost of timeouts) that a naïve ratio leaves undefined; and (iv) specifies a repeatable protocol with a per-run data schema.

1. Introduction

Large language models are increasingly deployed as agentic systems that receive a goal, inspect an environment, invoke tools, take actions, observe results, and iterate until a task is completed or abandoned. This is true across software engineering (issue resolution, refactoring, CI triage) and, increasingly, across knowledge-work domains such as legal and forensic analysis.

This shift changes how cost and productivity should be measured. In stateless API usage, models are naturally compared by input and output token price. In agentic execution, the relevant economic unit is no longer the isolated prompt or token; it is the completed task.

In agentic workflows, the cheapest model is not necessarily the one with the cheapest token. The economically preferable configuration minimizes the quality-adjusted cost per accepted task — and much of that cost is operational overhead (tool-output verbosity, context re-reads, turns) that is independent of unit-token price.

1.1 Contributions

  1. It proposes Cost per Completed Task as a system-level metric, arguing that CPCT is a property of the {model + scaffold + prompt + tools + repository/data + tests + reviewer} system, not of a model in isolation.
  2. It decomposes task cost into token, cache, tool, runtime, retry, and human-review components, with an operable rubric for the review-cost and quality terms.
  3. It provides a repeatable protocol — a controlled-variable specification and a per-run data schema — for comparing configurations.
  4. It presents a real, measured cross-domain case study (forensic report generation, not code) whose central result — an 8.6× cost swing on an identical model and task driven purely by operational overhead — provides clean evidence that token price is a weak proxy for cost.

2. Background

2.1 From code completion to agents

Early LLM coding tools suggested completions; controlled studies of GitHub Copilot reported gains such as 55.8% faster completion on some standardized tasks (Peng et al., 2023). Later work complicated this: a randomized controlled trial by METR on experienced open-source developers on mature projects found early-2025 tooling increased task time by 19%, even as developers perceived a speed-up (Becker et al., 2025).

2.2 Benchmarks for real tasks

SWE-bench evaluates whether a produced patch resolves a real GitHub issue without breaking existing behavior, across 2,294 problems in 12 Python repositories (Jimenez et al., 2024). AutoCodeRover pairs LLMs with program-analysis search at low dollar cost per task; Agentless shows a simpler localize–repair–validate process can be competitive at low cost; OpenHands provides an open platform with sandboxed execution and evaluation.

2.3 Token consumption and cost-aware evaluation

Efficient Agents frames an efficiency–effectiveness trade-off and introduces cost-aware ideas such as cost-of-pass (Wang, N. et al., 2025). Work on token consumption in agentic coding reports that agentic tasks can consume far more tokens than chat or reasoning, that higher consumption does not imply higher accuracy.

3. Positioning CPCT Against Existing Cost Concepts

3.1 Relationship to TCO and Activity-Based Costing

CPCT is a domain-specific specialization of two established ideas. Total Cost of Ownership (TCO) holds that acquisition price is only part of a resource's cost. Activity-Based Costing (ABC) assigns indirect and overhead cost to the activities that consume them. CPCT applies both to agentic work: the "purchase price" (token price) is one component, and the framework attributes operational overhead — turns, cache, tools, runtime, retries, review — to the activity that consumes it (the task).

3.2 A ladder of cost-aware metrics

MetricUnitCapturesMisses
Cost per token1M tokensNominal unit priceEverything operational
Cost per callAPI requestRequest-level priceMulti-turn loops, retries
Cost per generated patchpatchOutput producedWhether it works
Cost per resolved issuepassing taskTest-passing resolutionReview, quality, maintainability
Cost-of-passpassing taskExpected cost to a passHuman review, quality grading
CPCTaccepted taskAll operational + review cost(time — see TCPCT)
Quality-adjusted CPCTquality-weighted taskAbove + reviewer quality(rubric subjectivity)

4. Research Problem and Questions

Central problem. How should teams evaluate the economic efficiency of AI model configurations in agentic workflows where the objective is not token generation but accepted task completion?

  • RQ1. Does lower token price reliably imply lower end-to-end cost?
  • RQ2. Which operational variables most influence total task cost beyond token price?
  • RQ3. How can teams normalize comparison around completed tasks rather than model calls?
  • RQ4. What governance decisions can a CPCT metric support?

5. The CPCT Framework

5.1 Task and success definitions

LevelCriterionNotes
L1 — Technical passTarget output produced; automated checks passSWE-bench-style fail-to-pass / pass-to-pass
L2 — Review passA human expert accepts the outputMinimality, security, correctness, coverage
L3 — Production passIntegrated / filed without observed regressionThe economically meaningful outcome

5.2 Per-run cost decomposition

Cᵢ = Cᵢ(tok) + Cᵢ(cache) + Cᵢ(tool) + Cᵢ(run) + Cᵢ(rev) + Cᵢ(retry)

where the terms are, respectively, direct token cost; billed cache read/write cost; external tool/sandbox/CI/hosted-execution cost; wall-clock or compute cost; human review cost; and failed-attempt/rework cost.

5.3 Basic CPCT and the failure-mode fix

CPCT = Σ min(Cᵢ, Cᵢ(max)) / Σ Sᵢ, together with the acceptance rate Σ Sᵢ / n

Always report CPCT with the acceptance rate. When Σ Sᵢ = 0, CPCT is reported as "undefined" — no accepted tasks — with total sunk cost stated, never as a finite number that hides total failure.

5.4 Quality-adjusted CPCT

QCPCT = Σ min(Cᵢ, Cᵢ(max)) / Σ (Sᵢ · Qᵢ), with Qᵢ ∈ [0,1]

QᵢReviewer decision
1.0Accepted without changes
0.7Accepted after minor human-suggested refactor
0.3Requires partial rewrite
0.0Rejected

5.5 Time-adjusted CPCT

TCPCT = Σ [min(Cᵢ, Cᵢ(max)) + Vₜ · Tᵢ] / Σ Sᵢ

Tᵢ is elapsed time; Vₜ is the value assigned to time. The case study confirms that time is a separable axis: parallelizing independent stages cut wall time ~40% at identical dollar cost.

6. Experimental Protocol

6.1 Minimal unit and repetition

The minimal unit is a tuple (model, scaffold, task, run). Because agentic execution is non-deterministic, at least three runs per task per configuration are recommended, reported with dispersion.

6.4 Per-run data schema

FieldDefinition
task_id / stageTask or pipeline-stage identifier
modelModel name + version
agent_scaffoldAgent/tool harness
input_tokens / output_tokensGeneration and context cost
cache_read / cache_writeContext dependence
runtime_secondsOperational and time cost
success_levelL1 / L2 / L3 / fail
review_minutes / review_scoreC(rev) and Qᵢ
total_costAggregated Cᵢ

7. Measured Case Study: A Cross-Domain Forensic Pipeline

7.1 Domain and apparatus

The case study is deliberately outside software engineering, to test CPCT's generality. The domain is agentic forensic expert-report (laudo pericial) generation on the DraIara/OpenClaw harness: the agent reads an OCR'd evidence dossier, calls forensic tools via MCP, produces maps and crops, writes a sixteen-section report, and self-reviews. Two named models were used: Claude Opus 4.8 (premium, ≈US$15/M output) and Claude Sonnet 4.6 (budget, ≈5× cheaper per token).

7.2 The isolation experiment: same model, same task, 8.6× cost

Configuration (identical model & task)Cost (US$)TurnsInput tokens
Sonnet, no I/O cap7.9320766,715
Sonnet + I/O cap1.181024
Sonnet + I/O cap + tool-lean0.92616

Cost fell 88% (an 8.6× swing) with no change of model. The driver was tool-output verbosity: nine IP-geolocation maps returned ~766,715 input tokens that were re-ingested into context. This is the paper's strongest single piece of evidence that token price is a weak proxy for cost.

7.3 Stage-decomposed pipeline: the irreducible cost floor

StageModelCost (US$)TurnsRuntimeOutput tokCache read
extractSonnet 4.60.6728181 s9,344792,255
geoSonnet 4.61.819203 s4,454308,518
integritySonnet 4.60.4515135 s7,348394,178
osintSonnet 4.60.4621125 s7,043374,040
biometricsSonnet 4.60.215145 s768156,058
synthesisOpus 4.83.8221720 s50,3982,238,068
Totalhybrid7.4299~19 min79,3554.26 M

The single premium synthesis stage is 51.5% of total cost. It behaves as an irreducible floor: no amount of cheaper tokens on the other five stages can push total cost below it.

7.4 End-to-end scenarios and the time axis

ScenarioCost (US$)Wall time
Original (all-Opus monolith)~9.83~30 min
Decomposed naïve (all-Opus, estimated)~12–16 (est.)~30+ min
Optimized (Sonnet stages + Opus synthesis)6.29~25 min
+ Parallelized6.29~15 min

Moving 5/6 stages to the budget model cut cost only 36% (9.83 → 6.29), not ~90%. Parallelizing independent stages cut wall time ~40% (25 → 15 min) at identical dollar cost.

7.5 Findings mapped to the research questions

  • RQ1 — No. A ~5× per-token discount on most of the pipeline produced a 36% end-to-end saving; and the same-model 8.6× swing shows token price is not even the main lever.
  • RQ2 — Operational overhead dominates. The measured dominant drivers were tool-output verbosity fed back into context, the turns it induced, and cache/context re-reads — all above unit-token price.
  • RQ3 — Decompose and account per stage. Per-stage cost accounting localizes the irreducible floor and the movable work.
  • RQ4 — Route by stage, cap I/O, parallelize. Reserve the premium model for the accepted-work-producing stage; run cheaper stages on the budget model; cap tool output; parallelize independent stages for time without cost.

8. Discussion

8.1 The false economy of cheap tokens

Cheap tokens suit simple, high-volume tasks. For complex work, the measured floor result (§7.3) is decisive: the cost of the stage that produces accepted work sets a floor that per-token discounts elsewhere cannot break.

8.2 CPCT is a system property, not a model property

CPCT = f(model, scaffold, tools, data, prompt, tests, reviewer)

The isolation experiment proves this quantitatively: the same model spanned US$0.92–7.93 on one task purely from scaffold-level I/O discipline.

8.3 Operational overhead is the primary cost lever

The practical corollary is that context hygiene is a first-class cost control: capping tool-output verbosity, trimming re-ingested context, and reducing turns produced an 8.6× saving with no model change and no quality loss at L2.

9. Threats to Validity

Single-run measurements. Section 7 reports single runs on a few real cases; figures need repeated runs and dispersion before being treated as general point estimates.

Cross-domain transfer. Evidence is from forensic report generation; transfer to coding agents is argued (via the domain-independent isolation mechanism) but not directly measured here.

Model and price volatility. Models, prices, context windows, and caching policies change quickly. A CPCT figure is time-stamped and records model version, date, pricing snapshot, and scaffold configuration.

10. Practical Recommendations

  1. Measure cost per completed task, not per token.
  2. Audit tool-output verbosity and re-ingested context first — it can dominate cost more than model choice.
  3. Decompose pipelines and account cost per stage; find the irreducible floor.
  4. Route each stage to the cheapest tier that still passes L2; reserve the premium model for the accepted-work-producing stage.
  5. Parallelize independent stages for time without cost.
  6. Track failed attempts and include their capped cost.
  7. Log model version, scaffold, prompt version, data snapshot, and acceptance level.
  8. Use a published quality rubric; run ≥3 repetitions and report dispersion.
  9. Include standardized human-review time.
  10. Re-run evaluations periodically as models and prices change.

11. Conclusion

Agentic workflows change the economics of AI-assisted work. Token price is an incomplete and sometimes misleading proxy. This paper proposed Cost per Completed Task as a system-level measurement framework, resolved the failure-mode edge cases a naïve ratio leaves undefined, and specified a protocol and data schema.

A real, measured cross-domain case study grounded the framework: holding the model fixed, operational discipline alone moved the cost of one task by 8.6×; a premium synthesis stage formed a 51.5% irreducible cost floor that per-token discounts elsewhere could not break; and parallelism cut time ~40% at equal cost.

The broader conclusion is that AI configuration selection should be governed by completed-task economics: the true cost of AI is the cost of correct, accepted, and maintainable work — most of which is operational, not per-token.

Data Availability

The per-stage and per-configuration measurements in Section 7 are reported in full in this manuscript. An anonymized per-run dataset following the schema in Section 6.4, together with the figure-generation scripts, accompanies the archival deposit.

Conflict of Interest

The author declares no conflict of interest. The case study uses commercial models (Anthropic Claude Opus 4.8 and Sonnet 4.6) accessed via a third-party router; this is a factual description of the measurement environment and no sponsorship or commercial relationship influenced the results.

Suggested Citation

Filho, O. J. (2026). Cost per Completed Task: A Measurement Framework and Cross-Domain Case Study for Evaluating AI Model Configurations in Agentic Workflows. Zenodo. https://doi.org/10.5281/zenodo.21198419

References

  • Bai, L., Huang, Z., Wang, X., Sun, J., Mihalcea, R., Brynjolfsson, E., Pentland, A., & Pei, J. (2026). How Do Coding Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks. Preprint.
  • Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089.
  • Chowdhury, N., et al. (2024). Introducing SWE-bench Verified. OpenAI.
  • Dong, Y., et al. (2025). A Survey on Code Generation with LLM-based Agents. arXiv:2508.00083.
  • Jimenez, C. E., Yang, J., Wettig, A., et al. (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv:2310.06770.
  • Jin, H., Huang, L., Cai, H., et al. (2025). From LLMs to LLM-based Agents for Software Engineering. arXiv:2408.02479.
  • Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv:2302.06590.
  • Wang, N., et al. (2025). Efficient Agents: Building Effective Agents While Reducing Cost. arXiv:2508.02694.
  • Wang, X., et al. (2025). OpenHands: An Open Platform for AI Software Developers as Generalist Agents. arXiv:2407.16741.
  • Xia, C. S., Deng, Y., Dunn, S., & Zhang, L. (2024). Agentless: Demystifying LLM-based Software Engineering Agents. arXiv:2407.01489.
  • Zhang, Y., Ruan, H., Fan, Z., & Roychoudhury, A. (2024). AutoCodeRover: Autonomous Program Improvement. arXiv:2404.05427.

Sobre o Autor

Osvaldo Janeri Filho

Osvaldo Janeri Filho

Perito Digital especializado em perícia digital, nulidade de provas digitais, contratos digitais e cibersegurança. Executivo em Tecnologia & Segurança da Informação com mais de 15 anos de experiência no mercado.

Perícia Digital LGPD/GDPR Cibersegurança Direito Digital

Precisa de Consultoria Especializada?

Conte com minha expertise em perícia digital, direito eletrônico e cibersegurança. Agende uma reunião de 20 minutos para discutir seu caso específico.