Home · Writing · Research

The ROI Denominator Problem: Hours Saved are not Benefits Captured

A practical method for converting AI time savings into schedulable capacity, verified operating outcomes and finance-owned value, without treating every minute returned as cash.

TLDR

  1. A practical method for converting AI time savings into schedulable capacity, verified operating outcomes and finance-owned value, without treating every minute returned as cash.
  2. This paper is written for chief financial officers, transformation leaders and operating executives. Read Parts I and II for the decision logic, Part IV for the worked example, and Part V for the field artefact.
  3. Now try to spend it. The ten minutes arrive at different times, inside different roles, after different tasks.
  4. The distinction matters because productivity evidence is real, yet it does not license a direct jump to cash.
  5. The familiar ROI formula is simple: net benefit divided by investment. The disputed number is usually the benefit in the numerator.
Gross-to-captured-capacity waterfallAn illustrative waterfall starts with one hundred gross hours and removes adoption decay, fragmentation, demand gaps and quality rework, leaving thirty-seven usable hours. Where 100 reported hours go 100 hgross 85 hsustained 62 haggregated 47 hdemanded 37 husable
Figure 1. Gross-to-captured-capacity waterfall. Illustrative values: 100 gross hours become 37 usable hours after measured frictions. Status: synthetic scenario, not a benchmark.
On this page

This paper is written for chief financial officers, transformation leaders and operating executives. Read Parts I and II for the decision logic, Part IV for the worked example, and Part V for the field artefact.

Glossary

  • Gross time returned: task minutes avoided before testing whether people can aggregate or use them.
  • Usable capacity: time that is available in blocks large enough to schedule against real demand.
  • Capacity-capture coefficient, κ: the share of gross time returned that becomes usable capacity after measured frictions.
  • Conversion action: a management change that absorbs capacity into throughput, service, quality, risk reduction, redeployment or cost removal.
  • Operational readback: evidence from the system of record that the intended workflow effect actually occurred.
  • Realised value: a benefit accepted by its accountable owner under an agreed evidence and attribution rule.

Part i: the ten-minute dividend that nobody can schedule

Imagine a 3,000-person organisation in which an AI assistant saves every participating employee ten minutes each working day. At 220 working days and 70 per cent sustained use, the arithmetic produces 77,000 hours a year. Multiply that total by a loaded labour rate of £42 and the spreadsheet announces £3.23 million of annual benefit.

Now try to spend it.

The ten minutes arrive at different times, inside different roles, after different tasks. Some employees use the gap to improve the next document. Some answer email. Some absorb an interruption. A few leave earlier. Demand remains unchanged, no shift loses a post, no contractor invoice falls, no queue accepts more cases and no product reaches a customer sooner. The organisation may be better off, but its cost base and throughput have not moved by £3.23 million.

That is the thought experiment at the centre of this paper. Gross time returned is an engineering observation; realised benefit is an operating outcome. The missing link is a management mechanism that makes the returned time schedulable and assigns it to an outcome that somebody owns.

The distinction matters because productivity evidence is real, yet it does not license a direct jump to cash. A 2026 BIS working paper used data on more than 12,000 non-financial firms and estimated a 4 per cent increase in labour productivity associated with AI adoption after its identification strategy. It found no adverse firm-level employment effect in the short run and linked stronger gains to complementary investment in software, data and training. Those findings describe higher output per worker, not an automatic reduction in payroll. The authors also warn that longer-term effects remain uncertain and that benefits are uneven across firm sizes (BIS working paper).

The OECD reaches a compatible conclusion from a different evidence base. Its review of experimental research finds that effects vary with task and user experience, that human-AI collaboration matters, and that evidence on long-term business effects remains thin (OECD AI Paper 39). A task can become faster while the surrounding process absorbs none of the gain. Local efficiency and enterprise value live at different levels of analysis.

The proof gap is visible in executive surveys. Deloitte's 2026 survey of 64 US healthcare finance leaders classified 44 per cent as AI scalers, yet only 18 per cent of that group reported mature financial attribution. The earlier-stage group reported 31 per cent. The sample is small and sector-specific, so it cannot establish a universal rate. It does expose a useful tension: deployment can outrun the ability to attribute outcomes (Deloitte methodology and findings).

Vendor surveys often tell a more optimistic story. Google Cloud reported that 74 per cent of surveyed executives said they achieved ROI within the first year. It also reported large productivity gains among a subset of respondents (Google Cloud's 2025 survey summary). Those are source-reported perceptions from a vendor-sponsored study. They are useful market evidence, but they do not reveal how much saved time became cash, capacity, quality or revenue in each organisation.

BCG reported the opposite side of the adoption gap: its global survey found that 60 per cent of companies were not generating material value from AI despite substantial investment. Its interpretation centres on shallow use and limited work reinvention rather than a lack of tool access (BCG's 2025 adoption analysis). This is consulting survey evidence, not a causal estimate. It still supports an operating question that usage dashboards cannot answer: did the organisation change how work reaches an outcome?

The finance error begins when a proxy changes category without evidence. Minutes avoided become hours saved; hours saved become capacity; capacity becomes wage value; wage value becomes a booked benefit. Every step may be reasonable. None is automatic.

The denominator inside the roi formula

The familiar ROI formula is simple: net benefit divided by investment. The disputed number is usually the benefit in the numerator. The deeper denominator problem appears one level below it. A time-based benefit estimate divides a population's potential saving across eligible tasks and assumes that each returned minute has the same economic yield. It rarely does.

Task selection creates the first denominator. A pilot may measure only clean, repeated cases while the business case applies the result to every case. Active adoption creates the second. A licence assigned to an employee does not mean the tool was used for the eligible task. Accepted completion creates the third. Draft generation is not finished work when checking and repair follow. Capacity capture creates the fourth. Returned time has to form usable blocks against real demand. Benefit recognition creates the fifth. Finance needs an agreed rule for when a capacity, service, quality, risk or cash outcome can be accepted.

Each denominator can shrink the eligible base. That is why a programme can report an accurate nine-minute task saving and still publish a misleading annual benefit. The error comes from applying the measured difference to a larger or economically different population.

A defensible calculation preserves the chain. Report the number of eligible task events, the share actually assisted, the share sustained after novelty, the share accepted without extra correction, the share aggregated into usable capacity and the share converted through a named action. This allows a reviewer to reproduce the claim and challenge the precise assumption that matters.

The counterfactual belongs in that chain. If demand, staffing, seasonality or policy changed during rollout, a before-and-after comparison will attribute some of that movement to AI. A retained cohort, matched queue or staggered release will not remove every confounder, but it makes the alternative visible. A value claim without an explicit counterfactual is a forecast wearing the clothes of a result.

Five benefit classes, five different proofs

The cleanest remedy is to stop forcing every gain into one currency. A time-saving intervention can create at least five benefit classes.

First, it can create employee welfare: lower after-hours work, fewer interruptions or more time for learning. Second, it can create usable capacity: schedulable time against existing demand. Third, it can change an operating result: more completed cases, faster cycle time or fewer defects. Fourth, it can alter cash: lower overtime, contractor spend, headcount, leakage or cost-to-serve. Fifth, it can expand opportunity: more sales activity, more experiments or earlier market entry.

These benefit classes can coexist. They still need separate measures. A reduction in after-hours work should be reported as such. Calling it a salary saving makes the employee benefit less credible, not more.

Benefit class Proper unit Minimum evidence What does not prove it
welfare or resilience after-hours minutes, interruption rate, absence, sentiment baseline plus repeated employee and work-pattern measures loaded salary multiplied by reported time
usable capacity schedulable hours in the relevant queue or role task telemetry, adoption durability and rota or queue readback prompt count or one-off survey
operating outcome accepted units, cycle time, quality, service level system-of-record outcome with comparison model accuracy alone
cash or margin spend avoided, margin added, loss reduced finance reconciliation and owner sign-off theoretical wage value
option value experiment speed, learning milestone, decision date explicit hypothesis, expiry date and decision rule indefinite strategic narrative

Part ii: the capacity-capture mechanism

The capacity-capture model begins with a narrow system boundary: one task class, one population, one period and one operational outcome. It then follows six transformations.

Let gross time returned be G. It comes from eligible task volume, measured time difference and sustained adoption. Let the usable-capacity coefficient be κ. For a practical first model:

Here, N is the eligible population, f is task frequency, Δt is the measured time difference, a is active adoption and s is sustained use. The second line contains four conversion terms: g for aggregation into schedulable blocks, d for available demand, q for quality retention after rework, and m for the share attached to an executed management action. U is usable capacity.

This multiplicative form is deliberately unforgiving. If nobody owns a conversion action, m approaches zero even when the tool is fast. If quality falls and rework returns, q falls. If ten-minute fragments cannot be pooled, g falls. The coefficient is a measured explanation of leakage, not a discount selected to make a business case look prudent.

Estimating κ without inventing precision

Start with intervals rather than a single point estimate. For every component, record a lower bound, central estimate and upper bound. The range should come from observed variation across teams, weeks or task types. If there is no observation yet, label the input as a planning assumption and give it an expiry date.

Estimate aggregation (g) from work patterns. Analyse calendar or queue events in a privacy-preserving way and count the returned minutes that form blocks large enough for the destination work. The threshold must fit the role. Ten minutes may be usable in a call queue and useless for a credit review that needs uninterrupted attention. Test several thresholds rather than choosing one that flatters the estimate.

Estimate demand (d) from the constrained queue. Use backlog, abandonment, service-level misses, overtime or deferred work to show that additional capacity has somewhere to go. Demand should be measured at the point where the capacity becomes available. A backlog elsewhere in the organisation does not prove that the relevant team can absorb the time.

Estimate quality retention (q) from accepted completion. Include review, repair, reopened work and downstream correction. When the intervention changes the definition of a task, retain both the old and new measures long enough to bridge them. Otherwise the programme can improve the reported rate by quietly narrowing what completion means.

Estimate management conversion (m) from observed action. This is the least technical input and often the most consequential. Did the rota change? Did a service target tighten? Did overtime fall? Did named work move to the released block? Count only the share covered by an executed decision. A steering-committee intention is not an operating event.

The product of four uncertain terms can create false confidence. Use a simple simulation or sensitivity table to show how the result changes across plausible ranges. Correlations also matter. Weak aggregation may reduce demand absorption because capacity arrives at the wrong time. Quality problems may reduce adoption. Where dependencies are material, model them together rather than treating every coefficient as independent.

The value of κ is diagnostic. If its central estimate is 0.28, the question is not whether 0.28 is respectable. Ask which component constrains it, what intervention could move that component and what evidence would show the move occurred. That turns the model into an operating agenda.

Microsoft's current agent value guidance makes two helpful measurement choices: establish a baseline before rollout and, where possible, retain a comparison group. It also recommends combining telemetry with self-reported time and tracking where reclaimed time goes (Microsoft Learn). The capacity-capture model extends that advice by making the conversion frictions explicit.

AWS prescriptive guidance similarly starts from current process cost, links measurement criteria to autonomy and error tolerance, and calls for decision points that terminate non-performing agents (AWS Prescriptive Guidance). That is useful discipline. A benefits case needs an exit rule as much as an upside estimate.

Fragmentation curve for returned timeThree illustrative curves show that the same gross saving creates different schedulable capacity when minutes are dispersed across many people or concentrated in repeated roles. Fragmentation changes the yield many roles, irregular tasksstable team, repeat queuepooled work, rota redesigned gross minutes returned per person per day →schedulable share →
Figure 2. Fragmentation curve. Concentration, repeatability and pooling determine how quickly returned minutes form schedulable blocks. Status: illustrative relationship; local telemetry is required.

Why a coefficient is better than a haircut

Many benefit models apply a flat 30 or 50 per cent realisation haircut. The adjustment may reduce overclaiming, but it teaches the organisation nothing. A coefficient built from observable causes can guide action.

Suppose aggregation is the main loss. The intervention is work pooling, protected focus time or role redesign. Suppose demand is the loss. The intervention may be to remove a bottleneck elsewhere or stop counting the time as near-term capacity. Suppose management conversion is the loss. The fix is an accountable decision, not a model upgrade.

The coefficient should be estimated by cohort and work pattern. A contact-centre queue with minute-by-minute demand can absorb small fragments. A legal team with long, interdependent matters may need half-day blocks. A finance close process can absorb capacity only during a narrow period. A single enterprise-wide rate hides the mechanism.

Work pattern Likely aggregation yield Main measurement Plausible conversion action
continuous pooled queue high idle time, queue depth, accepted completions revise staffing curve or service target
repeat casework by stable team medium to high block availability, backlog, rework redesign rota and case allocation
knowledge work across many roles low to medium calendar fragments, task switching, deliverable cycle create pooled work or protected blocks
episodic specialist task low event frequency and avoided delay treat as resilience or option value
seasonal close or campaign variable capacity inside the critical window move deadlines, contractors or scope
Capacity conversion chainSix connected stages run from an observed task event through gross time, schedulable capacity, management action and operational readback to an accepted benefit. Value appears only after readback taskevent grosstime usablecapacity ownedaction operatingreadback acceptedbenefit telemetrytime studyκ estimateownersystem of recordfinance sign-off
Figure 3. Capacity conversion chain. Each stage needs its own evidence and owner. Status: Applied analysis.

The decision test

Ask one question before assigning monetary value: what observable management action would be different if the saved minutes were real? If no queue, rota, budget, service target, risk control or opportunity allocation changes, classify the gain as unconverted time until stronger evidence arrives.

Part iii: where the business case breaks in production

Self-report and telemetry answer different questions

Self-report can reveal cognitive load, hidden work and where people believe time went. It is vulnerable to recall, novelty and expectation effects. Application telemetry records use and event duration, but it can miss preparation, checking, waiting and work displaced to another system. Neither should dominate by default.

Use paired evidence. Sample tasks before and after. Record accepted outputs rather than generated outputs. Ask users where the reclaimed time went. Link the task event to the downstream case or service outcome. Where feasible, retain a comparison group or use a staggered rollout. The stronger design triangulates time, quality and outcome rather than perfecting one proxy.

Fragmentation has a geometry

Ten minutes returned to one hundred people is not equivalent to sixteen hours returned to one role. The arithmetic total is similar. The scheduling possibilities are different.

Calendar structure matters. So does queue topology. Continuous queues can absorb small increments; project work often cannot. Shared service teams can pool capacity; highly specialised roles cannot do so without changing who owns the work. If benefits depend on aggregation, the business case must include the organisational design that creates it.

Ten-minute thought experiment calendarA calendar grid shows small saved fragments scattered across employees and days, contrasted with a pooled half-day block created through rota redesign. Seventy fragments are not one shift distributed work MonTueWedThuFri pooled blockone owned queue · ½ day rota redesign converts fragments into an assignable unit
Figure 4. Ten-minute calendar. The same nominal total can remain dispersed or become a pooled block through work redesign. Status: illustrative thought experiment.

Quality can consume the dividend

Fast first drafts often create review work. A task timer that stops when the draft appears overstates the saving if an expert later checks evidence, repairs formatting or reconciles a downstream record. Measure accepted completion, including retries and review.

Quality retention (q) should cover more than model accuracy. It should include first-time-right rate, material error, policy adherence, downstream correction and reopened cases. A five-minute task saving paired with seven minutes of diffuse checking has negative usable capacity even if the model's answer looks fluent.

Demand may not be available

Some teams have more demand than capacity. Others do not. If the work queue is already empty, a faster task creates slack rather than throughput. That slack can have value as resilience, learning or faster response to peaks. It is not automatically monetisable.

Demand also moves across bottlenecks. Accelerating document preparation can increase waiting at approval. Faster case summarisation can increase the investigator's review queue. The unit of value is the constrained workflow outcome, not the accelerated local task.

Authority determines whether capacity moves

Even a clear block of available time remains theoretical if a manager cannot change work allocation. Employment agreements, skill boundaries, budget ownership, union rules, service commitments and risk controls can all limit conversion.

This is why the owner matters. The technology lead can measure time. The operating owner can redesign the queue. Finance can accept or reject the benefit class. People leaders can test workload and welfare effects. Without that chain of authority, a benefits dashboard becomes a commentary surface.

Six production failures that survive a polished dashboard

The first is denominator drift. The pilot measures a narrow task, then the portfolio model applies the saving to all users, all tasks and every working day. The dashboard retains the original task label, so the expansion is hard to see. Require an eligibility query that can be rerun from the event data.

The second is novelty decay. Early users are selected, supported and observed. Their adoption falls once the intervention becomes ordinary or the tool encounters less curated work. Measure sustained use after the support period and preserve cohort curves rather than reporting only cumulative users.

The third is shadow checking. People learn that the draft is usually right but still verify every fact because accountability has not changed. The checking time moves to email, a second browser or an informal peer review and disappears from the product telemetry. Periodic work sampling and downstream correction data expose it.

The fourth is bottleneck migration. The assisted step accelerates, but review, approval or fulfilment becomes the constraint. Local cycle time falls while end-to-end cycle time does not. Trace the item to final readback and report work in progress between stages.

The fifth is rebound demand. Cheaper or faster work causes more requests. That can be valuable, but it can also raise total cost. A legal summarisation tool may encourage teams to submit more material; a service assistant may lower the threshold for opening a case. Separate unit efficiency from total volume and ask whether the extra demand is desirable.

The sixth is double counting. Capacity is credited to the AI programme, then the higher throughput is credited again to a process programme, while avoided contractor spend is claimed by both. A common benefit identifier, one accountable owner and an allocation rule are mundane controls. They prevent the same outcome from appearing in several investment cases.

These failures have different remedies. More model accuracy addresses only some quality leakage. The others require measurement design, work allocation, authority and finance governance. A production review should therefore inspect the capacity chain with the same seriousness used for latency, error and cost.

Failure surface for capacity captureA two-dimensional heat field maps aggregation and demand absorption to capacity capture, with a marker showing that strong time savings can still sit in a low-capture region. The capture surface aggregation into schedulable blocks →demand absorption → high gross savinglow capture redesignedwork
Figure 5. Capture failure surface. High gross savings can remain in a low-value region when aggregation and demand absorption are weak. Status: synthetic decision surface.
Depth: attribution design and the negative control

The minimum negative control is a comparable population that receives no intervention during the same measurement window. Where randomisation is practical, randomise at the team or workflow unit to limit contamination. Where it is not, use phased rollout, matched queues or interrupted time series, and record concurrent changes.

Do not compare only task duration. Compare accepted completion, quality, queue movement and downstream outcome. The AI-assisted group may finish drafts sooner while the control group clears more cases because review burden differs. A defensible result reports both.

The strongest weakening observation for this paper's thesis would be repeated evidence that distributed, unowned time savings reliably predict finance-accepted outcomes without work redesign or demand absorption. If that pattern appears across contexts, κ may be unnecessary. Current evidence and operating experience point the other way, but the mechanism should remain falsifiable.

Part iv: a worked public-service rota

Consider a synthetic benefits-processing team of 48 case officers and six reviewers. The team handles a persistent backlog. An AI drafting assistant prepares the first version of routine correspondence and summarises evidence already present in the case file. It cannot approve a benefit, alter entitlement or send correspondence without review.

The baseline sample covers four weeks. Eligible officers complete 6.2 assisted tasks per day. Accepted completion takes 31 minutes at baseline, including preparation and checking. During a six-week controlled rollout, the assisted cohort averages 22 minutes. Quality and reopen rates remain inside pre-agreed bounds after excluding an early prompt defect. Active adoption settles at 78 per cent.

The gross calculation is attractive: 48 officers × 6.2 tasks × 9 minutes × 20 working days × 78 per cent equals roughly 699 hours returned per month.

The first operating review finds that only 29 per cent appears as blocks of at least thirty minutes. Officers are still assigned cases individually, and saved fragments land unpredictably. Reviewer capacity is also constrained. The queue completes only 3 per cent more cases, far below the gross-time implication.

The team then changes the operating design. Routine cases enter a pooled morning queue. Officers rotate through two protected drafting blocks. Reviewers receive a risk-tiered batch rather than interruptions throughout the day. The AI assistant writes an effect receipt that links each accepted draft to the case identifier, review result and final send event. Staffing does not fall. Overtime and agency cover are unchanged.

In the second period, aggregation g rises from 0.29 to 0.71. Available demand d is 0.92 because the backlog is persistent. Quality retention q is 0.95 after rework. The executed management-action share m is 0.86 because some capacity remains protected for training and complex cases. κ becomes 0.53 after rounding. About 370 of the 699 gross hours become usable capacity.

The estimates are reviewed as a range. In the cautious case, lower sustained use and aggregation produce κ of 0.39. In the central case, κ is 0.53. In the upper case, it reaches 0.62, but only if reviewer capacity expands with the officer blocks. The steering group uses the cautious case for funding and the central case for operational planning. It does not present the upper case as a commitment.

The negative control remains on the old rota for the comparison period. It uses the same assistant, receives the same training and serves a similar case mix. Its gross task saving is close to the redesigned group, yet aggregation and completed-case movement remain low. This matched contrast isolates the work-design mechanism more clearly than a pre-rollout comparison alone. It also prevents the programme from crediting the assistant for seasonal backlog changes shared by both groups.

Reviewers inspect equity and conduct effects. Faster routine processing must not create slower service for complex cases or pressure officers to classify borderline work as routine. The evidence pack therefore reports cycle time and error by case complexity, vulnerability flag and channel. A portfolio average could conceal harm concentrated in a small group.

The team also records what it chose not to monetise. Training blocks consume some returned capacity. Officers use part of the time for complex-case discussion. Those choices reduce short-term throughput and may improve future quality. They remain visible in the conversion ledger as deliberate investment, not leakage or cash saving.

That result still is not a cash saving. The team records it as capacity and tests the operational readback. Monthly accepted case completions rise by 11 per cent relative to the matched queue, median correspondence time falls by 1.7 days and reopen rate remains statistically indistinguishable inside the programme's practical threshold. These numbers are synthetic, designed to show the method. A real programme would publish denominators, uncertainty and concurrent changes.

The value owner chooses demand absorption: reduce backlog and improve service rather than cut roles. Finance therefore values verified additional completions and avoided backlog escalation under an agreed service-cost model. It does not book 699 hours × salary. The management decision selects the value path; the technology does not.

Rota redesign before and afterThe upper flow shows scattered officer drafting and interrupt-driven review. The lower flow shows a pooled queue, protected work blocks, risk-tiered review and case-system readback. The gain arrived after rota redesign before individualcase listscattereddraftinginterruptreviewweak queuemovement after pooledroutine queueprotectedblocksrisk-tieredreview batchcase-systemreadback same assistant · changed work design · different captured outcome
Figure 6. Rota redesign. Pooling, protected blocks and review batching convert scattered savings into usable capacity. Status: synthetic worked example.
Measure Before AI Assisted, old rota Assisted, redesigned rota Evidence status
accepted task time 31 min 22 min 22 min synthetic measured input
gross time returned 0 h 699 h/month 699 h/month calculated from scenario inputs
aggregation yield n/a 29% 71% synthetic operating estimate
usable capacity 0 h 139 h/month 370 h/month coefficient output
accepted completions baseline +3% +11% illustrative matched-queue readback
roles or payroll removed none none none explicit boundary

The counterargument: welfare is still value

A strict capacity model can undercount benefits that operational systems do not record. Employees may use returned time to think, recover, coach colleagues or improve work before anybody redesigns the queue. Those effects can matter deeply. The response should be better measurement, not forced monetisation.

Track after-hours work, work intensity, error recovery, absence, engagement and learning time where they fit the intervention. Keep welfare and experience alongside financial measures in the scorecard. Do not hide them inside a wage-rate multiplication. Honest plural value is stronger than fictional cash precision.

Celonis offers a useful first-party example of the distinction between elapsed-time acceleration and final value. It reports cases where agentic tooling produced initial process models in a weekend or under 24 hours, allowing teams to move earlier to optimisation (Celonis account). The time-to-first-model result is an operating milestone. The eventual economic benefit still depends on what optimisation decisions follow and whether their effects are verified.

Part v: the field method

The field method has four ledgers: task, capacity, conversion and outcome. Keep them linked by stable identifiers.

The task ledger records eligible task, actor, start, accepted completion, assistance mode, retries, review and quality state. The capacity ledger aggregates returned minutes by team, role, queue and schedulable block. The conversion ledger records the management decision, owner, effective date and intended value class. The outcome ledger reads back throughput, service, quality, cost, risk or opportunity from the system of record.

The ledgers do not need one large platform. They need stable keys and clear semantics. A task event can remain in the workflow system, a conversion decision in the change register and the financial readback in the finance data product. The evidence layer joins them through case, cohort, intervention and benefit identifiers. Access should follow the sensitivity of each source; a value analyst does not need unrestricted access to employee-level work traces.

Define the grain before implementation. Task-level evidence supports time and quality analysis. Team-week evidence is usually better for capacity and rota decisions. Monthly finance evidence supports benefit recognition. Mixing grains creates false joins, such as attaching one monthly cost figure to every task and then summing it thousands of times.

Version every definition. If accepted completion changes after a policy update, retain the prior definition and the bridge. If the assisted population expands, create a new cohort. If the model or workflow changes materially, open a new intervention version rather than extending the old result. This makes refresh and withdrawal possible.

Privacy is part of the measurement design. Work-pattern data can become employee surveillance if collected without a narrow purpose, aggregation rule and retention limit. Prefer queue and team measures where individual detail is unnecessary. Involve employee representatives and people governance early. A benefits system that damages trust can reduce adoption and invalidate its own estimates.

Finally, assign decision rights. Operations owns the conversion action. Finance owns recognition and double-count prevention. Technology owns task and system telemetry. Risk or quality functions own relevant boundaries. The programme office maintains the evidence chain but should not unilaterally approve its own benefits. Independent challenge is especially useful where an investment decision depends on a small number of optimistic coefficients.

This structure prevents a common failure: a dashboard shows hours saved while finance looks at a different population, operations has changed the process and nobody can reconcile the numbers.

Choose the conversion path before assigning money

Conversion path Management action Finance treatment Guardrail
absorb demand increase accepted throughput or reduce backlog value additional accepted units or service improvement check the next bottleneck and quality
redeploy move capacity to a named higher-value activity value verified output of destination activity confirm the work actually moved
reduce variable spend lower overtime, contractor or transaction cost reconcile actual spend avoided exclude shifted or deferred spend
remove fixed cost change funded roles, facilities or licences recognise after the cost base changes include change and transition cost
invest in quality or resilience increase checking, coaching, recovery or learning report outcome in quality or resilience units do not relabel as cash without agreement
create option value accelerate an experiment or decision record milestone, expiry and follow-on decision retire options that never inform action
ROI sensitivity to capacity captureFive curves show illustrative net annual value at different capacity-capture coefficients and management conversion choices, with a break-even band. κ drives more variance than the demo timer break-even band throughput valueredeployment valuewelfare onlyno conversion action capacity-capture coefficient κ →net annual value →
Figure 7. ROI sensitivity to κ. The value case is highly sensitive to the capture mechanism and selected conversion path. Status: illustrative sensitivity plot.

A runnable capacity-capture calculator

The following Python artefact uses synthetic inputs. It calculates gross hours, usable hours and evidence-appropriate value. The positive case absorbs demand into accepted completions. The negative case leaves minutes fragmented and claims no cash benefit.

Depth: run the calculator and inspect both cases
from dataclasses import dataclass

@dataclass(frozen=True)
class CaptureCase:
    people: int
    tasks_per_day: float
    minutes_saved: float
    workdays: int
    adoption: float
    sustained_use: float
    aggregation: float
    demand: float
    quality_retention: float
    managed_conversion: float
    minutes_per_accepted_unit: float
    value_per_accepted_unit: float
    annual_run_cost: float

def calculate(case: CaptureCase) -> dict:
    gross_minutes = (
        case.people * case.tasks_per_day * case.minutes_saved *
        case.workdays * case.adoption * case.sustained_use
    )
    kappa = (
        case.aggregation * case.demand *
        case.quality_retention * case.managed_conversion
    )
    usable_hours = gross_minutes * kappa / 60
    extra_units = gross_minutes * kappa / case.minutes_per_accepted_unit
    gross_value = extra_units * case.value_per_accepted_unit
    net_value = gross_value - case.annual_run_cost
    return {
        "gross_hours": round(gross_minutes / 60, 1),
        "kappa": round(kappa, 3),
        "usable_hours": round(usable_hours, 1),
        "verified_extra_units": round(extra_units, 1),
        "net_value": round(net_value, 2),
    }

positive = CaptureCase(
    people=48, tasks_per_day=6.2, minutes_saved=9, workdays=220,
    adoption=.78, sustained_use=.93, aggregation=.71, demand=.92,
    quality_retention=.95, managed_conversion=.86,
    minutes_per_accepted_unit=31, value_per_accepted_unit=18,
    annual_run_cost=86_000,
)

negative = CaptureCase(
    people=3000, tasks_per_day=1, minutes_saved=10, workdays=220,
    adoption=.70, sustained_use=.80, aggregation=.12, demand=.35,
    quality_retention=.98, managed_conversion=0,
    minutes_per_accepted_unit=30, value_per_accepted_unit=20,
    annual_run_cost=420_000,
)

print("positive", calculate(positive))
print("negative", calculate(negative))

Expected behaviour: the positive case produces usable capacity and outcome value because all conversion terms are non-zero. The negative case can report many gross hours, yet κ is zero because no management conversion exists. Its monetary benefit remains zero while run cost remains visible.

The 30–60–90 day sequence

During the first 30 days, select one constrained workflow and record the baseline. Define accepted completion, quality boundaries, demand signal and benefit owner. Instrument task and outcome identifiers before asking for a savings estimate.

By day 60, run a controlled comparison. Estimate each κ component. Look for the dominant leak. Test one operating intervention, such as queue pooling, protected blocks or review batching. Keep a negative control or phased cohort where practical.

By day 90, reconcile the operating readback. Decide whether to scale, redesign, narrow or stop. Finance accepts a benefit class only after the owner provides its evidence. Update the business case with observed run cost, review work and change cost.

Thirty, sixty and ninety day implementation sequenceThree large stages progress from baseline and instrumentation to controlled comparison and then to finance reconciliation and a scale, redesign, narrow or stop decision. Ninety days to an evidence-backed decision day 0–30baseline accepted workname the outcome ownerlink task to readbackdefine quality boundary day 31–60controlled comparisonestimate κ componentsredesign one work ruleretain negative control day 61–90reconcile outcometest full economicsaccept benefit classscale · redesign · stop
Figure 8. Ninety-day sequence. Measurement and work redesign precede benefit recognition. Status: practitioner operating method.

Release and review controls

The benefits case should carry an explicit expiry. Adoption, task mix, model behaviour, demand and labour cost change. Recalculate κ after a material workflow, model, prompt, policy or staffing change. Review core operating measures monthly and financial assumptions quarterly.

Treat uncertainty as part of the decision, not a footnote. Publish a value interval and show which assumption drives it. A board can accept a wide range when the investment is reversible and the learning value is high. It should demand stronger evidence before a large, irreversible scale commitment. The model becomes more useful when it reveals what must be learned next.

Link funding stages to that learning. The first tranche buys measurement and a controlled intervention. The second buys work redesign where κ is constrained. The third buys scale only after operating readback survives the comparison. This sequence protects useful experiments while preventing pilot estimates from hardening into permanent benefits.

A benefit should leave the ledger if the conversion action reverses, quality falls outside bounds or demand disappears. Keep gross time as a diagnostic, not as a permanent booked value. A value ledger earns trust by withdrawing claims as readily as it records them.

The final control is portfolio comparison. A project with modest time saving and high capture can outrank a project with spectacular task acceleration and no route to an outcome. That changes investment selection. It also changes the conversation with vendors: ask for the evidence chain, not the largest reported percentage.

The decision is a work-design decision

AI can make work faster. The evidence increasingly supports that proposition for particular tasks and, with more caveats, at firm level. The board and finance question begins one step later: what changed because the time returned?

The answer is rarely found in model telemetry alone. It sits in the relationship between task events, calendars, queues, demand, quality, authority and financial ownership. A credible business case therefore treats work redesign as part of the AI investment, not as an adoption activity that happens after the technology ships.

This framing also improves the human conversation. Employees no longer hear that every minute returned is a hidden redundancy target. Managers can state the intended conversion path before rollout. Finance can distinguish welfare, capacity and cash. Technology teams can optimise for accepted work and workflow movement rather than activity. The result is a value case that is both more demanding and more believable.

Do not book the minute. Book the verified change that the minute made possible. For one workflow, measure gross time honestly, estimate κ from observed frictions, execute a named conversion action and read the result back from the operating system. Scale when that chain holds. Narrow or stop when it does not.

That is a stricter test than an hours-saved dashboard. It is also the test that lets a serious AI programme distinguish employee relief, usable capacity, operating improvement and cash without diminishing any of them.