HR automation / Business case
HR chatbot ROI: fewer tickets can mean employees gave up
- Fictional pilot
- 60 of 100 queries avoid immediate handoff; only 35 are verified resolved without HR.
- Full cost
- Count answer upkeep, evaluation, integration and rework alongside the subscription.
- Stop expansion
- Materially wrong policy answers or obstructed escalation fail the pilot.
In this guide
Start by asking what the percentage counts
Does “resolved” mean the employee confirmed the outcome, the conversation ended, an article was selected, or no live-agent transfer occurred? Does the denominator count messages, queries, conversations or unique employee problems? Several numbers can be honestly calculated from the same pilot and still describe different things.
ServiceNow's Australia-release documentation, for example, describes conversation deflection using resolution statuses on query responses. That is a defined product metric. It should not silently become an estimate of all HR work removed, particularly if one employee makes several queries or later contacts a person.
Keep the supplier's metric, with its definition, on the report. Add a buyer outcome measure alongside it. We recommend a verified resolution count for a bounded topic and follow-up window, plus repeat contact, abandonment and unknown outcomes. Do not relabel unknown as success or failure to make the chart cleaner.
The 100-query example: 60% contained, 35% verified
All figures in this example are invented. Assume 100 distinct initial routine questions, one per employee, so the denominator is simple. The pilot team follows the same topic across approved channels for seven days. Seven days is a proposed measurement window, not a universal standard.
| Outcome | Count | How to use it |
|---|---|---|
| Immediate human handoff | 40 | Measure any triage benefit; no full resolution saving assumed. |
| No immediate handoff; later human contact on the same problem | 15 | Count the actual combined bot and human work, including repetition. |
| No handoff; outcome remains unconfirmed | 10 | Keep as unknown. Do not book a full avoided-case saving. |
| Resolved without HR, confirmed within the chosen window | 35 | Candidate for measured capacity released on this topic. |
| Total | 100 | All initial questions accounted for. |
Sixty avoided an immediate handoff. Thirty-five have verified outcomes without HR. Neither number establishes the correct result for next month's different mix of questions. Nor does the example imply any supplier delivers these results.
For sensitive subjects, avoid collecting unnecessary detail just to join analytics records. Agree a privacy-conscious measurement approach with the responsible team. Where follow-up cannot be safely or reliably linked, show the evidence gap rather than manufacturing a resolution rate.
Turn verified time into a useful budget comparison
Suppose the 35 verified cases each replace an observed average of eight minutes of routine handling. The provisional capacity released is 280 minutes, or 4 hours 40 minutes. If this pilot also needs 100 minutes of answer maintenance and quality review, net released capacity is three hours. At an explicitly assumed loaded rate of $50 an hour, that capacity is valued at $150.
Now make a deliberately simple monthly planning scenario: the same volume and mix repeats, and the supplier charges $300 a month for the scoped service. The $150 monthly capacity valuation does not cover that subscription, even before setup or unmeasured rework. Keep unconfirmed outcomes out of the savings calculation.
Quicker after-hours answers or employee convenience may still justify the cost. Record those benefits separately from cash savings.
Do not add “time saved,” “ticket cost avoided” and “headcount capacity freed” as three independent benefits when they value the same minutes. Keep cash reductions, redeployable capacity and employee experience in separate lines, with evidence for each.
Your automation may not apply to the channel people use
Read the operating documentation before multiplying a feature by the entire case volume. Workday's Case Agent instructions distinguish classification, assignment and summarisation. The documented assignment skill applies to specified new-case entry routes, excluding email, business-process and Create Case Advanced cases. It also depends on Workday data for leave-aware assignment.
If most of your requests arrive by email, applying that assignment benefit to every case overstates the eligible workload. Those email cases are outside the documented assignment scope. Count eligible cases before estimating the minutes saved.
A summary-writing tool needs a different test from a self-service answer bot. Time how long an authorised agent takes to read, correct and use the summary, compared with the existing task. Do not claim a whole case was automated because one step became shorter.
Test wrong policies, missing context and escalation
Start with a narrow topic whose approved answers have clear owners and effective dates. Give the pilot the awkward variants: a US and UK worker asking the same question, an outdated policy with a newer replacement, a question with insufficient context and a request that needs a person. Define the expected response before reviewing the output.
Record answer correctness, relevant source, applicable employee population, successful human handoff, repeat contact and maintenance effort. For a personalised answer, test permission boundaries using authorised fictional accounts. Stop the pilot if an account can retrieve another worker’s restricted information.
We would stop expansion if the assistant invents a material entitlement, repeatedly selects the wrong local policy or obstructs escalation. We would keep a narrower successful use case if its benefits survive the complete cost model. Rollout does not need to be all-or-nothing.
Buy an outcome you can keep measuring
Request the definitions behind usage charges, any minimum commitment, the relevant employee and channel coverage, model or feature allowances, integration work and what happens when usage exceeds the estimate. Ask who maintains approved content and how changes reach the assistant. Include routine evaluation and incident handling in the operating budget.
Keep the pilot's measurement sheet after signature. A product can improve while the knowledge base decays, or adoption can rise while the query mix becomes harder. A renewal decision should use current verified outcomes and cost, not the launch-day deflection slide.
Expand to another topic only after measuring the current pilot’s correct resolutions, repeat contacts, maintenance time and total cost.
Questions buyers ask
Is HR chatbot deflection the same as resolution?
Not necessarily. Deflection is a product-specific metric. Inspect the event and denominator, then separately track verified outcomes, later human contact and unknown results.
Can time saved be reported as cash saved?
Only when spending actually falls. Time released from existing staff is usually capacity that needs a useful destination; value it separately from a reduction in cash expense.
What costs belong in the chatbot business case?
Subscription, implementation, connections, usage charges, content upkeep, answer evaluation, incident handling and rework. Do not count the same saved minutes in several benefit categories.
How big should the first pilot be?
Large enough to cover the actual variants in one bounded topic and observe the chosen follow-up window. The 100-query example is a fictional teaching case, not a statistically validated sample-size recommendation.
Sources and research scope
ServiceNow metric documentation and Workday Case Agent operating instructions inspected 2 October 2026. All pilot figures and USD cost inputs are fictional. No model, chatbot or customer deployment was tested.
- ServiceNow conversation deflection definitionAustralia-release query-response resolution metric and example.
- Workday Case Agent skills and reportsSkill-specific entry channels, assignment inputs and reporting.