The headline rate is the least interesting number
Outcome pricing for support agents converged fast in 2026. HubSpot moved its Customer Agent from $1.00 per conversation to $0.50 per resolved conversation on 14 April 2026, halving the rate and narrowing the billable base at the same time. The pitch is straightforward: you stop paying for spam, abandoned sessions, and conversations the agent fumbled to a human.
That is a real improvement over per-attempt metering. It is also, structurally, a shift of definitional power to the vendor. Under per-conversation pricing you could count your own conversations and predict the bill exactly. Under outcome pricing, the count you need is one the vendor computes, on their clock, using a predicate they can restate at renewal. The rate is public and stable; the predicate is where the money moves.
Credits are a currency layer, not a discount
Most of these systems do not bill you dollars directly. HubSpot meters through “HubSpot Credits” at $0.010 each, sold in packs of 1,000 for $10. A resolution costs 50 credits. The arithmetic is deliberately one step removed: 50 × $0.010 = $0.50.
The currency layer does two things. It lets the vendor reprice per-action consumption without touching the exchange rate — the April change was 100 credits to 50, not a change in what a credit is worth. And it pools consumption: the same balance funds enrichment, data answers, prospecting and workflow AI actions. A support-agent budget is therefore not isolated from the rest of the account, and when the pool empties the agent stops accepting conversations across every connected channel until the reset date.
The resolution predicate is a disjunction
This is the part most write-ups get wrong, including by describing it as a checklist of conditions that must all hold. It is not a checklist. There are two independent ways to bill:
(A) the agent posted at least one reply that shares a content source or performs an action, and there is no handoff to a human within 72 hours of the last agent response
— or —
(B) a lead is marked qualified, partially qualified, or not qualified — evaluated immediately; the 72-hour window does not apply
Read what each clause does. Branch (A) sets a low bar — citing a knowledge-base article counts whether or not it answered the question. And the only escape hatch in (A) is an actual human handoff inside the window.
Note what is absent: negative feedback. A thumbs-down does not prevent a resolution, and the docs are explicit that later messages, transfer requests, or negative feedback will not change the status once it has been set. If you assumed an unhappy customer meant an unbilled conversation, that assumption is worth discarding now.
Which gives the structural asymmetry worth internalising: silence resolves. A visitor who reads a wrong answer, sighs and closes the tab has satisfied branch (A) in full. The predicate cannot distinguish a satisfied customer from an abandoned one, because both look identical in the event stream. Any outcome-priced agent you evaluate will have this property in some form — check how the vendor handles the null case before you check anything else.
Why the status freezes at 72 hours
Under branch (A), resolution is evaluated 72 hours after the agent’s final response, and the status is then immutable.
The window cuts both ways, and it is worth seeing both edges. A conversation the agent handled almost entirely but escalated at hour 70 costs nothing — genuinely generous. A conversation the agent answered confidently and wrongly, which a human picks up on day four after the customer emails again, bills $0.50and stays billed. If your team’s follow-up cadence is slower than three days, you are systematically on the wrong side of the freeze.
Branch (B) has no such window at all — a qualification verdict settles immediately.
The second trigger nobody models
Deflection is not the only billable event. Marking a lead qualified, partially qualified, or not qualified also counts — including the negative case. The predicate is about the agent completing its job, not about you gaining a customer.
This is the clause that decides whether an agent is cheap or expensive, and it turns on deployment context rather than volume. A deflection bot sitting on a documentation site fires the predicate maybe half the time — plenty of visitors ask for a human or bounce before the agent cites anything. A pre-sale qualification bot on a marketing site fires it on nearly every real conversation, because qualification is the job and every outcome of that job, including “not a fit,” is a billable event.
The same product, the same rate card, two completely different unit economics.
The same traffic, three bills
Take 1,000 chat conversations a month on a Professional-tier plan, which includes 3,000 credits — enough for exactly 60 resolutions before overage. Run the same traffic through three plausible predicates.
A conservative deflection deployment at 55%: 550 × 50 = 27,500 credits, minus 3,000 included, leaves 24,500 × $0.010 = $245. At 65%: 650 × 50 = 32,500 − 3,000 = 29,500 → $295. With qualification enabled and 95% of conversations reaching a verdict: 950 × 50 = 47,500 − 3,000 = 44,500 → $445.
Identical traffic, identical rate card, an 82% spread. Triple the volume to 3,000 conversations and the same three configurations land at roughly $795, $945 and $1,395 a month. Outcome pricing has not removed the variance; it has relocated it from a number you control to one you do not.
A published resolution rate is a benchmark, not a forecast
Vendors publish aggregate resolution rates, and individual customer testimonials range widely — HubSpot’s own product page quotes deployments at 70–80% and 75%. Competing vendors tend to publish lower figures for the same products. Both directions carry an obvious interest, so weight them accordingly.
Neither forecasts your bill, because a resolution rate is a ratio over a population and your population is not theirs. Any published aggregate is dominated by mature support deployments with real knowledge bases answering repetitive, factual questions. If your content is marketing pages rather than FAQs, or your questions are pre-purchase and consultative rather than “where is my order,” you are sampling a different distribution entirely — and the direction of the error is not obvious, since a thin knowledge base lowers useful resolutions while the citation clause keeps billing them anyway.
Included allowances round to zero
Bundled credits look generous on a feature grid and evaporate under real traffic: 500 on Starter, 3,000 on Professional, 5,000 on Enterprise. At 50 credits a resolution that is 10, 60 and 100 resolutions per month respectively — two a day at the mid tier, shared with every other AI action in the account.
Treat the included pool as a trial allowance, not a floor, and model your cost as if it were zero. Testing consumes it too: 50 test conversations during configuration is 2,500 credits, most of a Professional month gone before the agent sees a customer.
Monthly expiry punishes seasonal demand
Credits refresh monthly and unused balance expires with no rollover. For steady inbound this is a minor annoyance. For seasonal businesses it is a structural mismatch: you forfeit the pool in the trough and pay overage through the peak, so the annualised effective rate is strictly worse than the flat-demand case every pricing calculator implicitly assumes.
The mitigation is boring and worth doing anyway — set an account-level spend cap and disable automatic credit-tier upgrades, which several platforms enable by default. Automatic upgrade is the mechanism by which a $295 month becomes a $900 month without anyone approving it.
What the markup buys over the token floor
Price the same interaction against raw inference. Assume a six-turn retrieval-grounded chat: history replay plus retrieved chunks puts cumulative input around 40,000 tokens and output around 1,500. At commodity mid-tier rates of roughly $1 per million input and $5 per million output, that is $0.040 + $0.0075 ≈ $0.048. (Those turn and token assumptions are ours, stated so you can vary them — not measured from a vendor deployment.)
So $0.50 per resolution sits at roughly ten times the token floor, and the multiple widens as model prices fall. That gap is not a scam — it buys hosted retrieval, CRM state, omnichannel plumbing, handoff routing, analytics, and no engineering headcount. The honest framing is that you are renting integration, not inference, and the integration is the expensive part.
Which means the build-versus-buy line is drawn by volume against engineering cost, not by the per-unit price looking high. At 300 resolutions a month the markup is a rounding error against one engineer-week. At 3,000 it is a five-figure annual line item and the arithmetic changes.
Auditing an outcome-priced agent before you sign
Four questions get you most of the way. What is the exact predicate, in the vendor’s own documentation rather than their pricing page? What happens in the null case — does an abandoned conversation bill? Is there a settlement window, and can the status reverse after it closes? And which non-support events, qualification especially, also trip the meter?
Then instrument before you commit. Most of these products ship a trial period; the number worth extracting from it is not satisfaction but your predicate hit rate on yourtraffic. Everything else in the model is arithmetic once you have that ratio, and nothing in the vendor’s published rate can substitute for it.
Takeaway
Outcome pricing is a genuine improvement on per-attempt metering, and it deserves credit for aligning spend with delivered work. But “you only pay when it works” is a claim about a boolean, and you did not write the boolean. Before you compare $0.50 against a competitor’s $0.99, find out what each one is counting — the predicate spreads costs further than the rate does.
The worked example behind this piece: EyesInAI vs HubSpot Customer Agent →