You will recognise some of this
- Someone spends their morning reading documents and typing what they say into a system
- Supplier statements are reconciled by eye, line by line
- The support inbox is triaged manually before anyone can answer it
- A manager asks the same question every week and someone rebuilds the answer
- You have tried a chatbot and it confidently gave a wrong answer to a customer
- You suspect AI could help but cannot tell which parts are real and which are demos
Where does AI actually pay for itself?
In the places where a human is currently acting as a translator between two systems: reading an invoice and typing it into a ledger, reconciling a supplier statement, triaging a support inbox, turning a week of transactions into a summary a manager will read. These are high-volume, low-judgement tasks. That is exactly where the technology is strong today.
Where does it not?
Anywhere a wrong answer is expensive and hard to detect. We will not put a model in front of a pricing decision, a clinical judgement or a compliance filing without a human check in the loop, and we will say so during scoping rather than after go-live. If AI is the wrong tool for what you are describing, that answer is free.
How do you keep it from making things up?
By constraining what it can do. Retrieval grounded in your own data, structured outputs validated against a schema, confidence thresholds that route uncertain cases to a person, and logging on every decision so you can audit what happened and why.
How it runs
- 1
Automation audit
One to two weeks
We measure where the hours actually go, not where they are assumed to go, and rank candidates by hours saved against effort and risk. Some tasks come back marked "do not automate" with the reason attached.
- 2
Pilot on real data
Two to four weeks
One task, your data, running alongside the human doing it. We compare outputs and measure the disagreement rate — that number, not a demo, decides whether it ships.
- 3
Ship with a human in the loop
From day one
Confidence thresholds route uncertain cases to a person, every decision is logged, and cost per run is monitored so spend never arrives as a surprise at the end of a month.
- 4
Tune and widen
Ongoing
The disagreement log is the backlog. We tighten the cases it gets wrong before we extend it to the next task.
When not to hire us for this
- Putting a model in front of a pricing decision, a clinical judgement or a compliance filing without review
- Replacing a person on day one — pilots run alongside the human until the numbers earn the change
- A general-purpose chatbot on your website because a competitor has one
- Anything where the running cost approaches the labour it replaces; we will show you that maths first
What this rests on
We build AI into our own product rather than only for clients — including a metered token wallet, because we pay the running costs ourselves and had to make them visible.
See the platform it runs inQuestions
- Does our data get used to train a model?
- No. We use commercial APIs under agreements that exclude training on your data, and we tell you exactly which provider processes what, and where. If data residency matters for your market, raise it during scoping — it changes the architecture, not the outcome.
- What does it cost to run?
- It depends on volume, and we measure it per run from the pilot onward so you have a real number before you commit. Any automation whose running cost approaches the labour it replaces is one we will recommend against.