← Back to Blog
AI agent for expense review: policy checks and receipts
AI Agents

AI agent for expense review: policy checks and receipts

byBruno Galo · Published on 22 Mar 2026

Last updated 12 Aug 2026

Available inCatalàEnglishEspañolPortuguês

Expense approval in most mid-market companies is a manager clicking approve on a list of amounts. The policy exists, is several pages long, and is not being checked — not through negligence, but because checking it properly would mean reading each receipt, comparing it to a limit that depends on grade and destination, verifying it is not a duplicate of a claim submitted six weeks ago, and confirming the VAT treatment. Nobody has that time for a claim of forty euros, so the control is nominal.

This produces a specific pattern: high compliance with the process, low compliance with the policy. Everything is approved, submitted on time, correctly coded — and the actual rules go largely unenforced. When something is eventually found, it has usually been happening for a long time, which makes it awkward for everyone.

An agent changes the economics of this. Reading every receipt, checking every claim against a policy that varies by grade and destination, and detecting duplicates across months is exactly the work that is trivial in principle and impossible in practice at human cost. What is left over — the genuinely judgemental cases — is small enough that a manager can actually consider it.

Why this matters

The direct financial exposure is usually modest and worth quantifying honestly. Expense leakage in a mid-market company is real but rarely material against revenue. If the case for this rests on recovering leakage, it is a weak case.

The stronger arguments are elsewhere. The first is VAT recovery: reclaimable VAT on employee expenses is frequently under-recovered because it requires a valid receipt with the supplier's tax identifier and the correct treatment applied per category, and manual processes do not reliably capture that. In Spain and Portugal, where documentation requirements are specific, the recoverable amount left on the table is often larger than the leakage the process was designed to prevent.

The second is control quality. An expense process that is nominally controlled and actually unchecked is a control weakness, and it is the kind auditors notice. Being able to demonstrate that every claim was checked against policy — with an audit trail — converts a soft control into a real one.

The third is manager time. Expense approval consumes a small amount of many senior people's attention, in fragments, on decisions they are not really making. Removing it is worth more than the leakage.

At a glance: what the agent checks and what it must not decide

Check Agent authority Note
Receipt present and legible Full Reject and return to claimant with a reason
Amount matches receipt Full Discrepancies routed, not silently corrected
Within category limit for grade and destination Full Rule-based, however many permutations
Duplicate of an earlier claim Full — flag, never auto-reject Duplicates are often innocent; accusation is not the agent's role
Valid tax documentation and correct VAT treatment Full for validation, propose for treatment Determines recoverability; jurisdiction-specific
Correct cost centre and account coding Full, with proposal Learns from historical corrections
Missing required detail — attendees, business purpose Full — request from claimant Most common cause of an invalid claim
Policy exception requested by claimant None — route to approver The exception process is the point at which judgement is required
Pattern suggesting deliberate misuse None — route confidentially to a named owner An accusation, however soft, is never an automated output
Claims by senior staff or the approver's own chain None — route by a defined independent path Self-approval risk needs structural handling

The last two rows deserve care. An agent that surfaces a pattern in an individual's claiming behaviour is producing something with employment consequences. That output must go to one named person by a defined route, not into a general queue, and the policy for handling it should be written before the agent is switched on.

What works, and what to be honest about

What works:

Checking at submission, not at approval. The agent validates while the claimant is still present and can fix the problem. A claim returned two weeks later for a missing business purpose is a cost to everyone; one flagged at submission is a small correction.

Approval by exception. Compliant claims are approved automatically; the manager sees only what failed a check or requires a judgement. This is what makes the remaining review genuine — a manager looking at three flagged items reads them properly.

VAT validation as a first-class function. Checking that the receipt carries a valid supplier tax identifier and applying the correct treatment per category is where the measurable financial return usually sits. It is also the check least likely to be performed manually.

Duplicate detection across a long window. Duplicates are usually accidental — a claim resubmitted after an approval delay, a receipt claimed by two attendees. Detecting them requires comparing against months of history, which no human reviewer does.

Coding proposals that learn. Expense coding is repetitive and mostly predictable from the merchant and the claimant's role. An agent proposing the code, corrected occasionally, removes a persistent low-grade irritation.

What to be honest about:

This is employee monitoring, and it has a legal dimension. Automated assessment of individual behaviour engages data protection obligations, and in Spain and Portugal there are specific employment-law considerations around monitoring and works council consultation. Get this reviewed before deployment rather than after — it is not usually an obstacle, but discovering it late is a poor position.

The financial case is often weaker than the control case. Be honest in the business case. Overstating recoverable leakage produces a deployment that is judged against a number it will not reach, when the real benefits — VAT recovery, demonstrable control, manager time — are perfectly defensible.

Policy will need rewriting before it can be encoded. Most expense policies contain ambiguity that human approvers resolve silently. An agent cannot, and it will surface every ambiguity at once. That rewrite is real work and it improves the policy regardless.

A strict agent generates resentment quickly. An agent rejecting claims on technicalities, in an organisation where the policy was previously unenforced, feels like a change in the deal. Phase enforcement, communicate before switching on, and start with flagging rather than rejection.

Receipt reading is good, not perfect. Poor photographs, thermal receipts, handwritten additions and foreign-language documents all reduce accuracy. Design for a proportion requiring human reading rather than assuming full automation.

Decision framework: where to start

Run in order. Stop at the first match.

1. Is your expense policy unambiguous enough to be encoded?
If not, rewrite it first — limits by grade and destination, what documentation is required, how exceptions are requested and who may grant them. This is the actual project and it is not technical.

2. Have you taken data protection and employment advice on automated review?
Do this before building. Automated assessment of employee claims engages obligations that vary by jurisdiction, and in Iberia there are consultation considerations worth confirming.

3. Is your current approval a genuine control or a formality?
If claims are approved without checking, say so internally before deploying. The agent will produce a visible change in rejection rates, and an organisation that has not acknowledged the prior state will experience that as the agent being unreasonable.

4. Are you capturing valid tax documentation and applying VAT treatment per category?
If not, this is where the measurable return is. Prioritise it over policy enforcement.

5. Do you have a defined confidential route for behavioural patterns?
Define it before go-live: one named owner, a documented handling process, no such output in any general queue.

6. All of the above in place?
Deploy in flagging mode only for the first two months — the agent checks and informs but does not reject. Publish what it would have rejected, adjust the policy where it turns out to be unreasonable, then enable enforcement.

7. Running well and claim volume still consuming time?
The remaining cost is likely submission friction rather than review. Look at how claims are captured — mobile capture at the point of spend removes more total effort than any review automation.

Indicative cost and effort

Workstream Typical elapsed time Effort profile
Policy review and rewrite for encodability 3–5 weeks Light effort, needs decisions
Data protection and employment review 2–4 weeks Light — external advice
Receipt capture and reading integration 4–7 weeks Medium
Policy rule configuration 3–5 weeks Medium — permutation-heavy
VAT validation and treatment logic 3–5 weeks Medium — jurisdiction-specific
Duplicate detection across history 2–4 weeks Light to medium
Exception routing and confidential escalation path 2–3 weeks Light, must be deliberate
Flagging-mode pilot 8 weeks Light — monitoring

Assumes one entity or a small group with a single expense policy. Multiple entities with differing policies and jurisdictions extend this. Get a quote for a scoped estimate.

Frequently asked questions

Can the agent reject claims outright?
For objective failures — no receipt, amount mismatch, missing mandatory detail — yes, with a clear reason and an easy resubmission path. For anything involving judgement about whether an expense was reasonable, no. That distinction should be written into policy rather than left to configuration.

What about duplicate claims that are innocent?
Most are. The agent flags and asks; it does not accuse. The message a claimant receives matters here, and it should be drafted by someone who understands that the majority of flags are honest mistakes.

How do we handle claims from senior staff?
By a defined independent route, decided in advance. Self-approval and approval within one's own reporting chain are the classic weaknesses in expense control, and an agent enforcing them consistently is one of its more valuable properties — provided the route exists.

Will this pay for itself?
Usually, but check where you expect the return. If it depends on recovering leakage, model that conservatively. If it includes VAT recovery, demonstrable control and reclaimed manager time, the case is generally comfortable.

How do employees react?
Better than expected when the change is communicated honestly and phased, worse than expected when an unenforced policy is suddenly enforced without notice. The communication is more determinative of success than the configuration.

How quickly can this be connected to our real expense and ERP data?
The connection itself is usually quick — a certified partner like Stacksync can have real-time sync between the expense system and the ERP running within weeks. The policy checks described above take longer to get right, because they require someone to actually decide the thresholds and exceptions the agent will enforce.

Closing — Next steps

Expense review is a control that most organisations perform in form rather than substance, and everyone involved knows it. Automating it makes the control real, which is valuable, and also visible, which requires managing — because the first honest report will show that the policy was not being followed, and that finding is about the prior process rather than the people.

A concrete first step: take last quarter's claims and check a sample of thirty properly against policy — receipt validity, limits, duplicates, tax documentation. The failure rate in that sample is your baseline, and it is usually the most persuasive item in the business case.

About the author

Bruno Galo is the founder of Atypical Tech, a NetSuite consultancy serving mid-market clients across Iberia. He specializes in connecting CRM and ERP systems for seamless order-to-cash workflows, building automated order management pipelines that eliminate manual data entry between sales and finance teams. As an official Stacksync implementation partner, Bruno designs and deploys AI agents on integration platforms to handle exception routing, document processing, and reconciliation — turning fragmented order flows into reliable, self-monitoring systems.

LinkedIn: https://www.linkedin.com/in/brunogd

Sources

URLs are publisher-level and should be verified before publication. Employment and data protection aspects are jurisdiction-specific — take local advice.

Comments

No comments yet.

Leave a comment

Your comment will be reviewed before publishing.

An unhandled error has occurred. Reload 🗙