Skip to main content
A fraud-risk scoring framework for expense operations

A fraud-risk scoring framework for expense operations

How to score, prioritize, and sweep transaction-level risk without building a fraud department you can't afford

Most SMB expense fraud programs die the same way. Someone finds a bad transaction — a duplicate reimbursement, a card used at a vendor nobody recognizes, a mileage claim that doesn't match the calendar — and the finance team reacts hard. New rules get written. Approvals pile up. Everyone's inbox fills with receipts. Six weeks later the rules are quietly ignored because they slowed everyone down and caught mostly nothing.

The problem isn't that small businesses lack controls. It's that they treat every transaction as equally risky, which means they either check everything (impossible at scale) or check randomly (useless). What's missing is a way to rank risk so the limited hours your team has go toward the transactions most likely to actually be a problem.

That's what a scoring framework does. It turns "we should probably review expenses more carefully" into "these 14 transactions this week deserve a human, and here's why." This post walks through how to build that scoring model, how to prioritize which controls to invest in, and how to run scheduled sweeps with real deadlines attached — all sized for a team that might be one bookkeeper and a part-time controller.

Why "review more" always fails

Ask a small finance team how they catch fraud and you'll usually hear some version of "we look at anything that seems off." That's not a system. It depends entirely on who's looking, how tired they are, and whether they happen to recognize a vendor name.

The operational reality: a 12-person company might run 400–600 expense transactions a month across cards, reimbursements, and vendor invoices. If your controller spends even three minutes on each one, that's 20–30 hours a month gone — and they still won't catch the clever stuff, because the clever stuff is designed to look normal. Meanwhile the obvious problems (a $9 duplicate, a rounded-up mileage claim) eat attention that should go elsewhere.

The deeper issue is that risk isn't evenly distributed. In real operations, a large share of fraud and error exposure sits in a small slice of transactions: new vendors, manual reimbursements, transactions just under approval thresholds, and anything touching an employee who both requests and approves. If you can score transactions on those dimensions, you can stop reviewing 500 things and start reviewing 30 — the right 30.

This is the same logic behind separating detection from investigation. If you haven't already, the SMB detection-to-investigation playbook covers the response side; this article is the layer that sits in front of it, deciding what's worth investigating in the first place.

What a transaction-level risk score actually looks like

Forget anything fancy. A workable score for a small business is a weighted sum of a handful of signals that each nudge a transaction up or down the risk ladder. The point isn't precision — it's consistency. The same transaction should score the same whether your controller reviews it Monday morning or Friday afternoon.

Here's a starting model you can adapt. Each transaction gets points across a few dimensions, and the total decides how it's handled.

Risk signalWhy it mattersPoints
New vendor (first payment ever)Highest fraud vector for SMBs; fake vendors start here+25
Amount just under an approval thresholdClassic structuring behavior+20
Requester = approver (self-approval)Broken segregation of duties+30
Round-number amount ($500.00, $1,000.00)Correlates with fabricated claims+10
Weekend / holiday transaction dateNot automatically bad, but worth weighting+5
No receipt or receipt uploaded lateDocumentation gap+15
Vendor category mismatch (office supplies at a restaurant)Miscoding or misuse+15
Repeat vendor, consistent pattern, on-time receiptTrusted behavior−20

A transaction scoring 0–20 flows through untouched. 21–45 gets a quick documentation check. Anything 46 and up goes into a manual review queue with an actual person assigned. Tune the thresholds after a month of watching where real problems land — every business has a slightly different distribution.

The mistake people make here is over-engineering the model on day one. You don't need 40 signals. You need 6–8 that you'll actually maintain. A model nobody updates is worse than a simple one, because it slowly stops reflecting how your business actually spends.

The control prioritization matrix

Scoring tells you which transactions to look at. It doesn't tell you which controls to build first — and for a resource-constrained team, that's the harder question. You can't implement everything, so you rank controls the same way you rank transactions: by impact versus effort.

  1. High impact, low effort — do these now. Example

    blocking self-approval in your expense tool. Usually a settings change, catches a whole category of risk.

  2. High impact, high effort — plan these. Example

    a full new-vendor verification workflow with callback confirmation. Worth it, but it needs process design.

  3. Low impact, low effort — do them when convenient. Example

    flagging round-number claims for a light check.

  4. Low impact, high effort — skip or defer. Example

    manually reconciling every sub-$25 transaction. The math doesn't work.

Small teams instinctively start in the wrong quadrant — usually low-impact, high-effort busywork like re-checking tiny receipts, because it feels thorough. Meanwhile the high-impact/low-effort fixes (turning off self-approval, requiring receipts before reimbursement) sit undone for months.

If this control had existed last quarter, how many real problems would it have caught, and how many hours would it have cost?

Where the score and the matrix connect: your controls should target the signals driving the most score. If "new vendor" is your heaviest weight, then vendor onboarding controls jump to the top of the matrix. The two tools reinforce each other instead of running as separate initiatives.

Scheduled risk sweeps: making review a routine, not a fire drill

Scoring transactions in real time catches things as they happen. But some patterns only show up over time — a vendor that gets paid a little more each month, an employee whose claims creep upward, duplicate payments that land two weeks apart. That's what scheduled sweeps are for. Recurring, time-boxed reviews of the higher-risk slice, run on a calendar so they never depend on someone "getting to it."

A sweep cadence that works for most small teams without burying anyone:

  1. Weekly (15–20 min)

    Pull every transaction that scored 46+ during the week. Confirm each was either legitimately reviewed or resolved. This is your fast-moving safety net.

  2. Monthly (1–2 hrs)

    Run three specific checks — new vendors added this month, top 10 largest transactions, and any employee whose total claims rose more than roughly 30% versus their trailing average. Trend stuff, not single-transaction stuff.

  3. Quarterly (half day)

    Test the model itself. Did high-scoring transactions actually turn out to be problems more often than low-scoring ones? If not, your weights are wrong. Also review any control that's been triggering constantly with no findings — that's alert fatigue building.

The quarterly self-test is the step almost everyone skips, and it's the one that keeps the whole framework honest. A scoring model that hasn't been validated against real outcomes is just a set of assumptions with numbers attached. Pull last quarter's flagged transactions and ask what actually happened to them. You'll usually find one or two signals doing all the work and a couple contributing nothing but noise. Cut the noise.

Here's a quick visual of the sweep cadence and how flagged transactions flow into weekly, monthly, and quarterly reviews.

Process diagram

A sweep cadence that runs on the calendar makes reviews reliable instead of opportunistic, and ensures patterns — not just one-offs — get the attention they deserve.

Remediation SLAs: putting a clock on findings

Finding a suspicious transaction means nothing if it sits in someone's queue for three weeks. What turns scoring into an actual control is attaching a deadline and an owner to every finding — not a vague "we'll look into it," but a specific person and a specific date.

  1. Critical (suspected fraud, self-approval on a large amount)

    Owner assigned same day. Initial assessment within 48 hours. Card frozen or payment held immediately if warranted.

  2. High (missing documentation on a large claim, unrecognized new vendor)

    Owner within 2 business days. Resolved or escalated within 5.

  3. Standard (round-number flags, minor category mismatches)

    Batched into the weekly sweep. Cleared within 10 business days.

The measurable outcomes worth tracking are narrow on purpose: percentage of findings resolved within SLA, average days-to-resolution, and the false-positive rate per signal. Three numbers. If findings resolved within SLA drops below around 80%, your queue is overloaded and either the scoring is too aggressive or you need to route more to automated checks. If a single signal's false-positive rate climbs toward 90%, that signal is costing more attention than it earns.

This SLA discipline is the same muscle used across expense operations generally — it's the difference between a control that exists on paper and one that actually runs. Teams that already have expense governance with clear ownership and lightweight audits tend to adopt remediation SLAs almost effortlessly, because the accountability structure is already there.

Where automation carries the weight

None of this scales if a human has to compute scores by hand. That's the honest limit of a spreadsheet-based approach — it works for a month, then someone gets busy and the scoring quietly stops.

Expense platforms with built-in scoring and flagging are where this actually becomes sustainable. The scoring model runs on every transaction automatically the moment it hits the system; new-vendor flags, threshold detection, and self-approval blocks fire without anyone remembering to check. Scheduled sweeps become saved views that populate themselves, and remediation SLAs run as timed reminders that escalate when a finding sits too long. The human stays in the loop for judgment — deciding whether a flagged transaction is actually a problem — but the detecting, ranking, and clock-watching happen without manual intervention.

The point isn't to replace your controller. It's to make sure the 30 transactions that need her attention are the right 30, surfaced automatically, with context already attached. Getting the data clean enough for scoring to work is its own prerequisite, which the pre- and post-transaction controls matrix covers in more depth.

A real scenario

A regional architecture firm, around 18 people, was running expenses through cards and monthly reimbursements with no scoring at all — the office manager eyeballed the card statement and approved reimbursements in a batch every month. Nothing dramatic had gone wrong, which was exactly why nobody worried.

When they applied a basic scoring pass to six months of history, two things surfaced. A design vendor had been paid three times for what looked like the same invoice across two months — roughly $2,400 in duplicate payments nobody caught because the amounts landed in different statement cycles. And one employee's reimbursements had crept from about $200 a month to closer to $700, all just under the $750 threshold that would have required a second approver.

Neither was elaborate fraud. The duplicate was a genuine vendor error that would've been repaid if flagged; the reimbursement creep was one employee getting loose with the rules because nobody was watching. Both had gone unnoticed for months. After setting up scoring with a weekly sweep on anything 46+ and a monthly trend check, the office manager's review time actually dropped — she stopped scanning every line and started looking only at what got flagged, maybe 20–25 transactions a month instead of all 300-plus. They recovered the duplicate, tightened the threshold behavior, and cut monthly review time by more than half.

When this framework makes sense — and when it doesn't

When it makes sense: You're processing enough transactions that manual review is unreliable — roughly 200+ a month is where it starts paying off. You've got at least one person who owns finance and can act on findings. And you've had, or worry about having, the "how did we miss that" moment.

When it's overkill: If you're a five-person shop running 40 transactions a month, a scoring model is more machinery than you need. Just require receipts, block self-approval, and have someone glance at the card statement. Build the framework when volume outgrows attention, not before.

Who should not do this yet: Teams whose underlying data is a mess. If vendor names are inconsistent, categories are unreliable, and half the receipts are missing, scoring will produce garbage flags and everyone will stop trusting it within a month. Clean the data foundation first — consistent vendors, enforced receipts, reliable categories — then layer scoring on top. A model built on bad data doesn't just fail quietly; it actively trains your team to ignore alerts.

Closing thought

The goal of a fraud-risk scoring framework isn't to catch every bad transaction — that's not realistic for a team this size, and chasing it burns people out. The goal is to spend your limited review hours where the risk actually concentrates, put deadlines on what you find, and check quarterly whether the model still reflects how your business spends.

Do that consistently and fraud-risk assessment stops being a periodic panic and becomes something that runs quietly in the background, surfacing the handful of things that genuinely need a human. Fewer transactions reviewed, better ones caught, and a team that trusts the flags because the flags have earned it.

The goal of a fraud-risk scoring framework isn't to catch every bad transaction — that's not realistic for a team this size, and chasing it burns people out. The goal is to spend your limited review hours where the risk actually concentrates, put deadlines on what you find, and check quarterly whether the model still reflects how your business spends. Do that consistently and fraud-risk assessment stops being a periodic panic and becomes something that runs quietly in the background, surfacing the handful of things that genuinely need a human.

Built for Businesses Tailored for streamlined expense tracking & budget management
Save Time Automate expense entry and reporting workflows
Gain Control Track budgets and spending with real-time insights
Increase Profitability Identify cost-saving opportunities and optimize expenses