#ai#finance#fraud detection#machine learning#fintech#risk management

AI Fraud Detection in Finance: Models, Thresholds, and the Scams Models Miss

Reported US fraud losses hit a record $15.9B in 2025. How fraud models work, how to set decline thresholds with cost math, and why scams need other defenses.

๐Ÿ“… January 8, 2026โœ๏ธ Updated: September 27, 2026โฑ 7 min readโœ Web3 Listicle Editorial Team

A secure, modern financial data center with glowing blue neural network lines overlaid, symbolizing AI protecting digital transactions.

People reported losing a record $15.9 billion to fraud to the US Federal Trade Commission in 2025, up from $12.5 billion the year before. Investment scams accounted for nearly half of it ($7.9 billion). The FBI's Internet Crime Complaint Center, which counts differently, recorded $20.9 billion. Both are reported losses only; the FTC's own testimony calls its figure just a fraction of consumers' actual losses.

Those numbers include two very different problems. Unauthorized fraud is when someone other than the customer uses the account: stolen cards, account takeover, synthetic identities. Machine learning has been effective against this for decades. Authorized fraud, or scams, is when the real customer is tricked into sending the money themselves. Most of the growth in losses is here, and it is much harder for models because the customer's device, login, and behavior all look normal.

This guide covers how fraud models work, how to set thresholds with actual cost math, and what defenses work against scams.

How transaction fraud models work

The workhorse is a supervised model, usually gradient-boosted trees (XGBoost, LightGBM, or similar), trained on past transactions labeled as fraud or legitimate, typically from chargebacks and confirmed cases. It scores each new transaction in milliseconds. The features matter more than the algorithm:

  • Behavioral deviation. How far this transaction is from the customer's normal amount, merchant types, time of day, and location.
  • Velocity. Number and value of transactions in the last minute, hour, and day; number of new payees added recently.
  • Device and session. New device, emulator or remote-access software detected, password recently reset, session behavior unlike the customer's usual pattern.
  • Merchant and counterparty risk. Fraud rates at this merchant or receiving bank; newly opened receiving accounts.

Two other techniques complement it:

  • Graph analysis connects accounts, devices, phone numbers, addresses, and payees. Twenty "unrelated" accounts sharing three devices and a mailing address is a ring, even if each account looks fine alone. This is the main tool against synthetic identities and money mule networks.
  • Anomaly detection (isolation forests, autoencoders) flags unusual activity without labels. It finds new patterns but produces more false alarms, so it usually feeds an investigation queue rather than blocking payments.

Rules do not disappear. Hard rules for sanctions, known compromised cards, and regulatory limits sit alongside the model, and a small number of fast rules can respond to a new attack while the model is retrained.

Sophisticated machine learning algorithm visualization on multiple screens, showing complex pattern recognition models analyzing financial transaction data streams.

Setting thresholds with cost, not accuracy

A model outputs a probability. Where you draw the line between approve and decline is a business decision, and it should come from the cost of each kind of mistake.

  • Letting a fraudulent transaction through costs the amount lost plus fees (for example, a chargeback fee).
  • Declining a legitimate transaction costs the lost margin plus some chance the customer uses another card or leaves.

Decline when the expected cost of approving exceeds the expected cost of declining:

p ร— fraud cost > (1 โˆ’ p) ร— false-decline cost

which simplifies to declining when p > false-decline cost รท (false-decline cost + fraud cost).

A worked illustration for a $200 card-not-present purchase:

  • Fraud cost: $200 + $25 chargeback fee = $225
  • False-decline cost: $30 (lost margin plus an estimate of goodwill)
  • Threshold: 30 รท (30 + 225) = 11.8%

So the merchant should decline if the model says the fraud probability is above about 12%. For a $2,000 purchase with the same false-decline cost, the threshold drops to about 1.5%, which is why high-value transactions face more friction. For a $20 purchase, it rises to 40%.

In practice there are three zones: approve below a low threshold, decline above a high one, and step-up authentication in between (3-D Secure for cards, a push notification or biometric check in a banking app). Step-up costs a little friction and catches a lot of fraud, so the middle band is where most of the tuning effort goes.

Also check the base-rate math before going live. If 1 in 1,000 transactions is fraudulent, even a model with a 1% false positive rate flags about ten good transactions for every fraudulent one it catches. Our AI financial risk management guide works through that calculation.

Scams: why models struggle and what helps

In an authorized push payment (APP) scam, the customer logs in on their usual phone, passes every check, and sends money to a "safe account," a romance partner, or a fake investment platform. From the bank's side, only the payment itself looks unusual, and sometimes not even that.

Defenses that work better than device and login signals:

  • Payee verification. UK Confirmation of Payee and the EU's Verification of Payee, mandatory for euro credit transfers under the Instant Payments Regulation since October 2025, check that the name matches the account before money moves.
  • Scoring the receiving account. Mule accounts share traits: recently opened, sudden inbound volume, immediate onward transfers. Receiving banks can detect them, and inbound scoring is often more effective than outbound scoring.
  • Context-specific warnings. A generic "beware of scams" message changes nothing. A warning that says "Banks will never ask you to move money to a safe account" at the moment someone pays a new payee flagged as "transfer to my other account" is more effective.
  • Delays for high-risk first payments. A few hours' hold on a large first payment to a new payee gives the customer time to reconsider or call someone. With instant payment rails such as FedNow and RTP in the US, money is gone in seconds without it.
  • Signals from the conversation. Screen-sharing or remote-access apps running during a banking session, or an active phone call while making a payment, are strong scam indicators some banking apps now detect.

Liability is shifting too. Since October 2024, UK payment firms have had to reimburse most APP scam victims up to ยฃ85,000, split between the sending and receiving firms. That has made scam prevention a direct P&L issue for UK banks, and similar debates are underway elsewhere.

Deepfakes and document fraud

Generative AI has made fake identity documents, voice clones, and video deepfakes cheap. FinCEN issued an alert to US financial institutions in November 2024 on fraud schemes using deepfake media, particularly in account opening. Countermeasures include liveness detection in identity verification, checking document images for signs of generation or editing, and never treating voice alone as proof of identity for high-value instructions. Our AI threat intelligence guide covers impersonation attacks on businesses.

Governance and fairness

The EU AI Act classifies creditworthiness scoring of individuals as high-risk but explicitly excludes fraud detection from that category. That does not mean anything goes. Fraud models can still discriminate if features act as proxies for protected characteristics, and wrongly frozen accounts cause real harm. Test outcomes such as decline and account-closure rates across customer groups, give customers a way to resolve false positives quickly, and document the model under your supervisor's model risk expectations (SR 11-7 in the US).

A diverse financial analyst team collaborating, reviewing AI-generated fraud risk assessments on a large screen, emphasizing human-AI partnership in investigation workflows.

Building or improving a program

  1. Measure the baseline: fraud losses in basis points of volume, false decline rate, step-up rate, and investigator hours per confirmed case.
  2. Fix labels. Models are only as good as the fraud labels. Link chargebacks, customer claims, and investigation outcomes to the original transactions.
  3. Set thresholds from cost, separately for different transaction types and amounts.
  4. Add graph features for account opening and payees to catch rings and mules.
  5. Treat scams as their own problem, with payee checks, inbound mule scoring, contextual warnings, and friction on risky first payments.
  6. Monitor drift weekly. Fraudsters adapt within days of a new control.

For insurance-specific fraud, see AI in insurance risk assessment. For the compliance side of transaction monitoring, see AI regulatory compliance.


This article is for informational purposes only and is not financial, legal, or security advice. Evaluate fraud controls with qualified compliance and technology advisors.

Frequently Asked Questions

People [reported losing a record $15.9 billion](https://www.ftc.gov/news-events/news/press-releases/2026/03/ftc-testifies-joint-economic-committee-agencys-efforts-combat-fraud) to fraud to the US Federal Trade Commission in 2025, up from $12.5 billion in 2024, across about 3 million reports. Investment scams accounted for $7.9 billion. The FBI's Internet Crime Complaint Center [recorded $20.9 billion](https://www.ic3.gov/AnnualReport/Reports/2025_IC3Report.pdf) in losses for 2025. Both figures count only what victims reported, so actual losses are higher.
Most production systems use gradient-boosted tree models trained on labeled past transactions, scoring each new payment in milliseconds using features such as amount, merchant, device, location, velocity of recent activity, and how far the transaction departs from the customer's usual behavior. Graph analysis links accounts that share devices, addresses, or payees to find rings, and anomaly detection helps surface new patterns.
Compare the expected cost of each mistake. If letting fraud through costs the transaction amount plus fees, and wrongly declining a good customer costs lost margin plus goodwill, decline when the fraud probability exceeds the false-decline cost divided by the sum of both costs. Scores between approve and decline thresholds can trigger step-up authentication instead.
No. The AI Act lists AI used to evaluate the creditworthiness of individuals as high-risk but [explicitly excludes](https://artificialintelligenceact.eu/annex/3/) AI used to detect financial fraud from that category. Fraud systems still fall under GDPR, anti-discrimination law, and banking supervisors' model risk expectations.
Only partly. In these scams the real customer, deceived by the fraudster, makes the payment themselves, so device and login signals look normal. Useful defenses include payee name checks, scoring the receiving account for mule behavior, warnings tailored to the payment's context, and short delays for unusual first-time payments. In the UK, banks have had to [reimburse most APP scam victims up to 85,000 pounds](https://www.psr.org.uk/publications/policy-statements/ps247-faster-payments-app-scams-reimbursement-requirement-confirming-the-maximum-level-of-reimbursement/) since October 2024.