# AI collections: which exceptions need human review?

- HTML page: https://www.billabex.com/en/blog/ai-collections-agent-human-exceptions/
- Language: en
- Version française: https://www.billabex.com/fr/blog/agent-ia-recouvrement-exceptions-humaines.md
- Updated: 2026-09-10
- Site index: https://www.billabex.com/llms.txt

An AI agent can prepare a factually accurate reply about an overdue invoice and still make the wrong next decision. Accepting an extension, acknowledging an invoicing error or announcing legal escalation does not commit a business in the same way as sending a reminder about a verified due date. Defining human exceptions means establishing those differences before the first sensitive case arrives.

The question is becoming practical as adoption expands. **20% of EU enterprises with at least ten people in the scope of Eurostat's survey used an AI technology in 2025**. That figure measures adoption, not the reliability or autonomy of payment reminders, and does not cover all microbusinesses. [Eurostat, 2025 ICT survey, published 11 December 2025](https://ec.europa.eu/eurostat/web/products-eurostat-news/w/ddn-20251211-2).

For a founder or finance director, the useful question is which actions the organisation authorises, on what evidence, and who decides when a required condition is missing. The answer must be understandable to accounting staff, sales colleagues and any provider helping with implementation. A policy that only its original author can interpret is unlikely to survive a busy week.

## Start with decisions rather than customer labels

A label such as “sensitive customer” does not define an authorised action. The same customer may receive a factual reminder while a requested debt reduction needs approval. Conversely, a small invoice can contain a significant contractual dispute. Value indicates possible consequences, but cannot replace an examination of the decision itself.

List what the team actually does: request documents, confirm information, suggest dates, accept concessions, amend balances, pause actions and resume contact. For each action, identify the necessary evidence and the person with the relevant authority. Permission to send a message does not automatically include permission to accept whatever the customer proposes in reply.

CNIL recommends defining permitted and prohibited uses of generative systems, training users and organising checks. It also explains that plausible outputs may be inaccurate. The matrix below applies those concerns to a collections workflow; it is not a universal legal classification of debt collection software. [CNIL, generative AI questions and answers, questions 2 and 6](https://www.cnil.fr/fr/les-questions-reponses-de-la-cnil-sur-lutilisation-dun-systeme-dia-generative).

## Four reasons to request a decision

**The evidence is missing.** A customer reports a payment, but no allocation is confirmed. The system should acknowledge the information without inventing a receipt. Investigating an [unallocated payment](https://www.billabex.com/en/blog/payment-received-invoice-still-open-tracing-unallocated-cash/) is an example of work that may need accounting input before the next balance is communicated.

**The information conflicts.** The invoice appears overdue while a salesperson mentions an agreed extension. The useful next step is to establish which instruction is valid. Using whichever record was updated most recently is inadequate if that change was never approved by someone with authority.

**The action changes a commitment.** A discount, waiver, unusual schedule or requested set-off requires the organisation to check decision rights. The agent can present the request and its context. The business must establish who can accept, refuse or offer an alternative.

**The context becomes sensitive.** Suspected fraud, a reported insolvency proceeding, an incorrect recipient or a substantial dispute calls for a specific route. The case must remain understandable to the person taking over, even when they have never seen that customer's history. A generic alert saying “needs attention” does not provide enough information.

## A practical matrix to adapt

| Observed situation                       | What can be prepared                     | Decision to assign                          |
| ---------------------------------------- | ---------------------------------------- | ------------------------------------------- |
| Purchase order requested                 | Invoice and available contract reference | Who obtains or validates the document       |
| Payment reported but not found           | Amount, date and reference summary       | Who confirms the allocation                 |
| Discount requested in return for payment | Request and verified balance             | Who may approve the concession              |
| New due date proposed                    | Previous commitments and current request | Who may amend the agreement                 |
| Unusual banking instruction              | Anomaly recorded without execution       | Who verifies through an independent channel |

This is an organisational proposal, not a claim that every product provides these functions. Use it to request a clear demonstration and allocate responsibilities internally. A row without an owner is more than an unfinished setting: it represents a decision that may remain unresolved when a real case arrives.

ANSSI recommends limiting automated actions when an AI system processes uncontrolled inputs and controlling its interactions with business applications. In collections, a customer's email remains information to evaluate, rather than authority to rewrite the company's rules. [ANSSI, 29 April 2024 guide, recommendations R26 and R27](https://messervices.cyber.gouv.fr/documents-guides/Recommandations_de_s%C3%A9curit%C3%A9_pour_un_syst%C3%A8me_d_IA_g%C3%A9n%C3%A9rative.pdf).

## Few alerts do not prove sound autonomy

A system that rarely requests help may be efficient, or it may miss important situations. Its escalation rate alone cannot distinguish those possibilities. Review cases it escalated and cases it handled without intervention. Otherwise, your evaluation only sees problems the system already recognised by itself.

Consider an **entirely simulated test of 100 conversations**. Independent business reviewers decide that 16 required human judgment. The system escalates 20: 12 correctly and 8 unnecessarily. Among the remaining 80 conversations, it misses 4 exceptions. It therefore identifies **12 / 16 = 75%** of the cases requiring escalation, while **12 / 20 = 60%** of its alerts are relevant.

These percentages assess no supplier and establish no acceptable performance threshold. They show why two measurements are necessary. Reducing the eight unnecessary alerts without understanding the four missed exceptions could improve apparent convenience while worsening the consequences of mistakes. Examine the nature of each case. A missing document and an unauthorised concession do not carry the same implications.

The business should also record disagreements between reviewers. If two experienced colleagues cannot agree on the expected decision, the problem may lie in an unclear internal policy. Treating either opinion as unquestionable ground truth would make the evaluation misleading. Resolve the policy question and retain the difficult example for subsequent checks.

## Make the handover useful to a decision-maker

An exception should contain the question to resolve, relevant amount, supporting documents, latest commitments and possible next steps. It should distinguish the customer's statement from information the business has verified. A lengthy summary that mixes those categories forces the reader to reconstruct the case and consumes some of the time automation was expected to release.

The owner should be able to return to the original exchange. A summary can omit a qualification, a negative or the scope of a proposed date. Ask how a human correction is retained: the decision should be visible during the next interaction, with its author and scope. A one-off approval must not become blanket permission across the customer base.

Plan for absence as well. When the normal decision-maker is unavailable, which action stays paused and who can substitute? Waiting should leave the team with a clear status, without producing a commitment merely because nobody responded. Continuity is easier to prepare before holidays than during a customer's dispute.

## Test before expanding the delegation

Create representative examples using authorised or anonymised data: missing documents, reported payments, verbal extensions, discount requests and ambiguous messages. Ask the people handling those situations to define the expected decision. Only then compare the system's result with the agreed reference.

Keep difficult cases in the review after launch. New customers, process changes and software updates may alter results. The purpose is not to accumulate approvals indefinitely. It is to know which decisions the organisation can delegate and which remain assigned to a competent person, with evidence supporting that boundary.

The guide to [AI-powered debt collection software](https://www.billabex.com/en/blog/ai-powered-debt-collection-software/) introduces the category, while [compliance and security](https://www.billabex.com/en/blog/compliance-and-security-in-debt-recovery-software/) covers other assessment questions. To evaluate [Billabex's collections approach](https://www.billabex.com/en/debt-collection-software/), bring an anonymised sensitive case to the demonstration. Ask where the agent stops, what information it presents and how your decision controls the next action.

## Sources

- [Eurostat: 20% of EU enterprises use AI technologies](https://ec.europa.eu/eurostat/web/products-eurostat-news/w/ddn-20251211-2), 11 December 2025, 2025 ICT survey and methodological note.
- [CNIL: questions and answers on generative AI systems](https://www.cnil.fr/fr/les-questions-reponses-de-la-cnil-sur-lutilisation-dun-systeme-dia-generative), questions 2, 6 and 8, accessed 7 September 2026.
- [ANSSI: security recommendations for a generative AI system](https://messervices.cyber.gouv.fr/documents-guides/Recommandations_de_s%C3%A9curit%C3%A9_pour_un_syst%C3%A8me_d_IA_g%C3%A9n%C3%A9rative.pdf), version 1.0, 29 April 2024, R26 and R27.
