---
title: "Testing collections automation: selecting a comparable pilot portfolio"
canonical: https://www.billabex.com/en/blog/collections-automation-pilot-portfolio/
lang: en
alternate: https://www.billabex.com/fr/blog/automatisation-relances-portefeuille-pilote.md
updated: 2026-09-20
index: https://www.billabex.com/llms.txt
---

# Testing collections automation: selecting a comparable pilot portfolio

An automated collections pilot can show more receipts simply because it was given the easiest accounts. If the manual group keeps the old disputes and unreachable contacts, the result mainly compares two different portfolios. Before starting, ask which accounts would have progressed comparably under the existing process. Without a credible answer, a polished dashboard can make a selection decision look like a software effect.

That question changes how the pilot is prepared. You need to choose a population, preserve its starting position and define what you want to learn. A pilot checking message quality does not use the same protocol as an evaluation of payment outcomes. Both have value, provided an operational demonstration is not presented as financial impact that has already been established across the wider business.

## Define the decision before selecting the accounts

Write down the decision the pilot should inform: extend automation to straightforward invoices, retain human approval for particular responses, or defer rollout until the underlying data is usable. Then choose a primary outcome consistent with that decision. For receipts, an example is cash actually received and allocated to invoices present at the start, over an observation window fixed before the pilot begins.

UK guidance on evaluating AI interventions recommends establishing a baseline and specifying what the comparison group receives. That methodological principle can inform a business pilot; the guidance is not a study proving the effectiveness of collections software. It helps distinguish a change in outcomes from a change caused by the intervention being assessed. [HM Treasury guidance, updated 15 May 2026](https://www.gov.uk/government/publications/the-magenta-book/guidance-on-the-impact-evaluation-of-ai-interventions-html).

Retain additional measures such as human handling time, incorrect amounts, reminders sent to the wrong contact and cases needing intervention. A favourable financial result is insufficient if the team must subsequently correct inaccurate communications. Our guide to [exceptions requiring human review](https://www.billabex.com/en/blog/ai-collections-agent-human-exceptions/) can help define operational boundaries before the first customer responses arrive, instead of making exclusions only after inconvenient cases appear.

## Capture the differences likely to affect payment

At the starting date, record days past due, open amount, disputes, supporting documents, verified contact details and any existing payment promise. Include language and, where relevant for a European portfolio, the country or payment route. Observe these characteristics before assignment. Information discovered by the software during the pilot cannot retrospectively become a baseline selection criterion without changing the meaning of the comparison.

Avoid creating so many categories that each contains a single account. Focus on differences that could materially change interpretation. Two groups with identical total balances can still be very different if one contains a few large customers and the other many small debts. Inspect the distribution of accounts, rather than relying on a headline average or total that hides concentration and different operational workloads.

The NIST describes blocked experimental designs: group situations that are similar on selected factors, then compare treatments within those groups. For collections, that suggests distinguishing recent invoices without blockers from cases requiring a prior resolution. This is an application of the method, rather than a guarantee that all hidden differences have been removed or that accounts inside a category are identical. [NIST guidance on blocking](https://www.itl.nist.gov/div898/handbook/pri/section3/pri332.htm).

## How portfolio composition creates an apparent improvement

Consider a simulation using no real customer data. Each group contains 100 customers with one €1,000 invoice per customer. The automated group has 80 recent cases without blockers and 20 blocked cases. The usual-process group has only 20 recent cases and 80 blocked cases. Assume that, during the observation window, 80% of recent cases and 20% of blocked cases are paid in full under both approaches.

| Simulated composition and outcome | Automated group | Usual-process group |
| --------------------------------- | --------------: | ------------------: |
| Recent cases without blockers     |              80 |                  20 |
| Payments from those cases         |              64 |                  16 |
| Cases blocked at the start        |              20 |                  80 |
| Payments from those cases         |               4 |                  16 |
| Total invoices paid               |              68 |                  32 |

The raw result shows €68,000 against €32,000, a €36,000 difference. Yet the settlement rate is identical between approaches within each category. In this example, the entire difference comes from portfolio composition. Describing it as additional collection caused by the software would therefore be wrong, even though every sum in the dashboard was arithmetically correct and the cash amounts reconciled to the simulated invoices.

With a common composition of 50 recent and 50 blocked cases, the same rates produce 40 + 10 = 50 paid invoices, or €50,000 in each group. This recalculation illustrates composition effects; it does not establish that retrospective weighting would repair a real pilot. Unmeasured differences can remain, including commercial relationships and other work performed by staff while the test is running.

## Assign customers without mixing the approaches

Where practical, randomly allocate eligible customers within the categories defined in advance. Keeping all invoices for a customer in the same group avoids exposing that customer to competing follow-up approaches simultaneously. Where multiple entities use the same accounts payable centre, check whether separate allocation would also create interference. The number of invoices is then different from the number of genuinely independent situations informing the analysis.

The usual-process group should continue receiving its normal follow-up. Deliberately withholding reminders would change the comparison and manufacture a more favourable difference. Conversely, asking sales to quietly rescue every pilot account can conceal the effort required. Record interventions on both sides, using a consistent way to measure the [cost of reminders across several teams](https://www.billabex.com/en/blog/payment-reminder-cost-three-teams/) without counting only the visible drafting time.

Avoid starting the groups in different periods without recording the consequence. Customers' monthly payment runs or holiday absences can move receipts independently of the tool. Use the same calendar period where possible and give each entry the same observation window. Invoices added later must not be judged as unsuccessful merely because they had fewer days in which a payment could be observed.

## Clean the starting data consistently

Before launch, check credits, previously received payments and duplicates in both groups using the same procedure. Finding [unallocated cash](https://www.billabex.com/en/blog/payment-received-invoice-still-open-tracing-unallocated-cash/) improves balance accuracy, but it does not represent a fresh receipt caused by a reminder. Preserve the actual bank movement date and the correction history. Otherwise a helpful data-cleaning exercise can accidentally become the largest claimed financial benefit of the pilot.

Decide what happens if a case becomes disputed during the test. Human intervention may be necessary, but removing the account from the denominator afterwards would flatter the outcome. Retain its initial assignment in the overall report and explain the intervention. You can show both the full assigned population and the work performed, without erasing difficult cases discovered after the pilot had already started.

## Make a decision at the precision the evidence supports

A small pilot may show that a workflow operates without estimating its average impact precisely. NIST explains that a **95% confidence interval** refers to a method whose intervals would cover the true parameter in approximately 95% of repetitions, under the relevant assumptions. It is not a probability of commercial success or a guarantee attached to the particular estimate in your report. [NIST on uncertainty around a mean](https://www.itl.nist.gov/div898/handbook/eda/section3/eda352.htm).

Size the analysis around the outcome, variability and grouping of customers; the 100 accounts in our simulation are not a universal minimum sample. If evidence remains limited, gradual expansion with stopping criteria may be more defensible than a numerical performance promise. [How Billabex works](https://www.billabex.com/en/product/how-it-works/) provides context for preparing that scope. Your protocol must then establish what the team observed, on which accounts and with which limits, before the result can support a wider deployment decision.

## Sources

- HM Treasury and Evaluation Task Force, updated 15 May 2026: [impact evaluation of AI interventions](https://www.gov.uk/government/publications/the-magenta-book/guidance-on-the-impact-evaluation-of-ai-interventions-html).
- NIST/SEMATECH e-Handbook, section 5.3.3.2: [randomised block designs](https://www.itl.nist.gov/div898/handbook/pri/section3/pri332.htm).
- NIST/SEMATECH e-Handbook, section 1.3.5.2: [confidence limits for the mean](https://www.itl.nist.gov/div898/handbook/eda/section3/eda352.htm).
