September 12, 2026
Case Study: AI in Debt Collection Put to the Test
Case study: AI in receivables management — how a defined workflow prioritizes dunning notices, makes exceptions auditable, and reduces team workload with clear oversight.

Case Study: AI in Receivables Management Starts with Process, Not Models
A case study on AI in receivables management rarely begins with a bad model. It begins with a process that doesn’t account for its own exceptions: partial payments recorded in the ERP, responses sitting in email inboxes, payment commitments noted in comments. The dunning notice goes out anyway. This isn’t an AI problem—it’s an operational failure with direct impact on liquidity, customer relationships, and effort.
Why Manual Receivables Management Fails on Exceptions
In an anonymized case of a B2B company with 12,480 open invoice cases per year, a three-person team processed a daily list from the ERP. The list included invoice number, due date, amount, and debtor. But for a usable decision, critical information was missing: recent payments, open complaints, agreed installment plans, and responses to prior dunning notices.
The workflow involved eight manual handovers. One employee exported receivables, a second checked bank transactions, a third drafted or stopped dunning notices. For disputed cases, the team had to email sales or accounting. Until a response arrived, the case stalled or proceeded based on a blanket rule.
The errors were predictable. A customer received a reminder despite a payment commitment. A customer with a partial payment was treated as fully overdue. An invoice missing a reference number wasn’t followed up on for too long. None of these situations require an autonomous decision by a language model. They require clear data reconciliation, documented rules, and a human where data is insufficient.
This is where many initiatives fail: A model is supposed to read and prioritize texts, while status fields go unmaintained, payment data is imported with delays, or responsibilities are unclear. The system then processes incorrect cases faster. Honestly, automation without exception handling increases the reach of a bad process.
Case Study: AI in Receivables Management—What Was Specifically Tested
The following case study doesn’t promise generic success but describes a transferable operational model. The starting point was 18 months of historical data: invoices, payments, dunning levels, email responses, and processing times. Before building anything, it was tested which decisions could be rule-based and which cases required expert review.
The goal was defined measurably: The team shouldn’t see fewer cases but should manually handle fewer standard cases. A case was considered standard only if invoice, debtor, payment status, and dunning level clearly matched. Any deviation went into a separate queue.
Prioritization Follows Payment Data, Not Linguistic Intuition
The first component was workflow automation between ERP, bank data, and case management. An API integration or scheduled data sync matched payments using invoice number, amount, purpose, and debtor. For clear matches, the workflow updated the case status. For multiple possible invoices or mismatched amounts, it didn’t close the case automatically but created a review task.
The system doesn’t prioritize solely by overdue status. It checks at least four factors: amount and age of the receivable, documented payment commitment, prior communication response, and open clarification flags. An invoice with a confirmed commitment from yesterday is treated differently than an equally old invoice with no response.
For email responses and attached documents, document processing is used. Data extraction looks for invoice number, payment date, amount, or terms like “complaint” and “installment plan.” The model doesn’t just classify—it provides a confidence score and the text passage it’s based on. If the invoice number is missing or the amount contradicts the ERP, exception handling kicks in.
Humans Decide Where Rules End
A human-in-the-loop isn’t a fallback. It’s the defined control layer for high-risk cases. In this setup, automated steps were limited to cases with clear data. Disputed receivables, installment requests, unclear payments, high amounts, and cases with escalation risk remained with qualified staff.
The expert doesn’t see a long email history but a case with evidence: invoice data, payment reconciliation, identified statement, suggested next action, and justification. They can confirm, modify, or reject the proposal. Every decision is logged with a timestamp and reasoning. This is clearer than an inbox where the last relevant message must be searched for.
The Operational Path from Pilot to Ongoing Processing
A useful pilot doesn’t start with all receivables. It starts with a narrowly defined group, such as undisputed B2B invoices in one dunning level and within a specific amount range. Over four to six weeks, it’s tested whether data reconciliation, text classification, and handovers work under real conditions.
Before the first automated dispatch, a decision logic is needed. Who can stop a dunning notice? Which events reset a dunning level? When is a case handed off to sales, accounting, or legal? Which wording is permissible per country, contract type, and customer group? These rules belong in an approved process description—not just in the head of an experienced clerk.
The technical implementation connects existing systems instead of creating a shadow register. The workflow reads receivables from the leading system, processes incoming payments, handles messages, and writes the case status back. The daily reconciliation then checks two questions: Was every relevant payment processed? And does the number of processed cases match the ERP data?
This cross-check is critical. If an import fails, the system must not silently continue sending dunning notices. Monitoring tracks data flows, error rates, unassignable payments, and unusual spikes in escalations. For critical processes, uptime and SLAs are defined: Who responds to errors, within what timeframe, and how are incorrect actions corrected?
An AI agent can play a useful role in this setup, such as pre-reviewing incoming responses or drafting the next message. However, it shouldn’t decide on legally or commercially relevant exceptions without rules and approval. This depends on receivable type, volume, jurisdiction, and risk tolerance. For low volume, clean rule automation may be more economical than a complex model.
How to Measure Operational Success
A demo shows a message can be classified. Operations show whether the result is reliable. That’s why a baseline and fixed metrics are recorded before launch. In this case, these included manual processing time per standard case, the share of correctly stopped dunning notices after payment, the number of unresolved exceptions after 48 hours, and time to first qualified response.
A key metric is error rate: How many automatically prepared actions required correction? This rate is evaluated separately by case type. An overall rate might look good but still be unacceptable for partial payments or complaints.
Two weeks aren’t enough for judgment. Payments and customer responses follow their own cycles. A meaningful comparison requires at least one full dunning cycle—multiple cycles for seasonal business models. Only then does it become clear whether prioritization actually resolves cases earlier or just shifts work to another queue.
Smooth Operations Mean: Responsibility Remains Visible
Receivables management isn’t a project that runs itself after go-live. Payment formats change, ERP fields are adjusted, employees set new statuses, and customers respond differently than in training data. Without an operational concept, a good implementation becomes an unclear special process within months.
A pragmatic operation therefore has a named business owner, technical responsibility, and a fixed rhythm for exceptions. Weekly reviews check misclassifications and blocked cases. Monthly checks verify whether rules, texts, and permissions still match the current process. Changes are versioned and only deployed to production after testing.
This avoids surprises—not because every exception is known in advance, but because every exception has a predefined path. CINDR.LA builds such systems not as slideware but so that data reconciliation, human approval, monitoring, and ongoing responsibility work together operationally. The most sensible first step isn’t asking which model to use. Test first on 100 real cases: Which decisions today take too long, come too late, or are made incorrectly—and why.