Machine Learning · Term 4 · PGDM 2025-27
* Receivables risk scoring
A collections team can only chase a fraction of its open accounts. Working the ledger top to bottom treats every account as equally likely to default. This model ranks them instead, so the same effort recovers more. It runs entirely in your browser: the trained trees ship as JSON and score as you type.
30,000
accounts in the training file
0.783
AUC on held-out data
51%
of defaults in the top 20%
2.56x
lift over ledger order
* Score an account
Account details
Eight inputs. The other 22 features sit at their training median.
Last month first. This one column carries 51 percent of the model's importance.
Risk score
Probability of default next cycle
Loading 200 trees
* In short
I took 30,000 credit accounts from the UCI repository and built a model that answers one question: of the accounts that owe us money, which ones should we call first? I added seven features of my own on top of the raw columns, and two of them turned out to matter more than almost anything in the original file. Nine models were compared on the same split, and gradient boosting won. The result is not a number a manager has to interpret, it is a work list. Review the top fifth of the ledger and you reach half of everything that is going to go bad.
I picked this problem because my ERP capstone recommended it. We proposed three AI use cases to sit on top of an Odoo implementation, and scoring open invoices by payment risk was one of them. This is that recommendation, actually built and deployed.
* How it was built
Default of Credit Card Clients
UCI Machine Learning Repository, 30,000 accounts, 24 features, no missing values
| ID | LIMIT_BAL | SEX | EDUCATION | MARRIAGE | AGE | PAY_1 | BILL_AMT1 | PAY_AMT1 | DEFAULT |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 20000 | 2 | 2 | 1 | 24 | 2 | 3913 | 0 | 1 |
| 2 | 120000 | 2 | 2 | 2 | 26 | -1 | 2682 | 0 | 1 |
| 3 | 90000 | 2 | 2 | 2 | 34 | 0 | 29239 | 1518 | 0 |
| 4 | 50000 | 2 | 2 | 1 | 37 | 0 | 46990 | 2000 | 0 |
| 5 | 50000 | 1 | 2 | 1 | 57 | -1 | 8617 | 2000 | 0 |
| 6 | 50000 | 1 | 1 | 2 | 37 | 0 | 64400 | 2500 | 0 |
| 7 | 500000 | 1 | 1 | 2 | 29 | 0 | 367965 | 55000 | 0 |
| 8 | 100000 | 2 | 2 | 2 | 23 | 0 | 11876 | 380 | 0 |
First 8 of 30,000 rows. DEFAULT is the column being predicted.
What the columns mean
- LIMIT_BAL
- Approved credit limit in NT$, the total exposure on the account
- SEX / EDUCATION / MARRIAGE
- Demographics. Education and marital status carried codes outside the documented range
- AGE
- Age in years
- PAY_1 to PAY_6
- Repayment status for each of the last six months. Minus one means paid in full, one means one month late, and so on
- BILL_AMT1 to BILL_AMT6
- Bill amount for each of the last six months
- PAY_AMT1 to PAY_AMT6
- Amount actually paid in each of the last six months
- DEFAULT
- The target. One if the account defaulted the following month
What I found before modelling
345 records carried an education code outside the documented 1 to 4 range and 54 carried an undocumented marital code. Left alone, the model would treat each stray code as its own category and learn noise from a handful of rows. I folded them into "other".
The target is unbalanced: 22.1 percent of accounts default. Predicting "nobody defaults" scores 77.9 percent accuracy and recovers nothing, which is why accuracy is not the metric this model is judged on.
Data: Default of Credit Card Clients, UCI Machine Learning Repository. The model scores in your browser, so nothing you type is sent anywhere.