CredWise: A Controlled Agentic Decision-Intelligence Framework for Explainable and Auditable Credit-Risk Assessment
Organizations: Department of Mathematics, Indian Institute of Technology Kharagpur, Kharagpur 721302, West Bengal, India
Abstract
Credit-risk prediction is important in banking, but a prediction alone does not explain why an applicant is risky or how it should be combined with other evidence. This paper presents CredWise, a decision-support framework that integrates credit-risk prediction, probability calibration, explainable artificial intelligence, policy retrieval, SQL analytics, and controlled agent-based workflows. An XGBoost model is trained on Lending Club data (1,345,310 loans, 18 features) using a temporal split: 2007--2016 for training, 2017 for validation, and 2018 for testing. On the 2018 test set, the calibrated model achieved a ROC-AUC of 0.7109, PR-AUC of 0.2993, F1-score of 0.3714, and accuracy of 65.44%. Calibration reduced the Brier score from 0.2157 to 0.1273 and the expected calibration error from 0.2862 to 0.0585. SHAP explanations were temporally stable, with a Spearman correlation of 0.9959 between 2017 and 2018 feature rankings. On 28 labeled queries covering nine policy sections, FAISS achieved the best Hit@1 (0.929) and MRR (0.964), while all three retrieval methods reached Hit@5 = 1.0. Agent routing achieved 95.6% accuracy (43 of 45 cases), and the SQL benchmark scored 1.0 on exact-match, execution-success, and result-match across six cases. These results show that CredWise can combine predictions, explanations, policy evidence, and structured analytics in one controlled workflow. It is an academic research prototype, and final decisions remain with a human reviewer.
Figures & tables
| Period | Purpose | Samples | Default Rate |
|---|---|---|---|
| 2007–2016 | Training | 1,119,699 | 19.70% |
| 2017 | Validation | 169,300 | 23.12% |
| 2018 | Test | 56,311 | 15.75% |
| Total | 1,345,310 |
| Feature | Meaning |
|---|---|
| loan_amnt | Loan amount |
| term | Loan term |
| int_rate | Interest rate |
| installment | Monthly payment |
| sub_grade | Loan sub-grade |
| emp_length | Employment length |
| Metric | 2017 | 2018 |
|---|---|---|
| Accuracy | 0.6403 | 0.6544 |
| Precision | 0.3539 | 0.2602 |
| Recall | 0.6724 | 0.6484 |
| F1-score | 0.4637 | 0.3714 |
| ROC-AUC | 0.7091 | 0.7109 |
| PR-AUC | 0.4075 | 0.2993 |
| Metric | Feature Set A | Feature Set B | Change |
|---|---|---|---|
| Accuracy | 0.6462 | 0.6497 | +0.0035 |
| Precision | 0.3113 | 0.3215 | +0.0102 |
| Recall | 0.6370 | 0.6799 | +0.0429 |
| F1-score | 0.4182 | 0.4366 | +0.0184 |
| ROC-AUC | 0.6985 | 0.7198 | +0.0213 |
| PR-AUC | 0.3650 | 0.3842 | +0.0192 |
| Metric | Before Calibration | After Calibration |
|---|---|---|
| Brier score | 0.2157 | 0.1273 |
| ECE | 0.2862 | 0.0585 |
| Metric | Estimate | 95% CI |
|---|---|---|
| Accuracy | 0.6544 | [0.6505, 0.6584] |
| Precision | 0.2602 | [0.2544, 0.2661] |
| Recall | 0.6484 | [0.6380, 0.6579] |
| F1-score | 0.3714 | [0.3644, 0.3783] |
| ROC-AUC | 0.7109 | [0.7050, 0.7164] |
| PR-AUC | 0.2993 | [0.2904, 0.3080] |
| Measure | Result |
|---|---|
| Spearman rank correlation | 0.9959 |
| Spearman -value | |
| Top-5 overlap | 100% |
| Top-10 overlap | 90% |
| Top-15 overlap | 100% |
| Feature | 2017 PSI | 2018 PSI |
|---|---|---|
| revol_util | 0.0967 | 0.3305 |
| int_rate | 0.0741 | 0.1531 |
| revol_bal | 0.0200 | 0.0915 |
| loan_amnt | 0.0267 | 0.0600 |
| installment | 0.0259 | 0.0439 |
| dti | 0.0024 | 0.0392 |
| Predicted Non-default | Predicted Default | |
|---|---|---|
| Actual Non-default | 31,102 | 16,342 |
| Actual Default | 3,118 | 5,749 |
| Component | Metric | Result |
|---|---|---|
| Policy RAG (BM25) | Hit@1 / Hit@2 / Hit@3 / Hit@5 | 0.714 / 0.857 / 0.857 / 1.000 |
| Policy RAG (BM25) | MRR | 0.818 |
| Policy RAG (FAISS) | Hit@1 / Hit@2 / Hit@3 / Hit@5 | 0.929 / 1.000 / 1.000 / 1.000 |
| Policy RAG (FAISS) | MRR | 0.964 |
| Policy RAG (Hybrid) | Hit@1 / Hit@2 / Hit@3 / Hit@5 | 0.821 / 1.000 / 1.000 / 1.000 |
| Policy RAG (Hybrid) | MRR | 0.911 |