Loading...
Loading...
Credit risk data cleaning and variable screening pipeline for pre-loan modeling. Use when working with raw credit data that needs quality assessment, missing value analysis, or variable selection before modeling. it covers data loading and formatting, abnormal period filtering, missing rate calculation, high-missing variable removal,low-IV variable filtering, high-PSI variable removal, Null Importance denoising, high-correlation variable removal, and cleaning report generation. Applicable scenarios arecredit risk data cleaning, variable screening, pre-loan modeling preprocessing.
npx skill4agent add github/awesome-copilot datanalysis-credit-risk# Run the complete data cleaning pipeline
python ".github/skills/datanalysis-credit-risk/scripts/example.py"| Function | Purpose | Module |
|---|---|---|
| Load and format data | references.func |
| Organization sample analysis | references.func |
| Calculate missing rate | references.func |
| Filter abnormal months | references.analysis |
| Drop high missing rate features | references.analysis |
| Drop low IV features | references.analysis |
| Drop high PSI features | references.analysis |
| Null Importance denoising | references.analysis |
| Drop high correlation features | references.analysis |
| IV distribution statistics | references.analysis |
| PSI distribution statistics | references.analysis |
| Value ratio distribution statistics | references.analysis |
| Export cleaning report | references.analysis |
DATA_PATHDATE_COLY_COLORG_COLKEY_COLSOOS_ORGSmin_ym_bad_samplemin_ym_samplemissing_ratiooverall_iv_thresholdorg_iv_thresholdmax_org_thresholdpsi_thresholdmax_months_ratiomax_orgsn_estimatorsmax_depthgain_thresholdmax_corrtop_n_keep