← IKHSAN.M.N
CASE FILE UNIT-01
THRESHOLD 0.35
--:--:--
経路
01Problem02Dataset03Class imbalance04ML pipeline05Benchmark06Error analysis07Explainability08Inspector09Stack
CASE FILE — UNIT 01不正検知機構・解析記録
市場取引・不正検知

Credit Card
Fraud
Detection

An imbalanced binary classification system over market transaction data. This file documents the full workflow: the data challenge, the pipeline, model benchmarking, error analysis, interpretation, and a live inference interface.

解析経路 / ANALYSIS PATH—— / 09
○01·······························Problem○02·······························Dataset○03·······························Class imbalance○04·······························ML pipeline○05·······························Benchmark○06·······························Error analysis○07·······························Explainability○08·······························Inspector○09·······························Stack
Transactions
1.04M
取引件数
Fraud rate
0.18%
不正比率
Best PR-AUC
0.874
最良性能
Recall @ 0.35
92.5%
検出率
01 — 問題定義

Problem

Detect fraudulent transactions while minimizing missed fraud and unnecessary false alarms.

Fraud detection is an imbalanced classification problem: fraudulent transactions represent a very small fraction of all activity. The objective is therefore not to maximize accuracy — a model that labels everything legitimate would already score above 99%.

The real objective is to control the balance between two error types, and to justify where that balance was set.

誤差の代償 / COST OF ERROR
FNMissed fraud
Fraud classified as legitimate. Passes through undetected and becomes a direct financial loss.
FPFalse alarm
Legitimate transaction flagged. Costs review time and inconveniences a real customer.
02 — 資料概要

Dataset

Transactions
1,048,575
post-deduplication
Features
31
24 engineered
Fraud cases
1,842
positive class
Legitimate
1,046,733
negative class
Fraud rate
0.176%
1 in 569
03 — 不均衡問題

Class Imbalance

階級分布 / CLASS DISTRIBUTIONn = 1,048,575
Legitimate99.824%
1,046,733 transactions
Fraud0.176%
1,842 transactions
A constant "legitimate" predictor scores 99.82% accuracy and detects zero fraud. Accuracy is excluded from model selection for this reason.
対処法 / HANDLING
SMOTE (train fold only)APPLIED
Synthetic minority oversampling inside cross-validation folds to avoid leakage into validation.
Class weightsAPPLIED
scale_pos_weight tuned on the training split for the tree-based models.
Threshold tuningAPPLIED
Operating point selected from the precision–recall sweep rather than the 0.50 default.
Random undersamplingREJECTED
Discarded majority-class signal without improving PR-AUC.
04 — 処理系統

ML Pipeline

Raw Transactions
Data Validation
Preprocessing
Feature Scaling
Class Balancing
Model Training
LOGISTIC REG.
RANDOM FOREST
XGBOOST
BiLSTM
Evaluation · PR-AUC / MCC
Threshold Selection
Fraud Prediction
Explainability · SHAP
05 — 性能比較

Benchmark

Held-out test split, threshold fixed at 0.50 for comparability. Ranked on PR-AUC — the metric that stays informative when the positive class is rare.

ModelPrecisionRecallF1MCCROC-AUCPR-AUC
Logistic Regression0.6120.7410.6700.6730.9680.702
Random Forest0.8120.7690.7900.7900.9810.831
XGBoost0.8470.8350.8410.8400.9870.874
BiLSTM0.7880.8120.8000.7990.9830.848
SELECTED — XGBOOSTRANKED ON PR-AUCACCURACY EXCLUDED
06 — 誤差解析

Error Analysis

The decision threshold is a design choice, not a default. Move it and watch the confusion matrix and every downstream metric respond.

判定閾値 / DECISION THRESHOLDt = 0.35
混同行列 / CONFUSION MATRIX
PRED LEGIT
PRED FRAUD
ACTUAL LEGIT
TN
1,044,490
correctly accepted
FP
2,243
false alarm
ACTUAL FRAUD
FN
120
missed fraud
TP
1,722
fraud caught
Precision
0.434
Recall
0.935
F1
0.593
MCC
0.636
閾値掃引 / THRESHOLD SWEEP
ThresholdPrecisionRecallF1
0.200.1730.9850.295
0.300.3340.9560.495
0.350.4340.9350.593
0.400.5390.9080.676
0.500.7260.8350.777
0.600.8520.7350.789
0.700.9210.6040.730
0.800.9540.4400.603
Operating point selected at t = 0.35: recall is prioritised because an undetected fraud carries a larger cost than a reviewed false alarm.
07 — 解釈性

Explainability

Mean absolute SHAP value per feature across the test set — what drives the model globally, before looking at any single transaction.

Transaction frequency
0.42
Amount deviation
0.31
Hour of day
0.27
Location change
0.22
Merchant category risk
0.16
Device fingerprint
0.11
Account age
0.07
08 — 取引検査装置

Transaction Inspector

Adjust the transaction characteristics. The scored probability and its local attribution update live, then are compared against the operating threshold set in section 06.

取引入力 / TRANSACTION INPUTID #48291
Transaction amount取引金額
Transaction time取引時刻
Transaction frequency取引頻度
Location change位置変化
New device新端末
Merchant risk加盟店危険度
判定結果 / MODEL OUTPUTXGBOOST · t = 0.35
Fraud probability
92.4%
0.00THRESHOLD 0.351.00
PREDICTION — FRAUD / 不正
局所寄与 / WHY?
Transaction frequency
+++
Transaction time
+++
Location change
+++
Transaction amount
+++
Merchant risk
++
New device
·
09 — 技術系統

Technical Stack

PythonPandasNumPyscikit-learnXGBoostTensorFlow / Kerasimbalanced-learnSHAPMatplotlibGit
Access Source ↗← Return to Index
版数情報 / BUILD MANIFESTDEPLOYED
Release版数·······································································v1.1.0
Interface画面系統·······································································NERV CONSOLE / REV 01
Channel配信区分·······································································STABLE
Statically generated Next.js build. Project records and reported figures reflect the state of each source repository at compile time; interactive models on the case files are illustrative reconstructions of published results, not live inference. Superseded builds are replaced in full at deploy.
© 2026 IKHSAN MOCHAMMAD NOOR記録は閲覧者の要求に応じて展開するEND OF CASE FILE