Machine learning benchmark for detecting pump-and-dump patterns around IDX Unusual Market Activity (UMA) events. This file documents the data challenge, the classical-ML versus deep-learning comparison, the threshold trade-off, model explainability, and an interactive model lab.
Does deep learning actually outperform classical ML for IDX market-manipulation detection?
Two model families are trained on the same engineered features and the same temporal split, then compared under one threshold-tuning procedure before any surveillance decision is made.
IDX OHLCV + UMA Events
Feature Engineering
Classical ML
LR · Random Forest · LightGBM
Deep Learning
BiLSTM · CNN-LSTM · Transformer
Threshold Tuning
MCC · F1 · PR-AUC · Recall
Surveillance Decision
02 — 資料難度
Data Challenge
Test samples
165,900
2025+ held-out split
Positive rate
~2.8%
manipulation-window samples
Event window
±5
trading days around UMA
階級分布 / CLASS DISTRIBUTIONn = 165,900
No event97.2%
161,255 samples
Manipulation window2.8%
4,645 samples
A constant "no manipulation" predictor scores 97.2% accuracy and detects zero events. Severe class imbalance makes accuracy an unreliable measure, so MCC and PR-AUC are emphasised throughout this file.
標識の性質 / LABEL PROPERTIES
UMA-derived labelsPROXY
An IDX Unusual Market Activity notice is a regulatory flag, not a confirmed manipulation verdict. The label is the closest available public proxy.
±5 day event windowAPPLIED
Samples inside five trading days either side of a UMA announcement are treated as positive.
Temporal splitAPPLIED
Training data ends before the test period begins, so no future regime leaks backwards.
AccuracyEXCLUDED
At a 2.8% positive rate a constant negative predictor already scores 97.2%. MCC and PR-AUC are used instead.
03 — 特徴設計
Feature Engineering
Market-behaviour features derived from daily OHLCV bars. Each one describes a different aspect of the price and volume signature that surrounds a pump-and-dump episode.
price_accel_5d
Price acceleration
The change in price momentum across the 5-day feature window — the rate at which returns themselves are speeding up, rather than the size of any single move.
価格加速度
vol_ratio_20d
Volume activity
Traded volume relative to the stock's own 20-day average volume, so each ticker is compared against its normal level rather than against the market.
出来高比率
volatility_20d
Volatility
Rolling standard deviation of daily returns over the trailing 20 sessions.
変動率
rsi_14
RSI (14)
The 14-day relative strength index — a bounded measure of how one-sided recent price movement has been.
相対力指数
ma_ratio_5_20
Moving-average ratio
Ratio of the 5-day to the 20-day moving average — how far price has extended above or below its own short-term trend.
移動平均比
ret_1d
Daily return
The single-session percentage change in closing price.
日次収益率
temporal_dow
Temporal features
Calendar position of each observation — day of week and relative position inside the ±5 trading-day event window.
時間的特徴
04 — 実験設計
Methodology
The split is time-based, not random. A random split would let future market regimes leak into training and inflate every score reported below.
時系列分割 / TEMPORAL SPLIT
TRAIN
2021 — 2024
TEST
2025 +
NEVER SEEN DURING TRAINING →
訓練経路 / TRAINING
StandardScaler
Fitted on the training period only, then applied to the test period unchanged.
SMOTE
Synthetic minority oversampling applied to the training data so the rare positive class is learnable.
Model training
Six architectures trained on identical features and identical windows.
評価経路 / EVALUATION
Time-based split
2021–2024 for training, 2025 onward held out entirely.
Validation threshold
Each model's operating point is chosen on validation data, not on the test set.
Final test evaluation
MCC, F1, PR-AUC and recall reported once, on data the models never saw.
05 — 主結果
ML vs DL Benchmark
Matthews correlation coefficient on the held-out 2025+ test split, each model at its own tuned threshold. MCC is used as the headline metric because it stays honest under a 2.8% positive rate.
MCC 順位 / RANKINGTEST SPLIT
Random Forest
Classical ML
0.3405
CNN-LSTM
Deep Learning
0.3061
LightGBM
Classical ML
0.2545
Transformer
Deep Learning
0.2387
Logistic Regression
Classical ML
0.2154
BiLSTM
Deep Learning
0.1756
Random Forest achieved the highest MCC. CNN-LSTM was the strongest deep-learning model — and still finished below it.
Deep learning did not automatically outperform classical machine learning on this dataset. Sequence models had access to the same feature window and the same labels, and three of the four highest-capacity architectures ranked below a tree ensemble. That result is the reason the rest of this file examines cost rather than accuracy alone.
BEST — RANDOM FORESTBEST DL — CNN-LSTMACCURACY EXCLUDED
06 — 適合率再現率曲線
PR Curves
Precision against recall as the decision threshold sweeps. With a 2.8% positive rate the random baseline sits at 0.028, so the whole usable region of this chart is close to the floor. Select a model to highlight its curve.
Random ForestCNN-LSTMLightGBMTransformerLogistic RegressionBiLSTM
Marker shows the tuned operating point for Random Forest. Curve shapes are reconstructed from each model's tuned precision/recall pair and are illustrative between measured points.
07 — 複雑性と代償
Complexity vs Quality
Does the additional complexity of deep learning provide enough performance improvement to justify its computational cost? Plotting detection quality against model complexity turns model choice into an engineering decision rather than a leaderboard.
検出品質 ↑ / 複雑性 →MCC vs COMPLEXITY
Random Forest
CNN-LSTM
LightGBM
Transformer
Logistic Regression
BiLSTM
CLASSICAL ML→ COMPLEXITY →DEEP LEARNING
運用費用 / OPERATIONAL COST
Logistic Regression
VERY LOW
LightGBM
LOW
Random Forest
LOW–MED
BiLSTM
HIGH
CNN-LSTM
HIGH
Transformer
VERY HIGH
Relative scale only. Training and inference cost were not benchmarked in wall-clock terms in this study — the ordering reflects parameter count and architecture class, not measured latency.
The narrative is not "I trained six models". It is: model architecture was evaluated as an engineering decision, and on this dataset the cheaper family won.
08 — 判定閾値
Threshold Trade-off
A classifier does not produce "manipulation" or "no manipulation". It produces a probability, and the decision threshold determines how sensitive the surveillance system becomes. Model selection and threshold selection are separate decisions.
Random Forestt = 0.35
PRECISION
31.5%
RECALL
41.5%
More balanced detection. Roughly one in three alerts is a real event, at the cost of missing over half of them.
LightGBMt = 0.95
PRECISION
13.5%
RECALL
64.3%
More aggressive surveillance: higher recall, but substantially more false positives for the review desk to absorb.
判定閾値 / DECISION THRESHOLD — Random Forestt = 0.35
0.05TUNED 0.350.95
混同行列 / CONFUSION MATRIX
PRED CLEAN
PRED EVENT
ACTUAL CLEAN
TN
157,062
correctly cleared
FP
4,193
false alarm
ACTUAL EVENT
FN
2,717
missed event
TP
1,928
event caught
Precision
0.315
Recall
0.415
F1
0.358
MCC
0.340
閾値の効果 / DIRECTION OF EFFECTAS THRESHOLD ↑
Precision↑
Recall↓
False alarms↓
Missed events↑
Raising the threshold buys precision and pays for it in missed events. Which direction is correct depends on how much analyst review capacity the surveillance desk actually has — not on the model.
Sweep behaviour is reconstructed from each model's published tuned precision/recall pair. Only the tuned operating points are measured values.
09 — 模型実験室
Interactive Model Lab
Now explore the model yourself. Pick an architecture, change how many features are shown, and select any feature to read what it measures and why it matters for manipulation detection.
模型選択 / MODEL
特徴重要度 / FEATURE IMPORTANCE
01price_accel_5d
0.215
02vol_ratio_20d
0.186
03volatility_20d
0.147
04rsi_14
0.124
05ma_ratio_5_20
0.098
Feature names are the implementation's. Importance weights shown here are illustrative and used to demonstrate the comparison interface.
特徴検査 / FEATURE INSPECTORRandom Forest
price_accel_5d
Price acceleration
価格加速度
WHAT IS IT?
The change in price momentum across the 5-day feature window — the rate at which returns themselves are speeding up, rather than the size of any single move.
WHY DOES IT MATTER?
A pump does not look like a large return. It looks like a return that keeps getting larger day over day. Acceleration separates that pattern from an ordinary rally on news.
IMPORTANCE — Random Forest
RANK 01 OF 070.215
10 — 限界と技術
Limitations
This is positioned as a research-oriented surveillance benchmark, not a production-ready market-surveillance system. An MCC near 0.34 is a meaningful signal on a 2.8% positive rate, and it is also far from a deployable decision.
Noisy UMA labels
A UMA notice flags unusual activity; it does not confirm manipulation. Some positives are legitimate volatility.
Concept drift
Manipulation tactics change. A model fitted on 2021–2024 behaviour degrades against later strategies.
Survivorship bias
Delisted and suspended tickers are under-represented in the sample.
No order book
Only daily OHLCV is available. Layering, spoofing and quote-level patterns are invisible to these features.
No social signals
The promotion phase of a pump often happens off-exchange, on channels this dataset does not observe.
Statically generated Next.js build. Project records and reported figures reflect the state of each source repository at compile time; interactive models on the case files are illustrative reconstructions of published results, not live inference. Superseded builds are replaced in full at deploy.