← IKHSAN.M.N
CASE FILE UNIT-02
ResNet50 + linear
--:--:--
経路
01The finding02Architecture03How we got here04Decisions05The 3×3 grid06Pareto frontier07Where error lives08Why RBF loses09Dataset10Limitations
CASE FILE — UNIT 02脳腫瘍検出・要因実験記録
脳腫瘍検出装置

Brain Tumor
Detection

Brain Tumor MRI — What Does Accuracy Actually Cost?

A full 3 × 3 factorial benchmark of frozen CNN feature extractors × SVM kernels, measured on both accuracy and the clock. Every published hybrid CNN–SVM pipeline for brain tumor MRI reports accuracy. Almost none reports what it cost to run.

The kernel the field reaches for by default turned out to be the worst cell in the grid: slower and less accurate on all three backbones.

解析経路 / ANALYSIS PATH—— / 10
○01·······························The finding○02·······························Architecture○03·······························How we got here○04·······························Decisions○05·······························The 3×3 grid○06·······························Pareto frontier○07·······························Where error lives○08·······························Why RBF loses○09·······························Dataset○10·······························Limitations
Configurations
9
構成数
Best accuracy
94.75%
最良精度
Runtime spread
5.3×
時間差
On frontier
3
有効構成
01 — 要点

The Finding

Nine configurations. Six are Pareto-dominated — some other configuration is both faster and at least as accurate. Only three survive.

🥇 RECOMMENDED8.89 s
ResNet50 + linear
94.58%
94.58 of a 94.75 ceiling at 69% of the cost. The operating point this study recommends.
⚡ CHEAPEST3.91 s
DenseNet121 + linear
93.00%
Cheapest cell in the grid — 5.3× faster than the slowest, and 1.22 points more accurate than it.
🎯 CEILING12.89 s
InceptionV3 + polynomial
94.75%
The accuracy ceiling, but it buys only 0.17 points over ResNet50 + linear for 45% more time.
主結論 / THE RESULT THAT MOTIVATED THE STUDY

The RBF kernel is Pareto-dominated on all three backbones.

It loses 1.5–2.4 accuracy points while costing 1.9–3.3× more time than the linear kernel on the very same features. There is no accuracy argument for it here, and a large runtime argument against it.

02 — 系統構成

How The Pipeline Earns Its Shape

The backbone is frozen. That single decision reorganizes the entire cost structure, and it is the reason this study exists.

5,712 MRI slices · 4 classes
PAID ONCE — OFFLINE, AMORTIZED TO ZERO
Resize + canonical preprocessing
Frozen ImageNet CNN — head removed, avg_pool kept
Feature cache · CSV · 1,024–2,048-d
PAID EVERY TIME — RETRAIN · RECALIBRATE · CROSS-VALIDATE
SVM — linear · poly · RBF
Accuracy · macro F1
confusion matrix
Wall-clock timer
fit + predict

Because gradients never flow back into the backbone, feature extraction happens once, is cached, and is reused forever. Every recurring cost — refitting on new institutional data, site-specific recalibration, hyperparameter cross-validation — lands entirely on the SVM stage.

So the dominant tunable knob for recurring cost is the kernel, not the backbone. Yet the efficiency literature optimizes backbones — precisely the component the frozen regime amortizes away.

03 — 経緯

How We Got Here

The from-scratch baseline is kept in the repository deliberately. It is the empirical justification for the transfer-learning route, not a failed experiment to be hidden.

基準線の失敗 / BASELINE FAILURE
Custom CNN, trained from scratch
Conv2D(32) → Conv2D(64) → Conv2D(128) → Dense(128) → Dropout(0.5) → Softmax(4)
Train accuracy~86%
Validation accuracy24–40%
Test accuracy~52%
A 34-point train/test gap that survives augmentation and regularization is a data-quantity signal, not a hyperparameter signal. 5,712 images is not enough to learn good low-level visual features from zero. The fix is to import those features, not to keep re-deriving them.
決定木 / THE ROUTE TAKEN
Custom CNN from scratchFAILED
86% train / ~52% test. Textbook overfitting.
More epochs, augmentation, CLAHEFAILED
Still ~52%. Too little labeled data for the fix to be a hyperparameter.
Transfer learningTAKEN
Import the low-level features instead of re-deriving them.
Freeze the backboneTAKEN
Cache features once, so the classifier stage is all that recurs. Fine-tuning would return the recurring GPU cost and the overfitting risk.
Which classifier, which kernel?STUDY
The question this study answers.
04 — 技術判断

Engineering Decisions

Each of these was a fork in the road. The reasoning matters more than the outcome. Select a decision to read the record.

判断記録 / DECISION LOG
01Abandon the scratch CNN
02Freeze, don't fine-tune
03Cache features to CSV
04Backbones by inductive bias
05No extra preprocessing
06No per-cell tuning
07Paired, not merely matched
08Narrow runtime definition
09Internal split, stated loudly
決定記録 / RECORD01 OF 09
Abandoning the from-scratch CNN instead of tuning it harder
自作網の断念
DECISION
Training accuracy climbed to ~86%; test accuracy sat at ~52%, with validation bouncing around 24–40%. Epoch sweeps of 15 → 100 → 1000 with EarlyStopping, ImageDataGenerator augmentation, and a custom tf.data pipeline adding Gaussian noise, contrast adjustment and CLAHE-style equalization — none closed the gap.
WHY
A 34-point train/test gap that survives augmentation and regularization is a data-quantity signal, not a hyperparameter signal. 5,712 images is simply not enough to learn good low-level visual features from zero. The fix is to import those features, not to keep re-deriving them.
CONSEQUENCE
The baseline is kept in the repository deliberately. It is the empirical justification for the transfer-learning route, not a failed experiment to be hidden.
05 — 要因表

The 3 × 3 Grid

All figures over the same 1,143 evaluation images, identical row ordering, one paired comparison. Click any cell — it drives the Pareto plot, the runtime replicates, and the per-class breakdown below.

LINEAR
POLYNOMIAL
RBF
SPREAD
DENSENET1211,024-d
FRONTIER93.003.91 s
DOMINATED93.007.53 s
DOMINATED91.1612.88 s
-1.84
RESNET502,048-d
FRONTIER94.588.89 s
DOMINATED93.8811.79 s
DOMINATED92.00 †17.01 s
-2.58
INCEPTIONV32,048-d
DOMINATED93.2610.64 s
FRONTIER94.7512.89 s
DOMINATED91.7820.87 s
-2.97
PARETO FRONTIERDOMINATEDSELECTED
The whole grid spans just 3.6 points — no configuration is catastrophic, so this is genuinely a trade-off rather than a correctness question. But note the pattern in the last column: RBF is the worst kernel on every single backbone. † ResNet50 + RBF's confusion matrix was not retained; its ≈92% is the classifier's own two-digit report.
06 — 費用境界

The Pareto Frontier

Accuracy against Run 1 wall-clock. A configuration survives only if nothing else is both faster and at least as accurate. Selected cell: ResNet50 + linear.

5s10s15s20s91%92%93%94%95%RUN 1 WALL-CLOCK →ACCURACY ↑DenseNet linDenseNet polDenseNet rbfResNet linResNet polResNet rbfInception linInception polInception rbf
実行時間 / RUNTIME — TWO REPLICATES
RUN 1 / RUN 2
LINE
POLY
RBF
DENSENET121
3.91 / 4.05
7.53 / 4.89
12.88 / 14.09
RESNET50
8.89 / 6.99
11.79 / 8.85
17.01 / 11.60
INCEPTIONV3
10.64 / 8.38
12.89 / 10.63
20.87 / 14.21
In all six backbone × replicate cells, without exception: linear < polynomial < RBF. Absolute times are not stable — shared-tenancy contention moved them by up to 32% between replicates — but the ordering is perfectly stable. Every runtime claim here is made at the level of ordering, none at the level of absolute latency.
選択構成 / SELECTED CONFIGURATION
ResNet50 + linear🥇 RECOMMENDED
Accuracy
94.58%
Run 1
8.89 s
Run 2
6.99 s
Errors
62
The recommended operating point — 94.58 of a 94.75 ceiling at 69% of the cost. The glioma→meningioma cell shrinks to 23, but is still the largest error in the matrix.
07 — 誤差の所在

Where The Error Lives

Aggregate accuracy hides the only thing that matters clinically. Per-class F1 for ResNet50 + linear — change the selection in the grid above.

GliomaOK
神経膠腫
0.94
F1 SCORE
MeningiomaBOTTLENECK
髄膜腫
0.89
F1 SCORE
NotumorSOLVED
腫瘍なし
0.96
F1 SCORE
PituitarySOLVED
下垂体腫瘍
0.98
F1 SCORE
最大誤差セル / LARGEST OFF-DIAGONAL CELL
Glioma→Meningioma23
Still the single largest error in the matrix, even at the recommended operating point.
構造の一様性 / UNIFORM STRUCTURE
Notumor and pituitary are solved
F1 never drops below 0.95 and recall exceeds 0.96 in every one of the nine configurations.
Meningioma is the bottleneck in all nine
Without exception, spanning 0.82 to 0.90. Best anywhere is InceptionV3 + polynomial at 0.90; worst is DenseNet121 + RBF at 0.82.
Glioma → meningioma is always the largest cell
20 to 46 cases depending on configuration. No backbone or kernel choice removes it.
Meningioma leaks in two directions
Into notumor (12–20) and pituitary (5–21) — consistent with meningiomas being extra-axial lesions whose appearance on one axial slice can resemble both normal tissue and sellar-region masses.

Swapping the backbone moves meningioma F1 by at most 0.08. Swapping the kernel, at most 0.06. Neither eliminates the confusion.

The residual error is a property of the representation, not the classifier — so progress requires attacking the glioma–meningioma boundary directly: tumor-region preprocessing, multi-slice or volumetric context, targeted fine-tuning. Not more backbone or kernel search.

08 — 証拠

Why RBF Loses

The usual case for RBF assumes the data isn't linearly separable. Global-average-pooled CNN features are 1,024–2,048-dimensional for only four classes, and the backbone was already trained to make a thousand-way distinction linearly decodable from exactly this layer. The premise doesn't hold.

学習曲線 / LEARNING CURVE — RESNET50 + LINEAR
Training accuracy is exactly 1.000 at every training-set size.
Including the smallest subsample. A linear hyperplane already fits the training data perfectly — so the extra capacity of a polynomial or RBF kernel has no fit left to improve. It can only change the inductive bias, add variance, and cost more. Cross-validated accuracy climbs from 0.885 to 0.948 across the same range.
assets/resnet_learningcurve_linear.png
検証曲線 / VALIDATION CURVE OVER C
Accuracy varies by under 0.002 across four orders of magnitude of C.
This validates leaving C at its default, and is itself further evidence of easy separability — a hard problem would show a pronounced regularization optimum. Accuracy is flat for every C above 0.1.
assets/resnet_valid_linear.png

The persistent gap between the perfect training score and the 0.948 cross-validation score marks a variance-limited, not bias-limited regime. The CV curve has not plateaued: more labeled data would still help. More kernel search would not.

RBF is not a bad kernel. It should not be the unexamined default when the input is a high-dimensional pretrained CNN embedding. Testing the linear kernel costs seconds and, here, won on both axes nine times out of nine.

09 — 資料

Dataset & Setup

Training partition
5,712
publisher's train folder
Fit / evaluation
4,569 / 1,143
internal 80/20 split
Reserved test folder
1,311
not consumed
Classes
4
glioma · meningioma · notumor · pituitary
License
CC0-1.0
Kaggle · masoudnickparvar
特徴抽出器 / BACKBONES — CHOSEN BY INDUCTIVE BIAS, NOT DEPTH
BACKBONE
INPUT
FEATURES
STRUCTURAL ASSUMPTION
DenseNet121
224²
1,024-d
Dense feature reuse
ResNet50
224²
2,048-d
Residual shortcuts, structural detail
InceptionV3
299²
2,048-d
Parallel multi-scale convolution
These three embed genuinely different structural assumptions. A depth ladder — ResNet18 → 50 → 152 — would only tell us that capacity helps. This choice means any difference in feature-space geometry is attributable to architectural inductive bias rather than to size alone.
実験規約 / PROTOCOL
Random state42
Identical row ordering across all three feature matrices — a paired comparison, not merely a matched one.
SVM defaultsC=1 · γ=scale
degree=3, coef0=0, one-vs-one multiclass. No per-configuration tuning.
PreprocessingCANONICAL ONLY
Bilinear resize plus each backbone's own ImageNet preprocessing. No augmentation, no skull stripping.
Runtime scopeFIT + PREDICT
Image loading, backbone forward passes, CSV I/O and plotting are all out of scope.
Extraction batch size1
~0.27–0.31 s per image on the 224² backbones. Known inefficiency, deliberately unfixed — it sits outside the measured quantity.
10 — 限界と展望

What These Figures Do Not Support

Stated plainly, because a benchmark that hides its caveats isn't one.

Internal split
The 1,143 evaluation images are an 80/20 split of the publisher's training partition. The official 1,311-image test folder was never consumed, so absolute accuracies are not comparable to published figures.
Possible patient leakage
3,064 slices in the collection come from just 233 patients. Slices of one patient can fall on both sides of an image-level split. Not ruled out.
Shared-tenancy timing
Colab, uncontrolled neighbour load, ±32% between replicates. All runtime claims are ordering claims. A dedicated machine with repeated runs would turn them into statistical ones.
Mechanism inferred, not measured
"RBF retains more support vectors" follows from solver complexity and fits the data, but support-vector counts were never logged.
One missing matrix
ResNet50 + RBF's confusion matrix was lost. Its ≈92% is the classifier's own two-digit report.
One dataset, one modality
T1-weighted contrast-enhanced axial slices from one public collection. Whether RBF domination generalizes is an open empirical question.
展望 / ROADMAP — ORDERED BY VALUE, NOT EFFORT
01
Migrate to the publisher's held-out partition
Extract features for the 1,311 reserved test images, refit on all 5,712, report on untouched data. This converts every accuracy above into an externally validated figure. Runtimes must be re-measured, not reused — n grows 25% and solver cost is superlinear in n.
02
Log support-vector counts per configuration
Converts the mechanistic explanation of the RBF penalty from inference into measurement. Cheapest high-value addition in the list.
03
Feature fusion + PCA
The notebook already implements horizontal concatenation across backbones (3,072-d pairs, 4,096-d, 5,120-d for all three) and PCA compression, producing 24 further confusion matrices. Implemented but not yet formally measured — and given the cost structure established here, the question is not just whether complementarity buys accuracy, but at what price.
04
Attack the glioma–meningioma boundary specifically
The only remaining source of meaningful error.
05
Repeat timing on a dedicated machine
Several runs per cell, reporting mean and standard deviation.
技術系統 / TECHNICAL STACK
PythonTensorFlow / Kerasscikit-learnDenseNet121ResNet50InceptionV3NumPyPandasMatplotlibGoogle Colab
Access Source ↗← Return to Index
版数情報 / BUILD MANIFESTDEPLOYED
Release版数·······································································v1.1.0
Interface画面系統·······································································NERV CONSOLE / REV 01
Channel配信区分·······································································STABLE
Statically generated Next.js build. Project records and reported figures reflect the state of each source repository at compile time; interactive models on the case files are illustrative reconstructions of published results, not live inference. Superseded builds are replaced in full at deploy.
© 2026 IKHSAN MOCHAMMAD NOOR記録は閲覧者の要求に応じて展開するEND OF CASE FILE