DevLog 4: PD Biomarkers
Today was an absolutely massive workday! I spent 2 hours and 31 minutes fixing two layers of data leakage.
🛠️ Bug Fixes & Refactoring
- Data Leakage Elimination: Previously, CPM normalization and the gene filter were running on the full dataset before the train-test split.
- Pipeline Consolidation: Fixed this by nesting CPM normalization, gene filtering, and Boruta feature selection inside the cross-validation fold loops.
-
Notebook Cleanup: Consolidated everything into one clean notebook:
03_nested_cv_evaluation.ipynb.
📊 Model Performance (Leak-Free Pipeline)
When I ran the leak-free pipeline, Elastic Net blew everything else out of the water:
- Elastic Net: 0.992 ± 0.010 AUC-ROC
- Random Forest: 0.962 AUC-ROC
- XGBoost: 0.938 AUC-ROC
🧪 The Proof (Permutation Testing)
To make sure the 0.992 wasn’t a ghost score, I ran a permutation test with shuffled labels.
- AUC Collapse: Dropped to chance level (0.43 - 0.48).
- Feature Reduction: Boruta features dropped from 205 down to just 9 per fold.
- Conclusion: This proves the biological signal is 100% real.
Next Steps & Deliverables
- Batch Effects: Skipped ComBat for now since Phase 1 is just a single cohort. I will check the GEO metadata for internal batch run dates tomorrow.
-
Diagnostic Plots: Wrapped up the day by generating a clean set of plots, including ROC (real vs. shuffled), PCA, and a SHAP plot flagging
hsa-mir-23aas the top biomarker.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.