You are browsing as a guest. Sign up (or log in) to start making projects!

3h 37m 59s logged

DevLog 4: PD Biomarkers

Today was an absolutely massive workday! I spent 2 hours and 31 minutes fixing two layers of data leakage.

🛠️ Bug Fixes & Refactoring

  • Data Leakage Elimination: Previously, CPM normalization and the gene filter were running on the full dataset before the train-test split.
  • Pipeline Consolidation: Fixed this by nesting CPM normalization, gene filtering, and Boruta feature selection inside the cross-validation fold loops.
  • Notebook Cleanup: Consolidated everything into one clean notebook: 03_nested_cv_evaluation.ipynb.

📊 Model Performance (Leak-Free Pipeline)

When I ran the leak-free pipeline, Elastic Net blew everything else out of the water:

  • Elastic Net: 0.992 ± 0.010 AUC-ROC
  • Random Forest: 0.962 AUC-ROC
  • XGBoost: 0.938 AUC-ROC

🧪 The Proof (Permutation Testing)

To make sure the 0.992 wasn’t a ghost score, I ran a permutation test with shuffled labels.

  • AUC Collapse: Dropped to chance level (0.43 - 0.48).
  • Feature Reduction: Boruta features dropped from 205 down to just 9 per fold.
  • Conclusion: This proves the biological signal is 100% real.

Next Steps & Deliverables

  • Batch Effects: Skipped ComBat for now since Phase 1 is just a single cohort. I will check the GEO metadata for internal batch run dates tomorrow.
  • Diagnostic Plots: Wrapped up the day by generating a clean set of plots, including ROC (real vs. shuffled), PCA, and a SHAP plot flagging hsa-mir-23a as the top biomarker.
0
1

Comments 0

No comments yet. Be the first!