[2026-05-29] Solving an elusive bug: process and lessons learned
Leonardo Collado Torres
0:00 / 0:00
[2026-05-29] Solving an elusive bug: process and lessons learned
9 просмотров · 2 нед. назад
Leonardo Collado Torres
822 подписчика
9 просмотров · 2 нед. назад
For more information on Nick Eagles, check https://bsky.app/profile/nick-eagles.....
For more information about the materials discussed in this video, check https://docs.google.com/presentation/....
For more information about the LIBD rstats club, check https://bsky.app/profile/libdrstats.b....
Summary:
🐛 Debugging a Tricky Data Analysis Bug | Spatial Transcriptomics Case Study
Video Summary
🔬 **The Problem**: While analyzing HD spatial transcriptomics data comparing transcription inside vs. outside cells, unexpected results emerged - correlations were surprisingly low and spatial registration statistics seemed off.
🔍 **Investigation Journey**:
📊 Initial plots showed suspicious negative correlations where positive ones were expected
🧪 Pseudo-bulk PCA revealed an extreme batch effect by sample
❓ Confusingly, the clustering results looked visually correct!
💡 **The Culprit**: The `duckplyr` package! Unlike standard `dplyr`, it doesn't preserve row order after joins for performance reasons. When using `pull()` after a join, the cluster assignments got scrambled!
⚠️ **Key Lessons Learned**:
1. 🧠 Don't ignore your intuition - sometimes "unexpected" results are actually bugs
2. 📖 Read the documentation carefully - especially for package compatibility differences
3. 🔎 Question your assumptions when debugging
4. ✅ Always verify data alignment with visualization checks
🛠️ **Tools Discussed**: R, dplyr, duckplyr, spatial transcriptomics analysis, pseudo-bulking, PCA
Perfect for data scientists, bioinformaticians, and anyone working with spatial -omics data!
#DataScience #Bioinformatics #Debugging #RStats #SpatialTranscriptomics #CodingTips #DataAnalysis #FICTURE #spatialLIBD