Backend
Finding Exoplanets in Noisy Data with Machine Learning
Tanmaya sree Chirra DEV Community
2 views
We Built an AI to Hunt Earth-Like Planets — Here's How
Finding planets around other stars is hard. Kepler gives us raw light curves — brightness measurements over time — and buried inside that noisy data are tiny dips caused by planets
crossing their star. We built Astrobit 1.0 to find them automatically.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
The Problem
Kepler's Simple Aperture Photometry (SAP) flux is messy. Instrumental systematics, cosmic rays, and quarter-boundary artifacts all look like signals. A naive threshold approach
misses real planets and flags false positives constantly.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Architecture
Raw SAP Flux
↓
[Cleaner] — mask bad cadences + sigma-clip outliers
↓
[Detrender] — per-quarter Savitzky-Golay filter
↓
[BLS Search] — 50k coarse grid → fine refinement → alias check
↓
[Feature Extractor] — SDE, depth, SNR, odd/even, secondary eclipse
↓
[Random Forest Classifier] — trained on 269 labelled stars
↓
[Platt Scaler] — calibrates scores to probabilities
↓
[Vetter] — secondary eclipse, odd/even, recurrence, systematics
↓
Ranked Candidates → submission.csv
Each stage is independent and cacheable — BLS results are cached to CSV so you can interrupt and resume without recomputing.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
What We Built
Cleaning — mask bad cadences, sigma-clip outliers from raw SAP flux
Detrending — per-quarter Savitzky-Golay filter. Recovers ~90% of true transit depth vs ~33% with a running median, and avoids edge artifacts at Kepler's quarterly roll boundaries
BLS Period Search — 50k log-spaced coarse grid + 600-point fine refinement around each peak. Alias checking at 0.5x, 1x, 2x, 3x catches period harmonics. ~100x cheaper than full-
resolution search
Feature Extraction — SDE, transit depth, SNR, odd/even depth ratio, secondary eclipse depth
Random Forest Classifier — trained on 269 labelled stars, replaces brittle single-SDE-threshold with a multi-feature decision boundary
Platt Scaling — calibrates raw model scores to actual probabilities on the dev set
Vetting — secondary eclipse check, odd/even depth consistency, per-quarter recurrence, known systematic period filtering
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Results
Stage
Impact
SG detrending vs running median
~90% vs ~33% transit depth recovery
Coarse-to-fine BLS
~100x faster than full-resolution search
Alias checking
Catches period harmonics at 0.5x–3x
RF classifier vs SDE threshold
Multi-feature boundary, fewer false positives
Platt scaling
Calibrated confidence scores on dev set
Vetting layer
Filters secondary eclipses, systematics, odd/even inconsistencies
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Lessons Learned
• Detrending matters more than the classifier. A bad detrend corrupts every downstream feature. We spent more time on the SG filter than the ML model — worth it.
• Cache everything. BLS on a full Kepler star takes time. Caching to CSV saved us hours during iteration.
• Single thresholds break. SDE alone is a terrible classifier. The moment we switched to a multi-feature Random Forest, false positive rate dropped significantly.
• Calibration is underrated. Raw model scores are not probabilities. Platt scaling on the dev set made our confidence scores actually trustworthy for ranking candidates.
• Vetting is not optional. The classifier catches most false positives, but secondary eclipse checks and odd/even consistency are cheap and eliminate a whole class of eclipsing
binary contamination.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Stack
Python · scikit-learn · lightkurve · scipy · numpy
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Repo
🔗 github.com/25wh1a6678-art/Astrobit_1.0
Read original: https://dev.to/tanmaya_sreechirra_d568b/finding-exoplanets-in-noisy-data-with-machine-learning-25d0
← Previous
Cloudflare Zero Trust: Enterprise Access Security Guide
Next →
AI Agents That Verify Their Own Output: The 30-Minute Validation Loop That Beats 30 Days of AI Code Review
Related
The benchmark everyone cited for "fastest backend" was archived in March, and the reason is the interesting part
Backend
0
Dev.to (EN Zone)
How to automatically find the batch size when using Accelerate with FSDP2? [D]
Backend
1
Reddit r/MachineLearning
Cross-Chain Bridge Risk Assessment: Aave V3
Backend
3
DEV Community
My guard asserted the output path was a scratch directory. The row landed in the permanent log anyway.
Backend
4
DEV Community
Comments0
No comments yet — be the first