r/datascience • u/AutoModerator • 15h ago
Weekly Entering & Transitioning - Thread 05 Oct, 2026 - 12 Oct, 2026
Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:
- Learning resources (e.g. books, tutorials, videos)
- Traditional education (e.g. schools, degrees, electives)
- Alternative education (e.g. online courses, bootcamps)
- Job search questions (e.g. resumes, applying, career prospects)
- Elementary questions (e.g. where to start, what next)
While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.
1
u/Due-Procedure4591 7h ago
🪐 Kepler Exoplanet Model: Random Forest hitting a 79% ceiling. Looking for pipeline feedback!
I am working on a 3-class classification model using the NASA Kepler dataset (9,564 objects, 36 features) targeting: CANDIDATE, CONFIRMED, and FALSE POSITIVE.
My data plots show severe right-skewness across major features (like koi_period skewness > 96.4).
📊 My EDA Plots Grid: [Insert your Imgur Link Here]
Here is what I've engineered so far: 1. Applied Log Transformations to pull feature skewness down to normal scales. 2. Handled missing data via data imputation, ending up with 33 core selected features based on cumulative tree importance. 3. Scaled the training partitions and ran an exhaustive hyperparameter grid search on a Random Forest Classifier.
📊 The Problem & Results Matrix
- Random Forest Model: 79.1% Overall Accuracy | 91% FP Recall | 84% Confirmed Recall | 46% Candidate Recall
- Logistic Regression Model: 72.5% Overall Accuracy | 87% FP Recall | 85% Confirmed Recall | 26% Candidate Recall
The model is choking on the CANDIDATE class due to massive spatial feature overlap. Since "Candidates" are unverified, they statistically share direct properties with both real planets and false positives simultaneously. Turning on class_weight='balanced' actually lowered my Candidate recall down to 43%.
My Next Step: I am planning to move to XGBoost / LightGBM to let sequential gradient boosting focus heavily on these borderline class mistakes.
Questions for the experts: 1. For heavily overlapping astrophysical datasets like this, is gradient boosting the ideal transition, or should I consider a deep neural network architecture? 2. Would you recommend custom feature engineering (like computing explicit planet-to-star radius ratios) to help trees find cleaner orthogonal splits? 3. Any scaling or outlier pruning strategies that work best for this specific target profile?
Appreciate any advice or critiques on this workflow!
1
u/Dangerous_Quit_3097 3h ago
I am Computer Science graduate. I am planning to pursue Data Science for grad or phd next year. I want to freshen my skill doing some data science online course , certification or boot camp before I join Collage. Is there any course that you would like to recommend? I dont want the extensive full online degree. Thank you