Pipeline mind map
Standard pipeline template for a tabular competition. Not yet specialised by an editor.
Data processing
- Read the CSVs, find id and target
- Encode categorical columns
- Handle missing values
- Scale / transform skewed numbers
Model design
- Logistic/linear baseline
- Gradient boosting (LightGBM, XGBoost, CatBoost)
Training
- K-fold cross-validation
- Early stopping
- Save out-of-fold predictions
Post-processing
- Match the sample_submission format
- Blend models by CV
- Choose finals by CV
Metric
Categorization Accuracy. Accuracy: share of rows predicted exactly right.
Public baselines
Most-voted public notebooks, refreshed daily from the Kaggle API (last 2026-10-10).
- Titanic Tutorial ▲61430
- Titanic Data Science Solutions ▲40137
- Introduction to Ensembling/Stacking in Python ▲15486
- A Data Science Framework: To Achieve 99% Accuracy ▲14000
- Exploring Survival on the Titanic ▲11042
- Titanic competition w/ TensorFlow Decision Forests ▲9041
- A Journey through Titanic ▲7522
- An Interactive Data Science Tutorial ▲6378
- EDA To Prediction(DieTanic) ▲6346
- [py] T1-1. 이상치를 찾아라(IQR활용) Expected Questions ▲6135
- upura-kaggle-tutorial-01: first submission ▲6022
Discussion 0
No discussion yet. Ask the first question or post a team-up for this competition.
Sign in with PocketPlay
Sign-in opens soon. Until then the forum is read-only.
Same account as pocketplay.win (Google or username). Posting needs a Google-linked account.
Markdown: **bold**, *italic*, `code`, ``` code blocks, - lists, > quotes, [text](https://link). No HTML.