Pipeline mind map
Standard pipeline template for a tabular competition. Not yet specialised by an editor.
Data processing
- Read the CSVs, find id and target
- Encode categorical columns
- Handle missing values
- Scale / transform skewed numbers
Model design
- Logistic/linear baseline
- Gradient boosting (LightGBM, XGBoost, CatBoost)
Training
- K-fold cross-validation
- Early stopping
- Save out-of-fold predictions
Post-processing
- Match the sample_submission format
- Blend models by CV
- Choose finals by CV
Metric
Root Mean Squared Logarithmic Error.
Public baselines
Most-voted public notebooks, refreshed daily from the Kaggle API (last 2026-10-10).
- Comprehensive data exploration with Python ▲33699
- House Prices Prediction using TFDF ▲15661
- Stacked Regressions : Top 4% on LeaderBoard ▲14383
- Handling Missing Values ▲7893
- Welcome to Data Science in R: Workbook ▲7833
- Feature Engineering for House Prices ▲6679
- Regularized Linear Models ▲5522
- [py] T1-4. 왜도와 첨도 구하기 (로그스케일) Expected Questions ▲4016
- Submitting From A Kernel ▲3869
- Selecting and Filtering in Pandas ▲3755
- XGBoost ▲3734
Discussion 0
No discussion yet. Ask the first question or post a team-up for this competition.
Sign in with PocketPlay
Sign-in opens soon. Until then the forum is read-only.
Same account as pocketplay.win (Google or username). Posting needs a Google-linked account.
Markdown: **bold**, *italic*, `code`, ``` code blocks, - lists, > quotes, [text](https://link). No HTML.