Pipeline mind map
Specialised to this competition from the overview, data page and public notebooks.
- Read train/test; target = satisfaction (True/False → 1/0)
- 4 categorical columns → one-hot or native categories
- 13 rating columns: keep as ordered numbers
- Delays are right-skewed: try log1p
- Flag missing Arrival Delay; try Arrival − Departure delay
- Sanity baseline: logistic regression
- Main baseline: LightGBM / XGBoost / CatBoost
- Later: RealMLP or TabPFN (seen in top public notebooks)
- StratifiedKFold, 5 folds
- Early stopping on AUC
- Save out-of-fold (OOF) predictions
- Compare CV with the public leaderboard
- Submit probabilities, not 0/1
- AUC only cares about ranking: rank-average blends
- Check header: id,satisfaction
- Pick final submissions by CV, not public LB
Key facts
- Timeline
- Started Oct 1, 2026; entry, team-merger and final submission deadline Oct 31, 2026, 11:59 PM UTC ↗
- Task
- For each id in test.csv, predict a probability for satisfaction ↗
- Metric
- Area under the ROC curve between predicted probability and the observed target ↗
- Data
- train.csv, test.csv, sample_submission.csv (102.54 MB, CSV, CC BY 4.0). Synthetic, generated to resemble the Airline satisfaction dataset ↗
- Size
- Train 699,635 rows × 23 columns, test 299,844 × 22; 21 features; 292 missing values, all in train ↗
- Target balance
- 55.64% not satisfied, 44.36% satisfied (train) ↗
- Prizes
- Kaggle merchandise for places 1–3, awarded once per person in the series; no points or medals ↗
- Participation (Oct 9)
- 2,768 entrants, 1,261 teams, 10,137 submissions ↗
Problem breakdown
One passenger's trip: who they are, the flight (class, distance, delays) and their ratings of 13 services.
A probability that satisfaction is True, for each id in test.csv.
Clean CSVs, one target, a standard metric, many public notebooks, no medals at stake, and a fresh episode every month to try again.
292 missing values, skewed delay columns, and the temptation to tune on the public leaderboard instead of your own cross-validation.
The 21 features
Age, Flight Distance, Departure Delay in Minutes, Arrival Delay in Minutes
Inflight wifi service, Departure/Arrival time convenient, Ease of Online booking, Gate location, Food and drink, Online boarding, Seat comfort, Inflight entertainment, On-board service, Leg room service, Baggage handling, Checkin service, Cleanliness
Gender, Customer Type, Type of Travel, Class
Column names from a public EDA notebook ↗
Metric: ROC AUC
AUC is the probability that a randomly chosen satisfied passenger gets a higher predicted score than a randomly chosen unsatisfied one. 0.5 is a coin flip and 1.0 is perfect. Only the order of your predictions matters, so rescaling or calibrating probabilities does not change the score, while rounding to 0/1 throws ranking information away.
For scale: one public XGBoost notebook reports 5-fold CV AUCs of about 0.959–0.961, and top public notebooks show leaderboard scores around 0.961–0.962 (Oct 9). Gains near the top are in the fourth decimal place. ↗
Public baselines
Snapshot of the competition's Code tab sorted by votes, taken Oct 9, 2026. Scores are public-leaderboard scores shown on Kaggle.
| Notebook | ▲ | LB | Why open it |
|---|---|---|---|
| PS|S6|E10: RealMLP · PyTabKit | 65 | 0.96105 | Most-voted notebook: a single neural tabular model (RealMLP via pytabkit). |
| S6E10 LightGBM CV 5 folds 0.96111 | 46 | 0.96069 | Plain LightGBM with 5-fold CV: the best first fork. |
| EDA, Feature Engineering & XGBoost | 30 | 0.95932 | Readable EDA (shapes, target balance, feature groups) then XGBoost with StratifiedKFold. |
| S6E10 | One LightGBM | LB : 0.96017 | 24 | 0.96024 | One model, no stacking: good to see how far a single GBDT gets. |
| Single CatBoost 🐱 EDA | PS6E10 | 11 | 0.96046 | CatBoost handles the categorical columns natively. |
| Airline | LGBM/CatB/XGB/RealMLP| Baseline | 21 | 0.96075 | Side-by-side baselines of four model families. |
| S6E10: What Each Step Was Worth | 14 | 0.96153 | Shows the score change of each step: useful for learning what matters. |
| Cleared for Takeoff: LB 0.9606 + What Failed | 13 | 0.96056 | Also lists what did not work. |
| Stacking TFMs with Public OOF | 37 | 0.96176 | Stacking of out-of-fold predictions: step after you have 2–3 models. |
| S6E10 | TabPFN + Route Categories | LB 0.96160 | 29 | 0.96164 | Tabular foundation model (TabPFN) plus engineered route categories. |
Most-voted notebooks right now
Refreshed daily from the Kaggle API (last 2026-10-09). Not reviewed by an editor; the table above is the reviewed list.
- PS|S6|E10: RealMLP · PyTabKit ▲65
- S6E10 | Are You Satisfied ▲48
- S6E10 LightGBM CV 5 folds 0.96111 ▲46
- Can't Get No Satisfaction. Delays I try and try 😂 ▲43
- Stacking TFMs with Public OOF ▲37
- ✈️ Flight Log S6E10 | Captains + RealMLP Stack ▲36
- Secrets to Satisfaction (It's Stacking!) ▲31
- EDA, Feature Engineering & XGBoost ▲30
Step-by-step plan
- Day 1 — join and read
Sign in, accept the rules, read Overview and Data. Open sample_submission.csv to see the exact format.
- Day 1 — first submission
Fork a simple LightGBM notebook (for example the 5-fold LightGBM one below), run it, and submit. Getting any score on the board is the goal.
- Days 2–3 — your own CV
Rebuild the baseline yourself: StratifiedKFold with 5 folds, AUC per fold, saved OOF predictions. Write down CV and public scores for every run.
- Days 4–7 — features
Try one idea at a time: log1p on delays, a missing-delay flag, delay difference, rating sums. Keep a change only if CV improves.
- Week 2 — second model
Train CatBoost or XGBoost with the same folds. Rank-average its OOF with LightGBM and check CV.
- Week 3 — read and learn
Read “What each step was worth” and a stacking notebook. Copy one idea, measure it, and note what failed.
- By Oct 31, 23:59 UTC
Select your final submissions by CV score, not public leaderboard rank, and write a short notebook of what you learned.
Discussion 0
No discussion yet. Ask the first question or post a team-up for this competition.
Sign-in opens soon. Until then the forum is read-only.
Same account as pocketplay.win (Google or username). Posting needs a Google-linked account.