Pipeline mind map
Standard pipeline template for a text competition. Not yet specialised by an editor.
Data processing
- Clean and tokenize text
- Check label balance
Model design
- TF-IDF + linear baseline
- Pretrained transformer
Training
- Folds by label
- Small learning rate, few epochs
Post-processing
- Calibrate / threshold for the metric
- Ensemble seeds
Metric
Categorization Accuracy. Accuracy: share of rows predicted exactly right.
Public baselines
Most-voted public notebooks, refreshed daily from the Kaggle API (last 2026-10-10).
- Tutorial Notebook ▲1820
- KerasNLP starter notebook Contradictory DearWatson ▲515
- Text-Representations ▲435
- Watson :: XLM-R & NLI :: inference ▲256
- Text Analysis + Topic Modeling with spaCy & GENSIM ▲247
- Basics of BERT and XLM-RoBERTa - PyTorch ▲218
- TPU Sherlocked: One-stop for 🤗 with TF ▲212
- All about Bert you need to know ▲180
- Using Google Translate for NLP Augmentation ▲136
- Contradictory Watson: Concise Keras XLM-R on TPU ▲107
- More NLI datasets - Hugging Face nlp library ▲95
- Watson - KFold XLM-R + Translation Augmentation ▲82
Discussion 0
No discussion yet. Ask the first question or post a team-up for this competition.
Sign in with PocketPlay
Sign-in opens soon. Until then the forum is read-only.
Same account as pocketplay.win (Google or username). Posting needs a Google-linked account.
Markdown: **bold**, *italic*, `code`, ``` code blocks, - lists, > quotes, [text](https://link). No HTML.