Part A — Debug. The strong move is the order of operations:
- Reproduce each failing test in isolation and capture its actual vs expected output.
- Classify each failure as loss-mismatch, label-mismatch, encoder-head wiring bug, or split/leakage bug.
- For the "known" bugs, look at the docstring / git history first.
- For "novel" bugs, write a minimal failing example before editing the model.
- Add a regression assertion per fix that genuinely distinguishes the bug from the fix (e.g., a test that asserts the label encoder returns
long not int).
Gotchas that sink candidates: chasing the "novel" frame and missing that a known bug is misdescribed; rewriting the encoder instead of testing whether the head consumes [CLS] vs pooled mean; failing to run all tests after each fix.
Part B — Train/evaluate. Show EDA that drives a decision: class counts, untruncated length histogram, truncation rate, duplicate rate. Pick a leakage-free split (stratified K-fold if K is small); compute class weights from train only; use linear warmup + cosine decay; gradient-clip at 1.0; dynamic padding; select on macro-F1 (resists imbalance) with a confusion matrix; report ROC-AUC when K=2. Pin seeds and torch.use_deterministic_algorithms(True) for reproducibility — and call out the run-time trade-off.