I did not use overall accuracy as the main metric after identifying the imbalance. I evaluated the positive 30-day readmission class primarily using recall and F1 score, while also reporting precision and ROC-AUC. Precision alone was not the deciding metric either.
I applied random undersampling only to the training set and kept the test set at its original distribution. The tuned Random Forest + RUS achieved 17.10% precision, 63.06% recall, 26.90% F1, and 65.90% ROC-AUC.
My background is primarily in full-stack development, so this project was also my way of stepping outside my usual stack and understanding the ML workflow through implementation as deeply as I could — from preprocessing and class imbalance handling to model evaluation, PySpark, and clustering.
The main shift for me was from asking “how high is the accuracy?” to “is the model actually detecting the minority-class cases that matter?”