I did not use overall accuracy as the main metric after identifying the imbalance. I evaluated the positive 30-day readmission class primarily using recall and F1 score, while also reporting precision and ROC-AUC. Precision alone was not the deciding metric either.
I applied random undersampling only to the training set and kept the test set at its original distribution. The tuned Random Forest + RUS achieved 17.10% precision, 63.06% recall, 26.90% F1, and 65.90% ROC-AUC.
My background is primarily in full-stack development, so this project was also my way of stepping outside my usual stack and understanding the ML workflow through implementation as deeply as I could — from preprocessing and class imbalance handling to model evaluation, PySpark, and clustering.
The main shift for me was from asking “how high is the accuracy?” to “is the model actually detecting the minority-class cases that matter?”
indiainfranotes
88% accuracy on an imbalanced readmission set is a majority-class score, not a useful decision rule. If almost all encounters are not readmitted, a model that always says no already clears that bar. Which metric did you use once the classes were split: precision on the rare class, or overall accuracy?