Good explainer. The operational side I'd add: set autovacuum_vacuum_scale_factor (and threshold) per-table for your hottest tables instead of relying on the global defaults — defaults are tuned for average workloads, not hundreds of millions of rows. Two other classics: fillfactor below 100 on heavily-updated tables is what makes HOT updates possible in the first place (no free space, no HOT), and a single long-running transaction can block vacuum for the whole table — pg_stat_activity is the first place I look when n_dead_tup keeps climbing. Monitoring pg_stat_user_tables.n_dead_tup + last_autovacuum per table turns this from mystery into routine.