CLUSTERED BY vs ZORDER BY in Databricks
Understanding Data Distribution vs Storage Optimization
In large-scale Spark workloads, performance problems often lead to one of two suggestions:
“Let’s bucket the table.”
“Let’s ZORDER it.”
Both
bijudevassy.hashnode.dev5 min read