Học data platform engineering (10)
PHASE 20 - Batch & Distributed Compute
Môi trường: Spark (local + một cluster nhỏ), Flink (từ P19), Trino, DuckDB + Polars, Ray (local); dữ liệu lớn thật (TPC-DS, hoặc dataset công khai vài chục GB);
yentrinh.hashnode.dev53 min read