ADAniketh Deshpandeinani-db.hashnode.dev·4d ago · 58 min readWhat Actually Happens When You Call spark.read? One Line of Python, a Thousand TasksTL;DR: spark.read.parquet(path) reads almost nothing. The real work starts at the first action, when Spark turns your DataFrame into four plans, slices your files into tasks, ships those tasks to exec00
ADAniketh Deshpandeinani-db.hashnode.dev·4d ago · 65 min readThe Distributed Compute Playground: Everything You Can Plug In and Tune in an Apache Spark JobTL;DR: An Apache Spark job is not one thing. It is a stack of nine layers, and almost every layer has parts you can swap: where executors come from, how the code is written, which engine runs the plan00
ADAniketh Deshpandeinani-db.hashnode.dev·6d ago · 47 min readThe Database Playground: Everything You Can Plug In and Tune in PostgreSQL and MySQL (Including AI Workloads)TL;DR: PostgreSQL and MySQL are not black boxes. Both let you swap storage engines, index types, durability guarantees, where data physically lives, and even add vector search and ML, often per table 00