FFFred Fenginfredfeng.hashnode.dev·3d ago · 7 min readVortex TSDB: Live Metrics in One CommandHTTP Writes, Minute Aggregates, Sliding Windows, Replicated in Memory 1. Overview Lightweight. Replicated. Real-time. Vortex TSDB is a distributed time series database for live metrics. Push numbers00
VSvikas sharmainvkdigiservices.hashnode.dev·Sep 22 · 4 min readBest Data Lake for Modern Cloud Data ManagementThe amount of data generated by modern businesses continues to expand. Applications, websites, business platforms, connected systems, and digital services can all produce information that may become u00
VSvikas sharmainvkdigiservices.hashnode.dev·Sep 22 · 4 min readBest Data Lake Solutions for Growing Business DataEvery growing business eventually reaches a point where its data becomes too diverse to manage through isolated storage systems. Customer information, application logs, documents, media files, backups00
RKRupak Kulkarniindata-to-decisions-tech.hashnode.dev·Sep 12 · 5 min readFrom Data Engineering to AI: How Modern Data Platforms Enable Smarter Business DecisionsArtificial intelligence is changing the way organizations make decisions, automate processes, and interact with customers. But behind every successful AI initiative is something much less glamorous an00
SPSudhanshu Prajapatiinblog.altimate.ai·Sep 11 · 21 min readBeyond the Databricks Cost Calculator: 6 Levers That Actually Cut Your Bill The official Databricks price calculator provides you information on the DBU rate times the hours you plan to run, the instance type for cloud providers, the costing of VMs (basically compute types) f00
Aaslangamzenur079ingamzenuraslan.hashnode.dev·Sep 11 · 4 min readI Finally Understood Why We Need OLTP, OLAP, Data Lakes and LakehousesLately I have been spending more time on data engineering and one question kept coming back to me Where should the data actually live At first this sounded like a simple storage decision. Then I reali00
AVAtul Vishwakarmainatulcodes.hashnode.dev·Aug 25 · 9 min readMapReduce Explained: How Google Parallelized Computation Across Thousands of MachinesThis post is a summary and discussion of "MapReduce: Simplified Data Processing on Large Clusters" by Jeffrey Dean and Sanjay Ghemawat (Google, Inc.), presented at OSDI 2004. All credit for the origin00
DRDor Rafael Avrahamofinoptipipe.hashnode.dev·Aug 8 · 4 min readWhy I stopped guessing at Spark and dbt config valuesI've spent more than a decade building data pipelines, and the part nobody warns you about isn't the pipeline logic. It's the tuning. Executor memory, shuffle partitions, cluster size, thread counts. 00
AAAfzal Ahmedincodelabuk.hashnode.dev·Jul 25 · 7 min readArrow Flight Serialization: Bridging High-Performance Computing with Data AccessWhat is Arrow Flight? Arrow Flight is a high-performance data transport framework built on Apache Arrow, designed to efficiently move large datasets between systems, processes, and services. Unlike tr00
FIfor IT theinieee-projects.hashnode.dev·Jul 26 · 3 min readBig Data Projects for Final Year Students: A Quick GuideIf you're a final year Computer Science or Engineering student looking for a project that's both academically strong and genuinely relevant to industry, Big Data is worth serious consideration. With o00