Databricks’ Latest Database Innovation: Fact or Fiction?
When Databricks asserts it has solved an enduring database challenge, it backs it up with a striking promotional statement: “One data, zero compromises, zero copies.” This naturally ignites curiosity among tech enthusiasts. Supposedly, the company has integrated OLTP and OLAP without duplication. Databricks, the creator of the Apache Spark framework, labels its invention LTAP—lake transactional/analytics processing—featuring Reyden, a new compute engine, and Lakebase, its serverless PostgreSQL hosted on open object storage.
Addressing the Database Puzzle
Databricks seeks to resolve a fundamental database issue. OLTP (online transactional processing) focuses on small, row-oriented reads and numerous writes, whereas OLAP (online analytical processing) emphasizes large, column-oriented reads and batch writes. Combining these two within a single system poses a significant challenge. With AI applications on the rise in business contexts, addressing this issue is more pressing than ever.
What’s the Assertion, Though?
The promotional narrative indicates that Databricks does not attempt to stuff OLTP and OLAP workloads into a single engine or conceal the underlying complexities. Rather, it integrates data at the storage level, consolidating transactions, analytics, streaming, and operational data into a single storage instance within their data lakehouse—Databricks’ phrase for a blend of data lakes and data warehouses.
Zero Copies – An Exaggeration?
Does Databricks genuinely provide “zero copies” of data, as highlighted in promotional materials and a Forbes CEO discussion? Not exactly. The transactional aspect of LTAP builds on Databricks’ first managed PostgreSQL database, Lakebase, utilizing technology from Neon, which Databricks acquired last year for copy-on-write branching and autoscaling serverless computing.
Insights from Databricks Engineers
During a PostgreSQL conference in May, Databricks engineers disclosed insights. Under the topic “Analytics directly on OLTP data,” the pageserver handles storage while Spark’s analytics executor retrieves complete page images from object storage’s image layers. Some insiders at Databricks concede that, technically, there are two copies: pageservers acting as a cache or materialization layer, and object storage for OLTP and OLAP.
SingleStore’s Counter to Databricks’ Claims
SingleStore, eager not to be overshadowed, quickly rebutted Databricks’ HTAP failure assertions. “You can’t label HTAP as a failure and then rave about its advantages. LTAP is merely a fresh coat of paint on the same old engine,” remarked SingleStore’s CTO, Nadeem Asghar. “Three engines on top, each with its own cache, freshness validation, and possible point of failure. Different data formats necessitate maintaining their synchronization.”
Other Competitors in the Arena
Many companies have attempted to align analytics with transactional systems. MongoDB provides column-store indexes for analytical queries, Oracle’s HeatWave for MySQL allows transactional applications to conduct analytics without needing specialists like Teradata or Snowflake, and SAP has been advocating real-time analytics since 2011 with its HANA in-memory database.
Databricks’ View on Copies
Databricks maintains that their “zero copy” assertion is valid by eliminating the requirement for multiple synchronized authoritative data copies. “In LTAP, users manage only one authoritative data copy. It serves as an Iceberg (open source table format) source of truth filled with Parquet files. Yes, akin to any database, there are numerous intermediate internal copies, from memory caches to the storage hierarchy,” explained a Databricks representative.
Remarkable Engineering Amongst Bold Assertions
Despite the promotional hype, Databricks has demonstrated commendable engineering prowess with its Reyden execution engine. As noted by Andy Pavlo from Carnegie Mellon, Reyden efficiently processes PostgreSQL pages. “The engine interprets PostgreSQL pages, a sophisticated task, enabling faster analytics without the waiting period associated with S3.” While critics may emphasize overhyped claims, Databricks’ engineering efforts are indeed impressive.
Conclusion: Zero Copies or Zero Logic?
Databricks may have developed some advanced technology, effectively merging transactional and analytic workloads. However, with marketing claims like “zero copies,” they navigate a precarious path. Promotional narratives may obscure these engineering accomplishments. Stay tuned to gadgetlad.co.uk for additional insights!