Denodo Architectures: Data Lake and Lakehouse
You can translate the document:This document shows an architecture pattern for a Data Lake / Lakehouse.
Data Lake and Lakehouse
The Data Lake/Lakehouse plays a central role in the data management architecture in this pattern, and it is being used in many cases to replace the data warehouse.
- Access to All data:
The semantic layer in this case connects to the data lake engines (for example Spark, Impala, etc.) and federates the access to data outside the data lake, either in the data warehouses, in enterprise applications or even in SaaS applications.
This extends the scope of the data lake, because thanks to the Denodo semantic layer the user can get access to all enterprise data, being a universal semantic layer on top of the lakehouse. The user is not forced to copy all data into the lakehouse, something that, in practice, in many cases it is very difficult or impossible to do (for example different jurisdictions, legacy applications that can't be replaced, etc.). With this semantic layer the user can leverage all data regardless of the data origin. - Metadata-based Semantic Layer for accelerating the generation of data products:
The user can benefit from using a metadata based layer, having more agility for building data products thanks to its declarative approach and also because of its better reusability, as it is much easier to reuse existing data products than building new ones, compared to traditional data pipelines where usability is very hard if not impossible.
In addition it`s much easier to accommodate any change to deal with new business requirements with much shorter data product development cycles.
The semantic layer becomes the data foundation for AI applications such as AI chatbots and agentic apps. - Virtual gold/platinum zones for last mile customization:
We recommend to follow a virtual approach especially in the gold data zone in the lakehouse, or in other layers on top of it, such as the Platinum zone, for providing the last mile customization to accommodate and fulfill any business requirement.
The user can fully substitute the physical gold layer or apply a mixed approach with some tables persisted and other virtualized, or leave the gold zone persisted and then use the virtual approach for subsequent layers for further personalization to accommodate specific business requirements in a more agile way. - Truly Zero-copy Architecture:
The Denodo semantic layer provides a truly Zero copy architecture, Denodo federates the access to operational applications, software as a service applications, mainframes, data lakes/data warehouses in other clouds without copying the data. - Data Materialization Support:
Even when data products can be built in a virtual way they can be persisted at any time as Parquet files or as open table formats such as Delta lake or Iceberg. - ELT for data uploading to the data lake and generation of lake zones:
Denodo can be used to apply any required transformation for building new data zones in the lakehouse. Denodo supports ELT transformations, fully delegating the data processing to the data lake.
Denodo can also be used to populate data into the Data Lake. - Unified Security and Governance layer:
Another key advantage of the Denodo semantic layer is that it provides a unified security and governance layer that expands not only to data in the data lake but also in any other data source in the company.
Denodo leverages policies already defined in data Lake catalogs such as Unity or Polaris by pushing down user credentials to them, and it's able also to govern data outside the data lake, being an universal Security and governance layer that can be used across the whole organization. When policies are defined in external enterprise government tools Denodo can import those policies and enforce them in the Denodo semantic layer. - Embedded Lakehouse Engine:
The Denodo semantic layer includes a lake house engine based on Presto, the Denodo Lakehouse Accelerator, that can be used either as a standalone solution, or to complement other data lake engines, with the goal of reducing the total cost of ownership.
A typical approach is to off-load the traditional analytical workload from other data lake engines, given that Presto is very efficient for executing this type of workload, leaving heavy lifting data processing based on spark data pipelines in other engines such as Databricks, MS Fabric, etc., which are specialized on Spark processing. The Denodo Lakehouse Accelerator includes Velox for high performance data processing in modern vectorial SIMD processors. - Performance boost for analytical workloads in the data lake:
The user can leverage Denodo sophisticated query optimization capabilities for improving performance with workloads requiring data within and outside the data lake. Smart query acceleration based on summaries (partial aggregations of data that are reused by the optimizer to accelerate the execution of future queries) can achieve sub-second execution times. Denodo includes recommendations based on artificial intelligence to help the user identify the best summaries to persist. - Self-Service:
Data marketplace for facilitating self-service, discovery, exploration and collaboration between business users through advanced GenAI-powered interfaces.
When to use this Architecture Pattern
This architectural pattern is commonly found nowadays in modern cloud data architectures where the lakehouse is the central piece of the data management strategy.
Apply a semantic layer for achieving a Zero-copy architecture and gaining more agility for building and delivering data products to users. Building the gold or platinum zones as virtual views in the semantic layer offers significant advantages in terms of agility for quickly reacting to accommodate new business requirements. Reduce costs lowering data engineering efforts and infrastructure costs by offloading part of the workload to the Denodo Lakehouse Accelerator engine. The Denodo Data Marketplace facilitates self-service, data discovery and exploration to business users through advanced GenAI-powered interfaces.
References
Logical Data Management Reference Architecture
Denodo Architectures: Data Lake and Lakehouse
Denodo Architectures: Operational Workloads
The information provided in the Denodo Knowledge Base is intended to assist our users in advanced uses of Denodo. Please note that the results from the application of processes and configurations detailed in these documents may vary depending on your specific environment. Use them at your own discretion.
For an official guide of supported features, please refer to the User Manuals. For questions on critical systems or complex environments we recommend you to contact your Denodo Customer Success Manager.
