Modern Data Platform | DataOps Day
The goal of modern data platform is to facilitate an organization’s decision-making process based on analytical, transactional, raw, clinical & non-clinical data and to enable data scientists and analysts to extract significant insights from the data. A modern data platform also provides its stakeholders with the opportunity to consume and add value to data.
In this session, Sameer Paradkar will demonstrate how to design a user-friendly, cost-effective, highly available and scalable modern data platform to provide orchestration and monitoring, data governance, deployment automation, access control and auditing and orchestration of data pipelines.
Transcript
Hi, uh, this is Samir. Uh, welcome to the DataOps Day. Uh, I'm gonna present a session on Modern Data Platforms.
Uh, just to give, give you my introduction. I'm Samir. I'm an enterprise architect, uh, digital group of, um, I provide architecture and technology leadership, uh, for, um, our strategy customers and large deeds.
Um, I'm based out of Mumbai. Uh, so, uh, let's begin the session. Uh, I'll, I'll, I'll go to the next slide.
Uh, this is basically the agenda for, uh, the modern data platforms that I'm gonna basically, uh, discuss. Uh, so, uh, the initial set of things would include, um, uh, things around the data-driven economy, uh, the overall limitations of, uh, or the traditional limitations of the deployment models that we have, uh, in, on the, uh, as part of the legacy platform. Uh, then what are the typical, let's say, drivers, uh, that, uh, drivers for the modern data platforms?
Uh, obviously they're the ones that, uh, is a precursor to building, uh, all the analytics capabilities that you have, uh, as part of your, as part of your digital landscapes, uh, digital, uh, architectures that you have. And then, um, a comparison of, uh, three key, um, uh, architectures. Uh, the data warehouse versus the data lake versus the data lake house.
Uh, these are the ones that are predominant in, um, the current era. And, uh, just to give you a, let's say a brief, um, overview of what these architectures are, and then the various patterns that we have as part of, uh, the data architecture. So there's alarm data versus Kapa versus Delta and, and where they would basically be leveraged as part of your landscape.
And then finally, um, data analytics platform, uh, uh, depicting the reference architecture. This is basically a platform which leverages Azure services. Uh, and probably, uh, just to give you a view on finally, when a data, uh, modern data platform is built, I think that that becomes the foundation of, of your analytics platform.
And how that basically is, is is then, let's say, expanded into, uh, into a data, data analytics platform. Is this, uh, reference architecture? So let's begin.
So I think this is the, these are the four vs, right? I think that, that we are aware of, um, which basically constitutes the volume, velocity, uh, variety and, and the var, right? So definitely, uh, these are pretty much still pre prevalent, uh, in the current, uh, era in terms of, uh, what they are.
Volumes still running in terabytes and petabytes because billions and billions of devices, iot devices and your age devices are also getting integrated into your digital landscape. Um, so volume is ever increasing, so is a velocity. Um, so I think while we have, uh, architectures that, uh, typically consist of, um, the, the batch processing mode and the steam processing mode, so, uh, definitely the, the data tion is something that, uh, becomes one of the key parameters to, to make sure that you have architectures that are, that are able to handle that kind of, uh, velocity in terms of your data tion.
Uh, then the variety of, obviously, uh, you have structured unstructured and semi-structured data, uh, coming in from your, uh, various touchpoint, uh, and how that data is then, let's say, uh, uh, ingested into your data platform is probably what the variety is about. And then the diversity, right? Uh, as a complexity of information grows, organization must improve the level of trust, uh, the users having their information, ensuring consistency across organizations.
So establishing a competency is essential, uh, for the business, for the, for driving a better business business. So that's basically the four vs. Uh, the volume, velocity, variety and the velocity, uh, which basically is, uh, the key to any of the modern data platforms.
Um, then, then we come to data driven economy, right? Uh, uh, the role of, uh, data, data analytics is ever expanding, uh, and it's becoming more and more key, crucial and strategic, uh, permission critical systems. Uh, so if you see, um, data and analytics is a strategic asset, uh, where we basically have a solution which, uh, with which you can actually sort of leverage to, uh, do various things around your, um, uh, around your, um, landscaping, ensuring that you have, uh, the right, uh, business models, ensuring you have the right cost optimization done, ensuring that you have the right set of, um, capabilities to do the predictive, uh, analytics and the prescriptive analytics, uh, around it.
I think all this is, is a key I asset. It's not, it's not your, uh, tactical solution. Uh, and, um, more and more of these are becoming mainstream.
It's not an add-on, uh, let's say, uh, capability that you'll have as part of your landscape and higher expectation, right? I think, obviously, when you have a strategic asset, you'll need to have, uh, increasing velocity and scale of analytics in terms of embedding more and more capabilities into, into our platform is what, uh, the data-driven economy is about. So, my making sure that, uh, the data becomes crucial, and I think the, the insights, uh, from those that data is, is leveraged to, uh, make sure that, uh, your business remains competitive and you basically have the right set of elements to drive growth and flexibility.
So I think that's, that's about the data even economy. Um, so I think we'll go to the next slide. Um, and this is basically a comparison between the legacy stack and the modern stack, right?
What, what constitutes the traditional data warehouse versus the modern data platform? So if, if you see the traditional data warehouse, and that's where the comparison is, uh, which, which I'm gonna cover as part of the data side. So, traditional data warehouse is where you, you'll typically see a lot of inflexibility in terms of the structure, right?
We can only handle one type of data. The architectures are pretty complex. They're not able to scale.
Uh, while you would typically have, um, uh, data warehouses, you'll basically also need to align your, your initiatives from different business units and business lines, which probably becomes more complex to, to let's aggregate as part of that. Uh, traditional data warehouse performance is also not up to mark, usually because considering that those earlier platforms were not built on, on cloud technologies or cloud services, and obviously automated technology, because some of them are, are the ones that were, let's say, developed, uh, uh, leveraging things like, uh, the, the, the SISs and, uh, the dimensional modeling, like, which, which are basically not really leveraged as part of the, the, the modern platforms. So visavis, the modern platform is the one which are pretty flexible, uh, where you can actually have a lot of, um, complexity and a lot of iterations that we can, let's say, let's say create, to build the minimum viable product and scale your solution.
So that's probably what the flexibility is. And also in terms of making sure that you can actually incorporate more, uh, use cases into the platform is what these platforms are. And I'll basically have a reference architecture covering that.
So you can run multiple, let's say, use cases, multiple scenarios on the same platform, considering that the underlying platform has the, uh, right elements of your right elements like capabilities in terms of the building blocks, an open and simplified architecture. Obviously, it's based on things like restful APIs and, um, things around open source, uh, technologies, which are basically not, not really proprietary and, and tied to a particular, let's say, uh, I S V platform or pool. So that's where the simplified architecture, or open and simplified architecture comes from high performance through in memory concluding.
So, uh, there are some of these solutions provide the in memory processing, uh, which makes it more efficient and more, um, um, uh, reliant, uh, or more scalable, uh, or more performant, uh, due to the internet nature of the in memory processing, and that's where the high performance is delivered. And then future oriented technology. So, um, uh, leveraging, uh, let's say methodologies like agile and scale, uh, you would still be able to, um, uh, let's say embed the, the set of new capabilities and, and functionalities into your platform in an incremental way.
Uh, traditional model limitations, um, are the ones which, uh, basically includes, uh, I think if you see on the right side, right? Uh, I think if I, if I have to cover these three elements, uh, the big data appliances or the big data domain, uh, they have their own set of limitations considering that there's no elasticity, they have performance issues. And also there's a high acquisition and operational cost, considering that you'll have to manage and run various clusters and, and notes to basically orchestrate the entire report.
Uh, data lake and Hadoop clusters are the ones which basically is, let's say, tradition model that we have, uh, which is inadequate for many use cases due to performance issues and operational costs. And then we have the data warehouses, right? Where, where you have inherent performance limitations, there silos, um, and, uh, they basically lack scalability.
I think this is referring to the fact that usually when you have a complex organization, uh, you don't, you won't actually limit yourself to only a single data warehouse. You'll have various, uh, federated warehouses, uh, created, uh, for various business banks and, uh, various, uh, various units. And probably that's where the data warehouse, uh, architecture is not able to scale.
So that, that's the traditional limitations that we have, uh, driving. I think the next slide is about, um, drive better business outcomes. Um, and I think that constitutes, uh, organizations want to modernize their data analytics platform to drive business outcomes, uh, by using data as a differentiated asset for innovation and competitive advantage.
So again, uh, as they say, data is an oil, uh, or, or the, or how it basically is leveraged as, uh, as an asset, uh, to make sure that, uh, you leverage that, uh, to the fullest to, to make sure that you, uh, in, in incorporate those insights and, and, um, your, uh, your, um, uh, your various, uh, let's say, uh, things that you, um, that you get as part of your reports and, and dashboards in a way that you make it, uh, make it, uh, make your business, make your business grow big, make your business competitive, uh, ensuring that, uh, you have the right set of, uh, capabilities built on top of the platform to, to do that. So, um, so this is basically constitutes, uh, why we leverage, um, uh, data, data analytics platform to drive business outcomes. And that's the, those are the three key elements.
Um, so I think you have valuation of new types of data. I think obviously it doesn't remain, uh, the platform is not really, uh, limited to a certain type of data. You have, uh, data, uh, various types of datas that can be harnessed, that can be interested into your platform.
And then, uh, basically, uh, converted into insights and dashboards, um, dashboard. That's basically, and also, uh, it actually has multiple integration touchpoints where the data is ingested from, be it your iot device or your business critical applications, or, or it might be your streaming data from, uh, some of the other, uh, let's say, uh, integration touch points that you have in, in your application, uh, stack, right? And then the negatively next generation approach is what it it constitutes.
So we basically are looking at, uh, a high-end, uh, data platform, um, enabling organizations to make fast PAC based decisions that need to better business outcomes. So that's probably what the next generation platform is about, something which is scalable performance and, and something that can basically be adopted to run multiple scenarios, uh, as per the business needs, right? And then the cost optimization of the current usage of big data.
So definitely it basically provides you, uh, with a way to have the right, uh, let's say stack built in right deployment model in place so that you have your performance optimized, you have your cost optimized. So that's about the, uh, driving better business outcomes through data and analytics platform. And I think this is a very interesting slide here.
Uh, it actually talks about, uh, data platform is embedded in the digital business platform. So if you see at the center, that's where your data analytics platform is, that's which is actually built on the modern data platform. And then on, on the, on the, uh, on and around it is probably what you have as the various capabilities and the building blocks of your digital business unit, right?
Be it the customer interaction platform, uh, be it your connected, uh, connected things, uh, or, or your iot platform, or it can be your, uh, business support system, right? Information systems, systems platform, or it can be your ecosystem platform to your partners and, and vendors and things like that. So I think it's, while, while it's at the center, um, it actually has all these integration touch points, um, with the various, let's say, uh, applications and, and solutions in your landscape so that it can basically get that data, process, it, run the machine learning algorithms and get you the, get you the insights that you need to make the right business decision.
So this is basically a very interesting side. Again, as it, as it says, it's a modern data platform, is embedded into the larger digital platform, um, in most of the, uh, fortune hundred organizations, uh, uh, which, uh, which basically is explained here. Uh, then I think the journey, uh, the, the journey is something that, um, uh, if you see, um, it actually consists of, uh, if, if you see the overall, um, stages that it goes to, uh, obviously it starts with the data warehouse.
It can be your on-premise data warehouse. Then you have the next, uh, deployment, which is your earlier version of the data warehouse on the cloud. Uh, and then you have the data lake as the next, uh, evolution.
And then you have the cloud data platform, which will constitute of constitute of lakehouse, um, uh, kind of a, let's say, uh, deployment architecture or deployment pattern. So this is basically what the journey looks like. Um, if you see on the table on the left side, it, it basically compares elements of the data users, uh, the past answers and, and there's still databases.
So what, uh, there are limitations in terms of the earlier, uh, generations of your data platform. And obviously that is the reason why you have the cloud data platform and a lot of new deployments and, uh, new user scenarios are the ones that are actually deployed as part of the cloud platform. Uh, it may constitute of a lakehouse, it may constitute of a data lake, or it basically can be a mix of these things.
So that's probably what this particular, let's say, highlighted gray, uh, column, uh, basically depicts. Uh, so there's basically the journey to, to the cloud data platform. And then I think that's probably what, uh, it's, I think if you see how, uh, the platforms are built, it's typically the incremental approach.
You have the agile based methodology where you, uh, you have, you have a set of, uh, uh, let's say, um, uh, requirements that are basically built for a given iteration, typically a team of about six to six to six to nine people, uh, with the, with the, the, the two weeks of, of, of, of, of the, of the, of iteration cycle that where you actually sort of pump out the, the requirements. And that's basically something that most of the organizations I've adopted to the Scrum methodology. Uh, so I think that's probably how the platform is built in incrementally these, uh, in, in most of these fortune hundred organizations that I've worked for.
And I think some, some of these platforms, I think I've just tried in embedding, uh, into the presentation. I'll have a workflow of that at, at, at a point in time, uh, in, in, in sometime now. Uh, so I've been actually also emphasizing on the data workloads, right?
So this basically is the one that, uh, uh, that depicts the data workloads, uh, for the cloud platform. So, uh, if you see there's, there are at least six different type types of workloads that we know of. So first is the data warehouse, which is the classical data warehouse.
Um, then it has the data engineering, which actually is the, uh, data pipelines that we have, which also is one of the data workloads. And then we have the, uh, data lake, uh, which is, uh, one platform for all data. And then there's a new version of that as well, uh, which is basically the lakehouse, which also is something that we have, uh, we, we cover as part of the presentation later.
Uh, and then there are these data applications, right? Model solutions, which are based on, based your data platforms, uh, constitute you constituting of, uh, which constitutes of business intelligence, which constitutes of reports, dashboards, which constitutes of your, um, predictive analytics, et cetera. So these are, these are the data applications, uh, that you have.
And then data exchange secured governed access to the data. Uh, so making sure that you would, would want to basically transfer data from a source system to a target system. That's the fifth workloads that we can run on the cloud data platform.
And then it comes. And the last one is a data science where there are machine learning and, and ai, which basically also, um, is one of the workloads. And, uh, that's, that's, that's basically the gamut of things that you, you have as part of the data platforms.
And this is basically covered as part of, uh, the various functionalities of a, of a large scale deployment. Uh, and this basically is something that is pretty much prevalent, uh, in systems and architectures where you have data analytics, uh, which run multiple use cases and, and provides you with insights, uh, to, to drive your business outcome. So, we'll, we'll cover that in a bit.
Um, so I think the next slide is about strategizing modern data platforms, uh, from addition to, to insightful action. So over here, what, what, what it, what, what, what happens is when you have a, when you have a data platform, right, uh, uh, it, it helps stakeholders to take evidence-based decisions on analytical, transactional, raw, and your raw data, right? So this is basically the key, uh, it makes, it helps you, uh, understand, uh, and get insights.
And based on those insights you can make, uh, the, the right decisions, uh, uh, which, which will drive the business outcomes. Um, then you have the data scientists and ANA analysis as well. They can consume an extra data from the platform.
I think there's an interesting use case for data scientists and analytics, and it says, um, it can use predictive and restrict analytic pipelines, identifying metadata insights and patterns in your structure and UN data. But I think the interesting part is, it's, it's something that is pretty much self-service. So you don't have kind of a thing that is, let's say, built and, and, and, and provided for you.
But data scientists, analyst are one of those, let's say, entities in the system, which actually leverages self-service analytics to get this predictive analytics. So they do the slicing and dicing, they do the application, they pro they, they extract the reports out of it to, to get the right insights. So that's different from the typical stakeholders, which will only be sort of, um, uh, leveraging it for, uh, their, their, uh, their insights and, and decision making.
Uh, the third one is, uh, uh, it, it exposes the outcome of these pipelines as human and machine interface. So, uh, human interface, yes, uh, where you can actually sort of have your, uh, analysts and scientists do the slicing, dicing aggregation and providing and such. And also machine interface wherein you can enable, uh, an integration touch points, uh, from your machines into, uh, another application which would want to consume this outcomes of data analytics activity.
So that's basically another, uh, let's say, uh, capability that is provided by, uh, the next generation data platforms. And then making use of advanced analytics such as text mining, nlp, and ai, uh, to pass and future such information to augment it with metadata that, that leads to, uh, datadriven decision. So, uh, again, that AI ml part, which can also be integrated and, uh, added to your platform so that you can actually do the national language processing, artificial intelligence cetera, to drive decision making.
Uh, and then obviously, these are the elements that you'll also build on top of the platform data governance policy and processes, uh, which is you to making sure that you have the right, uh, set of elements governing it. And then if you see the, the overall, uh, platform, um, logical view, right? It constitutes, uh, uh, various blocks is a load.
They store this process, serve and consume. So these are the blocks that are prevalent, and they also call it a data pipeline. Uh, and these are the blocks that, uh, basically are, uh, all the data pipeline, and they basically are integrated in such a way that data flows from, uh, the left side where it, it gets ingested to your platforms and assuming devices and your, uh, real time, uh, applications.
Uh, and then the conse consequence, uh, subsequently, they're actually sort of processed and aggregated. So I'll show you that. So this is basically a logical architecture.
I probably wanna just cover the key one. So this is basically the, the reference architecture I want to cover. So if you see the data analytics platform, right?
Uh, the modern data platform is built on, um, uh, the modern data analytics platform is built on this, this kind of, uh, data platform, right? So if you see, as I said, right, it constitutes of several blocks of data. So there's addition, there's storage, there's processing, there's analytics and cognitive services, and then there's visualization.
So these are services, uh, these are the blocks that are integrated. Um, data goes from, uh, the left to the right. And at the bottom are the horizontal services here, have office station governance, and then DevOps, monitoring, security, cost management, and your development, uh, teams or development environments.
But I think the interesting part is to understand that, uh, you have your data sources integrated, uh, into the platform at the data stage, and then you have the use cases that can be run. So, as I said, you can build the platform incrementally. You can have various use cases that you can actually build on this platform, and various personas can actually sort of consume those use cases.
So it can be your, uh, it can be your data scientist, it can be your general public, it can be your analyst, or it can be your business users, right? So these are the ones that probably are become your consumer, uh, persona, who, who would actually consume these use cases. So this is basically the reference architecture.
Um, and then, uh, what happens is the modern data platform, speed saving and, and secure governance. I think that's probably what constitutes, um, the, the overall data platform structure or, or the overall architecture. So it can, it can, so I think it's typically getting con compared to the earlier versions of the data platform, like data warehouses.
We service the modern data platform. That's where, uh, these numbers are coming in. So it says rapid ingestion, migration, and amalgamation of enormous volumes of structure and unstructured data for any legacy system up to four times faster, uh, add up to 60% savings.
So I think there's, there's basically, uh, with the E T l with the data pipelines in place, um, the ingestion, ingestion of that, uh, data becomes faster, it becomes more efficient. So is the processing and the, uh, the machine learning analytics on it, right? And then you have, um, it can, it can actually sort of, uh, provide a, a savings.
So these are the ballpark numbers in terms of what it looks like, um, uh, in terms of, uh, uh, savings. And, uh, these are the ones that, uh, are coming in from some of the, let's say, projects that have, uh, been executed in, in the similar domain. Um, then you have, uh, uh, one of the things that are pretty, uh, prevalent, and that's probably what we have observed, is the self-service part, right?
Uh, for some personas like the analyst, like the data scientist, they would want to do their own slicing and dicing. So the platform should provide you with that capability to do that slicing, dicing and, and aggregation and, and build and, and get insights, uh, on basis their own exploration. So that's the key element.
And then implement ethical data governance, including enabling sandboxes for development, experimentation and security. So I think the ethical sort, sort of part is also getting more, uh, importance make, making sure that it's done in the right way. And, uh, I think it follows the processes and, and protocols of, um, the, the specific domains, uh, that they're working on.
And also specific regions if, uh, it's getting, let's say, uh, built for, uh, a kind of a EU customer or, or a UK customer or a, or a, or a US customer, right? So this is basically where the protocols and processes and governance come, come in the picture. Uh, and obviously there are a lot of prebuilt, uh, frameworks that promote security, data privacy and trade.
So these are the ones that are there. I think each of the platforms have a prebuilt set of, uh, uh, uh, frameworks that can be leveraged so that it becomes out of the box, and you can actually sort of get the platform up and running in, in, in lesser time or, or quicker, uh, as quicker as, as, as, uh, it can. Yeah.
And then, uh, there's a data modernizing framework as well. So I think this basically provides you with a, with a list of, um, things that constitutes how do you want to, let's say, modernize your, your data, right? Uh, modernizing your data platform.
So, uh, these are the ones that I don't want to go into detail. I basically want to cover some of the elements, uh, pertaining to, um, the architectural, let's say, building blocks. So this is basically the, the one that, uh, uh, that constitutes the, the, the logical view of your platform.
Uh, so I think it starts with, um, the data integration layer that allows to ingest your data, um, structure, semi-structured and, um, unstructured, uh, in, into your platform. And then there are these cold storage and the hot storage, um, parts, uh, in the platform. So cold storage being, um, the batch mode of, of, of your, uh, platform.
And then hot is the streaming mode where you get your insights in real time. Uh, and then the common dataset as well, uh, like all products, customers or the master data. So these, the ones that are also provided as part of the platform.
And then probably a common business intelligence and collaboration spaces for power users. So this is basically the self-service capability that you wanna, would wanna build in. So this again, constitutes the, the logical architecture.
I've been, let's say, emphasizing on addition storage, your storage, you processing, uh, your analytics and data visualization. And, uh, I think I just want to quickly show you the, uh, the, the, um, the data pipeline architecture. So I think this is something that I've been, uh, I've been also, I have also explained, uh, so there are various, uh, uh, uh, elements as part of your data pipeline, be it the batch one, uh, be the streaming one.
So this batch is called the hot part. Sim is, uh, batch is called the cold part. Steaming is called Hot, hot part.
And then you have the cloud platform where you ingest that data into your data pipeline. It goes to the cloud platform. It's, it's basically where it, it gets processed, uh, based on ingestion, storage, processing, um, machine learning and visualization.
And then that's probably, so just a view, just a different view to show that depiction. Um, so I think I come to this part where I just wanna explore, explain more in terms of the data warehouse versus the data lake and Lakehouse. So if you see these three architectures, right?
Data warehouse is the one which was the tradition architecture for business intelligence analytics. Data lake is probably the next evolution of it, where, um, data, data warehouse was just a structured data. Uh, data lake was structured, semi-structured and unstructured data, but it still had, uh, a processing stage where you would have data warehouses.
And then those data warehouses, uh, would basically be then leveraged to get the reports and, um, uh, business reports and dashboards, uh, of, and of course do the data science and machine learning. But data warehouses was, was still there, right? As one of the components, maybe, uh, an incremental layer, maybe, uh, maybe it was more of a, uh, of a for intermittent logistic stage of your data lake, but they was still there, whereas Data Lakehouse completely managed that.
It's a, a kind of a, a very coherent platform, uh, where you ingest your data in the, uh, data lakehouse, and then that actually is what you'll leverage to get your reports, dashboards, data science and machine learning. So this is basically the data five data analytics platform that I'm mentioning. Delta Lake is, um, uh, a kind of a pattern where it actually has, uh, uh, stores the data in such a way that, uh, for all your data types, you can still run transactional queries on it, uh, just like you would do for an sql.
Uh, which is not the case for, uh, unstructured data, I think for unstructured data. I think, uh, that's something that benefit of transactional, uh, uh, pattern is something that cannot be leveraged. And data, uh, Lakehouse is the one which actually sort of welcomes that through this interface.
Obviously, it has its own very engine, it has its own, um, um, uh, transactional layer, but it actually makes sure that you get the same benefit of your, um, a transactional SQL databases, and at the same time, the search ability or the search capability that you have in unstructured data into an integrated, coherent platform. So that's, that's what the Databricks is about. So this is basically the, the comparison of, uh, uh, data warehouse versus data lake versus Lakehouse.
Uh, and I think I've just tried explaining, uh, some of the key ones. Uh, it's probably pretty, uh, uh, let's say, um, uh, it has a lot of data. Uh, it has a lot of data, so I'm sure I request all the participants to actually, uh, read through.
Uh, and I'm sure there's a lot of critical let's details available. But to touch upon the key ones, definitely as I said, um, warehouse is structured data, data lake is structured, unstructured and semi-structured data. And Lakehouse is combined, structured and unstructured, um, both schema and Read and sche on, right?
So this is basically something that, uh, uh, we can actually do a on read and on the right, uh, which you can see it, it, it picks up from both words, the data warehouse as well, and the data lake as well. And that basically is what, um, uh, constitutes a lakehouse. It basically gives you the benefits of, uh, data warehouse and data lake into one coherent platform.
So that's probably how the comparison is. And, um, also it's scalable both vertically and horizontally. Uh, so that's another key difference between them.
And, uh, uh, it does, um, uh, it does, uh, have the, uh, integration capability for BI supports in, uh, integrated BI tools, dashboards, and advanced analytics. So this is basically, uh, what all it can actually sort of run as part of the workload. So these are interesting table that I have included, I'm sure you would, would wanna refer to it at, at some point in time.
Um, I, again, interesting, uh, to know, uh, if you see the Delta Lake, right, uh, it actually has three stages internally. So steaming basket is ingested into the bronze table, uh, then it's cleansed and made it into the next stage, which they call it a silver table. And then it's aggregated and, um, um, it, it is got it, it is got into the right format, uh, which is basically the gold standard or, or the gold stage on which the analytics and machine learning is, uh, is, is then, uh, run.
And that's probably what constitutes the stages of the Delta Lake. Uh, and that's probably what is depicted here. Uh, I think I have just explained to you the Datadriven culture.
Uh, I, these are different roles that, that are there, architect, engineer, lead, steward, data, citizen, uh, con, uh, the data scientist, uh, and the data consumer. So these basically, what are the different roles? And there's a, there's a difference between, uh, how they do it.
Some roles just, uh, leverage the existing solutions. Some actually do a lot of self-service on top of it. Uh, the analyst and the data scientists are the ones which actually would do a lot of self-service, uh, self-service analytics, so that those, those are the typical evolution.
And citizen consumer is another role which does that. So they don't wait for the IT Alexa team to actually build those dashboards and reports for them. They actually do a lot of self service to get insights, uh, to drive business outcomes.
And this is basically what is explained, uh, in the next diagram. So it's just a reputation. Uh, I just don't wanna go through it, uh, these slides.
But yes, I think I've explained to you the realtime tion, uh, which is basically scheming and batch and realtime scheming. Uh, this is basically what constitutes a typical logical architecture in the analytics background. And then, uh, I just wanted to, uh, let's say compare the LA and the copper architecture.
Uh, so there, there are two of them, right? Uh, one that constitutes, um, hot part and the other that constitutes a cold part. Uh, Lambda is the one that actually has, uh, has both these parts, uh, configured as part of the architecture, uh, just to get the right benefits finally.
But the limitation being, uh, it does have a lot of complexity. It does need a lot of configuration and, and skill to manage it, whereas Kapa is something which only has one, uh, let's path, uh, it's easier to build and, and, and feed, but at the same time, there's a loss in terms of the functionality. So I think this, this is basically the, the table that compares those three architectures.
So Delta architecture is what we saw earlier, uh, in the Databricks, uh, uh, Lakehouse. So this is this comparing various, uh, various architecture dimensions of, uh, your, your architectures and where they are used. Um, and I think, as I said, the Lakehouse is, uh, is the one that actually has, uh, best of both worlds, just like, uh, what we had as part of the lakehouse as we, as we saw earlier.
So I think Data Lake Data Data is the one that basically cons, leverages best, best of both worlds so that you get the benefits of Lambda and copper. And at the same time, you definitely have, uh, uh, those, those things, uh, to achieve a kind of a unified processing with a, with a kind of a coherent platform as we saw it. So this is basically how, uh, it it is.
And I think lastly, I just wanted to, uh, uh, basically show you the data analytics platform as I, I was emphasizing, this is the data analytics platform built on all Azure Stack. Um, if you see on the left side, these are the structure, unstructured ran, semi-structured data sources. Also, your steaming data is right here that gets ingested through data pipeline, and then it goes into the data lake.
And then from there on, I think, uh, we have those. So I think if you see those view stages, um, the vertical docs inges, so this is the ingestion, then you store it in the data lake, then you process it, leveraging your, uh, your data, uh, data spark, the, the spark pools and serverless dedicated SQL pools. And, uh, those are the ones that are then enriched using, uh, the AI machine learning, uh, solutions or services that you have.
And then basically those are the ones that are then, let's say, industry into your, uh, visualization layer, which constitutes of either a power BI or the human machine interface that I mentioned, right? So it can be exported into, uh, a Cosmo db or, or it can be leveraged as part of cognitive search, or it can be shared as through a data, data share services of Azure. And then this is a visualization.
So this basically just maps your logical building blocks into the Azure services. And this is basically what constitutes a kind of a full blown, uh, into an architecture that can be leveraged for multiple use cases. Uh, and again, it can be built incrementally.
It can be we incre incre incrementally as per, uh, what the, what the solutions or what, what the business, uh, requirements or the objects are. And at the bottom are the horizontal layers. So there's, uh, uh, discover and go.
Then there's Azure purview, uh, which is again, your catalog. And then there are various elements of, uh, DevOps, security. Uh, you have, uh, Azure, keyword, Azure monitor, cost optimization, so on and so forth.
These are all the services that you would typically wanna leverage as part of your end. So I think that is more or less, uh, I've reached the end of, uh, my presentation. I think I've covered most of the things that I wanted to discuss as part of this presentation.
I hope, uh, you found this, uh, insightful and thought provoking. Uh, these are my Twitter handles. Uh, these is, this is my mail id, uh, this is my LinkedIn, uh, uh, let's say, uh, link details.
Uh, feel free if you have any queries, uh, I'll be more than happy to. Have a nice day. Bye.





