Unsung Heroes, Data Engineers – The AI Times EP 3
Artificial intelligence and machine learning get all the glory, but behind every great model is a great dataset and great data engineers to wrangle that data into shape. Data engineers get hired in a nearly two to one ratio at companies, but they often go unheralded. Their work is less glamorous, but without ingesting, cleaning, QAing, testing and modifying data so that it’s free of mislabeled data, redundancies, and dozens of other errors, there is no AI/ML. In the past, much of that data work fell to ML engineers themselves, but as organizations have grown, they’ve increasingly split the tasks into dedicated teams that focus on data and focus on algorithms. In this episode, hosts Mike Vizard and Lee Baker are joined by Shirley Wu (Juniper), Gaurav Pathak (Informatica) and Gopi Kokkonda (Alignity) to talk about the way of the data engineer and why it makes all the difference to your models in production.
Transcript
this week hello and welcome to another edition of the AI times our topic for discussion today. Is that unsung hero the data engineer? I'm your host Lee Baker.
I'm the General Secretary of the AI infrastructure Alliance, and I'm joined by my co-host Mike visit who is Chief content officer for Textron group. We've recruited a wonderful panel for you today. I will allow our guests to introduce themselves and due course, but but let's dive in let's take off with our with our first question and and in no particular order.
I'm going to pick on one of our guests and go be I'm gonna start with you. Um, do you want to talk a little bit about the essential skills and defining characteristics of a data engineer? Well, it's a pleasure to be in this panel with you all.
I am Gopi coconut CEO and founder of alliedt solutions where data engineering is one of our top practices and we help startups to all the way Enterprises by and understanding and solving their data problems through our Consulting and Outsourcing services and coming to your question. Lee what are the essential skills of a data engineer now, it could be as simple as the skills but more importantly also the attitude that is required for data engineer especially taking back. How data engineering a world since the last two decades or so.
I mean it would it was evolved as what we call as data warehousing and people doing menial tasks and just starting with writing SQL code and then starting with Hadoop and the architecture around that to how things have evolved in the recent years that they're having to play more strategic role than just more tactical aspects. So I would say a simple answer could be as as few specific techn. is like right being able to write SQL code and specific experience in data warehousing and the grasp of the working and building like data warehouses and then data architecture and coding with like python as such and then some experience with data Ops and operating system and then few things like machine learning but those are just the set of skills, but more importantly it's important to understand the business objectives of what the data is and where it's is ingested from and then how the data is perceived by the consumers and what is value for them.
And then where is the differentiation between revenue and profitability as how the chief Revenue officers are coming into the current realm of things considering the macroeconomic conditions and then being cognizant about that and then playing a more white role is how I would put the combination of skills and attitude of what in data engineer should be to be a very successful data engineer. Well that makes sense. I totally agree with the Govi said so from my background, I just give Two sentence about my background.
So I'm mainly my name is Shirley Wu and so now I'm currently I'm in the architect and also lead the measure, you know, air-driven self-driving Network Solutions. So we have a project name called Marvis. So it's you know, you probably Google Juniper or Miss Aid.
The Marvis is a product name. Actually that's marvelous is collectively in the cloud. We build a intelligent, you know piece of we give that our software like the AI identity called Marvis.
Um, so I'm kind of you know started as a starting the foundation work, you know to build the whole Cloud infrastructure working with our cloud devops. And as a today that we have, you know, so many devices it be Deport Network device deployed the strong world and to the data engineer is essential role. So I cannot strongly agree with Gobi data Engineers, you know involved from the data warehousing engineer up to today.
So the engineer role first most is a good software engineer and the besides you need to have a good understanding of data and the business concept. So it's a you know for any ml good Solutions in the data scientists for sure, but the data engineer it's a similar and requirements but more focused on the engineering side of the data scientists for more focused on the algorithm, right and you know algorithm side, but is collectively we need the data engineer and the data scientists to successful ml Solutions And if I could build on what Gopi and Shirley added right like and before that I'll reduce myself. Everyone got a Pathak vice president of product management here at Informatica.
I lead the AI division for data management providing productivity benefits for our data management users. So data engineers. And before we go into skills, let's also understand.
What do they do every day, right and you know, we think of them as these pipeline developers and you know working every day understanding what's happening with data about 40% of their time is doing ETL development elt development testing Buck fixes and so on but a very important part of their job is working with the stakeholders business teams and understanding the key requirements. So alongside technical skills all the understanding of data ETL python data warehousing and so on. Alongside of soft skills where they you know, they are able to communicate better and and have problem solving Max.
I think a very important part of a data engineer should be the domain knowledge. And that's something that we have seen there are various different reasons for which data Engineers work with data create data for ML use cases. For low stakes AI development.
Let's say we are I'll call it low stakes compared to high stick AI development where we are using AI for example cancer detection or for example finding out if somebody can go blind because their eyesight is starting to deteriorate. These are now areas where ml is starting to get used and being able to look at data being able to provide clean data to ml models for these high stakes use cases. Very very important.
And that can only happen Engineers understand. What are these bits and bytes that are getting fed into the ml models and not see them as just data points within you know within a table or a CSV file that's moving in. So I think data Engineers alongside Technical and soft skills having that understanding of the domain is very very important as well.
And don't like me add to what government said as per Nick kudaker from Gartner 85% of the data engineering projects fail. And the reason is not technology. So as I write as he writ.
The right communication skills and understanding much more than just the coding skills or software development skills is much more vital in achieving the business objective for what age engineering role is and then his coordination with the other roles like the data analyst and data scientists as surely has put in oh and up on there. Let me go to surely then. I'm sure you didn't Engineers is one of the hottest job categories going right now.
There's people out there are climbing over themselves to go find these people and my question is is how do we find ourselves in a position where we needed to create this function in the first place? Because is AI driving a lot of this do we need to find ways to clean up our data to make the AI models work and that's what that's a core of this issue or is it something bigger than that? And because my perspective is maybe it's a bias, but Hardly any organization I know would get a Good Housekeeping seal of approval for the way they've managed data for the last 15-20 years.
So are we just kind of cleaning up our own mess? Finally? You are right actually nowadays that every organization, you know, they claim they're the data-driven because you know, the internet nowadays generate a lot of data.
So any business they wanted to you know, make a decision smartly. So they have to be data driven so and the core and the business, you know, you are the traditionally you just need a web developer, right? But nowadays it's involved.
So like I said data engineer first of foremost is a software engineer, but it's not enough anymore. You need to understand the data and also this is row is involved from used to be the data warehouse engineer. So now this road is much more just a data warehouse and you need to be the good solid the software engineer and also they understand the data architect traditional from a traditional database to the months ago database architecture and distributed system.
And on top of that, you know, the Hadoop big data analytics skills. Actually. It's essential nowadays, right?
So this is Just a technical front, but on top of that, you know, like I said, the business is a key. All these big data is just the tools fundamentally. We need the uterus to to solve the business problem.
Right? So that's why the for every organization you have your unique business problems you want to address then that's the data engineer came into the picture. They need to understand where the data come from first of all and then with the data come from be able to collect the data then analyze the data then figure out.
Hey does the business problem really can be answered or addressed by those data? If we cannot then we have to go back to the source to find the more data then that you know, then we identify what type of problem I want to solve is going to be we need to do a prediction or detection or clustering or do the forecasting. So then that can help the business solve the problems and we can involve to talk about hey.
Need a you know data scientists or analysts. So this can you know, the data engineer is at the starting point in the help to collecting the data clean the data then to look at the data. Can they answer the questions, you know business have or can they addresses issues business have so this is why the data engineer at the first I think a lot of organization before you have a data scientist.
You should the first to get your data Engineering in place get the data collected and be able to just generate the general simple reports a dashboard before we can talk about what's going to be the next step. Yeah, I love I love that. I love that idea of dashboards.
Sorry garib. Let me allow you to LEAP in on that, but maybe I can steer you with your your comment on this a little bit. Surely mention kind of data quality consistency and also the source, you know, where is this?
Where is it being derived from? According to to create great data sets from from purely web scraping and what kind of problems did that create for Unique versus private dentists. And ultimately what are the tools and techniques that you use to achieve this That's a great question.
And and you know, it most organization the if I have to disc organizations adapters to describe the state of data today it I would describe it describe it as neglected, right and because most of the folks are worried about, you know, the cool things the machine learning the AI part of it and then that's you know that those are the things that get all the spotlight but all the work done to make that AI model great is in data in collect different collecting different sources of data and and feeding that in making sure that it standardized it's high quality and it's covering the scope of all that ML and AI needs to do in web scraping. For example, I I know a few cases where retailers have this automated system of scripts that go ahead and look at all the competition. And figured out what are the you know, what are the prices of products?
They they want to make sure that on any given day their prices are the lowest and and this happens every day. So if some somebody is on a website and they're doing price compared, they should go to x retailers website instead because of the price now, it sounds it sounds good on paper. But having to do this at scale at internet scale is very very difficult.
So you have scripts that bring in this data, but you know, the websites change the underlying data models of those websites change and getting this properly is is very very difficult. So one week later live you they had a team of 300 people sitting and just looking at that data and making sure that the data today was as good as yesterday and and then providing human input to making sure that the data was was going properly into the system. Well follow up on that Gopi.
Um How automated can all this get it does seem like we're very manual eccentric here. We got a lot of effort required on the data Engineers. All right, surely take a crack and we'll have to say so actually I can't share the specific use cases.
We are developed. Unfortunately. I see how great with a graph there's no golden.
You'll know the tools that account magically automate for every organization every business use cases. It's just a peruse cases what type of data you are dealing ways. So for example for us and you know back to the previous question, you know, can we you generally the great data from a webs web scraping actually, you know Chad gbt a lot of training data is by scripting the internet, you know data right for us.
Actually we have one of the product that is the feature. We need to generate, you know, so for our customer the utilize Journey missed a product if they have questions go to our web page we wanted to provide. A good quality response.
So Juniper, you know missed we have all this documents product Document Services meeting materials, right all this on the internet and so in instead, you know, internally, we're trying to gather all the resources all the materials which is a scrape our you know public accessible web page and you know download those, you know, the little Snippets and index into our own database and you use you know, natural language, you know, universal encoders one of the language model in code the all this knowledge is entire database. So when somebody have a support questions and if we know this question inside our days database we have the index which is directly show to the user instead of you know provide, you know, they have to do the keyword search they have to do Google on the internet to find the response. So so, you know all this process actually, what do we do?
We kind of even though it's a web scraping we try to modularize that for different type of Source and we utilize the airflow that's when the open source, you know the scheduler so that we can reindex the web sched scraping process to the week every week. And if we have identified as a Content together modified, we also update our database inside the bank end. So this is kind of you know, like I said, it's a case by case scenario.
It's collectively the engineer team we work together identify always, you know, my, you know, the principle was not though. You don't want to write the code for specific one use case you always trying to gather, you know, more the business requirements and it's right to the code for this type of problems. Not just write the code for this specific one problem.
Yeah, I know you guys have been doing a lot of work with AI and data engineering. So maybe you know, what's the state of the art? Absolutely, and and before we talk about automation, let's first understand.
How do these data quality problems manifest themselves and data and as they make it into ML and AI there are various different reasons, right? There are environmental reasons people reason. There are more you know that we have seen over the course of working with thousands of data quality users.
If we are looking at for example data set that helps with traffic violations and there was this data set that we were working with and what we realized and this was data set coming from India. What we realized was that even a slight drift in the way the way the camera was positioned would create data quality problems and India, it is notorious for traffic violations. I've lived there for most of my life and every time there was a strong When the camera would move just a little bit and then that would create all sorts of data quality problems within the data sets as they come in.
There were people drifts that were happening the way you know, for for various different reasons people wouldn't report on certain things that they should report for the ml model to be able to provide the right predictions. And then there are things that happen Downstream, right? You can create a very pristine AI model looking at pristine data that's coming into it.
But the real world is messy when this you know, for example, if you're trying to predict if somebody has an eye disease from the scans of the retina if there are like some specs of dust on that, you know on that film The ai go. All right, so there are various different reasons in which data quality problems get used. Not everything can be solved through automation.
We can use automation to figure out for example, if dates are standardized better if the data set is complete or not, if two columns that are coming in our correlated, you know father's height and and the child's height and and if there are deviations our AI models well will flag it will say there are outliers in this data set. This child's height is not correlated with this Father's Side compared to what you know, what models are given so those kind of issues can be captured through AI a lot has been done in that area in Informatica tools and and others but I think human effort is still required if you have to solve the data quality problem, so I actually say attention is all you need. This is you know from our you know, GPT paper and transform it Transformer people that came out in 2017.
We've solve the data quality problem. You need human attention to that problem as well. Back to you if you said that there's no there's no shortcuts.
Is that what you said? Yeah, absolutely. You know there is there is no if there was a magic button that we could press and and all data quality problems go away.
I think that that would be but no there is you know, you have to care for the problem specially in high stakes use cases understand their data is coming from and then only we can get to pristine data quality. And again, I'll say, you know, we have to go from goodness of fit which is what Ai and ml models look for to goodness of data. Right and then that requires attention that requires a lot of work.
And I just want to add rather ask girls because he mentioned Chad GPD and AI. this whole new world of what the future of data is going to be if I mean I get gaurava or surely you because your your Or more Hands-On technical. What are the biggest challenges you see with the AI coming in and taking over the world and you see well, what do you see the immediate immediate one or two years with the world will being so murky with this AI thing coming in and how it takes over or disturbs the the data world.
You see any specific things over there coming up? Yeah. So it's a good question.
I think it's a it's exciting and also it's a little uneasy rights uncertainty about what's going to be you know for the next one year three years five years, right so gonna be a lot of startups, you know new ideas people trying to explore it. So one of the thing I can see right now is a chance to be T. Just like, you know, the Precast previously in love mention about the human still is gonna be in the loop, but probably gonna require different sets the skills that you know, give different type of skill sets for The engineers or Engineers for example for data quality.
That's a thing. There's no, you know Silver Bullet can solve all the data quality issue same for jgbt. I think everybody played with ggbt depends what type of data you fit in the change gbt, he can give you all kinds of, you know, different whatever problem they provide that GP can go wild right so then humans still going to be in the loop, but what type of skill set requires new next generation of the engineer.
I wish I know for now. But yeah, I'm just, you know still learning for me as well. So it's great that we have this conversation.
I would like to love to hear you guys, you know, everybody's saying those thoughts. What? If I understand the essence of this, you know, we're trying to balance the Need for Speed and Agility in days your engineering with the need for data security governance compliance.
Go are we getting that balance rights and and to Shirley's Point? What are the new skill sets that we're going to need to ensure we've got kind of that stability. Sure, and and this is an exciting new world that we are living in, you know, we call them these new large language models as Foundation models.
Everybody is excited about how we can you know, not only work with Chad GPT and make ourselves more productive. But every company in the world is now looking at how can we use these Foundation models and make their products and their services better. So it's exciting new area.
And as surely said, this is also this will also bring in the need for developing new muscles for data Engineers data stewards data Security Professionals. And and so on as gpt4 was announced there were a lot of good example. There was a Morgan Stanley example where they have created this wealth management chat for their internal folks.
So this this chat is supposed to have the knowledge of 60 80 100 Years of what Morgan Stanley has done and any intern Will wealth manager could go and ask it questions. Now when I hear use cases like that, I'm very excited. I think this is really the you know, the the Forefront of Technology right now, but as you delve deeper, you understand that these things are not Plug and Play yet.
They Morgan Stanley put 300 people, you know long before gpt4 was announced and this example was released to the world to make sure that the right data was first fed into these models, right so that you know, these models don't go around and start professing their love for the internal wealth managers and and so on right about hundred thousand documents were created by these data Engineers to make sure that the wealth management chatbot was you know was up to what Morgan Stanley wanted and I think that's what we will see we will need to scale along with how Ai and models and how Ai and ml models are scaling how technology is scaling in the world as well. Yeah having said that the question is about the skills and then well how the role itself would evolve the fundamental skills. I guess would still have to be the right foundation which includes the basic.
Programming in SQL and Python and then into a knowledge of like Hadoop and architecture and things like that now. That being skills and an attitude. There may be a new definition of roles and instead of being called a data engineer but itself it could be more specialized roles like a data reliability engineer who would ensure data quality and if you did a product manager who would like boost adoption and monetization data Ops, I think that would be the new thing which would replicate the roles of software engineering in terms of adopting the best practices and things like that and even like they so I guess we just see need to see how the data engineering field and the roles would evolve and specialize more to just take on the opportunity of data and also take on the challenges which both opportunity and challenges that AI will present on the future of data.
Shirley how do I know when I have a good data set? I think a lot of people are kind of sitting around going. Well, this is great.
But how do I know when I get there? You know like I guess my kids are in the back of the car and my head are going Are We There Yet. How do I know?
Yeah. So this is you know, there's a generous things of what is the definition of good data sets. Also, it's a peruse cases.
What type of domain Urie right so specifically for us, you know, we are in the kind of networking iot device domain. So this domain the data quality what we're talking about is, you know, Quantity Quality quality, which means that hey every device send it to the cloud as I do have a missing data. Do we have a certain Fields required fields, which is empty.
Right? And the second is a quantity. How much is the data we are talking about.
So if you just multiple Is you know, it's multiple devices or innocence right now for us. Fortunately we have this journey, we have already have a device deployed throughout the world, you know cover pretty much every continents, right so different geolocation that gives a realistically what's happening around the world. But if you know, we have a new firmware release new hardware had been deployed which is only being available in a few labs.
This is a lot of good quality of data. We can utilize to build any kind of ml solutions to solve the problem. So the third principle of us is going to be variety.
So variety is at the cover the different aspect of the problems. We're trying to address, you know, do we if this is an industry criteria before we can take on any kind of a business plan problems to to solve so we look at the data. We have a good quality and the quantity and there's a variety of data before we can you know, say this data is good for us to implement any kind of ml Solutions or any kind of, you know product features.
I'm sure you know for you know, I think it's some of the I think the Kobe since you are from Consulting, so probably you're dealing much more variety of the use cases your customer facing so the criteria may be different from your perspective, right? Yes. No, I'm seeing various Trends and then I think from The phase of where a customer is whether it's a startup and then their journey in the data engineering cycle to Enterprise customers.
We are seeing like different challenges with the amount of like volume the data security data velocity and the tools that they are using and then we are learning as well. And then I think conversations like this or so helping us learn and share that all aspects of it. Yeah.
The interest is to develop on that a little bit because data sciences and with before I say this old due respects to my data science colleagues data science. This can infamously be a little little myopic and very very focused on kind of the model accuracy and performance. How does a data engineer differ in terms of their requirement to work effectively with other teams such as data Sciences analysts business stake.
Is to ensure that data is used effectively to drive business outcomes. Well, I would say the challenge being that they're like Various levels at 10,000 foot year where there's management level challenge and a thousand foot view where there's an architect level challenge and we get a hundred foot view. There's an engineer level challenge.
So a data engineer. for a percentage of situations which could be anywhere from 40 to 60 percentage is still trying to understand his real role and then their X number of times his role could be misconstrued or is overlapping with other roles like a data scientist or data analyst. Add it to that the Crux of this conversation obviously is a data engineer role being less glamorous and then all the great going to other roles as such so that being the case.
The the first thing to understand is just embrace the problem itself. number one and that is for the Senior Management to keep that in perspective and second thing is Keep more long-term objectives in terms of hey, what is the expectation? And what is the objective that I want to achieve and not put short term but medium or long term objectives and then have it clear definition of role and also understand where the overlapping responsibilities are with respect to what the data engineer data science and data analysts are and then we'll have good combination of open and transparent feedback and mutual setting of goals could be a very high level answer to your question.
Well, I know that there are a lot more nitty gritties to going to but that requires a bigger and more focused on conversation, but that's what I would put it for now and anything shortly and that would like to add to this. Yeah, I would like jumping it's a great question. I wanted to share from my perspective.
I think in my team, I would like to defy a successful. Ml project is by Street criteria. And first is efficacy.
This is a primary goal, you know data scientists, like, you know, the mic put it out data science is really focus on the efficacy of their model, but this is not one of the it's only one of the criteria to define the ml Solutions successful or not second. There's a criteria is that the latency, you know, you can predict something you can take something how long you gonna wait, you know need the after a month after a day after an hour. You can predict, you know, what's going to happen or you're gonna detect Something's Gonna, you know have already happened.
So latency is also another criteria. The third answer to the most important is computation cost. Um, do you really need those GPU to support your solution, you know how much it cost for the computation cost versus how much business value you can bring in for this ml solution.
So this is from a business owners perspective really that Define the success for the ml solutions for you can see now that the data scientists only focus on the efficacy and latency latency is actually the combined with the data scientists and the engineer infrastructure, right? The computation cost is primary actually really being neglected by the engineering doesn't matter the scientists or engineering or data engineering but this is a lot of time is a business owners. You need to Jumping.
Hey, you know, what is a business value for this ml solution so collectively as a team that we need to try to find the max kind of give the maximum business return minimal computation cost and also the Maximal efficacy and the minimum latency. So look at this is three. So that's the singular.
Ml project is optimization problem, right? You know all these three factors they conflict with each other. We need to try to find the right balance not just focusing on one particular one aspect.
So this is collectively with a either working with a business engineering and the data scientists to find the right balance that can bring the success to the business. So that's my personal take and as a and there's actually that's a good golden, you know the standard for us to deliver our ml Solutions inside my team. they're kind of leads us into our next question, which I'm going to throw it ago, or I'm So a lot of data sets are imperfect can I fix them because there's a tendency among data Engineers or anybody else who looks at anybody else's work and they're like, oh that's awful and I'll just start over again, which is in particularly productive.
But is there some way to look at a data set and say, okay and this is how we optimize it or augmented from where we are versus constantly Reinventing the same wheel over again. Absolutely, and this all goes back to I think what Gopi mentioned about organizational roles and incentives as as we develop these data sets. The goal should be to make sure that we can provide as much information metadata to tell what is this data set good for what do individual Fields here mean so that the end Downstream use cases whether it's ml AI development with reports or dashboard development.
They use the data set properly. We were working with this, you know with this Enterprise and you know, they were dealing with these kind of drifts because the research and Engineering teams were seen as two different teams all together. So engineering teams would make changes.
So I'll obfuscate their exact use case, but I'll try to tell the essence of what happened. They had this variable which would Tell whether a customer would like a particular feature or not. And at one point of time that got encoded as zero or one in the data set But as time progressed that changed to a scale of 1 to 10, how would you you know the talk to this talk about this product your friends and so on right recommend this product to your friends.
So when this kind of change happens the data set itself drifts. Right and and you need to document and you need to say now this is the new field which captures the human input about whatever product feature that we're building. It's no longer a scale from zero to one.
It's now scale from 1 to 10. But ml model that was using this as an input feature, right suddenly breaks down because now they're not getting the exact data set that they're looking for right? So being very open about metadata alerting the downstream end users of this data set whenever drift occurs is very very important when these changes happen.
And if I may add to that more at a higher level not just a data set level we have things related to like an architect level challenges. It could be like poor workflow orchestration like premature scaling and streaming and even things like late governance and that being an article level. We see even a data engineer while he's still trying to settle or evolve into his role.
There are quite a few times. We have seen that data engineer and one of the points that you brought at least with respect to they're meeting the business needs and how the business eases and we have seen quite a few cases that a data engineer is often tempted to create too many dashboards just out of excitement which often happens at the business either doesn't even look at it or they only out of only 20% of the dashboards, which are created are actually like seen by business. So it's very important to understand and streamline.
What is done versus what is really required? And what is a priority number one number two in terms of optimization as well. We have seen data engineer quite a few times.
maybe with his Not enough guidance or mentorship rights? like long queries, which could be like a lot more optimized and then research says that at least queries could be optimized. 40 to 60% in terms of their length and things like that.
So well, I just want to get into more General. Iron level challenges than just the data set. I've told that ad City value to this point of conversation.
but surely we exist in in an explosion of Open Source tools that are available to us to kind of help Wrangle these these data challenges whether it's versioning labeling management to deployment up. How do you keep up with the the latest developments and data engineering and one of the resources that you're relying on and how do you blend that mix of Open Source versus proprietary? Yeah.
This is a good actually a lot of companies right to use open source build yourself or you go buy right? So this is a continues of dilemma for my career in the different company related to different project. So first, I really appreciate the open source, you know that the, you know communities especially for Big Data, you know, now it's really matured, you know the space right?
We have a few Big Data tools. That's pretty much in mainstream. So at least of for a lot of startups, so I came from a big startup background.
Then now we are part of the acquired by Junior person where you know part of the big big company now, but that we kind of still maintain them and maintain that we want to First. Is always open source and but we are not just take you know the open source, but we also be brave enough we can forecast the repo and make our own changes. But we also keep buying on that what's emerging on the you know Horizon like any kind of big players, you know, we invite them to have our givea Brown Bag give a demo and to see we're missing and some yeah, of course, you know a lot of open source for the such a large scale for what we are operating in the cloud.
It's a problem. It's not sometimes it's not sustainable because it requires a lot of head count internally to maintain. So in that case we have to go ways of proprietary or maybe more assess offering, you know proprietary from a cloud offering so that we can don't need to spend our head count on those infrastructure projects.
Maybe you can more focus on the business use cases, so it's more on Which stage of the company you are if it's a early day startup, right? You don't have much of the resources go with the proprietary. And we also you don't have that much of the data warning and the complex business use cases probably should go with open source on your own until you figure out the you know, specifically use case and you can continue to grow your customer base when you reach a certain level, you know, that maybe you should consider a proprietary data or proprietary vendors product.
Maybe can serve you better. So it's you know, the part of the grow up from the startup to mature company. This is a question always need to be continuous being assessed.
Right? So, this is my personal journey and I feel from my He was a few, you know company experience. That's what I can you know come to this type of conclusion.
And if you should think of these data science needs like a pyramid right at the bottom, you have the data infrastructure the storage the database is and all of that at one layer up you're looking at integration capabilities building pipelines bringing in collecting data from a lot of sources and bringing it to the infrastructure. Then you go one more level up and now you're talking about data quality data preparation wrangling data making sure that data is fit for use for AI and ML. And once you have all these three bases of the pyramid, then you can start thinking about Ai and ml you know, being you know with the company that I am in we definitely have strong view points on what kind of tools to use for data management.
I suggest when you have mlai requirements when you have data and politics requirements, look at this pyramid look at where you are and what are the needs for your scientists, but also, at how all of these data management needs You're not looking at them as Point items because all of data management needs themselves can be made better by underlying metadata. For example, if you are able to find where the right data assets are you can then use that to suggest joins in while you're creating pipelines. You can automatically recommend where quality problems are.
So being looking at a tool set that's integrated that is in the cloud. Like Shirley said the help a lot of data projects. And how should how should internal teams work with a vendor an assess provider?
Like Shirley describes and Gareth identifies. How should an internal team work with a vendor to create a business case for for the c-suite? And say combination I first added more into what Shirley said and Garland and just come to this point.
So combination of the clients that we have. It's all over the Spectrum as to they're using a combination of Open Source and proprietary tools. So what are the reasons it could be a combination of number one the resources and the budget number one and number two, they're risk taking ability and the support what a proprietary tools may offer versus what an open source tool may not offer.
The third thing is an internal Champion for that tool whether it's a proprietary or open source, right? So the management makes decisions based on this and the confidence and the bandwidth and risk taking ability and the budget and resources they have so those all in some combination. form as the reasons or the parameters in decision making and making a case for What tool to use and then went to adopt it and then how often to revisit their decision to either like change or migrate or move from open source to a proprietary tool based on all the combination of regions now there could be several use cases based on several products based on hey whether you're using as what government mentioned the Pyramid of where that tool forms, but those could be the very parameters in your decision making process or chain.
All right, guys, I think we're coming to a close. I'm gonna have we take us out and folks. Thanks for sharing your knowledge and expertise was awesome.
Yeah, thank you guys. I think we had a we had a great panel today and we could go on and on but we have to bring this to a close. So thanks for listening to our show.
You can find this episode was show notes on the tech strong TV website and through the AI infrastructure social media channels like YouTube. You can also follow us on your favorite social media platform And subscribe to us on YouTube, too. We look forward to seeing you next time and just reminds me to say thank you to all of our panelists.
Thank you for connecting this and Shirley and Garo very nice meeting you lot to learn and share from this panel. Thank you. I appreciate it.
Samir. Go peace. Thank


