Edge-Oriented Data Patterns | DataOps Day
Edge applications are emerging as a way to build and deploy distributed solutions that span data center core, cloud and edge boundaries. However, there are various barriers to this pursuit. In this presentation, we will introduce a comprehensive edge-oriented data pattern model that consists of architectural and design patterns for edge-hybrid and edge-native data systems. We will further delve into the role of cloud-native capabilities in supporting edge data solutions and the trade-offs involved. We will also showcase some examples of edge applications with enabling technologies and tools. Best practices and lessons learned will be discussed as well.
Why Attend?
-Learn about the best practices via the pattern approach
-Get hands-on experience with edge-native technologies
-Understand the latest trends in edge computing and data patterns
-Get inspired to contribute to open source and the larger community
Key takeways:
-The different types of edge computing architectures
-The benefits and challenges of edge-hybrid and edge-native data patterns
-The role of cloud-native technologies in edge systems
-Best practices for developing and managing edge applications
-Lessons learned from edge implementations
Transcript
Hi everybody. Good morning, good afternoon, good evening. From wherever you're joining us from, my name is Richard Shan, and today I'm presenting edge oriented data patterns.
So first, a quick rundown of what I'm doing in today's presentation. First, I'll introduce patterns in edge computing, what they are, where we're, where we are today, and where we're going in the future. Then I'll introduce my pattern model specifically, which is a six part pattern, which is a six part model.
And then I'll drill down for the sake of time in today's presentation onto two of the sub, onto two of the six subcategories. Then we'll look at how you can apply this model onto real life under the practice section. And I'll end off with the cook summary.
So first the introduction. What is, where is edge computing right now? Well, C N C F tells us just this year that over the next five to 10 years, edge computing is about to grow, about 40%.
Edge computing is already a hot industry in today's world, and it's only poised to get bigger as the world goes on. But what is edge computing? Well, at its core, instead of sending data to a large data center to process it, edge computing allows you to process data on the edge near where the data is being generated.
Let's give, let's give you an example. I, whenever you go to your local grocery store and you go to the checkout line, whenever you scan your orange or your milk carton or whatever, it's not going back to the grocery store's data center to check what the price of that item is. All the prices are cashed locally in the cash register.
That's edge computing. Instead of sending, instead of checking back with their data center for help on the prices, it knows what to do at the cash register. But why should we be concerned about what edge computing is and what's happening to it?
A few reasons. First is reduced latency because we're processing data on the edge where the data's being actually generated at the cash register. There's no network latency.
You don't have to set, you don't have to send the barcode all the way back to your data center to check the prices. You know the price right there at the cash register. And next is bandwidth savings.
This ties in hand in hand with reduced latency because we're not sending all this data and we're not communicating with the data center data every second. We can save a lot of bandwidth and we can save a lot of networking because we're just all storing the data locally and we're processing the data locally. And third is security.
This is really important, especially if your team or your company is dealing with personal or confidential data. Because all the data is being handled locally and it's being processed locally, you don't have to send it over the network. You don't have to send it over the web to remain data center because it's all being stored locally.
It's gonna heavily increase the security of your data. And last is scalability. At the end of the day, your main data center is always gonna have a cap, even if it's really large.
But as your company scales and gets bigger, so are its edge assets is edge assets are always gonna expand as the company expands. If your team and your company can utilize these edge, these edge assets to be part of your edge computing cluster, you're never gonna run out of computational power. Next, we talked about how great edge computing is, but let's talk about why it can't get to its full potential today.
First is management complexity as opposed to your regular data center where everything is centralized and you have lots of maintenance procedures and a lot of, and a big maintenance crew to check over the data center 24 7. Managing an edge cluster is very difficult because you have a lot of small devices that spread over a large region and each device can individually fail without the rest knowing it's hard to manage. And next is limited processing power.
This should fundamentally make sense. Your edge computing devices perform a cluster, but that cluster comes nowhere close to your main data center. Of course, your main data center is always gonna be your heavy processing power, your heavy processing power area.
It's gonna be your powerhouse of com comput of computing, but your while your edge processing air, why your edge processing cluster is a lot faster and a lot safer. It's always gonna have limited processing power just 'cause it can't compete with your server racks in your main data center. And third is reliability.
Again, managing a data center is a lot easier than ha than conducting maintenance on the edge. If a single small device fails on the edge, it's gonna take you quite a while to figure out where the device was, what it did, and how to fix it. And lastly is data synchronization, especially if your edge implementation is, is designed for not every single edge device to be online and connected to your main data center 24 7.
There can be a large gap between where, what's happening in your data center and what gets and what gets pushed to your edge cluster devices. So we talk about what's stopping edge computing from becoming its full potential, but let's talk about how a pattern can actually solve these barriers. First is standardization.
Patterns are industry-wide best practices. They've been tried and true by industry taught experts. Whenever you adopt a pattern that the entire industry is using, you help standardization.
This is really good because everything, because everybody is on the same page, it's a lot easier for your team to collaborate with other companies and other groups and other projects. Next is re usability Patterns at their core are designed to be overarching generalities that are applicable to many problems. Patterns should be applied, should be applicable to everyday problems by everybody because we're, you're adopting a pattern that's able to be reused and start instead of setting, instead of starting at square zero every time you need to develop a new solution to a new problem, you can adopt similar solutions to, to problems that you've solved in the past.
Third is interoperability. When building an edge cluster, you're probably not gonna be using the same kind of machines. You're gonna have laptops, you're gonna have chips, computers, cameras, devices from all different kinds of vendors and different kinds of companies.
This is gonna be really hard to manage, especially as these companies aren't designed to work together. But if you adopt a pattern that takes these concerns into mind and solves for them before you even get started building your cluster, then you can ensure that your cluster has long-term sustainability and is very interoperable with its own parts. Lastly is faster development.
Again, instead of starting from nothing, every time you develop a new solution to a similar problem that you've already solved before, you can als, you can reuse the same framework using the pattern that you've already used before. This framework is gonna allow you to create faster development to create new solutions at a faster rate. 'cause you're not starting from square zero, you already know the framework and the pattern that you're going to use.
You just have to implement it slightly differently or maybe even the same as you have before. So let's go down into my specific, um, pattern approach model. My model for all patterns in ai in on edge computing is access, specifically AI, flow, composite, clustering, edge, native, central, smart and segmentation.
There's a brief overview of what those, uh, of what those module are on left. But for the sake of time in this presentation, I'll be drilling down on AI flow clustering. So let's start drilling down onto AI flow.
There are six key patterns inside of AI flow. I'll be going over those in the next few slides. Let's start off with basic app flow.
This is your typical and most basic AI application. It's gonna take an input, run it through a module, do some steps, and then give you an output. For example, whenever you unlock your phone, you probably use space id, uh, you probably use space id.
What happens is your phone will take a picture of your face and com and that, take that as the input, then it'll take that input of your picture and compare it using an AI model to pictures that you've trained your phone on that they, that it knows it's your face. After, after running it through the AI model, it'll decide, oh, is this you or is it not you? And if it is, it'll make a decision and that decision will unlock your phone.
Next is the serial flow. So let's say we don't just wanna recognize your face to unlock your phone, but we also wanna see if your face is happy, you're sad, and maybe if you're in a dark room or if you're in a bright room. And we can change the phone's brightness depending on that.
It's multiple sequential models. You're, it's multiple sequential models. After recognizing your face, then you'll feed it into a separate algorithm that'll decide whether you're happy or sad.
The output of one algorithm is the input for another one. And next is parallel. Sometimes you don't necessarily want the models to be one after another, unlike recognizing a face and then seeing it's happy because that means in order to check if your face is happy, it has to first be recognized by the phone to unlock your phone.
But let's say we're trying to analyze a picture of an animal and see if it's a dog or a cat that's gonna be using two separate models. One model is gonna see if it's a dog and output a percentage. One model is gonna, and the other is gonna see if it's a cat and operated percentage and your model and your algorithm is probably gonna choose the higher percentage and output that as a result.
That's an example of parallel processing where two models are being run at the same time with post-processing. Another example is say we have a picture of a bird that's being fed into your algorithm. That's not good 'cause none of your models are gonna know what to do.
That's where your outlier detection algorithm will come in. A lot of the times you're always gonna have an outlier detection model running in the background. That's always running parallel.
So when you're dog and cat algorithm don't know, algorithms don't know what to do with a picture of a bird, the outlier detection model picks up the slack and says, oh, this doesn't belong here. I'll classify it as an outlier. That's a really common application of parallel flow.
Next is cascade flow. Cascade flow is two models or more that are running together. Let's take an example.
In the financial sector of fraud detection. A bank may classify a transaction as fraudulent, but only after ca only. After taking into account a few things, first they might look at the amount of money involved in the transaction, then I might look at when the transaction happened during the day, then I might look at what type of transaction it was and a lot of other things.
And then after all of that, it'll classify, it has paid flow is pretty to series flow. There's one after another in algorithms, but cast played flow instead of discarding, each output keeps all past outputs. So it'll first look out how much money was in output, then it'll say, oh, the algorithm, or, oh, the transaction happened at maybe 9:00 AM and it was a credit card transaction.
It'll take all of that data. And then finally your model will, your fi, your model will take all that data into consideration and then conclude the transaction was either fraudulent or not. As opposed to series flow where the input into one algorithm is the output and then another output and then another output.
You only have one final output and all the rest is discarded Cascade flow. Cascade flow takes in all of the outputs and looks at them together holistically to make a decision. Next is ensemble flow.
Ensemble flow is really important because a lot of the time your AI algorithm is, your AI algorithm is not going be perfect. One model is not always gonna be the best as it possibly can be, and it could even be incorrect if it's based on a bias training pool or the training pool is not large enough. That's where ensemble flow comes in.
Ensemble flow runs the same input through multiple different models and mathematically averages them together to give you a more accurate result in the end. And last is blend flow. Let's go back to the face.
Id unlock your phone example here with blend flow, instead of just relying on a visual picture of your face to decide if it's you or not. A blend flow, uh, algorithm may also require audio. So you have to show your face to the, to the camera, but also say something this can, this can be combined for better security as well as not only will it analyze your face, your facial features, but also your voice.
And if you combine them, it might analyze you with your lips and your facial structure as you're speaking to ensure it's you. Now let's move on to the second. Now let's move on to the second part.
The second sub module, which is clustering. Clustering has three subcategories. First is the multi edge cluster, second is the edge node, and third is the edge multi resource.
Starting off on the multi edge cluster, let's say that you are working with Google on the edge and you have to process thousands if not millions of queries every second. Of course, your home PC is not gonna be enough to process all of these queries. That's where clustering comes in.
Clustering allows each individual processing unit to be a node inside of the larger data clustering system. Especially in today's world, Kubernetes is getting really hot. There are a lot of open source projects that can help your team build a multi edge cluster for the first time, like K threes and micro K eights.
It's an equivalent to building your data center on the edge. It's building a cluster on the edge, but you don't have enough computational power, like you don't have all of your server racks and stuff. And it's kind of like building a data center in the kitchen.
You're having all these smaller devices work together in a cluster to mimic the effect of your data center, but it's working on the edge for larger computational power and all of the benefits that come from working on the edge. Next, let's talk about an edge Node and edge nodes are really important. Edge nodes allow devices to be federated as part of a cluster, but they aren't a cluster in and of itself.
It's like a smart light bulb. A smart light bulb can be a node on a cluster. It's a small device.
You obviously can't load a cluster onto a smart light bulb because it's not even close to advanced enough, but it's still part of a node. Edge. Nodes allow devices that can't support a cluster to still be part of a cluster and support its computational power.
And lastly is multi resource. So not only do you have the device in and of itself, but each device has computing power. Let's look at A U S B drive.
U S B drives don't have much computational power by themselves, but they can store data. They're good at storing data, they can store a lot of it. These simplistic devices and others like the U S B drive aren't actually nodes per se on the cluster, but they are resources.
Multiple small data storing devices can work together to become a node of the cluster. That's where they can be re that u, that's where the resource term comes from. The whole cluster can utilize these U S B drives as a resource to stor data.
So while they're not actually using it to perform any computation, any computations, they use it to store data there. The U S B drive is a resource of the larger cluster. People have already started working with these concepts and building projects with these such as RY and working with K threes.
So now let's discuss how this can be used in practice. First is design considerations. This is just a generality.
You should apply these when you're working with any pattern, not specifically just access first. Good design is built to only solve hard problems because these patterns are general. They're all because these patterns are pretty general.
They also have to be pretty complex in order to be able to solve all of these problems. So that's where your judgment comes in. You shouldn't be applying a complex pattern to solve a really simple problem.
If you know you can solve the problem pretty easily, then just go and do it. You shouldn't apply a complex model to it that's just wasting your comp, your computational resources, and it's probably wasting your team's brainpower trying to figure out how to apply a complex pattern to a really simple problem. And next is malleability.
Patterns at its core are designed to be broad generalities that everybody could use for everyday problems. That means when you're working with a pattern or you're even making your own pattern, you have to make sure that it's applicable to a lot of different problems, even if it's only applicable to one subset of problems. And you need to make sure that it's, that it's able to be a general solution.
Next is unintended consequences. We all know that AI is getting really hot today, and with that comes a lot of unintended consequences. With ai, AI can generate false information and the like, your, your, your algorithm and your project and your team need to make sure that you're not just gonna cross these bridges when you get to them.
You need to make sure that you can preplan these unintended consequences and that if, and that when not, if the AI starts to generate a wrong prediction, then you know that it's not gonna have a long-term effect. And lastly is integration. AI is already being used in every aspect of human society.
It's being used in medicine, it's being used in education, and it's being used in the financial sector. You must realize that AI is not always gonna be correct. And, and, and we've seen lots of news about AI being incorrect and causing lots of damage.
You need to ensure that when you're implementing AI and and these different patterns into your project, you must ensure that you're considering the effect on a larger ecosystem. What happens if the AI messes up? What happens if it goes wrong?
You need to make sure that you're implementing fail safes and you need to make sure that if the AI messes up, nothing bad will go wrong. Next, we'll talk about how we should be using patterns effectively. Again, this is a generalization that can be applied to using all patterns, not specifically just access.
First is to keep it simple. Remember, patterns are designed to solve hard problems. You shouldn't be applying a complex pattern to a really simple problem.
If you can solve the problem by itself, go ahead and do it. But at the end of the day, if you have concluded that you really do need a pattern to solve the problem, then you should keep it simple. If you, if that doesn't work, you can take one, you can take one step up the ladder of complexity and go to the next simplest pattern.
But don't overdo it. Don't just get, grab a really complex pattern to solve a really simple problem. Again, that's just wasting away your resources and your team's brain power.
Next is optimization. It's really key to continue optimizing these patterns. Remember, patterns are key.
Patterns are created to be general solutions to everyday problems. What that means is that patterns are applicable almost everywhere. But what that also means is patterns may not specifically be a good fit for your problem or your project.
This means is after you find a pro a pattern that can fit your problem really well, you should con you, you should continue adapting that pro that pattern. You shouldn't stick with the pattern that you know is going to work, but you should adapt it. You should fine tune it.
So if it's your project really, really well. But lastly is to adapt your problem. So not only do we want to fine tune the pattern, so you also fine tune your problem.
You wanna find a middle ground where you're adapting your problem and you're adapting your pattern. Because we know that these patterns have been tried and true by industry actors and they know that they're all, and we know that these patterns are working. What this means is that if you don't, if you don't feel comfortable changing the pattern up too much to fit your problem, what can happen is you might write part of a solution for your problem.
You might write a solution to your problem such that it fits the pattern a lot better. So again, find the middle ground where you're able to adapt your problem to the pattern and you're able to fine tune the pattern to your problem. Now let's move on to best practices.
First is data minimization. Remember, edge computing is really fast and really secure. You should be using edge computing for these kind of critical decisions that need to be made in real time.
You should not be wasting your edge computational power on something that isn't very important. Remember, edge computing power is limited. Use it for real time and important decisions.
If you have a, if you have a, if you have a computational algorithm, that's, that's pretty heavy and that's also not very important. You should continue running it either in the e on either in the cloud or in your data center. You should only be running computations on the edge that really matter that you need to have done in real time and on site.
And next is latency awareness. Remember, you should be processing real-time tasks at the edge and regular tasks and heavier algorithms can be done in the cloud. Keep your heavy algorithms and keep ev and keep your non-essential algorithms in the cloud while your real time decision making algorithms are kept in the edge.
And lastly, distributed computing. Again, edge computational power is limited. You should be spreading your tasks across multiple mode, especially if you have a really important task that needs to be done instantly.
You shouldn't shove that entire computation on one edge node. You should be distributing it across your edge, across your entire edge cluster. You should be distributing it across multiple edge nodes so that the task can be done faster and without and with less stress.
So now let's move on to a quick summary. At the top, we explored what edge computing was. We explored where it is right now and where it's poised to grow into the future.
We explored what is edge compute and how patterns can be applied to it. I introduced the six part access PA pattern and we drilled down specifically into AI flow where we talked about different kinds of AI flow patterns. We talked about, uh, we talked about uh, basic serial parallel cascade ensemble and blend AI flow patterns.
And then we drill down into the cluster part. Cluster is ma is the mo is the meat of this presentation and is probably the most important thing you need to take away. Clustering tells you how exactly you should help build your own cluster through multi edge clustering, through edge nodes and through multi resources.
This will, this should let your company and your organization and your project get started by building a Kubernetes cluster on the edge. Then we explain some best practices. Remember, these practices are not just to do with the access pattern.
All of the subcategories, these practices should be done with all patterns. In general, you should make sure that your pattern is adv is malleable, but it's not over specific. You should make sure that you are planning for pre, that you're planning for unplanned contingencies that you need when your pattern goes wrong.
And at the end, remember you should always try to find the middle ground between your pattern between adapting your problem and adapting your pattern. That's really key and that's probably what you should be taking away from this presentation. Patterns are generalities.
Your problem is nuanced, so you should try to find the pattern that works, but at the same time addresses your problem, adapt your problem to the pattern so that way your pattern fits really well. But also adapt the pattern. Don't settle for a generality.
Continue fine tuning your pattern and continue fine tuning your problem until they fit really well together. Thank you guys all for coming today and I hope you guys learned something from this presentation. Continue working on developing your own patterns and, and continue working on working on other people's patterns.
I hope one day that you might be able to help make another pattern and contribute to the community. I put my contact information at the bottom and there's a QR code that you can scan, uh, with your phone to get my contact information. Once again, thank you guys for coming to my presentation and I hope you guys have a good rest of the conference.





