Nasuni’s John Capello and Tetra Tech’s Adrian Grancea on Accelerating Edge Data
John Capello of Nasuni and Adrian Grancea of Tetra Tech discuss Nasuni’s recent Nasuni Edge for Amazon Simple Storage Service (S3) availability announcement. Nasuni Edge for Amazon S3 is a cloud native, distributed solution that allows enterprises to accelerate data access and delivery times while ensuring low-latency access that is crucial for edge workloads, including cloud-based artificial intelligence and machine learning (AI/ML) applications, all through a single, unified platform.
Transcript
This is Textron tv. Hi. I have the great pleasure of being joined by a couple gentlemen here.
We're gonna be talking about hybrid, uh, cloud network hybrid storage. You know, as we live in this world of hybrid clouds, how do we manage our data across all those environments? I'm joined by John Capello, who is with Nasuni.
I had an announcement recently that we're gonna be discussing, and Adrian, cia, and I believe Adrian's a customer, right? Or user of the technology, uh, working with Nasuni, so, uh, with TetraTech. So, uh, John, why don't you start off, just tell us a little bit about your role in the company and, uh, overview just of what Nasuni is, and then Adrian, if you just mention what you do.
Sure. I'll start off with a, just a, a quick intro. I'm John Kiel.
I'm Field CT at Nasuni. Um, I get the pleasure of having to work with a lot of our largest customers. Um, TetraTech being one of our, um, uh, of our biggest customers that really pushes what we do at Nasuni and really takes advantage of a lot of things that we do at Nasuni.
Um, a lot of the innovations that we've been working on recently, um, I've been lucky to be able to engage TetraTech on that. Um, I also work with a lot of our cloud partners, so, um, a key part of our solution is that you have an object storage underneath that you have, um, that you placed Nasuni on top of. And so we work with all of our major cloud partners, AWS, Amazon being, um, obviously one of our biggest ones too.
Um, but I'll hand off to Adrian for, um, intros, and then I'll say a little bit about Nasuni after that. Great. Sure.
So, um, I'm Adrian cia, um, with TetraTech, uh, part of the IT system engineering group. Uh, and, um, my main involvement with Nasuni was, um, designing the architecture and, um, doing the project management for the transition from the typical Windows file systems to Nasuni ecosystem. Very good.
Well, tell us about the announcement. I think it's relative, relative to S3 and AWS environments. Sure, yeah.
So, um, uh, just to start, lemme kind of, uh, set the table about like what Nasuni is. Um, now I'll talk about the fact that we've added S3 as a protocol on top of our massive file system that does a lot of really good things for us. Um, so number one, you know, Nasuni really is a, as Adrian was talking about, kind of the, the evolution of file infrastructure within the enterprise.
So think of it as a modern approach to how to store your unstructured data. Um, up to this point, a lot of customers have been storing it on nas on, um, file servers. So whether it's a Windows file server, whether it's something like, um, you know, a NetApp or an Isilon, um, those are all fantastic data center based solutions for storing data and storing unstructured data.
But Nasuni came along and we said, what if you could have all the power and the benefits of a nas, but you are backed by some kind of, um, cloud scalable solution. Um, so what we did was we built a, um, infinitely scalable versioned file system on top of object storage. So you as a customer, you've got access to your AWS account or an Azure account or Google account, um, or maybe you have a private object store.
Um, and we, um, allow you to then create a file system within that object storage within your own, um, your own tenant. And on top of that, you'll, um, then be able to access your data through a software defined layer of these edges that we call them basically like, um, uh, look, look and feel a lot like NASA's. So for us, you know, we're taking a different approach to how you manage file infrastructure because we have a really, you know, unbelievably scalable foundation to it with an object store.
Um, we have built in security and how we actually encrypt the files from edges and move them into the object store. But we also have a versioning capability that is, um, really unparalleled in the industry. Um, people talk about restore points or access or, um, you know, backups or, um, uh, you know, uh, snapshots to be able to restore their data.
And those are usually hundreds or thousands of different restore points. We have literally millions to billions of restore points because you can restore files and folders. So security for us is built in, data protection is built into us.
And then the fact that's a, a software defined model where the access points are not the object store, that's not how you get to your data. You're actually getting it through these edges, these, um, software defined appliances that live as virtual machines, really, wherever you can deploy a virtual machine. So up to this point, those edges have been speaking files.
So it looks like a nas, acts like a nas, I can write to it through SIFs, right to it through NFS. Um, but what we just announced is we've added a new layer on top of that, a new protocol in which you can get into that, uh, file system, and that's using the S3 protocol. So, um, what you don't see anywhere in the market today is the ability to have this massively scalable file infrastructure, or let's just call it now a global namespace.
I can write to it through sips, I can read to it, um, through NFS. And now that same namespace, without migrating any data, is now available through S3. So in some ways, you can think of what we've developed, not just as a way of adding protocols into your file system or extending out the global namespace, but now up to this point, what you have thought of as almost like a, um, like a data lake, um, an object, um, based architecture for storing unstructured data is now available as caches, wherever you wanna put those caches.
So now think of it as the fact that like my S3 based, um, uh, access into a global namespace can now happen through caches that I can put anywhere. Um, super excited about that. And, um, TetraTech has long been the soon customer, um, uh, uh, Adrian, we've, we've worked together on lots of other projects including, um, a, uh, analytics and tier system.
Um, I know you guys have been a, um, huge supporter of some of the original data propagation analysis that we've done, as well as, um, what's now, um, uh, been announced called na, Sunni iq, but, um, you guys were also there to, um, really help us to think about S3. And so, um, if you wanna talk a little bit about sort of how S3 works within your environment within TetraTech, or how you're seeing the opportunities for it. Yeah.
So, uh, just to, um, emphasize a couple of points here that you mentioned, John. Uh, first of all is, uh, this idea of, uh, centralizing, uh, file system, the whole, the entire, uh, data stores for, for the enterprise, uh, without sacrificing the performance. So if you have a centralized file system, uh, you need access, quick access, and that is through the caching devices that you mentioned.
And second biggest, uh, advantage is, uh, the, the backup and especially the restored. Uh, everybody e every backup system has a big flaw when it comes to disaster recovery. And, uh, here it was, it's the main advantage because you can restore in minutes or probably hours instead of weeks and, uh, in, in case of a ransomware attack.
So those are the two main features that were attractive to us when we started. And after that, we discovered, uh, more and more features that you continuously add to this ecosystem. Na Sunni is a file services applications.
Uh, so it's a big distinction. It's not just a file server file system, it's file services. So services are added all the time.
And you mentioned the SUNY iq, um, a side of the basic, uh, anti, uh, um, virus scan and the ransomware protection and so on. This, uh, S3 is the newest edition, and, um, it, it, it, it evolves. So it's an evolving system.
It's exciting, it keeps up with the times. And, um, recently, uh, this S3 that we tested together is, is a, uh, a benefit for it. It may be a niche, uh, type of, uh, enhancement, but it has a future because more and more application will support natively S3.
And this is the, the key part, because uploading data into the cloud from the field, uh, it's a painful, uh, endeavor. Right now. We have, uh, lots of field engineers going in disaster areas and capturing, uh, huge amounts of data later data, uh, images that they have to be uploaded and processed, um, internally.
But right now what they do, they have, uh, they try to use Dropbox, uh, USB drives. Um, it's a multi-step, uh, process that will bring data from the field to the, uh, to the cloud. So by, by having S3, uh, protocol available, um, it streamlined the whole process.
Uh, it has a consistent, uh, transfer rate based on our testing is not like windows that goes up and down all the time. And, uh, the SIFs that has all the limitations, uh, this is, uh, it, it's a consistent transfer rate and it's very stable. It's, uh, resilient to, uh, internet disruptions.
And, uh, it, it, it's solid. It, it has a big future into, um, a, a a range of applications. We tested with one of two, and it's based on whatever our clients were required from us.
Okay. So I'm guessing that, uh, part of this, there are a couple things that stood out in what you talked about. One of them was consistency across environments.
Right? Now you're talking about a essentially what a network attached storage like type service in the cloud in, in S3, but also, um, you know, file systems are great for just putting information, but there's a lot of data management practices that go around that, like we talked about, you know, restore points and backup and restore and things like that. What, what were some of the most important things you needed that S3 didn't have now that you've got uni on top of S3?
So for us, um, it, it is the, uh, how you move data from the field to the cloud. Uh, and I'm talking large amount of data, hundreds of gigs, files that are into gigabytes that, uh, you have to process them on the field, put them on connect USB drive mainly, or put them into your Dropbox and go into the office. And two or three step later on, you, you have data where it has to be.
So, S3 is, uh, it works. Um, we, we are testing right now to use VPN less, um, and there are some security issues there, but, uh, nothing that cannot be solved. Uh, and so it'll be a one touch, uh, data movement from what you have on your laptop in the field to, uh, and the SUNY Edge appliance into the cloud and data will be readily available.
They can process it right away. Uh, there are, um, hundreds of, uh, of gigabytes, if not terabytes, of data uploaded daily, that that's the, the biggest problem. So we have lots of contracts with, um, uh, US agencies that goes into the disaster areas, and there are hundreds of people in the field collecting data for insurance and all this kind of, uh, applications Jump in.
John, I, you know, I, I'm, I wouldn't be the first person, I'm guess Adrian isn't the first either, has worked with the cloud and said, getting data to the cloud to and from the data, and then synchronize it in across environments, managing it as a common namespace between what's in Amazon AWS and as well as other environments. Um, that, that's, that's a big challenge for data ops groups, if I can use that term, just in general for people in the data business. Yeah, this is, this is one of the exciting things here is that, um, you know, the, um, the, you know, the promise of the cloud is it's, um, you know, it's capacity, it's, you know, performance.
Um, but one of the challenges is trying to work with, you know, a a a company that's been doing, um, you know, kind of amazing engineering work for decades. And a lot of the, you know, current processes don't fit within their, um, uh, within a cloud model. And, and specifically I was thinking about the, the, the lidar and maybe some of the data flow there, which is, um, it's great to be able to upload NS three, but if you know the way that you're gonna be analyzing that means that I need to mount it as a, um, NFS share or mount it as a, um, a SIF share somewhere, and I need to read it because of some other, or some other process, some other application, you know, needs file access.
Like, um, that's a really unique thing that Nasuni can do. Um, so rather than having this complicated as, as you were talking about this data ops flow, which is like upload to one location, you know, process in location, and then transform and move it to another location, and then maybe do some processing, then move it to another location. You know, our, our goal is to simplify that as much as you can.
You know, when we say a global namespace, we really mean it, like it's global, everything can access it. And access this doesn't just mean like, Hey, I have the ability to, um, you know, send a request from some other location. It's like, I've got the protocol I can access it with.
And that's been the one big thing that's been missing in the cloud, which is I wanna be able to write using whatever modern protocol I, I, I, I need. And S3 is kind of the modern object protocol I wanna be able to read as well, and maybe even write as well using these other file protocols. And, um, Adrian, if I heard what you're saying, like you, you're, you're up doing uploading, um, the LIDAR data, it's gonna get processed.
Is that process necessarily gonna be in S3 using S3 protocol, or could it also be used, um, using the SIPS protocol for that? Well, there are, uh, specific applications, uh, that, um, uh, that use them for modeling. And, uh, so it's mainly SIFs, uh, when they use them.
But the biggest problem is, uh, uh, on the operational part, it, there are multiple, uh, locations where data resides. So there are offices sitting on piles of USV drives and we don't know has been uploaded successfully. Yes or no.
SIFs is not very reliable on that. It's large amount of data, especially if you go with the VPN. So, um, just moving that data internally in a reliable fashion, that's the a, a big challenge.
So S3, as I said, it, it's a very, um, predictable mm-hmm. Upload, uh, transfer, um, you, you just start it and you can't forget about it. You don't have to watch it all the time.
Oh, the, the disconnected, or right now it's on zero megabits per sec, uh, kilobit per second, and then so on, and its starts going again and oh, how long it'll take, I don't know, maybe an hour, maybe a day. So, um, it, it's very, it, it simplifies the whole upload process, the whole, uh, uh, single source of truth for, for later processing. And we can't talk about data too long without talking about security.
And, uh, you know, there are, yeah, you know, some security capabilities that are in the cloud. Different cloud providers have, have their own model. Of course, we have ours within our environment.
Uh, how about how does this help with the, the security aspect of managing that in AWS Um, I could, I could take the first, first crack at that. Sure. Um, so, um, for us, we've sort of built a lot of our security model in terms of access control around active directory.
And so when we have a file system that's built on top of the object storage, we are creating all the metadata structures to be able to represent your active directory controls as well. Um, so when we developed, um, this integration into S3, we were then kind of confirmed with the idea of like, wait, there's two different types of access controls that are coming into play here. You've got your ad system, and then you have what is, you know, typically the, the secrets and the access keys that you use for S3 protocol.
So, um, we worked hard with our customers to figure out how can we best map these two things together. And what we've come up with is a way in which you can create your own secrets and access keys, and you can put them onto an appliance. And so that appliance will have those secured there.
Um, and then you can map those keys to any part of your global file system, any part of the volume. So you can both put at the top level and then give access to the entire entire volume. Or you can use those keys to be able to map it to lower within the tree, which itself is its own way of being able to, um, map out, uh, security protocols as well.
So we kind of think of it almost the way that like, um, NFS users sometimes work with, um, with SIFs users being able to have like admin access or, or, um, super user access across the tree, and then being able to, to map exports to that. Um, but our, our goal here was to make sure that you, you don't have to re migrate or re permission your current global namespace, but we want to be able to have that work in, um, conjunction with the IAM model, the access key model. And so, um, we feel like we come up with a really nice little nice solution here.
It's easy to use, easy to manage, easy to update those, um, uh, those keys. Um, but it doesn't mean that you have to get away from what you're using today, which is, you know, for most, you know, large file infrastructure, it's gonna be, um, active directory. Yeah.
So, um, just to add that, uh, from our perspective is, uh, a side of what Nasuni is doing in terms of security, our security team is it's extremely, uh, diligent into assessing new technology. So anytime we start with a new protocol or a, a open a, a, a a hole into the firewall, everything goes into a secure area is mitered for weeks. And, uh, there are reports and security teams give us the blessing at the end of it.
So, um, it's a, uh, defense in depth, uh, and this is what we try to, regardless of what the vendor is saying, we take our own precautions. So, um, this is pretty much what, uh, how you address security. I'm curious.
Um, so we have to bring up also ai, of course, and people are investing more and more in their, in their, in their products, having AI capabilities, but also in the software that we're developing, uh, including models and machine language algorithms and generative AI as well. But that, that involves pushing a lot of data around, uh, whether it's training, training models, or it's continually feeding new data into models. Um, that data management challenge, transferring that data, you know, models start to drift, they need to be replaced.
All of those kind of things presents a new set of challenges for a lot of data management teams, I would imagine. Talk a little bit about how this might help with that. You wanna start out, John?
Sure. Um, so this is the one I'm really excited about because I think there's, um, a lot of things Nasuni does kind of natively today that, um, you know, ai, um, AI engineers of the future will just start to, or are just gonna start playing with it. Um, and, um, one big one is the fact that we can version at any level of the global namespace and with a lot of granularity over time.
So I like to think of it this way, like if, if, if you think of, um, you know, uh, Adrian's, um, the file systems that, um, Adrian has on us at TetraTech, um, I'm gonna guess I haven't looked at the exact numbers, but I'm gonna guess they have somewhere in the order of a few hundred million restore points. So if we wanted to, we could actually go back to, um, let's say a year and a half ago, we could pick a random date a year and a half ago, and we would say, what if we trained a model off of the dataset that existed on that day? We could do that with nai.
And to do that, like it really requires you to figure out, well, which part of the tree do I want to be able to train it off of? Maybe it's the projects folder, maybe it's a specific subset of projects, folders, maybe it's a combination of different projects, folders. And then we would tell the system, okay, let's restore in a read only way, just a, an appliance to that point in time.
Well, if you started to train on that day, just that point in time, and then you move forward and you trained a different day, what we've basically allowed is for you to take a training set instead of starting on day one. You can go back in time and start training your data from as long as you've been on Nasuni. And I think the power of AI, as we've seen today, is the fact that I can keep retraining these models, but almost everyone starts, uh, from like, you know, T zero as like the day I start my training model.
You don't have to do that with Nasuni, start with whenever you start create your Nasuni volume. So now TetraTech is available to them, you know, literally like, you know, um, a thousand times more actual training sets that they can use to train their model just by the fact that we have this incredibly granular version file system. So that's the one I'm super excited about.
Um, the second thing is the fact that the S3 protocol is kind of a modern protocol. It's one that works within a lot of these, um, kind of more AI based workflows. If you have a data scientist that is, um, you know, um, is, you know, working off of their Jupyter notebook and wants to test something very quickly, um, boy, it's so much easier to say, okay, well just point it to this S3 endpoint now than, you know, point to another S3 endpoint.
Some of them do use files and you might wanna mount a, a directory, but, um, we just give you a lot of flexibility for how your AI engineers, your, your, um, ML engineers wanna work today by just giving them a protocol that, um, it works so easily with their tools. So I'd say those are, those are two of the things. Yeah.
So, uh, um, for us at, at TetraTech right now, the, the main emphasis is to provide more services to our clients. So, uh, traditionally we just gather data, process it, and submit it to the clients. They locate to it, they use it for a while, and that's it.
Um, using an AI will allow us to provide more services, give them, uh, the ability to search through years and years of collaboration and, um, extract more value out of the existing data. So it, it's definitely very an exciting field and, uh, we're just scratching the surface of the surface at this point. So it's a long way to go, but data is there.
And as you mentioned, there are so many restore points that we can go really granular and, uh, be very specific. And I think many organizations are, are starting to learn some new challenges with generative AI and the training and, and, uh, yeah, you know, prompt engineering and so that, so there's a lot of data that you need to both develop and test those environments before it ever makes it to production as well. 1 point a.
Well, what was that, right? Just to give a kind of ridiculously, but probably common example. Um, but any other things on the announcement that you wanted to make sure we cover?
Uh, just, um, uh, a couple of the, um, the points that I think kind of were made here before, which is, again, it's all part of the existing na Sunni system. You have existing na Sunni volumes, you can apply that to the, um, S3 protocol to that. Um, one other little thing to touch on, um, because we're a version file system, because we're based off of, uh, um, really the appliances are using XFS underneath, but we're creating our own Uni FS file system in the background.
We have our own way of storing metadata. And what that means is that when we implement the S3 protocol, we now allow you to add more metadata than you would just using your standard S3 service. Mm-Hmm.
And, um, we're, we've seen customers get excited about that already. Um, I think more and more as we're in this AI space more and more where metadata becomes, um, such a locus of innovation, um, just having the ability to have more than say like, you know, 12 pairs of 12 key value pairs or more than, you know, 4K of, um, metadata. It's, we, we've, we've tested well, well beyond that and, um, the system that supports very large metadata structures.
So we're excited because now you can, um, think about your S3 target, not just as data, but think of it as being metadata rich as well. Fascinating. Any other points you wanted to make Adrian?
Uh, no, I think we touched on pretty much on anything, uh, that was pertinent for, for this. And, uh, I, I'm looking forward to, uh, I, I'm pretty sure S3 has a future because it's something native to the cloud. So, uh, who are excited to explore more and more applica, hopefully more and more vendors will have applications that supports S3 natively.
Um, so, uh, this is something that, um, we're looking forward to as it is right now. They, you, you need a third party client. Um, scalability is, I mean, if you want to go to thousands of users, uh, you have to manage something extra, but I'm pretty sure that the future is there for S3.
So, uh, looking forward. Well, fantastic. Congratulations on the announcement.
Ready to, uh, have this out in market. Where can folks find out a little bit more about this? John?
com/ S3 Edge, in particular, S3 Edge, EDGE. All right. Excellent.
Well, thank you, gentlemen. It was great talking with you. It's always, uh, it's fascinating.
It's kind of another way the cloud kind of grows up, right? It's, it, it innovates in its own way and this is also makes it a little easier for all of us to use all that great storage we have up there in meaningful ways as what AI being a big important one. So thank, thank you, John.
Thank you, Adrian. Thank you. Thanks.
Bye.