Implementing Zero-Trust IT Architectures with Bloomberg’s Phil Vachon
Phil Vachon, head of infrastructure for the office of the CTO at Bloomberg, describes what’s really required for organizations to implement zero-trust IT architectures.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Phil Von, who is head of infrastructure for the office of CTO at Bloomberg, and we're talking about Zero Trust as we head into 2024.
Phil, welcome to the show. Thank you. Thanks for having me.
Mike. I think a lot of folks have been talking about zero trust. It's probably one of the most hyped turns in all the cybersecurity these days, but what is it exactly?
'cause I think a lot of folks get confused. There's vendors running around out there that say they have a zero trust platform, but to me it's a little bit more than that. What's your take?
Yeah, I I mean, the, the way I think about it is like all the vendors are trying to brand themselves as a zero trust solution, but how do you have a solution to something that's really a design philosophy? Like thinking about how you build your systems to be as secure as possible, and what is the, the mindset that your team should have as they're building new infrastructure and introducing new technologies or cleaning up legacy as it may be. So you guys are down this path, but how far, it seems like it's a journey.
So how do you know, um, as my kids would say, in the backseat of the car, are we there yet? Well, I think, I think if folks are asking, are we there yet, um, maybe they're asking the wrong question. Of course, my son isn't quite talking yet, so maybe I get, I, I haven't had that pleasure yet of being bombarded with that question.
But, uh, but it actually is a, a journey that doesn't have a clear destination because in the end, you're adopting a design philosophy set of design principles around, you know, validating identity and, and, and authenticating identity continually, or ensuring that you, you know, opt for principle of least privilege. People are only privileged, or systems are only privileged to do exactly what they need to. So, you know, getting onto that journey requires taking some steps around making sure you have the maturity for things like identity.
Like do you know the things that you need to be able to authenticate or who those people might be or where they're expected to do so? Um, but then also there's a certain amount of maturity we require around understanding that the solutions for this, for this type of approach, uh, vary a lot depending on if you are looking at your corporate IT estate, for example, your end user devices that, you know, your employees are touching day in, day out to do their jobs, or looking at the data center or public cloud or infrastructure you're using to deliver services to your customers. In fact, uh, a well-designed product offering might include, uh, functionality to enable customers to have their own zero trust type solutions built around what you're offering is, and, and using, leveraging that.
So, you know, that's a very common topic, of course, in the business to business in the enterprise space. Um, but yeah, it's, it's, you know, it's a journey and, you know, we're, we're well into that journey. We've adopted these philosophies of trying to isolate as many pieces of infrastructure and equipment as possible to make sure that, you know, the network isn't the only control we have.
We have authentication pervasive throughout our infrastructure, how people authenticate to our systems internally or, or even how services talk to each other or authenticate to a database, whatever it may be. Uh, so all those types of, of, uh, kind of first steps are necessary. But, um, you know, at the core of it, um, you know, it's knowing what, again, like I said, what it is that you're trying to authenticate to and, and what we'll be trying to authenticate.
Uh, because without that, that inventory of identities or, or entities you just are, are completely lost trying to figure out really what something should be authorized to do. On the face of it, it seems like it's kind of common sense. I would argue we've been checking identity since the first caveman grunted who goes there.
So my question to you is, what is the challenge? How did we get to this point and and how difficult is it gonna be to unravel at all? Yeah, and, and I mean the, the challenge, especially if you talk about the identity problem, I mean, you know, humans, we, we rely very heavily on, you know, interpersonal interactions.
You and I talking, or you and I meeting face-to-face, shaking hands, whatever, getting to know each other on a personal level, that is how we identify each other for a computer. It's a harder problem, right? And, and we've seen various iterations on this.
You know, first it was, well, if I'm a, a human that's in your list of entities that should be allowed to access the system here, you know, my voice is my passport, ha uh, my, you know, whatever my password may be was how I would authenticate to the system. Obviously, we learned that passwords are something that can be stolen. Uh, and, and over the years we've evolved to include, you know, additional factors of security for multifactor authentication.
Now, all of these try to give us some level of assurance as to who those, those entities are, or who those humans are. But around that, there's a whole governance and management regime you have to think about. And there's of course, the life cycle of your second factor authentication.
I mean, uh, we saw this with, uh, a vendor who shall remain unnamed, who's providing authentication services and identity management services to a lot of customers, uh, where their entire set of controls was bypassed by a customer's help desk because of, uh, you know, a successful social engineering incident. So those types of issues, you have to think about them and think around them, but also just getting into the realm of how do you manage that list of entities as it evolves. We are doing as a, as a part of doing business, people change jobs, people leave the company, their roles are, they have things added to them, whatever it may be, having a good strong regime for how you deal with movers, leavers, joiners in your, in your enterprise, but also, um, how do you manage that list of, uh, you know, attributes that relate to what a person's role should be.
Those are all very complex business problems. And in the end, you know, they're the type of problems that programmers really hate to deal with because they're, they're very, very mushy and human problems. Whereas, you know, like it's, it's not something that you can just write some conveniently closed function that says, oh yeah, you know, this person shut this privileges, whatever it may be.
It's, it's a reflection of the business itself and, and of the roles people play in that business. Seems to me we're a little overly obsessed with identity as it pertains to people, but as far as I can tell, the issue extends all the way out to individual software components, the hardware, everything has identity and everything can be abused. So are we underestimating the scope of the challenge?
0 ads governance in there. I can't remember if it's the first or the last, but that's very important. But identify, and you've gotta know what are the things, which could be software, services, equipment, you name it.
In fact, those things could be layered even. And, uh, where, where this gets interesting is, yeah, we, we have a problem in regards to, um, the service identity side of the story or, um, as the system identity side of the story even. We've got many different mechanisms, you know, from trusted platform modules and as, as secure elements and servers, uh, or many laptops.
Of course you have, um, you know, apple devices have their secure enclaves. So there, there are some very clever hardware solutions. This problem, in fact, you can even look at UBI keys for identifying humans, or Bloomberg has its own b unit technology and bsat technology that we use, uh, with similar capabilities.
And, uh, we, you know, have to sort of think about, um, how do we interchange those identities first and foremost? And second of all, like, how do we evaluate what they really mean? Like, should I be trusting an identity from an Apple device the same way I trust, uh, the identity from A TPM?
And I mean, those are problems that are extremely difficult to reason around, but also, uh, very foundational to how we have to actually think about authenticating systems and services. But then the other part of this is, is, you know, the interchange of those identities is, is actually not standardized very effectively. There's some emerging standards, uh, I I'll call them defacto industry standards, uh, the spiffy standard, uh, secure production identities, uh, for everyone.
Um, spiffy is, um, a fantastic first step into this realm, but it still has a host of problems you have to deal with. Like, what is the list of entities you want to be able to authenticate, or what granularity should you authenticate? Is it a per process on a server that's talking to another process, or is it, you know, per server?
And of course, um, the team behind Spiffy would, would acknowledge that it's, you know, you have this classic turtles all the way down problem where, um, you know, I want to authenticate the bedrock hardware. Well, how do I know to trust the bedrock hardware? Well, maybe I need to authenticate the manufacturer.
How do I authenticate the manufacturer? So it's just, it is turtles all the way down. You have to kind of pick that, that bottom turtle that that will get you there.
But, but kind of zooming out a bit like this is kind of indicating the fun fundamental problem we have, which is, uh, what are the things we want to authenticate? How do you authenticate them? So the mechanism being something like spiffy or other mechanisms that are adopted across industry.
And, uh, of course, um, you know, what is the list of things that you then wanna be able to pull out and make an authority decision about? Like what should something be authorized to do at point in time? So there's, there's a number of these, um, just foundational problems.
Uh, I I kind of, I, you know, I'm picking on Spiffy 'cause I actually find it, you know, it's a technology that we've adopted. I find it to be one of the most compelling means to convey identity in a data center environment. Um, but you know, there's a lot of infrastructure you have to build up to support it.
And of course, um, you know, supporting spiffy identities and environment that is, you know, bare metal, you know, service running on a single server versus, uh, a server running in Kubernetes and a container, uh, are very different scopes of problems as well. So there's, there's still a lot left to the implementer. Uh, so this is where, where we've been focusing a lot of our energy is like, how do we make sure that these production identities are available to, uh, everyone in the company, uh, by default so that they don't even have to think about at least where do they get the identity.
Now, how do we or get the authentication, uh, of course how, you know, how do you identify a particular service that's a whole other can of worms that you have to open. What is it that you and the rest of your colleagues at Bloomberg know now that you kind of wish you knew when you first got started down this whole path? 'cause there's a lot of folks that are still early in the journey.
I think the, the hardest lessons have been assuming that, uh, pieces will play nicely together and interoperate nicely together. Um, this is an emerging space, especially in the data center. And if you're building largely proprietary technology, so if you're a hyperscaler, let's say, and you're not building on top of third party components, or you're building largely on a closed ecosystem, um, these problems are, are usually a little more turnkey to solve.
Bloomberg, we're a big investor in developing our own technology and open sourcing it. You can see this like in our, our own database infrastructure, uh, relational database that we've actually opened, sourced. And, and, uh, of course we use it extremely broadly internally, um, through to our service mesh through to kind of databases, other databases and all that.
We've, we've been big adopters of open source, uh, and big advocates for open source. Um, that doesn't mean we don't also have commercial products in the mix. So getting all of those pieces to talk together nicely are coming up with design patterns for those is, is quite difficult.
Um, the one thing that I will say has made us the most successful has been taking a pragmatic eye towards it where there were no identities or where you were, you know, purely working with network controls or similar in the past. Uh, taking those areas and identifying the biggest swaths, the biggest wins, where you get 40, 60% of your services able to authenticate to each other, that's a huge win. So identifying those, uh, high impact and, and then of course, high value, uh, opportunities while also, you know, acknowledging you're gonna have a long tail, especially when you're, you have a technology company whose technology stacks evolved over 40 years.
Um, you have to come up with different strategies to support and, and manage those, um, those other kind of technologies, those legacy technologies. 'cause they're still very important to running the business. But if you start with that, that long tail or trying to address that long tail, uh, it's gonna slow your journey down and, and a lot.
And, uh, obviously it will also, uh, paint you into corners in a few areas where you can get these higher impact broader winds. In general. You cannot walk down the street these days without somebody leaping out to tell you about their awesome new AI thing.
Do you think AI will be applied to this space? And how might it help you in your view of the world? Just, just in to, um, uh, just in, in order to keep my colleagues who, who work very hard on, on, on these types of problems.
I, I'm gonna use the, use the machine learning moniker. I think, I think ai, like, like zero trust is a highly abused term. Uh, and obviously outta respect for them, I'll stick to what this is, which is machine learning.
Um, but, but obviously there, there are opportunities where any sort of pattern and, and patterns need to be extracted from a large data set. 'cause one of these hard problems for us, uh, has also been, we have, you know, 40,000 plus services that make up, you know, the core of Bloomberg's infrastructure. And those are all, um, you know, speaking our proprietary protocol and all that, we've, we've built up a micro concept of microservices before the word microservices were coined even.
Um, but, uh, we, we had to invest very heavily in, uh, developing ways to automatically infer what policies might be between services and identifying those patterns of communication where there are something that should be authorized. So service A should be talking to service B, uh, inferring that policy, but then also, you know, using a bunch of contexts to figure out should that policy actually exist or is that someone abusing that service? Um, you know, making sure that you can collect that context and then reason around it automatically.
I think that's, that's where machine learning has a lot of value in these types of problems. So that's been a very interesting area that we've actually been doing a lot of research in. How do we apply those techniques to help us infer policies more rapidly and, and scale that up because in the end, the problem becomes very untenable if you do it manually.
How do I establish relationships and trust with other folks who hopefully are pursuing a zero trust policy, but eventually everything's connected to everything else. So how do we kind of collectively approach this Problem? That is an excellent question because, uh, again, this is an area that's dominated by defacto standards and something like, uh, you know, picking on spiffy, again, spiffy is not a great answer for how you talk to like other entities.
So if I'm, if I'm talking about like, I have a third party service provider that I'm depending on, that's not really the use case. Spiffy is designed for, um, obviously we're big believers in the, you know, the human equivalent to something like Spiffy would be single sign-on, uh, you know, SAML and OAuth are your, your, your bedrocks that you build on for those. Uh, and of course that becomes the, a very convenient means for managing, for example, uh, employee access to SaaS services and and so forth.
But once you get to the data center, it gets a little murkier. How do you establish those trust relationships between parties? How do you rotate credentials, you know, on an, you know, very rapid basis?
How do you attest for the fact that, hey, this endpoint I'm talking to is actually the endpoint I expect to talk to. Those are problems. We, we've been, you know, piecemeal solving over the years.
You know, OAuth service workflows obviously are, are helpful for this, but they're not a silver bullet because there's still a lot of manual intervention required. And obviously if you're doing that right, those are still tied to a human identity one way or another. And of course there's a philosophical argument about, well, isn't everything, you know, ordered or requested by a human, uh, and within the company, therefore it should be tied to a human.
But there, there's also an operational continuity aspect of things. If someone leaves the company, you know, critical service, uh, connectivity between a, a client and, and, and the company should not be, uh, um, should not be something that no longer works, you know, because their credential expires or is revoked. But, um, you know, I, I think this, the answers for those questions are not great.
And, and it amazes me how often this just devolves to, hey, pin this cer this CI web PKI certificate, um, and, uh, use this long API key. And that's how we'll authenticate each other to me it's like, it's a, it's a space we haven't solved very effectively. All right, let's imagine you are king of all things, zero trust for a day and you have your staff.
What's that one thing you're gonna fix? What by decree? What is that thing that just kind of rankles at you and go, boy, we all better off if this didn't exist?
Can I just say like, logging telemetry In the end, that's what we need. 'cause I mean, look, we can come up with all these great, secure by construction concepts, but building that telemetry, how is it being used? Is, uh, you know, what policies are being, being used and when, or applied and when, what credentials were being used and, and look like this is, this is the bread and butter of what keeps a, a security operations center running.
So, you know, making sure that all these wonderful zero trust, trust concepts are still being, you know, backed by good telemetry that helps us understand what is going on, you know, on the various platforms that we have in these environments so that we can actually, you know, get insights, perhaps apply some of that machine learning actually to the, uh, to the telemetry and get insights into where there might be abuse or where we need to change patterns. All right, folks, you heard it here from Bloomberg itself. Zero trust is attainable.
It just takes a, a lot of work and it's continuous, so you're actually never done and it just exponentially gets bigger and bigger as you go along. But take heart, because if you don't do it, the bad guys are gonna be all over us anyway. So Phil, thanks for being on the show.
Thank you very much. Appreciate it. And back to you guys in the.