batCAVE - Calling All DevOps Superheroes to CMS


The Centers for Medicare and Medicaid Services (CMS) provides health coverage to more than 100 million people through Medicare, Medicaid, the Children’s Health Insurance Program, and the Health Insurance Marketplace. The CMS seeks to strengthen and modernize the Nation’s healthcare system, to provide access to high-quality care and improved health at lower costs. To achieve improved delivery of health and healthcare management, CMS is adopting DevOps as the core to enterprise transformation.
Chapters
Full transcript
The complete talk, organized by section.
Robert Wood
Hello, everyone. My name is Robert Wood. I am with the Centers for Medicare and Medicaid Services. CMS is one of those funny agencies that people are vaguely familiar with enough. They know Medicare, they know Medicaid, or they know somebody who uses it, or they have parents that use it or something. But ironically, it is one of those agencies that... So we sit under HHS, Health and Human Services, which of course does everything from FDA to CDC and NIH to much smaller things, Indian Health Services, for example. And so CMS is like a gorilla that really nobody knows about.
So we have about 20% of the country's GDP flows through CMS's budget, through these massive programs. And beneficiary-wise, we serve about 130 million-plus people. We have data on even more folks. We have a lot of different systems, all serving the U.S. population in the form of helping people get access to healthcare services.
So this is me. I am not a lifetime government person. Most of my time has been spent in tech startups and security consulting and such. I ended up coming into the government through a couple of different conversations with the CEO or CIO, and the motivation was really like they wanted somebody who was technical and they wanted a culture lightning rod, because security oftentimes stinks in organizations, especially in big enterprises. It's a drain, it's hard, it's friction-filled, and they wanted to try to change that, to his good credit. And I felt like I could take a swing at that. And so here I am.
Some of the challenges that we had, and these were some of the things that got laid out to me initially when I met with CIO, a gentleman named Rajiv Uppal, is that security and GRC, or governance, risk, compliance, had become these massive resource strains. So people who are familiar with the world of government work know the term ATO: just many-month process, lots of paperwork, lots of drain, lots of friction. And then on top of that, CMS is like many government agencies going through these massive transitions to cloud services, to try to modernize their technology footprint, move away from the 50-plus brick-and-mortar data centers that existed into AWS or Azure infrastructure as a service. And the problem with that was, despite making that move, things were still really jumbled. They were hard to manage. There was a lot of special snowflakes floating around the environment.
And of course, being government, everything is sort of almost institutionalized waterfall development, just in terms of the way contracts are structured and all the things. And then we have a very federated environment. So there's probably like 6,000 CMS employees, over 100,000 contractors writ large, a lot of engineers. So 6,500 is a rough estimate. There's a lot more if you start counting data analysts and DBAs and sysadmins and stuff like that. And then that trickles into this other problem of one contractor doesn't want to work with another contractor that doesn't want to work with another contractor. This even happens inside of my group, inside of security. And that ends up creating lack of integration and a lot of pockets of data and functionality and process, and those pockets probably and oftentimes don't work together without significant friction, or you have to overcome a lot of friction to make those things happen. And that of course makes security inherently difficult. You're investing a lot more resources to secure things.
So this whole effort is being championed out of the systems office, from yours truly. And really what we are trying to do is win the value game. So CMS is a very federated environment, right? I mentioned that. So people just get to do whatever they want to an extent. If they go through the ATO process, they have money, they will spin up their own IT shops. They'll spin up their own security teams. They'll spin up their own cloud accounts, buy their own stuff. They don't use centralized IT if they have the budget to circumvent it. And that makes, of course, creating a centralized or uniform governance process difficult.
And so to work against that, or try to get people coming back into the world of centralized IT, our hypothesis was that if we can actually create a 10x value scenario where it is valuable enough to drive down the burden and the friction associated with change, associated with ATOs, associated with getting your mission capability to production, then we can get people to come start adopting these things. This compounds when we look at these big sweeping changes coming out of the White House, like the executive order that came out in 2021, 14028, the zero trust one.
So our goals, we have a lot of goals. The big one is that we want to focus on getting people to production faster. We want to enable true continuous delivery. Something that we see very often is that even once a team reaches production, their ATO, they're operating in production, they have production workloads, everything's set up, they need to change something: there's a lot of tickets they go through, sometimes month-long, three-month-long deployment cycles. And so you look at that through the lens of security, and it becomes a really dangerous situation. Like Log4Shell hits, and in order to respond to that, you've got a one- to three-month lead time of delay where you are left exposed. That becomes really problematic from our perspective. And so we want to enable things to move as fast as possible. And so we're really seeing a lot of excitement inside of the agency, because it's security oftentimes the ones who are getting in the way of all of these things, driving this stuff.
Now I'm going to tell this story through the lens of COVID-19. Everyone lived through this, no matter where you were in the world. So March 2020, the White House at the time goes through and shuts everything down. Now at that time, your problem space looked a certain way. There was a situation like the number of cases and deaths were numbered in the teens, and everything was under a hundred at that time. Manageable. Maybe we do this for a couple of weeks and we can get ahead of this whole problem.
So people start coming up with ideas. What are we going to do as an agency to respond? CMS is in the healthcare business. We play a part in the response to this potential national epidemic. Keep in mind that at this point everyone just switched to remote, so everyone's scrambling just on account of that. So we come up with ideas. We go through procurement. We can't do anything in the federal government without procurement: six months, a pretty ambitious procurement cycle, but let's say it takes six months. Now we have to onboard people, maybe do some human-centered design if you build that into your contract. Then you set up your infrastructure, and then you actually build the dang thing, and you ATO it.
And what I want you to really focus on here is the difference between the very first line and the very last line on the right, and these are fairly ambitious timelines, unfortunately and sadly, in the federal government space. And the problem here is the problem you are trying to solve way back in March and April of 2020 looks a hell of a lot different than it does two years, a year and a half later. The assumptions that you made way back when are totally out the window, and you have to be able to quickly and swiftly revise these assumptions and adapt. And if you even have to start going back six months, and you have these month-long release cycles, that's going to be really problematic. So something has to change.
So this is our flywheel and how we're trying to drive adoption. I'm not going to dig deeply into all of this, but you'll notice, looking at this, that it takes a very product-centric perspective on how we're trying to get this idea out there into the agency. There's a focus on user and design research. There's a focus on collecting feedback. There's thinking about marketing and total addressable markets. And what we're really trying to do here is win the value game. We're not just coming out and saying, thou must adopt, thou must use, thou must move over, because that almost never works. That is the heavy-handed, ham-fisted approach the security teams have taken for years. And you can only play that card once, twice, before people start to ignore you, and then they discount you, and then the game is lost.
So we're trying to win that value proposition. These are some quotes that came back from one of our early customers in an office called OEDA. They do a lot of open source data analytics stuff with data use agreements and patient records requests and things like that. And the thing that's exciting about this is our hypothesis is starting to play out as we anticipated. Getting developers to focus on the application tier, focus on being able to build features and ship features without having to worry about all the junk, the scaffolding, all of that stuff underneath, is exactly what we are hoping for.
Now there's a little bit of an eyesore if you're just looking at it from wherever you're sitting. So don't try to read the individual boxes, but I'll explain this at a high level. What you have here is three main tiers, which you see in these columns on the left hand. You've got this batCAVE administrative cluster, a lot of tools and controls and solutions and services all built in here. Then you've got your customer-specific stuff in the second tier, and then you've got integration with other CMS and OIT capabilities, whether it's single sign-on, data lakes, monitoring services, whatever that happens to be. Because oftentimes if we were to go back to that development lifecycle graphic, each team is uniquely building a lot of that stuff every single time that they go ahead and launch a new product, bring a new product into the agency. All of that takes time. It takes time for them to build. It takes time for them to engage with the team that they're trying to integrate with. And it just creates this cumulative effect of strain on the overall system.
And one of the key things that we've been building, and this is all packaged up declarative Terraform and Helm charts and such, is what we're calling the utility belt. Here, this is really about taking a series of tools and pre-built integrations and stuff that we can deploy at once, in one shot, one go. We can deploy it to both the batCAVE admin plane and the customer application plane, and they're both getting all of these tools and all of these integrations built into their stacks.
Where we've opted to go is take the... So a lot of what we have to do in the government is focus on controls and satisfaction of NIST compliance, 853, all of that fun stuff. Instead of focusing on these vertical stacks of compliance where you take a given application from the top all the way to the bottom, down to policy and infrastructure and stuff, we're trying to shift the governance to the code layer, into these modular building blocks, so to speak. That way, the utility belt has a certain amount of control mapping built into it, and we can deploy that where we need to deploy it, and the control inheritance and the governance sort of stacks cumulatively.
Now this is what our pipeline stages end up looking like. And again, I'm not going to walk through everything specifically because these talks are rather short. But the thing that is interesting here is in software security and the application security space, there's been a lot of buzz around shift left, shift left, shift everything left. Everything has to go left. And that's great, because security has notoriously been focused on just doing pen tests and scans and things, which are very right activity. But really, you can't just shift left and throw runtime and throw build time and throw deploy time out the window, because then you're going to have other types of exposures. So really it's about shifting left, shifting right, shifting everywhere, making sure that security and governance are built into all of these different stages of your systems lifecycle.
Now a couple of different major milestones that we're excited about. We just hit the point where we have a full and independent batCAVE ATO. The reason that matters, the reason why that's important, is that now allows us to do full control inheritance from the batCAVE infrastructure. So everything that you see here, you get to inherit into your ATO package, and we hit about 80% of our overall controls that a given system would have to satisfy. Usually these things take months and months to document and months to prepare for in terms of assessment. You get 80% of that right out of the box. And now we're starting to build in other core things, like response and requirements that tie into the executive order, integration with our security data lake initiative, etc. Really what we're trying to do is all of that other friction that's not necessarily explicitly ATO-related, we can still build that into the batCAVE environment, so that by adoption you're decreasing all of these potential and very likely points of friction that are going to slow down your mission capability.
Now other major accomplishments. I mentioned 80%. Something that I'm personally really excited about in this last bit, so I'm going to highlight two of these things. The second one is taking other parts of policy, so SBOM generation and ingestion, putting that into the batCAVE just as an output of the pipelines. The zero trust enforcement through service mesh, micro-segmentation between environments, mutual authentication, things like that. The integration with the data lake, being able to manage and collect all of the logs in a declarative way. So we're not going back and forth with audit teams and pen testers and system developers and such. That's just when your infrastructure is provisioned, you just get that stuff out of the box. A lot of what we do in the federal government ends up mapping back to some White House directive, congressional initiative, FISMA, whatever it happens to be. And so the more of that that we can take away from the cognitive load of the development teams and these business owners, the better.
The last thing that I'll mention on this slide is, from day one, we integrated something that's referred to as a purple team. Now, I don't know how many folks are familiar with that term or have heard the term red teaming or blue teaming in the security space. Red teaming is typically associated with doing more attack-oriented activities, ethical hacking or social engineering, things along those lines. Blue teaming is associated with defending. So those who are experts or went through kindergarten are probably familiar that red and blue together make purple. And so the idea is that you have a group of people who are skilled in both attack and defense working hand in hand with the engineering teams, with the designers and the product managers, and building abuse cases, doing threat modeling, really trying to make sure that right out of the gates, and this is a very shift-left activity, right out of the gates we're considering how this thing is going to be attacked from any given point in the platform's lifecycle.
And then as we have findings, turn those things into regression tests so that we do not find ourselves just playing the whack-a-mole game that security teams and dev teams oftentimes find themselves in. We do a pen test. We find two instances of a finding that is really probably systemic. We fix those two things, and then lo and behold, you do a pen test again in two months and you find those same two things in a different spot, and then you have to play the whack-a-mole game all over again. That's really frustrating for developers. It's really frustrating for pen testers. It's not a cost-effective way to do security at scale. And so using technologies like Semgrep to build regression tests allows us to build this thing into the pipelines, into this declarative provisioning process, so that we don't have regressions back to known security defects.
But what that also lets us do is start to create a curated piece of shareable tests, shareable code, that can go out to others who are using similar types of things, similar infrastructure constructs. They can apply those Semgrep rules, modify them for their own benefit, and deploy them in their environments. And then they start to get this. So what we're trying to create is this platform-oriented benefit, not just platform in terms of technology layer, but platform in terms of: there's directed effects of producers of value, consumers of value, and then sometimes those players may be a producer and a consumer. So if they're pushing things back upstream, then others can take advantage of that.
Now where we feel we need help. We're one agency, right? We're just CMS. We're part of a bigger ecosystem in the healthcare system, in HHS. There are a lot of other agencies that are starting to get interest in this approach, and that was spawned off by some of the awesome work done in Platform One in the Air Force and then the Navy. But what we oftentimes see is these kind of siloed effects, and we don't see the kind of collaboration that at least we feel is really necessary to driving true open standards and a true government open source direction.
And so we're putting all our stuff out there, wanting to make our code, our standards, our control mappings, all of this stuff as open as it possibly can be. And so willing partners who are willing to either adopt, who are willing to stick their necks out and either take the code, take the control mappings, take whatever it is and apply those things inside of their agencies. Because, for example, you might have CDC. CDC is another HHS operating division alongside of CMS. They're a peer operating division, but they might have a slightly different control interpretation of 853 than we do. And if they're not willing to stick their necks out and inherit controls, or have their controls satisfied by what we're putting out there, or be willing to contribute back upstream such that we have a control mapping of the utility belt or the core platform to both standards, and here's why and how, then the thing starts to struggle at scale.
In terms of industry partners, I have this line in here, mission over mine. Something that we run into all the time in the ISPG or in the cybersecurity function is there's a lot of just-me players. It's somewhat understandable. Of course everyone's running a business. You need to make money. Everyone's got pressures on them from their boards, from their CEOs, stakeholders, whoever it is. And that's all great. We want everyone to benefit. I was in the private sector for a long time. I get it. I understand bottom lines. But what we really need are partners who are willing to go and focus on mission with us, who might be willing to have a smaller part of a bigger pie than a bigger part of a smaller pie. And so that's the kind of folks that we really want to partner with: those who are willing to push the edge and try to increase the overall size of the pie, because we feel like that's going to be better for the American people, that's going to be better for the partners and for us.
So that's all I've got. I have a minute and 30 seconds left according to my timer. So I just wanted to say thank you, and I can take a couple of questions.