RegisterEnterprise AI Summit — Oct 7–8 · Charlotte, NC
Video Library

Log in to watch

Log in or create a free account to watch this video.

Log in
Virtual US 2022
Share

batCAVE - Calling All DevOps Superheroes to CMS

Robert Wood
Chief Information Security Officer & Director of the Information Security & Privacy Group, Centers for Medicare and Medicaid Services

The Centers for Medicare and Medicaid Services (CMS) provides health coverage to more than 100 million people through Medicare, Medicaid, the Children’s Health Insurance Program, and the Health Insurance Marketplace. The CMS seeks to strengthen and modernize the Nation’s healthcare system, to provide access to high-quality care and improved health at lower costs. To achieve improved delivery of health and healthcare management, CMS is adopting DevOps as the core to enterprise transformation.

Chapters

Full transcript

The complete talk, organized by section.

Host Intro (Gene Kim)

00:04

Thank you Arun. Another one of my favorite presentations from the DevOps Enterprise Summit six weeks ago was from Robert Wood who serves as Chief Information Security Officer and director of the Information Security and Privacy Group for the Centers for Medicare and Medicaid Services. So the CMS provides health coverage to over 100 million people in the United States which includes some of the most vulnerable children the elderly and people who can't afford their own health care. Robert describes, the Amazing Story of what his group is doing to liberate the over 6,000 developers that support CMS from the onerous work of ensuring the security and compliance of their work in a very highly regulated environment and how this capability help them when it mattered the most during the COVID pandemic.

00:54

To me what makes this work? So admirable is that this work to elevate developer and Ops productivity is driven by the security group, which is probably the opposite of most of our experiences with security. as you likely know in the Phoenix project the security folks are portrayed as a shrill hysterical people focused on technical minutiae slowing people down and sucking the will out of everyone they touched I'm just so blown away by the vision and aspirations of the work being done by Robert and team in an ideal world. I think every engineering organization would have an effort like this to integrate information security objectives into their daily work.

01:29

Here's Robert.

Robert Wood

01:34

Thank you, Gene. All right. So today hello everyone today. We are going to be talking about a program that we are affectionately referring to as the batCAVE or the continuous authorization and verification engine because every government program needs a good acronym to get itself off the ground.

01:51

So I want to start off with some background about CMS for the Centers for Medicare and Medicaid Services. So we are effectively a payer which means that we are providing payment for medical services and everyone who has been alive in the in the last couple of years and gone through health crises such as COVID-19 and the opioid epidemics. No that the being able to access healthcare Services is incredibly important and the programs that CMS specifically offer some Medicare and Medicaid are serving some of the most vulnerable or at-risk populations across the US and overall. Our programs are benefiting over 130 million Americans Nationwide.

02:38

And so really we are looking to become we are looking to service all of these individuals and make sure that they have access to high quality healthcare Services. We need to make sure things are up. We need to make thing. Make sure that things are secure and stable for them.

02:58

so A little bit of background about myself. I joined CMS as the Chief Information security officer in November 2020. So I just hit my two-year point. I came in coming in from a couple of different startups.

03:08

I was at three startups prior to this and so I have a lot of exposure in heavily, like data Centric and developer-centric environments that is near and dear to my heart. The scale of CMS is is pretty large. So we have over 230 Business Systems. The physical system is is basically a piece of software or collection of of software that goes into a some kind of governance boundary that's providing some kind of mission.

03:42

capability and so jumping right in these are some of the challenges that we have in the federal space. Security governance risk compliance or GRC as it sometimes referred to puts a huge amount of resource drain on the software development lifecycle and the nature of FISMA and NIST 800-53 and all of these big heavy compliance mechanisms or requirements. They end up almost reinforcing this very waterfall or water-scrum-fall software development process. And for anyone that's built software in the in the modern world there.

04:24

They're probably pretty familiar that these sort of processes are slow. They're burdensome and they're often error producing or error-prone because you get to the end what you thought you thought going into it has changed quite significantly. One of the other big problems there are challenges that we have is that CMS is a very Federated environment. So typically we do not have a bunch of software engineers and data engineers and infrastructure engineers and platform Engineers on staff as federal employees.

04:55

These are contracted Engineers sometimes prime contractors sometimes subcontractors to a prime and those are engaging with a federal person who's operating in more of a program management capacity. um And so just the very nature of that means that systematically if you if you're familiar with something like Conway's law we have this very hierarchical structure across CMS and that creates a lot of silos around the agency that sometimes makes collaboration difficult. Sometimes produces data silos technology silos, and at that makes the whole process of governance risk compliance and just overall software security sometimes quite difficult. So the solution or what we hope is a contributing factor to to the problem is what we're referring to again is the batCAVE.

05:52

So this is a platform as a service abstraction that we are that we are building around cms's core Cloud program and the whole goal of this is to reduce the burden on developers and accelerate the time to value for mission owners. Most of what happens in the the federal software development space is repeatable and to the extent that we can automate away a significant part of the infrastructure setup the onboarding process and the security compliance needs the better and CMS has been on a cloud Journey for years now going all the way back to to healthcare.gov. And that was first launched to one of the early projects that I that I first engaged in intersected with CMS on called tmsis. This was bringing all of the Medicaid data across the country into the cloud for the first time across all of these disparate disconnected state-owned mainframes, bring them all in and giving CMS analytics capabilities.

06:59

So there's that many big modernization push at CMS for a long time. And so a lot of development teams are quite familiar with the cloud and so the extent that we can abstract away all of that repeatable work. That's really what we're for here so Some of our goals and this this really boils down to our mission our goal our vision for the program. So first and foremost is reducing the time spent on ATO or authority to operate and SIA or security impact analysis processes.

07:32

So these are a couple of big compliance things that go into bringing something into production for the first time and then making changes to a thing which is where the SIA comes into place. Basically what that means is if you're making a change you you determine the security impact of that change and then you make it and so to the extent that we can abstract away and reduce time spent on that because sometimes these can take months the better. We want to reduce the security and infrastructure burden on teams and allow them to focus on building software. It's one of our core talking points going into this whole thing is that you don't really innovating at the way at the layer of your software set or your server setup your networking your which database you choose things like that.

08:21

There's some amount of optimizations that happen there sure but real value real creation real Innovation is happening at the software layer at the features that you're building at the value you're creating for end users because most that infrastructure stuff is abstracted away from end users anyways I also want to be able to ensure true continuous monitoring and security compliance. And one of the big things here is is making sure that by default so that way it's reducing cognitive load on the developer that we are looking at all of the things that we really need to and making sure that that's happening consistently across every piece of software that we have. Um, the other the other big thing here is reducing the cost for end-to-end application development. So some of that is opportunity cost and time spent doing things but also active infrastructure costs.

09:13

So those who are who have built things in the cloud have probably Fallen victim at some point or another to the to the effect of turning something on for getting to turn it off. It's really easy to procure and and provision new resources in the cloud. You can do that really quickly. I can start up a thousand servers with a push of a button but sometimes you don't always forget you turn off those thousand servers with another push of the button and so the to the effect that we can build cost and operational guardrails as well as security guard rails into this platform abstraction the better for the agency as a whole Um, and so stepping stepping down through the rest of the the goals here.

09:59

The biggest thing that I really want to drive home is getting to the point where we're allowing and enabling and empowering true continuous delivery. We want to get value to the American people faster or just the American beneficiaries of cms's core services and Technologies faster. And I want to tell this story through the lens of the COVID-19 pandemic. So imagine.

10:26

Rewind your your mental space back to March 2020 when the federal lockdowns first went into place. There is about 17 or the cases were measured in the teens at this point in time. And if you think about all the steps that are going to go into creating a piece of software to contribute to solving the problem all of those things take time. And so there's ideation.

10:48

There's procurement which again we don't have our own software Engineers on staff. We have to go out and procure something put something on the market. You're gonna be onboarding those people you're gonna be going through human centered design. You're gonna be setting up your infrastructure and provisioning things to actually build and then you're gonna actually go through and build the thing that you thought you needed to build way back in March April 2020 and then get it authorized go through the ATO process and at that point in time the situation the environment the the landscape has significantly changed underneath your feet.

11:23

And so the assumptions you had going into it and the Assumption or that the realities that you are now currently operating in are totally different you went from teens to hundreds of thousands of cases. So the way that we're actually going about getting this idea out into the federal space. You'll notice on this on this slide terms like marketing and sales customers and feedback continuous improvements such and so this is a typical product flywheel and we're really using this to inform our thinking around how we spread and propagate the idea of the batCAVE into the environment. That is CMS.

12:06

We're thinking about the agency and really like HHS as a whole but of course specifically CMS as a total addressable market and we need to provide value. to those potential customers to attract them in or thinking about this through funnels and through Market potential Market capture because we really want to make sure that we do we don't lose the site lose sight of as a platform and security engineering team of how we create value who we're creating value for and why we're creating value into this into this platform into this product making sure that we're measuring the right things and using the right metrics to drive incentives and and behavior change internally as well as externally So here's some of the technologies that make up the batCAVE. So it is a platform. It is built on kubernetes as our core container orchestration layer.

13:02

The thing that I mostly want to highlight in this particular slide is the fact that you'll notice two main vpcs on this Slide the batCAVE VPC as well as an Ado or an application development organization VPC. That's that's the development team that's consuming or building something on the platform and the thing that The thing that is unique in this particular case is the fact that we have all of these pipeline tools. So there they're shared resources and consumed out by an Ado team, but the utility belt which is something you'll see in both VPC. Both VPC environments is something that we're using to both deploy the batCAVE core itself as well as to deploy customer specific resources, and one of the reasons that we're doing that is we want to make sure that there is consistency in functionality across our environments so both in the shared core Services as well as the as well as the customer environment and what those things actually look like if we zoom in just a little bit is you'll notice all of these these sort of functional capabilities that that work themselves in here.

14:18

So everything from like service mesh console and visibility and istio and service mesh and and visualization and dashboarding through grafana Prometheus metrics collection stuff like that. So that the thing that I want to really try to drive home on this on this slide is that it's not all just about security. Real really what we want to do is Empower development teams to focus on what matters most and not get distracted from all of this other security and operational stuff that they need. But also while we're at it give them things that are going to help them run more effective services.

14:56

So that's where stuff like, you know built in Prometheus and grafana comes into play. We can give them high quality Telemetry. It also happens to emit logs. They go back into back into our centralized security data Lake and such and and operations collectors to satisfy compliance and governance requirements.

15:14

But also let them do what they need to do and do it effectively fantastic. istio, we want to be able to make sure that they are like getting getting like role-based access control and and network segmentation through namespacing and such like make sure that that is really like carried through in the in the way that their environments are segmented the way that one application segments from another and such like that is important for them to be able to do like environment-based development. But also their security property benefits to that that matter for the governance or compliance process as well as for them. Things like Valero helping with like baking in the backup and disaster recovery process.

16:02

So, you know you think about something that is handling Health Care Services or supporting healthcare services for half the country you don't really want that to go down. You don't want to lose data. You don't want the like if something affects the Integrity of data in that in that that in a particular system that's supporting that many people you want to be able to roll back you want to be able to recover and so making sure that that's happening consistently. Is is really important to us.

16:29

And so that's where like this this pattern this this platform engineering pattern. That is the utility belt is really important to us. The other big thing that's really that really matters. So there's container orchestration the runtime then there's also the cicd process the pipeline stages.

16:48

And so what you'll notice here as you kind of go through this journey is you'll notice you things will jump out to you like like build and unit tests and stuff like that. Then there's a lot of other More security specific things in here. So sased static application security testing or static code analysis. So there we're doing things like checking for hard-coded passwords or API keys or checking for patterns of code, like use of insecure methods or functions or classes things like that that are going to contribute if if we sense those things those are going to contribute to insecure software at the other end of this.

17:33

Scanning the the images that are coming in the container images for potential malware or like typo squatted type of squatted inclusions of packages that stuff like that stuff happens in real breaches, especially in the federal environment as we're seeing more and more supply chain oriented attacks making sure that we're able to get on top of that stuff early in the process pre-build. Matters, um container hardening is another really interesting one. So prior to this project. I was not really I was not hit on this particular approach but something that I'm really excited about now.

18:10

So this this is actually like we've talked in the security industry a lot about attack surface and dead code and stuff like that like all of this like sprawling attack surface that contributes to risk but doesn't actually contribute to functionality and so the container hardening or a tax surface reduction process in the pipeline is actually shrinking containers down to that which is only required for its functional runtime like functional use and so what that's actually doing is if let's say you have a set of containers and they have 100 vulnerabilities, but 80 of them are like referencing or or built on stuff that you're not actually using you see this in the like the open source Library space all the time. I used to work at a company called sourceclear. That was now they've since been bought and one of the things that we they were in the software composition analysis space and something that they are that that we focused on a lot there was if you if you get flagged for using a vulnerable dependency, but you're not using the code that is actually vulnerable. Is it actually a risk or it's it's absolutely much less of a risk and so similar sort of thing here.

19:26

This takes it a Next Step farther where you're reducing and actually just removing the stuff that is not used. It might have vulnerabilities in it. And so you're you're so any findings that still come out on the other end of that are gonna be High Fidelity stuff that you should actually be worried about or building Bill Gates on or things like that. There's a gate check so This is a little bit of a Hot Topic or hot button Topic in and around the devsecops or devops scenes where you know developers don't necessarily want security sitting in their pipeline saying no, you can't deploy this.

20:05

Yes, you can deploy this, you know, basically being the rule the rule makers or or the the sort of Gatekeepers of getting code out there. We don't necessarily want that either so really this is this is intended for for absolutely critical things that we feel are going to prevent a piece of software from being functional or is gonna is gonna lead to some sort of imminent breach or imminent risk that has to be That has to be addressed. And so this is stuff that's going to come up through follow-on pen testing and scanning and stuff like that. Anyways, so you'd rather know about it now and get faster feedback for it.

20:46

So really this this whole process what we're trying to do and and just a kind of tie this all together, what we're really trying to do is certify each of these stages. So the building blocks that you saw on the prior part all of those technology components as well as the cicd stages that go into each of the each of our pipelines. If we map controls to each of these things each of these blocks then what comes out in the other side of it software that comes out in the other side. That's what secure that's what's compliant.

21:23

We don't necessarily need to go through the rigor of checking every piece of software for every single control of which there are hundreds. To make sure that we have a compliance system at the other side on the other side of it. And so we're certifying and making compliant the process and what the process produces will therefore be more secure. And so that also gives us more more flexibility more agility that we can pass on to the developers at the other on the other side of this.

21:54

So a couple of major Milestones I want to highlight first is the full and independent ATO. This actually just happened about two weeks ago. We we hit this Milestone. The reason that is important is if you if you're operating in an Enterprise environment that has a lot of governance risk compliance process.

22:10

There's an element there's this notion of inheriting controls. Typically developers have to go through and figure out whether it's developers or a security manager or somebody has to go and figure out on the development and this is in the product team what controls are applicable which ones are partially partially inheritable or which ones I can actually inherit. And so by having this full independent ATO we can just inherit the entire. The entire control set if you're using the batCAVE platform that saves a tremendous amount of time and it leads to consistent results in our GRC ecosystem.

22:51

Multi-tenancy is the next big thing. And this is for this is for portfolio owners. So for myself, I am considered a portfolio owner I have about Five six products that are running that are running in the cloud right now more to come that's everything from it as Farm ingestion enrichment tools to GRC tools to data Lake like ETL in Richmond, like data capture Services things like that and by running these things any compliance secure multi-tenant environment, it helps me cut down on costs as a portfolio owner because I'm only paying for shared resources once and then I can take those costs and pass it on to other Innovative efforts or just things that are gonna add value to our security program in general. The next big thing.

23:39

Is this Enterprise services and data Lake integration. So for anyone who's familiar with the Federal space you may you may have heard of executive order 14028. This was basically the zero trust executive order that came out in 2021, May 2021 to be specific and a couple of big things that were highlighted in that executive order were zero trust. So all federal agencies implementing some form of zero trust architecture at some prescribed time and then also software bill of materials producing a software bill of materials for software built in house and then consuming a software Bill and materials from software consumed out of the agency and storing all of that and using it in some undetermined way at that at the time of the writing.

24:22

And so these things were really informed by like the Log4j attacks and SolarWinds attacks that were happening around the time that these things are being drafted and then released and so what we're working on right now is basically taking these big policy changes that are that were And basically instigated from the executive order building them into the platform abstraction around our core Cloud offering and then those who are consuming it those development teams are consuming it get those properties get those benefits for free. They don't have to do extra work to do the micro segmentation to produce an SBOM. It's just another pipeline stage that gets that just falls out of the use of the batCAVE then goes into our data Lake and then we can when it when it comes time for the next Log4j vulnerability or you know, supply chain related attack so many on our security team is that instead of engaging with developers doing data calls and you know knocking on their door their virtual door, they're at a query or they're at a console. Writing a query and determining exactly what version of code is running what component who owns it who they need to engage with checking logs to see if it was exploited.

25:43

And so we have this whole end-to-end process this whole end to end life cycle security lifecycle. That's more operationally focused that also removes operational burden from developers. They don't want to be responding to spreadsheet data calls that's annoying. It takes away from things that they have going on at that time.

26:05

And and so the extent that we can take that away by making ourselves more of a data Centric data-driven organization. That's value add in the in the long tail of software development. So a couple of major accomplishments that we've hit in recent times, I mentioned the control mapping in the in the full and independent ATO. So where we're at right now is around 80% We're just shy of that about 79.8% last last I was told of mapping cms's implementation of 853.

26:35

So what that means is by adopting the batCAVE you've hit 80% of your ATO requirements right out of the gate means no documentation, no assessment. No, testing no evidence production for 80% of what you're dealing with. And you get all of those other benefits around deploying these necessary security stages like security gates and security steps in your pipelines and and deploying infrastructure that's going to be used for like your backups and your container orchestration and like managing a runtime and all of that stuff you get that stuff for free. Other big thing as we've already started to codify some of the key strategic policy moves that I mentioned SBOM zero trust Etc into the platform.

27:21

So that is like some of that's underway. It's continuing to improve and this is something we're really excited about because some of the forecasted costs here were tremendous from some of the development teams. There's a huge range of maturity levels across the agency. And then something that I'm really excited about and this is the last bullet on here is the from day one.

27:46

We deployed what's referred to as a purple team. It's on the spectrum of security testing or security Engineers. You have red teams that are doing more offensive adversarial emulation. Then you have blue teams that are doing more defensive security operations Incident Management sort of work, you take your color theory and you put those things together and you have your purple team.

28:07

These are security Engineers who are basically have a foot in both camps and they're also working right alongside the batCAVE engineers and so as they are finding things and building out these abuse cases doing threat models and stuff like that. We can take something that we feel is gonna be a risk at the core platform. And if something like with a lot of these platforms you have these tremendous opportunities to drive value saving or like cost savings and value creation and all of that for all of the consumers but shared services can also be a tremendous risk because Benefits are shared by all consumers but risks are shared by all consumers. And so we have to make sure that those core services are really well locked down we have to do that work up front.

28:44

So this is where the purple team comes into play and if we are if we find something So we built the a regression testing framework on top of Semgrep for those that are familiar with it. It's basically an AST based. It's like security unit testing and so taking an abuse case or a like a test finding turning that into regression test. So that way we do not fall back into a bad bad pattern or a bad habit or have something really resurface that's going to cause risk for us for our beneficiaries or for our developer constituents is really really important to us and having this like really tight feedback loop directly inside of the platform engineering team is something that I'm personally really excited about.

29:35

So this is It's the to the best in mind knowledge. Just the first time CMS has had a purple team. I'm not aware of any other mature purple teams operating across the federal government. And so this is something that I just I I love the model coming from this background myself, and I'm excited to see more of this happening.

29:54

So and don't you don't necessarily have to take it from me. So this is a success story from one of our early adopters. And and basically this was a team a development team that was transitioning from an on-prem brick and mortar data center environment into the cloud initially. And so they decided to evaluate the batCAVE and now are fully fully baked into it.

30:12

And basically that some amount of continuous integration support. They had an existing ATO but they were obviously going through this major re-architecture using Cloud native tools and such and what they found is that they were able so their initial migration time line was measured in years and we were able to cut that down to like weeks months in terms of getting them up and operational pipelines set up getting them up and working. So very very exciting to see the hard work that our platform engineering team has done translate into successful stories and and the feedback that we've gotten from this and other early stage adopter research and human centered design research is basically feeding into the roadmap. In the backlogs that all the development teams are working with to make everything better as we think back to that flywheel and how we're how we're thinking about our overall funnel and and adoption approach.

31:21

So I want to transition and sort of end on this last slide of where we need help. So the batCAVE is a collection of organizations be a federal organizations and Industry partners that are contributing to this effort. And so what we're really looking for are agency partners both within CMS, but also HHS at large or other federal agencies who are willing to go out on a limb and actively collaborate with us. When I say collaborate.

31:51

I mean to to share standards share code. We're happy to share. We we invite pull requests and stuff back into our back into our our Mains and we also want to be sharing data like sharing data around that'SBOM sharing data, like sharing pattern sharing control mapping stuff like that. We want to we want to see collaboration really raise the bar for the entire Federal ecosystem because like, you know, this is sort of an internal like inside joke sort of thing.

32:18

Like we're in the business of the government. The government does a lot of important work whether you're a fan of the government or not. It drives a lot of important stuff that all of us as Americans have to rely on at a day-to-day basis critical infrastructure. Versus that we rely on all that stuff.

32:39

It's important that we be as efficient and effective as possible and I think as taxpayers we can all agree. We want our government to do better. And as far as industry Partners, we're we're hoping to see industry partners and really lean into those who who want to work in an open source model we value I personally value those who want to prize mission and mission delivery and mission capability over a sense of "mine" and what I mean by that is industry Partners who are really leaning into like this is my code my program my data my this my that so everyone else sort of Heisman and and, you know, stay back stay out of the way like this is my program. I want to I want to work with those who are who are gonna put we're gonna put their necks out on the line and contribute back to the mission because that's what other folks are doing.

33:30

And so if we have this this Active network of owners Federal partners and Industry Partners working together in this more Badgers environment. I believe we can really make a difference and so I want to end that by saying thank you so much to everyone who attended today. I hope you learned something. If you if you want to reach out.

33:55

We do have a batcave LinkedIn page. You can hit me up personally on LinkedIn. I'm at Holy cyber Batman. You can email me what have you I'd love to talk with you.

33:58

I know our team would love to talk to you about ways that we could collaborate about ways that we could work together, or if you want to be a consumer of the batCAVE core services. Please do engage. Thank you so much. Hope you have a great rest of the event.