SRE for Enterprises
Enterprises are starting to realize that reliability is a core property they need to strive for to be effective at what they do. Once they come to this realization, a quick look around finds Site Reliability Engineering. This engineering discipline has been focused on this question for over ten years, evolving on a parallel track from DevOps.
Let’s talk about what SRE is and how Enterprises can start to explore this path to help enrich their operational thinking.
Chapters
Full transcript
The complete talk, organized by section.
David Blank-Edelman
Hi there. My name is David Blank-Edelman. I have no title. We'll talk about that later. And it's my pleasure to talk to you about SRE for enterprises.
I was asked to come give, believe it or not, a not-DevOps talk at a DevOps conference, because this seems like it could be a thing that people in the land of DevOps want to talk about.
The two questions that I'm hoping to answer in this particular talk are: what is SRE, and how? But if you're paying attention, and I know you all are, even if you are coming in late, and even the analysts all the way in the back, you might want to ask the question, and this is where I'm going to do this dramatic thing: can it happen? Is it true? Is it even possible?
And I want to leave that as an open question. I have some answers for that, but it may be that what I'm going to tell you can't actually be done. So let's talk about it and see if it is entirely possible.
I do want to say that there are some people that are trying to prove that it can be done. For example, the fine folks at the Google Cloud Platform have something called CRE, Customer Reliability Engineering, in which they're attempting to see if SRE and enterprises can come together like a candy bar.
So let's talk about this. Let's talk about the what.
Let's just do a quick poll here. Number of people in this room who are very familiar with SRE, raise your hand.
Okay. So that's a lovely percentage. So you people should just check me.
The people that feel like they kind of, sort of heard of SRE, raise your hand for me.
Good, because now what we're going to do is we're going to make sure you know exactly what SRE is and you can walk away with this.
Okay, so here's the question. SRE is an engineering discipline that is designed to see if it's possible to engineer failure out of a system so you can bring it to the level of reliability that you desire.
Now, you'll notice that I didn't say absolute reliability, that I didn't say 100% reliability, that I didn't say that sort of thing, because it turns out that that's a crucial thing. The things that fly in the air, planes, those things, they pretty much need 100% reliability. The thing that's ticking in your chest needs 100% reliability. But I would bet the majority of the systems that you encounter on a daily basis don't, and we're going to talk about why that's an important property.
Okay, so here's my first quiz. Who can tell me what this is? Raise your hand if you know.
A bunch of snow. It is not a bunch of snow, nor is it a cow in a snowstorm. Yes, I know these jokes. I have an 11-year-old.
What this is, is this is what a PHP app looks like when it errors out. Right? Have we all seen this?
Does it make you feel better if I bring up a different language? Or does that cause you to twitch? Okay, I'll go back to this, just because then we can all be nice and calm.
The idea that I want to point out to you is that at a certain point, you can have all the features you want in your application, in your services, et cetera. It can be chock-full of everything. It can go and get the laundry. It can do Gene Kim's hair. It can do lots of things. But if, in fact, it's not reliable, then it's useless, and it's not useful to you.
So focusing on feature velocity over the fact that it's up is something you should consider the importance of. If it's not up, it's not useful to people.
Now, the tricky thing is there's a tug-of-war that goes on. There's a tug-of-war between the people whose job it is to produce the features, to iterate the software, to write the stuff. Let's call them devs for fun. And the people out there whose job it is to operationalize the code, to make things stable, to make sure things continue to run.
And so they're in this big tug-of-war, to a certain extent. What I think SRE provides is one way of thinking about how to deal with that tug-of-war. There are several of them. DevOps comes from the same challenges.
The thing that I think is important before I tell you more about SRE is just to realize that one of the things that DevOps did early in its journey was to assert that it was the next evolutionary step from system administration.
Are there any people who identify as sysadmins in this thing? Right? So both of us, let's go get a drink later, okay?
What I want to say is that just like I don't think DevOps was an evolutionary step to sysadmin, SRE is not an evolutionary step to DevOps. It isn't that someday you're going to be a protozoa, and then later on, when you're actually standing up and filling out your taxes, now you're SRE. It's not that sort of set of evolutionary steps. You just have to realize that it's more of a parallel track dealing with the same concerns that DevOps has.
So one place you could start in this question, and you wouldn't do badly, would be to start to read this awesome book called Site Reliability Engineering, put out by O'Reilly. Warning: they're my publisher too, so I have a bias. And it's a really good book. It's a little light on character development and the plots could use some help, but it's a really good book. I recommend you look at it.
I'm going to be talking about stuff that comes from the Google SRE model. We can talk about other models because one of the questions that I then get asked after this is, like, "Pfft, that's just Google. Only Google can do SRE."
There are lots of people, including people in this room, that have SRE in their organization. I could list them off. You'd know them: Microsoft, Spotify, Huawei. There's just a ton of them that are out there. It's not just a Google thing.
But it is the case that not everybody does SRE the same way. And that's part of the reason why I went along and offered to edit a book about SRE that's going to be coming out sometime early next year, for some definition of early, but definitely not for next year, that attempts to sort of address this question of how can you do it differently.
But let's go back to the Google model, because that's an easy way to talk about it, to start there.
The best way I know to describe SRE comes from a conference that I helped found called SREcon. There was, at the very first talk at the very first SREcon, a talk by Ben Treynor, now Ben Treynor Sloss, about what he thought that SRE was. And given that he founded the group that sort of started this whole thing, he gets to say some of that.
What I'm going to do is I'm going to show you a slide from his deck. You don't have to read the whole thing. Don't worry, we're going to go through a couple of things here. This is from his deck. I call it a dramatic recreation because he didn't give me the slide per se, but through advanced exfiltration things like looking very carefully at the video, I was able to come up with this slide.
What I want to do is show you a couple of things in this slide that will give you some sense of SRE. So the first thing that I think is really important to understand about this slide are these things that say: have an SLA for your service, which is interesting because they talk more about SLOs, and we'll get to that in a moment. Measure and report your performance against this SLA. And then use error budgets and gate launches on them.
Anybody in this room not know what an error budget is? I'm hoping there's a lot of people raising your hands, because this is where I'm going to bake your noodle in a good way.
So here's the way error budgets work. If, in fact, you agree with me that your service does not have to be up all the nines, and that it can be up, let's just use a round number, 80% of the time, that means that there's 20% of the time where you have the ability to perturb that service in terms of reliability.
If you think it can be up 80% of the time, then you've got 20% left over that you could do things. Do things like launch new versions and stuff like that. That's your error budget. Okay? Does that make sense to everyone? Nod your head vigorously if it does. Excellent. That's really nice. It's a little bit like bobbleheads. It's kind of fun to watch from up here.
So, yes. That's what an error budget is.
Let's just sort of go over that. The first idea is that what you want to do is sit down with the people that are creating the service, the product, et cetera, and come up with a set of service level objectives. These are the "how do I know my thing is up or reliable?"
It can be: it's available on a web browser. It can be this amount of throughput. It can be lots of things. What an SLO is for your service and product, only you know. But you want to spend some time figuring that out in detail and understanding it carefully. And it turns out to be not the easiest exercise. It sounds like it's easy, like, "Is it up? Is it not up? We all know what up is." But it turns out that what all up is isn't always the same.
Is it possible that I could be shedding 20% of my traffic? Is that okay? It might be okay.
Once you have that, then what you want to do is take that definition and embed it in your monitoring system. And that's going to be important. What that monitoring system will be is that will be a source of truth for you, in which you can look at it, you the people who are doing the operations, you the people who are doing the development, everybody can look at the same monitoring system, trust it, and look at it to see how the service is doing.
Okay, so now comes decision time. Now somebody has a new release. They want to roll the release out. So now what they do is they look at the monitoring system.
If, in fact, the thing has been up 85% of the time, launch. Go for it. New version. Push it to production. Go for it.
But if, in fact, it's been up 70% of the time, maybe you want to gate that release. Maybe you want to stop it. If it's been up 50% of the time, for sure you want to go back and say, "Hey, maybe you need to work on the reliability of this."
So that's how error budgets function.
What's nice about this sort of thing, and what's nice about a lot of things in SRE, is that SRE is about these virtuous and reinforcing feedback loops. That's ultimately what it's about. Because if you think about how the error budget works, everybody starts working towards, in a nice way that everybody understands, the reliability that you're searching for. Those sort of feedback loops are great.
Another feedback loop that got picked up by the DevOps movement that thrills me no end is this one down here, which says that there should be a postmortem for every event, every significant event. We don't mean like your kid's birthday or Valentine's Day. But there should be some sort of postmortem. Though my kid's birthday... Anyway.
Postmortems need to be blameless. They need to focus on process and technology and not people.
Let me read that to you again because I like it so much. Postmortems are blameless and focus on process and technology, not people. If something went wrong, not how stupid was Bob or Susan, but what was it about our process, about our technology, about our monitoring, about our automation that allowed Bob or Susan to do this? It's not about that person.
Because here's another observation. It is impossible to fire your way to reliable. It isn't the case that if, in fact, every time something goes wrong and you have an event, you fire the person responsible for that event, that suddenly your team will be able to provide a reliable service. What it will be able to provide is that one person cringing in the corner with their cigarette, doing this, afraid to do anything.
Sorry, that's going to make a terrible picture. I apologize for shaking like that.
So that's the thing to realize. In order to get where you want to be, you can't just fire your way there.
Okay, so that's the what of SRE. If you have questions about SRE, that's a really nice skimming just to give you some idea of the flavor of SRE. If you want to talk more about this, I really would love to go deeper into this too. But now I want to make sure that I give you what I was asked to give you by Jason and other folks, which is the how. How do I take this into an enterprise?
So one way you don't do this is you don't take a job description and you don't mutate it to be that. Don't do that. Don't take the same job description and just do a search and replace. I have seen this. I have witnessed this. It will not work. You will not get what you want. You might get slightly happier people because maybe they think now they need a pay bump, but do not do this. This is not the way to get to SRE in an enterprise.
Can everybody raise their hand and say, "I promise not to change my job descriptions"? Yes, very good. Thank you for vowing that. I've encountered that. I just don't want the term to be diluted. It really means something.
Okay. So the way to get to SRE in an enterprise is this combination of the people and the machines.
Let's talk a little bit about the machines first, because they're a little easier. The thing is that in order to do SRE in an organization, you have to do somewhat of what Martin Fowler used to talk about for microservices. You have to be this tall, which means that your environment has to already have some value and use and emphasis on automation.
It will not work if you do this without doing automation, as far as I can tell. You must have your provisioning down pat. CI/CD, does this sound familiar? Maybe you've heard this, and that song, "One Thing Leads to Another," is my favorite CI/CD song, in case you're curious. That's why I play it beforehand.
Monitoring is really key, and we'll talk about that in sheer seconds.
And here we are. So this is a reworking of what Mikey Dickerson calls the hierarchy of reliability. He is going after Maslow's hierarchy, which Maslow actually never put anything in a pyramid. I don't know if you know that.
Mikey Dickerson is the SRE that is most well-known for leading the team that saved healthcare.gov once upon a time, and that was a serious big deal. One of the things that he found in this journey when he was doing this is that the first thing that they had to do, if we look at the bottom of this pyramid, is they had to get their monitoring solid. Because if you don't have your monitoring solid, you have no way of knowing whether you're doing a good job or a bad job.
I have encountered, and I've been to enterprises where when I say, "So tell me about your monitoring," they say, "Well, we have two guys in charge of it, but nobody listens to it because it's too chatty and it's not up most of the time."
You can only go down from there. You cannot go up. And I've seen that, and it's really bad.
And then you want to start thinking about things like incident response, which is, again, how do you learn from the experience? If you don't learn from your experiences, you don't learn from your downtime, you are not going to get better.
So that's what happens when it's happening, and then learning from it through postmortem root cause analysis. Now RCA stuff is starting to give way for more sophisticated versions. Testing release procedures, capacity planning is up there, development, and then products like what color blue is it, is not at the bottom of the pyramid. If you're spending all your time dealing with the functionality, you're like, "I only wish that button was bluer. We'd have an awesome product." Mm-mm.
So this is what you have to get set up in the land of the machines. Get your act together along this to get into the realm of SRE.
Okay. Now let's talk about the people, and this is not a Facebook reference.
There are certain requirements in my experience, having talked to a number of organizations who are doing SRE and doing SRE well, that you have to pay attention to.
The first question is, what are the values of your organization, and can they be aligned with reliability and this sort of thing?
So an example is, I have talked to at least one company that says, "Our value is if we say we're going to ship a product and a release on this time, we're going to ship it on that date, come hell or high water."
And that's not a bad value. When you hear that, it's not by itself problematic that they say they're going to do what they're going to say they do. But that kind of goes a little bit against the error budget idea. That says you might want to gate your releases on something.
And you might ask yourself, and this is a question I've had with Jason and others, you might ask yourself, what if I said to you, "Hey, do you ship if you find a security bug? Hell or high water?" So the question is, is reliability and security, are they somehow connected? Might be a question you might want to ask yourself.
So you have to have the right values around this. The right things have to be important.
I think it's safe to say you have to have somewhat of an engineering culture. Because a bunch of this involves some software engineering that's going to be going on. Or at least a bubble of it that you can start within your company.
And then you need to have some resources dedicated to this. You'd be surprised how SRE doesn't happen if you just put the Sea-Monkeys in the water and wait.
Are you all remembering Sea-Monkeys? Okay, can we do a quick Sea-Monkey statement? Are you aware that there were two packets you had to put in the water? One of them was supposed to be a clarifying thing, and the other one was supposed to be, I don't remember what it was. But there were two packets, and it turned out that the first packet was the actual Sea-Monkeys. And the idea of putting a second pack in was to wait till they got big enough so you could see them.
Just saying. It's an interesting fact. Go look it up. I don't know how we're on Sea-Monkeys now, but there's a really awesome documentary about Sea-Monkeys. I'm glad to show it to you. I can't remember where it is.
Okay. The other thing you have to have in your environment is a toil-intolerant culture. So toil are the activities that we have to do every day that do not yield any long-term useful benefit.
I'm not talking administrative overhead per se, but there are lots of things that you spend a lot doing in your environment that don't have long-term benefit and you don't actually get anything from. And if you think it's great to do that in your environment and it's awesome to have Bob push the button once a day and that's it, awesome. Go for it. But if you think, why does this have to be? If it bothers you, that's a good start.
I think there needs to be certain ownership structures. If you're not doing things that are around products and stuff like that, or at least not allowing it so that there is joint ownership around devs and ops, you can't really pull it off, at least in the experience of what I've seen.
And similarly, I want to talk about org structure. So I remember talking with Pedro, who's the head of production engineering at Facebook, and he says that one of the things that was important to him is that his boss, the person who's in charge of the operations of the world, and he's kind of in charge of that, is on the same level as the person who runs the dev organization.
The theory being that if you have this org chart where the operations side and the dev side have to chase up the chain and eventually come to meet at the CEO, that's not going to work. So having the right org structure is important. Otherwise, you get into these battles that are not cool.
The thing that's kind of important is SRE requires management support at the highest levels. You can't pull off SRE if you can't say no. That sounds so pithy, I'm going to repeat it again: you can't pull off SRE if you can't say no.
And if you don't have the support from above to say no, or to say yes at the right times, or yes and, then you can't pull it off. That's what I've seen.
And I also want to say that there were some discussions. Andy gave an awesome talk a moment ago in which he was talking about different org structures. I have seen how Google does it, how Facebook does it, how Spotify does it. These are some of the things. And they aren't all the same model. It is possible to do SRE with different org structures and different models. But there are wrong ways. And we can talk about wrong ways, and that would be a lovely conversation at the bar. I'll buy you a root beer and we'll have a chat.
Okay. So how do you start? This is kind of a key question that people ask me all the time.
I think it doesn't hurt to have a problem or some downtime, or maybe you just have an epiphany. Maybe you went and fasted in the desert and your spirit animal popped up and told you, "Concentrate on reliability." Maybe that's what you did. But at some point you have to wake up and say, "Reliability, aha." And that's really useful. And sometimes you do that when you have a downtime. Sorry to say. I know you guys never have downtimes, but let's just say you did.
And then it's not a bad idea to get management support lined up. Make it so they think this is important, too.
And then I recommend going off and reading the SRE book, but read it critically. Don't just say, "Ah, this is gospel, and let us all turn to page 32 of the Google SRE book and open the Google SRE book and turn to page 32, and we will all now sing that hymn." I want you to read it and say, "Do I agree with this? Does this make sense for us?"
But read it critically. And when my book comes out, trash it, please, and then tell me what's wrong with it, because I really would like that.
And then I think it's also super important to hang out with other people that are thinking about this stuff. I can't begin to tell you, and you've heard this before in DevOps as well, there are conferences, there are places, there are meetups and stuff like that. Spend some time talking to other people. They make your brain grow bigger. They make you think about things. They have been some places that you have been. Listen to these things. I think it's super useful.
And then the last step that I would say, and you should note this is the last one, is I think you should start to dip your toe into the land of service level objectives.
Just for fun, use dry-run mode or something like that, where you sit down and say, "This service right here," and I suggest picking one that you can wrap your head around, that is human-sized, "what are the service level objectives for it?" And sit down and go through that effort.
And then once you've done that, put it in a monitoring system, and then pretend you have an error budget that means something. You could go ahead and have finger of God and make it true from the beginning, but just pretend it's really true, it's something that you care about. And then see, how well did we do this quarter on that sort of thing? And that will give you some idea of how you're going to feel about it when you actually really do it for real.
So I think playing with that sort of stuff is a good start.
I would love to talk more about SRE. I love this stuff terribly. Here's how you can find me. You can find me at my name or Otterbook. That's the first book I wrote for O'Reilly. It was the Otterbook. That was back in those halcyon days when we didn't really know what Twitter was for, so we just made up handles. And you can find me on LinkedIn.
And just so that I complete the thing that I'm supposed to do that Gene wants, which is how do you need help: I can state that two weeks ago, because of a corporate acquisition, I was laid off, and so now I'm looking for a job. So if you want to talk a lot about this subject, I am certainly available to work for you.