RegisterEnterprise AI Summit — Oct 7–8 · Charlotte, NC
Video Library

Log in to watch

Log in or create a free account to watch this video.

Log in
Las Vegas 2019
Share
Download slides

What’s in the Way IS the Way - A Story of an Unlikely Successful DevOps Journey

You may not have the kind of resources you need when you start.


Your people have to spend a lot of time supporting the on-going development and release so they don't have the time to work on automation.


Your architecture may not be as clean and well defined as you would like.


You just started with test automation and your coverage is still very low.


Above all, your customer and delivery partners are doubtful of the new initiative and are not willing to change.


In this talk, I will share the story of a successful DevOps implementation in a traditional organization with many these constraints. My experience in leading this journey taught me that if you can take a realistic assessment of your situation and plan your implementation carefully, it is entirely possible for a traditional organization with large and complex computing environment to adopt DevOps practices, in order to deliver software faster with significantly better quality, and at the same time do it in a way that makes peoples? lives much easier.

Chapters

Full transcript

The complete talk, organized by section.

Wayne Wang

00:02

Good morning, and thank you for coming to this session. I'm really excited about sharing my DevOps experience with you.

00:13

What's in the way is the way. What do I mean? When we are implementing DevOps in large, traditional, complex organizations, your journey is not going to be like this, no matter how much you want it to be like this. Instead, it's going to be something like this. You're going to meet a lot of challenges and difficulties. At times, those challenges can be daunting and overwhelming, and we can get very easily discouraged.

00:57

But the challenges and obstacles can also be your creative energy. We need to take a little different attitude. We need to let go of the expectation that our journey should be smooth and nice. Like playing a Tetris game, we have no control over what we have, and we have no control over what comes to us. We simply focus on how to figure out how to deal with the situation as each piece is thrown at us. With this, what's in the way can actually point into the way of solution. This is what I mean: what's in the way is the way.

01:45

In early 2015, I finally got the chance to move from ops to dev. Prior to that, I was the IT operations manager in one of the most challenging IT operations in a global automotive company. My plant was 24 by 7. It made high-volume, high-value product. If my IT system was shutting down the production line, the company would lose $45,000 every single minute. The pressure was very high. So you can imagine when I finally was moving from ops to dev, I was telling myself, "Okay, my life is going to be so wonderful. I'm going to be like semi-retiring."

02:47

Of course, quickly I found that's not true. Dev sucks even more than ops.

02:56

My organization was following the traditional waterfall method, with a lot of handoffs, long release cycles. We could only release software every six to nine months. But it took us eight weeks just to set up the new environment for each release. The release weekend was a nonstop three days, three nights weekend. Once we released the software, it's like this little robot on the floor. It rarely worked. So we usually spent six to eight weeks just to mop things up. We called it post-launch support.

03:43

I was thinking to myself, "There has got to be a better way, right?" And so I spent the next one and a half years studying how the best software companies deliver their software, and in the process, I discovered and learned about CI/CD and DevOps, including reading every book from Gene.

04:06

Fast-forward 10 months. By May 2018, we were able to release software every two months or more frequently, and it took us about one week to set up instead of eight weeks. And we cut down the deployment time from three days, three nights to about 10 hours.

04:32

If the speed is good, quality is even better. On this chart, the blue bars represent the total number of defects plus deployment failures. The orange bar is pretty much the deployment failures. The first four releases were the traditional manual, and the last two were using CI/CD pipelines. As you can tell, we cut down the defects, the release failures by about 90%.

05:08

Releasing software faster with better quality is probably every company's dream that is undergoing digital transformation, which is also why I'm working at J.B. Hunt Transport now, leading its DevOps initiatives. J.B. Hunt is a transportation service company from the beautiful northwest Arkansas. It has the largest domestic 53-foot container fleet in North America. In 2018, we had about 28,000 employees with close to $9 billion of revenue, and we had 2,000 employees in IT, and we grow very fast, too. Our flagship digital product is called J.B. Hunt 360. It's a digital brokerage software.

05:59

But the work I'm sharing with you today actually is from my previous employer. Let's just call it Global Automotive Company. It's 10 to 20 times larger than J.B. Hunt, and it has global footprint with millions and millions of customers. The sub-organization my team and I worked to develop the CI/CD capacity has 250 developers divided into 26 squads, pretty evenly split between North America and Asia, with a couple of squads in Europe. We supported about 50 applications, mostly core-type applications with very heavy customization. The computing environment has hundreds of servers, all three operating systems: Windows, Linux, and Solaris. And we supported about 30,000 users globally. So it's a fairly complex computing environment. We had a lot of challenges that's unique to us, but the five lessons I share with you today are pretty common across all companies and industries.

07:27

Let's look at number one. In the beginning, we probably do not have the kind of DevOps expertise. Really, there are two ways you can get these people: hire from outside, or develop your own. In the last three years, I have conducted at least 50 to 60 interviews for DevOps engineers. I can tell you hiring from outside is very difficult, time-consuming, very costly, and the quality is not always guaranteed. So very quickly, I realized, you know what? A better choice would be developing our own people.

08:13

We all have good developers and manual packagers in our company who are passionate about software delivery. We also have good system admins who can code and who like to automate. Those are the best candidates for the DevOps engineer. The challenge is we need to convince our people they have to be automators, not the automated.

08:45

What I did was telling my team, "Okay, this is new to us. I don't know how to do it. You don't know how to do it. Let's work together, figure it out." We have three choices. Option one: let us step up in the next three to six months. Let's figure out how to do this thing together. Option two: you may say, "Okay, that's not for me. I'm too this or too that." That's fine. We are a big company. I can help you move on to something else. Option three: stay where you are. Really, that's not a good option, because if everybody is stepping up, that means you are slipping downward.

09:31

So in the end, I did have a few people choose to do something else, and a few people I had to take actions to help them find some different alternatives.

09:48

Once we had a team in place, we needed to create an environment so they could learn, help each other, and grow. That included changing the physical working environment, like from the left-hand side to the right-hand side. That's very critical for you to develop your people and also develop your high-performing teams.

10:19

So if I think back, if I were given a lot of money to hire a lot of people, I don't believe we would be as successful as we are, because in the end, some of the best-performing engineers actually are from our own team, the engineers we developed, not those ones we hired from outside.

10:35

The next challenge: how do you balance the ongoing support with automation? Many of us are not given a big budget to hire a new team just for the development. Instead, we have to do ongoing support, ongoing release. Even if you are given a new team just to focus on developing, you will find once your thing is working, you are going to face the same issue because people start to love your stuff. They depend on you. So you have to make hard choices. You have to prioritize. And we all know making choices is not easy.

11:30

When my daughter was four years old, we took our family to Disney for vacation. At the end of the first day, my wife was telling our two kids, "Each of you can pick one toy from the store." So we spent the next 45 minutes in a store until the store was closed. My daughter was sobbing. She was crying. She did not pick a single thing. She loves everything about Disney. So if you force her to pick one thing, that means she has to give up everything else. That's a lot of stress.

12:16

But think about it. If it's difficult for a four-year-old to make hard decisions, I think it's even harder for us grown-ups, for managers or executives, to make hard choices. I want to show a hand. Which of you are only working on one thing, one task, and are allowed to do that by your boss? Okay. That's common.

12:49

All right. So we have to make choices, right? A common way of prioritizing is you draw those four-quadrant charts: not important, important, not urgent, urgent. What's going to happen? Everything is in that quadrant. Obviously, this way of prioritization is not working. But why? Look at this. If you have been in the corporate world for a little while, you understand anything asked by your management has to be important, has to be urgent. So if you are asking them to put something in this quadrant, you are insulting their intelligence, right?

13:43

Here is my improved version of the quadrant. Believe it or not, I had a lot more success using this version with my stakeholders. So once you can prioritize it, then you can focus your resources on those top-priority things. So if I'm thinking back, if my team were not having to do the support for so many things at the same time, we probably would not really spend time on prioritizing, and we would not be able to deliver so quickly. But because of this, we were forced to focus on those most important things, most urgent things.

14:37

I'm going to talk about the next two challenges together. They both are very difficult technical challenges in CI/CD pipelines. Let's look at my architecture. In the very beginning, I was very jealous of those companies who have those microservices architectures. They can deliver smaller changes, continuously release those into production with full automation.

15:11

But unfortunately, this is what I was dealing with: legacy architecture. And there is no quick and easy way to modernize, although it has to be done. Obviously, that's out of my team's control. So the question to you as the leader of DevOps: what are you going to do before that's happening? What are you going to do in the meantime?

15:42

Similarly, with test automation, when we start test automation on our DevOps journey, we probably do not have good tests or automated tests to begin with. Even if you have some test automation, I think the chances are you will quickly find out those test automations are not good enough for CI/CD pipelines.

16:07

On this chart, this is real data from a very successful large software company, and this is from one of their biggest product lines. As the orange bar on this chart shows you, in the beginning, they already had a lot of test automation, but they found the same thing. They found those test automations are not fast enough or too flaky for CI/CD pipelines. So they had to spend the next two years to redo their entire test automation suites. So the question again to you as a leader of DevOps: what are you going to do in the meantime?

16:57

One thing my team and I realized very early on is 80% of our customer requirements actually are very simple, cosmetic, very easy to automate, very easy to release. You don't need to shut down the system even to apply those changes. But those are the features our customers always want to have yesterday, today, next week, instead of waiting for six months.

17:30

So that is an obvious opportunity for us to construct our CI/CD release trains with two tracks. One is the express train. We can release very quickly. We run the train more frequently. We could probably deploy 80% of our changes that way and leave the other 20% to the freight train for the big batch release.

17:58

This is how it works. Assume in the next three months, or 12 weeks, we have 300 changes. If we have this express train, we release 20 changes per week. For 12 weeks, that's 240, 80%. That leaves 60 for the freight train, for the batch release in the end. As you can see, by doing it this way, your big batch release becomes a lot smaller.

18:34

And when we started this express release program, we picked one simple thing: style sheet. And there were five changes for the first time. It was a success. As we gained more experience, we slowly increased the complexity and types of changes, and we also increased the total number of changes we can do each and every week.

19:00

And we got very good results. The first six releases for the six weeks, we released 61 changes into production with only one failure. And we fixed the one failure in 30 minutes. So basically, the failure rate was 1.6%; MTTR was 30 minutes.

19:21

So as you can see, the legacy architecture did not prevent us from using CI/CD and releasing more frequently. Instead, it taught us a lot of valuable lessons, like: we picked up the easy changes as a learning experience for our team. It taught us the value of a small-batch release, more frequent release. We cleaned up the backlog for the dev teams. And best of all, because we were doing it so frequently, we were forced to release during normal business hours. Because if we were asking our folks to do this kind of release every weekend or off-hours, people are going to quit on us. So it ended up very positive for us.

20:27

Similarly, with the lack of test automation, we did not have a lot of good test automation. So we developed a small set of critical integration tests, and we used those tests as a guide rail, complemented by manual validation in both the lower environment and in production. I'm sharing this snippet of IM between myself and my business partner. As you can tell, as the release was in progress, I was texting my business partner in real time: "Okay, this job is done. This job is done. This job is done. You can check." As a result, we developed a very close relationship with our business.

21:22

Now, down to my last lesson I want to share. This is probably the most difficult aspect of any DevOps initiative: people's resistance to change. This is just a fact of life. Our work habits, the processes, procedures were the result of many years of work. If we can get by, we don't want to change. I don't like to change. Nobody likes to change for no reason.

21:55

And yesterday, the talk about Compuware, the CEO came in and said, "Yeah, we need to change everybody." So that's a rarity. If you have a CEO like that, count yourself as lucky. But for the rest of us, you are going to have to be very patient and take a collaborative approach with what you do. Even very small changes can become a big deal.

22:27

Here is an example. In April 2018, after four months of using CI/CD pipelines for weekly release into production, I felt we were ready for the big batch release using CI/CD. And this is my hand-drawn release plan for that first CI/CD release. And one of the decision points is actually where are we going to have these checkpoint meetings during the release? Because traditionally, our organization would have one checkpoint meeting every four hours.

23:10

In the previous weeks, my team already conducted four rehearsals using the same CI/CD pipeline for the release. And starting from April 16, we were taking 34 hours to deploy, and with each iteration, we cut down the time. Two days before that production release, we actually were down to nine hours. So I was very confident that if we budget 12 hours, we should be able to deliver. Actually, the production release was seven hours on May 4th. So that was a pleasant surprise.

23:49

Back to the meeting, talking about whether we should have those checkpoint meetings every four hours. Here is the conversation between myself and my business partner, where I was suggesting, "You know what? We don't need to have these meetings every four hours. If something goes wrong in IT, we just fix it instead of waking you up in the middle of the night."

24:12

This is the response I got from my business executive. She said, "That's condescending. Maybe I know you are still fairly new to our organization, and you probably don't know: we are so used to getting up at night every four hours, even if it's three days, three nights."

24:37

I said, "Okay. What about we do this? Let's have the first checkpoint at 10:00 PM Friday. And instead of doing the next one at 2:00 AM, let's push to 4:00 AM." I said, "Because 2:00 AM, Jenkins is going to be busy installing software. That is really not for us to talk about." She said, "Okay, let's do the 4:00."

25:03

Here's my conversation with my business partner at 4:00 AM Saturday morning. I said, "We are done. If you want to call the business testers, start testing, please go ahead and do so." And of course, she said, "Wow, that's five hours ahead of schedule. No, I think it's a lot of hassle to call people this early. Why don't you go ahead and send your people home? We're going to stick with the 8:00 AM original time." I can tell you, that was my happiest moment of my DevOps journey.

25:43

So with the CI/CD pipelines, we pretty much got rid of what I called the post-launch release support in the six to eight weeks when I talked about that in the beginning. After the May release, we had a meeting on Monday morning, and the meeting lasted for about five minutes because all the issues were already resolved. There were no new issues. And in the June secondary release, we scheduled a meeting on Monday morning. No business people even showed up, so we basically got rid of that thing.

26:26

In summary, difficulties and challenges can be the source of your creative energy and your creativity, and you need to take a different perspective. You need to let go of your expectation of what a CI/CD journey should look like, what a DevOps initiative should be like. Expect a lot of challenges on the way, but also you can treat those challenges as an opportunity. Because what's in the way can actually point into the way of solution. With this, I think you should be able to surprise and delight your customers, deliver value to the business. Above all, your life is going to be much, much better.

27:27

Of course, the journey never ends. I know mine is not. I was talking about the architecture, the test automation. What we did is just workarounds, waiting for those to be changed. But the hard work is still there: how are you going to modernize your legacy system, and how are you going to accelerate the test automation? Of course, in the conference, we heard a lot about how to bring business to agile so we are not doing this big waterfall type of planning up front.

28:06

Thank you for your attention. I probably have time for two questions. Any of you?

Q&A

28:14

Audience member: So you mentioned that you had a lot of COTS applications. What was one of the biggest challenges that you faced trying to mold DevOps into utilizing the COTS space?

28:36

Wayne Wang: There are many challenges. A couple of those come to the top of mind. One is their developing cycle versus ours, because if they are using waterfall, then they're probably going to release to you every six months or 12 months. And how are you going to synchronize your release cycle with them and develop the features into release of feature into production? That's one. Number two is how are you going to manage those backward and forward compatibility, right?

29:16

Audience member: Thank you.

29:21

Wayne Wang: You're welcome. Anybody else? Okay, thank you so much for your attention.