The Art of Practices and How They Support the DevOps Transformation at Africa's Largest Bank
This presentation focuses on how we have used DevOps Engineering Practices to drive our Digital transformation at the Standard Bank Group. When we started this journey our aim was to answer a simple question:
Can modernized Engineering Practices (DevOps) create a competitive advantage for our Business?
Our aim was to answer this question so that we could showcase a direct correlation between the way we engineered our products and services and the benefit to our customer.
This presentation will aim to highlight our journey to date and the work that we are still doing to answer the above question.
This presentation will also provide insight into how we have integrated this work with some of our strategic Programmes such as our Cloud Programme and our Resilience Program.
It will also showcase how a 150 year old bank is trying to reinvent the way we look at Engineering and some of the pitfalls that we have experienced in our journey but also some of the great successes that we have seen.
Chapters
Full transcript
The complete talk, organized by section.
Natasha Anderson
Hi, everyone. My name is Natasha Anderson, and I am the Software Engineering Practice Lead at Standard Bank. Thank you for joining me today on my talk, "The Art of Practices: What Happens When a Tester Like Me Gets Hold of the Paintbrush".
My team and I are responsible for helping teams and individuals across the group get excited about, and better at, software engineering. So for some teams, we help them understand and implement modern engineering practices, like DevOps or SRE, for example. For others, we help them expand on the adoption journey.
We work with other practice teams to advance engineering by collaboration. My love for engineering started very early on in my life. As a child, I spent most of my time breaking things apart just to understand how they work, what they look like on the inside.
I love solving puzzles. I love building things. So the day I wrote my first piece or line of code, it was in Basic, of all things.
I remember going home and declaring with as much certainty as any 12-year-old could, this was it. This is what I want to do. And so for me, this was one of the defining moments in my life.
And from that simple childhood declaration, I've since been able to take on various roles in this industry. So I've coded, I've tested, I've trained, and I've coached. And the excitement and fulfillment, I suppose, that I get from doing the kind of job I do always stems from being able to solve problems, being able to be creative, and to connect with people.
So as the engineering practices lead, I wake up every day with this excitement and immense gratitude, I suppose, that I can dream up ideas and work with an absolutely extraordinary bunch of people to bring them to life. I get to engage with and partner with like-minded people from across the group. And I think it's important for me that I see my time here as an opportunity to contribute to the growth of my organization, to my country, and my continent in whatever small way I can.
So what I'd like to do with you guys today is, first of all, to share some of the foundations in establishing practices, the challenges we've faced and have overcome, the lessons we've learnt along the way. So in this section, what I'm going to do is just share what the practice is, how we've defined it, a short history in terms of how we got here, and an overview of our strategy, and a brief look at some of the items that we've created in our toolbox. Our transformation journey has been well underway, and some key highlights include the implementation of the scaled agile framework, journey to cloud, the huge amount of work that we've done on data science and machine learning.
And so when I joined in 2017, the bank had already been implementing DevOps, and pockets of excellence were sprouting up all over the place. Open source was used. Soapboxes and knowledge-sharing sessions were commonplace, even back then.
From a practices perspective, we had practices representing multiple domains of expertise and technologies, and specific engineering practices, for example, development of quality engineering. But by the end of 2018, the organization was ready to shift us up to the next gear, to embark on the next phase, and one which led to the reshaping of our organization to allow autonomy across the various lines of businesses. And it was at this time that we collapsed some practices into a single engineering practice.
And our role is to define what good engineering practices look like. So at a very high level, what the practice is responsible is for driving improvements of engineering methods and tools and techniques based on lean and agile principles, for developing metrics and measures to create visibility at an enterprise level for the state of engineering across the organization. For developing methods for attracting and growing talented individuals into the group.
For creating and building an engineering culture with engineers in the organization through continuous learning, through knowledge sharing, and through platforms that allow them to evangelize engineering. So my previous practice was the test automation and the research and development practice, and that practice was collapsed as part of the changes at the end of 2018. So it was at that point that I took on the new challenge of establishing the software engineering practice.
And with a brand-new canvas to start creating this masterpiece, I began to ask myself: How will this practice be different? How will we make a difference? How will we make an impact?
How will we move this organization forward?
And so in the first weeks, being back in early 2019, we set ourselves the exam question, how can software engineering be our competitive advantage? Going through the State of DevOps report, at the time, looking at the metrics, it was research based on Gartner, the Accelerate book, and DevOps for the Modern Enterprise. And this data, especially from the State of DevOps report, that operating in the low IT performance world was surely a route to irrelevance.
The distance between them and the elite in high IT performance was quickly becoming very far apart. Yet, we live in a world where the technology and the learning is available for us and already showed what was possible. And so with this exam question, we then set out to look at what industry leaders were doing.
And we asked ourselves, "Well, how are these guys moving? How are the elite performers able to respond to outages so quickly, being able to implement changes so quickly?" How is this rate possible?
And it was then that we looked at, well, it's continuity. So continuous integration, continuous deployment, continuous testing, continuous monitoring. Everything needed to be continuous.
For us to be able to do that as an organization, we looked at some areas that we needed to make some shifts in. Some were underway, but we needed to get away from these pockets of excellence to get into it being commonplace. So for making the shift to continuous, we needed to shift the focus for our team from just delivering features to delivering real customer value, to get them to continuously ask, "Is this going to deliver real impact, and how will I know?" We all live in a sharing economy, and we know that we go so much further with each other than if we try to go at this alone.
So we needed to shift from seeing our suppliers as just vendors to partners we can co-create and innovate with. We needed to make sure that the shifts in engineering, the improvement initiatives we're investing so heavily in, could be measured and could be visualized across the group. We needed to make engineering an aspirational, singular identity.
So rather than people in teams focusing on my piece for testing, on my piece for devving, on my piece for database, whatever, we needed Standard Bank engineers to have a single identity of what it means to be a Standard Bank engineer, to be individuals with T-shaped skills. And finally, where's the learning, where's the fun if you can't share and contribute back to communities that helped in giving you your edge? So we also wanted to shift the organization from being followers in the global industry to being active global industry contributors and thought leaders.
This was, of course, the dream of where we wanted to go, and it led us to putting together, on a single piece of paper, our true north. So we would realize our competitive advantage of translating business vision to value to our customer by driving improvement in speed, quality, and innovation. Our adoption of DevOps and SRE, cloud, and agile, we'll consider the three spokes in this wheel to get us there.
So we work closely with the agile coaches of programs such as those that drive resilience, those that drive cloud adoption, to move the entire group forward. Underpinning all of this is an engineering philosophy, engineering participation, and engineering playbooks that supports this vision. The engineering practices is made up of the software engineering practice, the knowledge management practice, the RTE, Scrum, site reliability, and engineering tools practice.
Us as a collective is called the engineering practice. Our philosophy that we put forward was make everything continuous, automate whatever you can. We work from a premise of you build it, you run it.
Quality is everyone's responsibility. Learn, experiment, grow, and repeat. It's all about being a learning organization.
So our participation for engineers comes from my team, for example, that coaches teams across the group. They need to be influencers. So as part of our perform-to-grow process, part of what we have is a piece that you need to be out there coaching teams.
You need to be creating playbooks according to these ones that we have over here, but also you need to be out there contributing to open source, becoming a thought leader, publish whitepapers, participate internally in soapboxes and guilds, influence the industry. Our playbooks, policies, and standards, et cetera, cover the areas where we are building our organizational knowledge in these, what we call playbooks, which is how do we win at. So how do we win at engineering through a standard, which is an outcomes-based document, a living document that we are continuously working at reviewing, updating, and incorporating new practices.
What KPIs can we share so we can create consistency across the group? What guidelines can we put together? Technical frameworks, tooling, training, engineering maturity, playbooks for health indicators.
What does the recruitment and onboarding process look like, et cetera? Then, of course, the engineering metrics and tools which underpin all of this.
This is how we went about establishing what the practice would be responsible for, some of the philosophies, and some of the things that we wanted to create. This was all still 2019, and by the, I want to say February or so, we still needed to recruit the team. We still needed to put the engineering coaches together, get all the other practice leads in the team.
And since then, a year and a half now, we've been able to put together our toolbox for software engineering practices, and it consists of these areas. So the one that I'll spend a lot more time talking to you about is our metrics dashboard. So for the last part of the year, we went in and we said, okay, if we're going to be implementing improvement, we need to be able to show that this is delivering value in the organization.
And so we created what's called the DevOps dashboard, and it really tells are we getting better, right?
Then in order to be able to identify areas of improvement, we have a maturity assessment. We have maturity assessments that then lend itself to improvement plans and roadmaps that teams can then leverage to be able to say, these are the capabilities I will work on to get better at engineering. And if I've configured my pipeline through the DevOps dashboard, then I will be able to visually see, does my deployment lead time reduce?
Does my release cycle time reduce? And to help teams to be able to get better and to leverage off either lessons we've learned or other teams have learned, there's a series of playbooks. First of all, the overarching one is the engineering standard.
And then we have playbooks for-... tools that we have. We have playbooks that speak to specific areas like SRE, DevOps, and a whole lot of learning and development which supports it.
Now, when we put this toolbox together, one of the important things that we tried to do was to make it self-service. In order not to be the bottleneck, we service the entire group. So in order not to be a bottleneck, what we did was we looked at the toolbox and we said, okay, it can be facilitated by a coach, or you can now pull it from the toolbox, which is on our group sites and be able to use the playbooks to be able to run an SRE deep dive in your area or a DevOps maturity assessment or onboard yourself to the DevOps dashboard.
So I want to spend some time on the topic of the dashboard. So when we were looking at doing the dashboard together and answering the question of are we getting any better, this work was based on the Gartner report, Measuring DevOps Value, as well as underpinned by Accelerate, like I mentioned, and Mirco Hering's book, DevOps for the Modern Enterprise. So at the beginning of 2019, what we did was I sat with a few teams, our SAP team, our core team, and I said, "Guys, when you look at delivering features, how does it work?
I mean, what's that line that you draw from when it's requested, to the time it's deployed into production?" And this was the value stream that we put together. A feature is requested, it's prioritized with the PI, it's pulled into the team backlog, you dev, test, repeat, it's marked available for release, and then it's deployed to production. So we then said, okay, for our DevOps score, what we will do is we'll measure two things.
How quickly can we implement changes into production? And when things go wrong, whether it's through changes or whatever, how quickly can we respond? Because if you're moving at an extremely fast pace, then chances are things will go wrong.
So what do we do when they go wrong? So, the 12 or so metrics that we had here, deployment lead time, release cycle time, change frequency, throughput, MTTD, MTTR, et cetera, these were all based on the value stream that we put together. And we then looked at how the organization was working, and we looked at, okay, so if we took feature requested to production, one of the things we realized is that we use a tool for planning our work in progress, and we use another tool to manage changes as well as incidents.
So how do we tie these two things together? And once we understood and we collected data, I think it was from around 350 or so teams in the organization, we were then able to create calculations behind that to be able to say, from the time a Jira is created to the time it's released into production, if we could map that change request number, we could calculate that difference in time based on days. And so at the outset, we created an MVP, just pulling data out of Jira and using Excel and Power BI, and we were then able to calculate quite a few of these metrics based on just Jira and our Remedy system.
We then said, well, we need to be able to take this and ensure that it's easy for us to be able to pull the data, to consolidate the data, and calculate the rest of the metrics that may not be so visible here, like for example, pipeline coverage. And so what we did then was we looked at tools and open source tools, and we had already been playing around with Hygieia, which is a Capital One open source project, for a while now. And so what we did is we took that and we said, okay, this would be able to get us data from various pipeline tools, so whether it be Bitbucket, Remedy, Bamboo, et cetera.
So Bitbucket, Bamboo, Nexus, et cetera. So we'd be able to pull that data and show in a single pane of glass some of these metrics. Some of the metrics that we've implemented to date include deployment lead time, in which we've both collected for Remedy and Jira, where we've sourced that information.
And based on the mapping the teams do within the Jira itself, we're able to then calculate things like deployment lead time, release cycle time. Change frequency and throughput we do get from just the change requests that go through. And in this initial view, what we then did was then we started to look at the average time to repair and mean time to recover, which was items already captured in our Jira, so I mean, in our Remedy incident management process.
So we were able to bring that in. But when it came to the percentage of the degree of test and release automation, that's where we changed it to pipeline coverage. And that's where we've done a lot of work to date.
And in our degree of test and release automation, what we do is we look at a simple pipeline from source code management to build, to testing, to monitoring, security, et cetera. And we are then able to calculate how much of your pipeline is configured. And soWhat we built on top of Hygieia is the ability for teams to then go in and collect data or configure their data sources, so once they configure their sources, Hygieia is then able to translate that into data points for their team.
What is interesting here as well is that we were then able to not only show these metrics at a group level, we could show it at a portfolio level or right down to a team based on their application or a component in their application.
So, more important than what we created, though, is what we learnt along the way. So what I'm going to go through is just some of the lessons we've learnt. And, let's start here.
Culture is your canvas for creating any engineering masterpiece. So here's some of the lessons we learned. Number one, the most important thing of developing culture in your organization is leadership.
So let's take a moment to just reflect on this picture here. In this picture, you have an artist representing a leader, and his image symbolizes varying emotions. It arouses different emotions in us, right?
Yet, in this picture, that emotion that comes through is based on the work the artist puts in to make the experience real for us. And as leaders, we are the artists of culture in our teams. We set the pace and the course for any initiative, be it transformation, delivery, anything.
We were in a very bad space, the team, my engineering team last year. We faced the textbook storming challenges. As a new team, we didn't know each other, we had very little trust in each other.
So even though we had a group of really brilliant people, as a team, we weren't cutting it, and we were struggling to deliver. And at the time, we engaged a coach, and the words Mel left me with, now for me, a brain tattoo, "Leaders get what they deserve." Leaders get what they deserve. Let that sink in for a moment.
This was a lesson on what truly being accountable as a leader starts to mean for me. I get what I deserve. So if I get strained, disengaged, stressed out, non-responsive people, it's what I deserve.
And what she said was, it's because you're either teaching it or you're tolerating it. Either way, you'll get what you deserve.
So in terms of getting the best in terms of implementing practices, here's some of the leadership traits that I value as a leader and those that I respect in those leaders that I respect at Standard Bank. So the first thing is that leaders inspire. During the interview process, when engaging with various senior leadership in preparation for this talk, it was in one 45-minute session that I felt challenged, I felt encouraged, and I felt energized.
And those three things is what makes an inspirational leader for me. That energy not only allowed me to get comfortable here today to tell you my story, but also to be brave, to put myself out there, to contribute outside of my regular commitments. And it made me want to research and to learn.
It made me want to share that feeling with my team and those that I came in contact with.
Leaders also connect. Standard Bank is nothing like I've experienced. It's a place made up of a network, and this network will help you get things done faster.
That I've learned, and you will clearly see it if you walk past a bank on any campus, is that there's an accessibility and an inclusiveness of this network. And there's always someone willing to help or knows someone who can help. And this extends to the senior levels of leadership.
They set the pace. They set the speed at which we go. Their inclusivity, their ability to connect with people, their ability to inspire people is what gets us there.
The interaction is easy and welcoming, and they do this because they know their stuff. They're serial learners, and therefore, because of their ability to learn and connect, their interactions offer us an opportunity to grow. Which brings me nicely into my last point, that leaders outlearn.
In his last days with us, my previous lead shared something with me that I'll offer to you today. He said, "When your team succeeds, you succeed, and it must all be their work." So I said, "Chris, but there's so much to do. How can I not help?
But more importantly, wouldn't I just be seen as a delegator rather than a contributor? Would I really be adding value if my team did all the work?" And he said to me, "Natasha, the best leaders I work for spend their time investigating, reading, developing themselves. And if you went to present any piece of work, they would challenge you because based on that learning, they'd hold strong points of view.
And as a result, they multiply their capability as opposed to if they just spent their time buried in delivery." So check your calendar. I implore you, check your calendar. As leaders, how much of time do you spend on expanding your thoughts?
So Peter Senge in "Art of a Learning Organization" defined a learning organization as a place where people continually expand their capacity to create the results they truly desire, where new and expansive patterns of thinking are nurtured, where collective aspiration is set free, and where people are continually learning to see the whole reality together.
Ninety-two percent of organizations are more likely to innovate. Ninety-two percent of learning organizations are more likely to innovate in their industry. Fifty-eight percent are more prepared to meet future demands.
Yet only 1% of an employee's workweek is dedicated to training and development. So the five steps of a learning organization was defined as personal mastery, mental models, shared vision, team learning, and system thinking. And how we've implemented this is, first of all, clarity.
Are teams clear on what they need to achieve, and do they have a realistic perception of how to get there? Continuous learning for us is embedded in our Perform to Grow, which is our performance management process, in the form of something called 200 Hours of Learning. So when we set out our goals for the year, my team and I use it as an opportunity to reflect on what we are working towards in terms of the organization goals and strategies, of course, but also what does success mean for each of us personally?
And this year, the question that I posed as we were setting the scene was: how will the organization know that we were here? How will they feel your mastery? How will they feel your impact?
And for my team and I, our personal development is intertwined and infused into the goals that we have, with the outcome being the impact that we want to have in the organization. Having a clear target of 200 hours and specific areas we want to target as a team in terms of development, such as our cloud certification, SRE, security, et cetera, provides us with a shared vision, clarity, and outcomes, and also a spread of capabilities across the team. Our teams are supported through platforms that support either self-learning, formal training, through training partners or classroom training, as well as informal learning and experimentation through hackathons or through guilds and master classes.
And lastly, experimentation is part of the learning process. As part of our test automation practice, for example, practice Fridays were days to show up and to showcase, to list out any kind of problem we were facing in our individual teams and work together. How many of your teams would jump at a chance of learning if you, as their leader, suggested the appropriate material and courses?
Their research suggests up to 75%. What degree of importance do you think your team places on learning and development as a factor of happiness for them inside the team? Their numbers show as high as 94%.
And so finally, from a culture perspective, the last piece that I want to touch on is safety.
So have you ever tried to get anything done when you felt anxious or uncertain or afraid? How would you rank your productivity or ability to improve? Dismal, right?
So just take a moment to think about where we are as a society, the pandemics that we are facing, COVID, the social injustices that we see around us. These are trying times, and we know that the downstream effects are going to be felt years from now. Our people need to feel safe now more than ever, and it's our job as leaders to make sure they feel included, they feel safe to challenge, to learn, and to collaborate.
Why? Because when people feel safe, that they feel that they are included, that they are considered, when they feel safe to try, to debate, and to challenge, they commit, they collaborate. And I've seen firsthand how engagement and outcomes achieved skyrocketed as a result.
For us, how did we get there? So some lessons that we learned, for example, is don't blame. Don't talk down.
Don't tolerate it either. Don't interrupt each other. If things get too long-winded without clarity on next steps, time box, but commit.
Be curious about feedback and other people's input. These will take on different levels of complexity depending on your team. And I must admit, working with a team made up of grads is a lot easier than rockstar subject matter experts.
The stakes, I guess, are higher for being vulnerable. But this is our job as a leader. It's to have our finger on the pulse of our teams, to watch for quietness and nudge for input, to make time to create clarity, to reassure when mistakes are made, and encourage the team to see it as a learning experience.
So the last piece that I want to share with you here is a piece on mastery. Right? So I first heard about or came across 10,000 hours when I read through Malcolm Gladwell's book, "Outliers." And one of the things that I found interesting was, are geniuses made because they're genetically predisposed, or are they created because they put the hours in?
And what he proposed was 10,000 hours to make you an expert in anything. So whether it's cloud, continuous integration, data science, Scrum, spending the time learning, spending the time experimenting, spending the time to understand what works for you and your team and what doesn't work for you. When you do this, you'll be able to make the headway and succeed at overcoming almost any challenge you set out to.
And I hope so far that you've seen that the importance of making time and space for the team and for individual learning is, again, made clear. And it can't just be about new features to production. It has to be about: what are you doing with your people to help them to be future-proofed for the challenges that lay ahead?
And in my final thoughts to you, I'm going to summarize with this. As a leader, it's our job to make sure that there's a spark in the eyes of our people. Right?
It's our job to make sure that they feel safe. The next thing is that there's no better time for engineering to rise up and take your teams and organizations to a new level. So leverage tech, use open source, report on metrics that matter, please.
Lastly, keep expanding your skills and restocking on your supplies through your network, through your learning, and through your curiosity. And finally, I love this quote by Ben Shneiderman, and he said, "Leonardo da Vinci combined art and science and aesthetics and engineering, and that kind of unity is needed once again." Thank you.