Our First Inverse Conway's Maneuver Was Not Enough

I guess you already know that there is a Conways Law.
During this session, I will not tell you too much about what it is but why we thought we should make use of it. I will share what and how we did our inverse Conways Maneuver.
Finally I will close with some insights where we still feel a lot of pain and why.
Chapters
Full transcript
The complete talk, organized by section.
Rene Lippert
[00:00:08.160] Hello, my name is Rene Lippert, and I welcome you to the presentation which I've prepared for this year's DevOps Enterprise Summit. It is about our first inverse Conway's maneuver and why it was not enough.
[00:00:24.160] My talk will split into four sections. First is about the Lido business domain within Lufthansa Systems and what the Lido Flight product is about. Second, where did we originate from and why did we take the journey? Third, our trip into agility describes the new structure and where we are. And in the fourth part, I will end with where our pain points are today.
[00:00:52.860] So I will start now with the Lido domain. We offer products for commercial airlines of all kinds to ease their operations. In Europe, we are market leader, and we have customers all around the world. We also offer data services, which means that we digitalize information which is only available in paper. Yes, that still happens nowadays. And we have around 850 employees to offer the software product and the data services.
[00:01:27.680] Just to give you a feeling, the product we have has more than 2,700 configuration parameters to configure and tailor our software to the customer's needs, because they are quite special with all the laws they have to fulfill and all the different policies they have to fulfill. Of course, that makes our testing a nightmare, because every individual installation looks like a software in itself.
[00:02:00.140] Here you see a short look at our customers. In the previous slide, you saw that there are 120 airlines around the world, our customers, at least that was the figures before the crisis, before the COVID pandemic. As we are 100% owned Lufthansa subsidiary, we understand and we feel what it means at the moment to be an airline. Nevertheless, we are looking optimistic into the future, and we call the new normal being our new chance. And we know that our software will be needed again when the flights are coming back and the airlines are back into the sky.
[00:02:49.220] Here we see a map covering the different business scenarios where our products are in use. It starts at the bottom, where we plan the trajectory of the flight, providing the flight crew with the required information about their flight event, support to use airport maps for taxiing on the airport itself, take-off parameter calculation, and in-flight details. Our product allows aircraft-ground communication, so even providing updates during the flight, and ground staff can monitor the flight event on ground while it's in the air. And finally, we support during the approach and landing and taxi back to ramp. The last step is then the collection of the data for post-flight analysis.
[00:03:39.380] In today's talk, I will focus about one of these products used in this cycle. It's called the Lido Flight software. Oversimplified, we can say it's the TomTom for the flight crew, knowing not only the optimal way the pilot has to go from A to B, but also a lot more like what's the weather on the road, what's the wind, for example, and all the other information relevant for the flight.
[00:04:07.540] So now to the second part, which explains the way the organization looked before our agile transformation, and also some insights about the architectural characteristics during this time. We had many specialized persons on the development side as well as on the business side. Persons with dispatch know-how exactly know what the customers are doing, as well as persons with deep know-how in IT and writing code. You can see them here as green shirts.
[00:04:41.840] The resource management, so-called RPMs, planned our green shirts, first in a tool called Planta and later in MS Project. You see them here on the next slide in light blue shirts. The team leader was responsible for business and disciplinary decisions, so we made the most skilled person from the business side become a team leader, as this was the career path for the people. You can see them with the black cylinder. And the release manager was in charge to coordinate and track everything that was done according to the release process and being responsible to deliver the release in scope and in time, seen here with the gray cylinder.
[00:05:32.360] So here you see the schema how the development organization looked. You can see the Lido Flight product is developed mainly at two locations, Gdansk and Frankfurt. As I said, the team lead was responsible for the decisions on what has to be done and what can be delivered with which release. With the capacity planning for this, the RPMs managed the work for the people long in advance, and we even shared the people between different teams, and that was planned there as well. In short, we had a very classical, traditional waterfall product way of development: two version releases a year, seven service releases a year, and if needed, special patches. It worked, but we fought with all the problems such a traditional way of working has.
[00:06:25.768] Now, let's look on the software side. We had a wide set of languages in use. There was C, C++, embedded C, shell scripts, Perl, Python, Fortran, lex, yacc, SQL, Tcl/Tk, and lots more. All together, we had 3.6 million raw lines of code. Stripped down, it was more than 2.7 million. Our components had more than 100 dependencies. It was overall 26, and you can see a picture of it on the left-hand side.
[00:07:00.728] I like to speak about our system as if it was a solar system with a database core in the middle and over 1,500 planets, binaries, shell scripts around it. Not only the database was used as a persistence layer, we also used a shared file system and a complex shared memory construct was built around the product to provide data consistency and performance in this galaxy. We had a stable build and release process, which was already a big achievement, but it was a huge effort and hard to maintain and even harder to enhance. So we could not stay like this, and it was clear we need to transition in a modern world. We wanted to do this with a step-by-step and not a big bang approach, because the big bang approaches in the past we tried and had failed with them.
[00:08:00.128] We started a journey into a new, bright, and shiny world of working the agile way. Before I start explaining what we did during our redesign of the organization, I want to spend a minute on Conway's Law and what the inverse Conway's maneuver is. Let me quote one of the conclusions from Melvin E. Conway's paper, "How Do Committees Invent?" "The basic thesis of this article is that organizations which design systems are constrained to produce designs which are copies of the communication structure of these organizations. We have seen that this effect has important implications for the management of system design. Primarily, we found a criterion for the structuring of design organizations. The design effort should be organized according to the need for communication."
[00:09:02.588] Let me give you an example so that you better understand what we intended to do. So if you have two software components, A and B, which are closely related and you want to merge into a single component, you better merge the teams first. Then there's a good chance that you have created a homomorphic force which reflects in your architecture, so that the software component really becomes one. As long as you have two teams, your software will always have a kind of fraction where the two teams' work meets.
[00:09:40.668] We designed the new architecture to launch the inverse Conway's maneuver. Our lead architects built the new domain model with a clear responsibility. The aim was clearly decoupled services from each other and all communication happened through a well-defined technical layer, a kind of interface bus. The services are well structured into components. We call them building blocks. We decided to follow the self-contained system paradigm, so the architectural patterns to use have been defined in macro and micro architecture.
[00:10:20.988] After we had this, we set up the teams according to this model and hoped that by this, the architecture will simply follow Conway's Law and all will become good because of the homomorphic force the teams will use and make all bright and shiny. Unfortunately, we didn't understand, or underestimated, how much the communication, especially the uncontrolled communication, and how much the pull of the old architecture against this homomorphic force from Conway's Law is going to harm us.
[00:11:02.848] As the architectural work was done, we could start with our agile transformation with analyzing what different roles we need and how a team should be structured. We also thought about which roles we need beside the delivery teams. Here you see one of the whiteboards from this time. We decided that beside the delivery teams we had identified from the architecture, we need: one, platform teams like infrastructure, monitoring, and logging; two, across the teams, we need chapters to work on definition of standards, spread into all these teams, and reach by this an alignment across the teams; three, have, for very special know-how, traveling experts enabling the teams, for example a security expert going to teams and joining them for a sprint or two and sticking with them to enable them understanding security; four, have coaches who coach the team or individuals into their new roles on the job, for example an agile coach.
[00:12:18.468] Here we did a couple of mistakes. First and biggest, from my point of view, we did not take care about our ops people. We understood how agile software development is working, but not how to do agile software operation. Second is that we did not manage to bring the concept of traveling experts and coaches into an efficient life.
[00:12:48.468] Here's an overview about the roles which we have introduced in our organizational redesign. We still have many team members, of course, so for the majority nothing has changed in what they do on their daily basis. What has changed is that they are now clearly assigned to one team and that they have to work in an agile way. They are still the green shirts. To support the agile way of working, we've introduced the scrum masters, here as the light blue ninjas, and the product owners, here to be seen with the black cylinder and the red shirt.
[00:13:34.208] For us, the product owner is very close to the team, or some of them even understand themselves to be part of the team. And we have introduced one role called architecture owner, who bears the responsibility of the architecture, makes together with the team the technical decisions, and supports the PO in all the technical aspects. You can say the PO decides what has to be done and when, and the AO decides how it is to be done and documented. Here, we had a steep learning curve on our assessment for the right people in the right shirt, but this is a talk in its own.
[00:14:25.728] Today we have more than 30 of such teams with 15 POs and service owners, 15 scrum masters, and 17 architecture owners. How does that match with more than 800 people, you might ask? It is, as I said before, only the part of Lido Flight which I'm here describing.
[00:14:48.728] During the transition, we managed the organizational change, some might call it a reorganization, by the use of quality gates. We had three of them. The first was the preparation phase. We called it ready for decentralization. A clear mapping of the people into the teams happened, the tooling was in place and ready to use, and the team was aware that they are no longer centrally planned by resource managers.
[00:15:21.228] Second was ready for agile. Assignments of the new roles happened, so the POs, the AOs, the scrum masters, they all have been found and placed into the teams. The agile way of working was understood and the team decided which method to use. And the last one, the third, was done with Lido 4D. That was about upskilling the people in the team so that the new technology could be used and start working towards the new architecture to decouple from the centralized core, starting forming the team's own galaxy, if you want to see it in the old analogy with the solar system.
[00:16:07.128] Here you see the principles and values we introduced when we switched to the new agile way of working. First, the team must have the feeling to own their code, so give them the autonomy to decide and the responsibility to bear with the consequences. Second, we aim for simplicity and fight complexity, especially the accidental introduced one. Third, we want to have a collaborative culture. At the customer side, the product is seen as a single piece, so we need to collaborate to make this happen in front of the customer. Four, rules and guidelines from the lead architects help the architecture owner to design their micro-architecture in their service. Five, teams have also the budget responsibility so that they can become cost-efficient. Six, you build it, you run it. That was the way we intended to think with the individual teams. Seven, infrastructure as code and automation should be in the DNA of each team. Eight, increase the bus hit factor to something higher than one.
[00:17:37.588] So far, that sounded like we have a perfect plan, and it was very simple to execute and reach what we've achieved till today. I can tell you that was not the case. We had plenty of discussions, hordes of skeptical employees, and quite often we had to argue and fight for the greater good. We started beginning of 2017, and now more than three years later, we still face a lot of pain. Here are some examples.
[00:18:14.596] Before we started our change, the value stream was driven along the release process, though there was not much communication required for new releases, and also the technology change was rather low. When we allowed the teams to work in an agile way and reach out for the new architecture, we started to introduce a lot of new technology. The amount of REST services increased heavily. Automation tools like Jenkins popped in also for operational tasks. Containerization in the form of Docker was introduced. We also allowed the teams to bypass the release process to move faster for new services which are in Docker. But the old release process still dictates the value stream, and by this, also the overall speed.
[00:19:16.376] Shifting left the installation process increased the cognitive load on installing and THD team, because they had now to work out how to run all this new technology, learn all the new automation, and the tooling around the containerization. Instead of letting delivery teams do all this, including in their releasing of the service, we asked the delivery teams to automate the installation only on the internal systems, and they did. But we still try to protect our customers from the delivery teams. So we just shifted parts of the release process left, but the pain is still handled by the same ops teams, only with a much higher effort we need for communication and a much higher cognitive load by all that new technology coming in. Compared with our values and principles, we have not achieved the you build it, you run it ideal.
[00:20:27.956] Based on the book Team Topologies, there are four types of teams. I will here mention only two of them. There is the so-called stream-aligned team. They should be loosely coupled via a clear defined versioned API and being responsible for the value stream of its product, component, or microservice, you name it. They have the end-to-end responsibility of their part.
[00:20:57.336] Our teams, compared to this, focus at the moment a lot on developing new features or shifting to the new technology, but with the old logic and with the old coupling. So even if we have renewed the software to the new technology, we do have a tightly coupled monolithic architecture. In the book Team Topologies, what we now have is called a monolithic release. This is the reason why we still need the release heartbeat for almost all our teams.
[00:21:37.276] As told before, we still run the ops-runs-it model. The way we do this at the moment is a nightmare in operations. As a single team running all the components, being first line of defense, they have to collect and provide all the feedback to the delivery teams, which even results in more communication. On top, there is no good tooling around to operate the new services. Monitoring, traceability, and logging is not enough developed and not tailored for the operational needs, as this is mainly done by the delivery teams themselves, who barely get in contact with the customer systems where the incidents happened and where this should take place.
[00:22:34.536] As I told you before, we did not run through the agile transformation with the ops teams, so we have no clear idea what kind of team type the help desk, for example, should be. Should they be understood as a stream-aligned team delivering incident resolution and software updates, and more as a service to the customer? Or should we design them in a form of X-as-a-Service team, offering the services to the teams and cover incident resolution on a 24 by seven, but with a clear API to be used by all the different delivery teams? Or maybe a completely different form and type of team.
[00:23:26.076] We also still fight with the Three Ways because of these topics. We could do a lot better on the management of the value stream and focus more on flow. We should do better in the shift left to avoid too much indirect feedback and foster more direct feedback. And last but not least, we have to become better in the way we do post-mortems for customer incidents to be better with our continuous improvement and learning.
[00:24:03.996] Let's now have another look on our teams, more specifically on the team size. Dunbar's number oversimplified says, you can only build deep trust with a small group, five individuals, maybe two or three more. You can only keep regular contact with a few, name it 15 individuals. You can only stay in touch with some 115 individuals. Recently, that number was increased to something like 180.
[00:24:39.556] More than half of our teams are quite big. Our old philosophy on that was: we do lack skill in a team, so add a person with this skill to the team. The team feels overloaded. Ah, add more people to the team. On top of this, we do not manage the communication between the teams. We just let it happen. That leads to a very high cognitive load in many of these teams because a lot of people have to communicate, and a lot of our people complain that by this, we have far too many meetings, far too many communication, far too many information needed.
[00:25:28.636] We've easily forgotten one simple mathematical law. If the team size is N, you calculate the number of communication channels by N times N minus one divided by two. So if you have, for example, six people, take the six dots in the little icon I've shown here. You have six times five makes 30, divided by two makes 15 communication channels if all these people need to communicate simply within a team. And if you have six teams, it's the 30 communication channels between the teams.
[00:26:16.096] Our batch size is driven by the release cycle. As we still have the waterfall-based release cycles of two version releases and eight service releases, our batches are big. Big batches bear high risk, and by this we still suffer on the quality of our software. So we have to reduce these batch sizes and release more frequently, so that we can manage the risk a lot better.
[00:26:50.136] We still have many handoffs in our end-to-end value stream. We should reduce that and allow the delivery teams really to deliver and be responsible for their product at the customer's end. Not saying that there should be no ops. Ops is still needed. I'm pretty sure our very short reaction times of 15 minutes cannot be handled by an on-call of the delivery teams. Here, we still have the need for a clear operations 24/7 unit, and it's more a question of what is the right team type of that 24/7 unit.
[00:27:33.536] Last but not least, our ratio of unplanned work is not tracked, especially in our ops teams, and it's for sure far too high. It's even understood we are here to work on the incidents, so there is no planned work for us at all. We have only unplanned work. It is the planned work which fights against the unplanned work. You need to know what kind of unplanned work you have. So we need to start, track, and understand the roots of this unplanned work, and then we can plan to fight against it and avoid it from happening. So planning is your weapon against the unplanned work. And we need for sure a shift in mindset in our operations departments.
[00:28:31.776] Before I close, I would like to say thank you to Matthew Skelton and Manuel Pais for this book. The Team Topologies book helped me a lot to understand what kind of teams we have, what kind of problems we are facing in our organization and our organizational design, and it helped me to raise the why questions and answering them. If you have to design an organization, I highly recommend you to read this book first.
[00:29:03.876] Let me close with a very short feedback. I think we did the right things, but we didn't do them consistently enough. What makes me look optimistic into the future is the fact that the collaboration and the way the teams are working together has changed a lot over the last three years. I do have more often the feeling that we are becoming agile and not only doing agile. I'm sure we are going to address our pain points, and if you are interested in how successful we will be, then let me know this in the feedback to the talk. If you're interested, I should come back next year and tell you the outcome in another talk. Thank you very much for joining me on this session. I'm Rene Lippert, DevOps evangelist of Lido.