The Observable Universe...Creating Observability at Scale

As a Fortune 50 company in the Aerospace & Defense industry with 60K engineers across x countries developing products in the digital age, you can imagine the complexity.
Effectively transitioning to a DevOps mindset required exploring the problem space by asking:
-How do we create value that supports the cyber to physical spectrum at scale?
-How can we create observability of the work system?
-Is it possible to give teams end-to-end visibility of flow?
To answer these questions, we developed an approach that integrated lean, agile, and DevOps with sense-making and the science of socio-technical systems for teams be able to peer into the unobservable.
Chapters
Full transcript
The complete talk, organized by section.
Tiani Jones
Tiani Jones: Hi everyone. I'm really excited to be here today to share some work that I was able to do with a couple of colleagues that really expanded our ideas of transformation. So I'm just going to get the slides started and we'll get going with the discussion.
So the corporation that I was working in had questions about how to face the challenges of serving their customers, and how to position themselves to take advantage of new markets and how to innovate. And as we know, the aerospace industry has seen disruption already, and so how could they get ahead of those potential threats?
So then transformation became the topic of the day, and this problem started with a question, first of all. And the question was, can we build and design engines faster than five years? So this is a complicated product that is developed in a complex environment, multi-year development projects, programs with sometimes 1,000 people can be working on one of these systems.
We observed a few things that were very typical in most large conglomerates. So if you're familiar with Conway's Law or the mirroring hypothesis, then you know that any organization that designs a system will produce a design whose copy is a structure of the organization's communication structure. So naturally, we saw silos, organization around disciplines and departments rather than value streams. There were boundaries for information sharing that were based on countries, which was exacerbated by heightened or tightened ITC constraints to share and protect technical data so data wouldn't be leaked where it shouldn't be. There was also inspecting quality in, lengthy reviews and multiple reviews where others were inspecting quality rather than it being built into the ways of working, and the list goes on.
Additionally, they had standard work and a lean operating system, which sort of worked for operations, but they were trying to push that into engineering, and it was causing friction in engineering because they didn't have the flexibility that they needed to do the type of work that is done in engineering and design. They couldn't unlock velocity or flow. Things seemed to be getting slower. So there was a lack of understanding of the interconnectedness of the system as well, how and where to try things, whether that's a new tool, new ways of working, or changes to the org structure.
There was one team that we coached. We helped them experiment with a new material for engine fan design, and we helped them change from a typical V-model and this lengthy design and sort of kicking things over the wall and kicking them back between disciplines. We changed to a somewhat co-located and collaborative team, multidisciplinary team that designed experiments. So we asked them, well, how would you run this experiment? What's the first question that you need to answer? If that goes well, what's the second question that you need to answer? If that goes well, what's the third question that you would need to answer? And they put that together in a plan: three tests based on their hypothesis about this material, what they would need, what were the dependencies, what were the actual physical things they needed to build in order to test this new material.
And one of the scientists on the team, he had 30 years of experience in this domain. I remember him saying to me this quote: "I remember how we used to work. We'd just start building until we got it right. But somewhere along the way, the red tape came in, standard work came in, and we started to slow down. I'm just not sure how that happened."
So the question was, could we duplicate this? Could we take a team and think of an outcome we want to achieve and sort of change the way we work? How could that happen? That was part of the essence of the transformation.
Additionally, there was this hypothesis that if we incorporate model-based engineering, which is sort of like DevOps for cyber-physical products, it'll make things better. That should make things go faster. And then, you know what else? We're going to roll out Agile. So we heard people say, "Roll out Agile, implement Agile, deploy Agile."
Outside of where engineering was happening, though, there was not much in the way of new ways of working. And there were systemic issues that kind of came outside the boundary of the primary work system that teams couldn't solve. Yet all of these address everything just close to the primary work system and not really thinking through the interconnectedness through other parts of the organization.
Additionally, SAFe was kind of what they call the apex predator in terms of Agile at scale. And then there was the question, if we have our lean operating system, we have standard work, isn't that going to conflict? Who's responsible for helping us do things differently? And then became this wanting to scratch the Agile itch because that was the buzzword again, and not really sure where to start.
And so my team was a very small team, but we started out making sense of everything that we were seeing, and we thought, what could we propose? How should we be thinking about transformation? How should we be talking about transformation to these other groups? What could have success given that we were only a team of two? And ultimately, we came to one clear conclusion: we understood that Agile would not be enough. This was a systems problem. It required a systems approach, and so that's where we started.
Another part of our reasoning process was how could we talk differently and bring to bear some theories and research around organizations that might resonate in the system and start to change the conversation from can we just roll out Agile and SAFe, and can we morph it with our lean operating system, to a little bit more thinking through and perceiving what the system, that it exists and have an appreciation for it and some of the theory behind it. One of them being the fundamental attribution error.
So there's a white paper, as we were doing this research about different ways of thinking about it and talking with different people: nobody ever gets credit for fixing problems that never happened.
So what we did observe is that something would happen, a problem would happen, everyone would rush to address work or people who are close to that problem, assuming that that was the cause. And yet true cause may be distant in time and space from the defect that it creates. There's so many different examples you can think of about if you have a maintenance procedure where there's a high level of defects and the manager presumes that the operator is at fault because they're close to this procedure. The true cause could be inadequate maintenance procedure, it could be poor quality of the training program, it could be any other series of causes. And so this is the fundamental attribution error. That was one thing we observed, and it was another way of talking about looking at the system and change.
Another thing we discovered was the working harder versus working smarter, the capability trap, which is from that same white paper. So suppose that managers conclude that people, not the process or the ways of working or the system interconnectedness and the causes being distant in space and time, are the source of low performance. So once they make that attribution to increase the production pressure, then you have two effects that happen. One of them is worker effort immediately raises, which closes the performance gap, but then workers are now less able to achieve objectives by increasing the time they spend working, and so they continue to hit ever-increasing targets.
To continue to hit ever-increasing targets, they eventually resort to shortcuts because the pressure is still increasing to produce. And what that leads is to cutting the time they spent on improvement. And so if they start the shortcuts loop, which is shortcuts because they're cutting the time spent on improvement, in the short run, they might have the desired effect, but in the long run, there's a side effect that happens where you have less effort dedicated to improvement, capability begins to decline, performance falls, which offsets the initial gains that you had. And so by increasing throughput objectives and pursuit of better performance, managers who mistakenly attribute low performance to the attitudes and dispositions of the workforce have inadvertently forced the system into this capability trap. So this was another way to talk about and reason about this with people in different levels of the organization.
Another thing that we surfaced was, and that we thought through, is how could we distinguish ourselves from the fields of Agile coaches that were talking about SAFe, story points, velocity, and some on journeys of five years, and then hitting this wall of systemic blockages. So from early on, we were asked to bring Agile to the teams, but we knew that, as I said before, that we had to go beyond Agile coaching and the lean operating system. So we leaned into how can we reason on and talk about the system with the people that want to hear about Agile.
So I found this white paper, Agile-Based Patterns in the Agile Canon, and it gave us some food for thought about how to step back from specific methods and techniques and think through how to apply patterns.
And so the first one is a balanced measure economic progress, but the interesting thing about this is the author, Daniel Greening, puts forth that you should have a balanced suite of well-thought-out metrics. And in this case, the company really only had bowler charts, choppy bits of data at aggregate that provided no insight. There was no data thread.
The second was to proactively experiment. So that sounds great, but how can you know where to start experimenting? How do you know where to put the feedback loops? In addition, you have this sunk cost fallacy. So after we invest time or money in a project, we have to continue investing. And additionally, I'm just too busy building the thing. I'm too busy producing. I'm in that capability trap.
Collective responsibility, limiting work in process. So how do we illustrate and how do we explain how much work do we have in process? Do we know? Do we know what a work item is? When we're talking about engine design, is it a model, a simulation? Is it code? Is it a document? We found all of those things as work items. So how do you represent the value that's flowing through the work system?
Collective responsibility. How do we share responsibility based on the organizational structure and those other factors at play? How do we work together? What does it really mean in solving systemic problems? So bringing that appreciation of the system, surfacing what a system is and how it works was really important here.
And so based on all of that, we started to gather up all these morsels of information, all these morsels of reasoning and ideas that helped us develop some shared language with people already familiar with and exploring Agile, but we were leaning into the complex systems thinking.
So we started to articulate the values of the system. And okay, so you want to build faster, so maybe if you're trying to optimize the system, you want to optimize it or help it to be disposed to produce more quickly. And leadership was talking about speed a lot, building products faster. But then my teammates and I, we added learning. So is our work system disposed to or optimized for learning? Your ability to learn is related to the ability to go fast.
And so in product development, the goal is to produce the product that meets the outcomes of the customer market and the customer's needs and desires, but the system may not be operating with that same goal. It may not have the same goal. And when that happens, you get that strategic gap that Steven Bungay talks about. And you can put in some countermeasures to reduce that gap. However, the system may not exhibit the desired behavior of the system, which is learning quickly. So then why aren't we learning quickly? The system isn't designed for it. What are we actually observing in the system? Ultimately, it's behaviors in the complex adaptive system. Can we modify the system to enable the behaviors of learning? Maybe Lean, maybe Theory of Constraints or Agile will help. These are the questions we surfaced and how we were able to kind of bubble up this topic.
Additionally, we were working with Jay Bloom at one point. And so this is a quote that he had: "The observability is about seeing what's happening, not about an answer." So what we landed on was, given we have a system for developing these complicated products, observability was needed to enable transformation. So Jay actually worked with us to start shaping how we would propose a way forward to building products faster. Is the system enabling learning quickly? Can we observe the system's dispositionality to learning?
So we know about flow metrics. So flow metrics give you insight into the primary work system. Flow metrics are about the flow of value through the system. You have work in process, throughput, variability of throughput, and other things you can measure that are related to flow: queues, lead times. And this is the space of figuring out some of the answers to the questions that we had about the work items and handoffs, boundaries, ownership for a large program of up to maybe 1,000 people. And it's interesting because the lack of the data thread is where this surfaced, and can we do anything about that so that we can even represent flow? But we know about this, and there's a lot of research and experimentation and tools that we can use to instrument this if we need to, or if we're able to.
But the thing that was interesting, and if we start talking about is the system disposed to learning, how is the flow of knowledge related to the flow of value? It's really about measuring the impact of what creates flow, the correlated impact of measuring what creates flow. So it even feels like it's more upstream from the flow metrics, in a sense. And that's what we were really interested in exploring.
So Lean says to eliminate waste everywhere. All waste is equal, painted with an equal brush. Theory of Constraints says to find where the hunting is good. So what we understood is that if you have a large, complex, adaptive system, you're just kind of hunting everywhere. Equal hunting across all parts of the system would be wasteful. But hunting is like going to the Gemba, it's like going to where the value is created. So what if going to the Gemba and observing the system is looking for where there are things that are out of balance? So we create observability by means of flow metrics. We're trying to understand if we can do so regarding the flow of knowledge.
And what if we could look and see where things are out of balance? Not looking for active problems, because by the time the problem happens, you're too late. But where are things off? Where are there trends? And what if we could use metrics to hunt the system for behaviors related to the flow of knowledge specifically? Like I said, we already know about flow metrics. We know there's this idea of going to the Gemba and looking and seeing where value is created. The Gemba is the place where value is created. And if it's connected to knowledge, connected to learning, can we use metrics to do that?
So the key concept was that if observability is not about the answer, observability is about finding where to hunt.
What is the hunting ground then? The sociotechnical system. So another white paper that we discovered, "The Evolution of Sociotechnical Systems" by Trist, described the sociotechnical system as the whole organization, the primary work system, and the macrosocial phenomena. So the hunting ground then begins here. We know what the primary work system is, or at least we think we do, where the complicated products are developed. And now we can reason on, I see there are macrosocial phenomena and elements in the whole organization that now touch on and affect. There's interconnectedness there. So talking about the sociotechnical system and distinguishing that when we're talking about the work system became language that we could use and reason about as we explored this topic of the flow of knowledge.
And so I thought about what are the skills of a hunter? And I found this cool website that showed these maps, these really neat maps for all different kinds of animals that if you're a hunter. An armadillo in this case, and the black bear. It shows you the clusters of where you'll likely find those animals. And then in this website, they also talked about the top skills of a hunter. So marksmanship, mental toughness, physical fitness. I bolded the ones that I thought were kind of interesting to draw this analogy: navigation, bushcraft, observation, and patience.
So navigation, you'll need to know the hunting ground. You'll need your map or your compass. That helps you develop bushcraft. So what is bushcraft? It's all about venturing into the wilderness, and here's how they described it: venturing into the wilderness, feeling comfortable and confident. This comes with time spent on the hill, observing others and learning. Setting up a correctly sited camp, where you'll be out of the wind and protected from the worst of the weather, whilst avoiding any pooling rainwater. Identifying native flora and fauna, reading signs left by animals, and surviving any of the weather changes that Mother Nature has to throw at you. And that all comes from experience. Experience leads to confidence. Confidence leads you to explore. And the more you explore, the more comfortable you feel. And don't be afraid to ask questions.
So if you think about it, the hunter within the sociotechnical system, they need to develop a type of bushcraft where they feel comfortable rather than having the typical management response, which is to just jump in and assume that any problems or potential problems and their solutions or their causes are very close together. But you begin to be a hunter using tools for hunting, metrics in this case as one example, that allow you to peruse the hunting ground and get familiar with it. Asking probing questions, finding where the hunting is good, finding where there are things out of balance which can lead you to potential problems or opportunities before they become problems to get things back in balance and manage your system.
So now that we've established that this is a hunting ground and hunting is the skill rather than finding an answer, what if we could hunt for a Brent everywhere that he might be? We know the most common bottlenecks are policies and people. People comprise the socio aspect of the work system. They prevent us from learning quickly, these bottlenecks. We noticed that we had the ITC constraints, the matrix org silos. All these things indicate there's POP. And the first-hand experience of that propeller team that found one person in the lab that had all the skills to do the lab setup, nobody else knew how to do it. And they had to send someone in there to document it, save it, and other people were then able to take advantage of that knowledge. They took that knowledge, that tacit knowledge, out of that one person and were able to spread it around and unblock themselves. So it's that de-Brenting.
Brents have the tacit knowledge. And what brings forth or what comes forth from this is this idea of intra-predictability. Rather than controlling the response and making sure all responses are the same when if X happens, one person thinks that the response is Y. If X happens and others think it's Z, what if they work together? They share the information. They increase sense-making. The thing is, you don't know what you really know until you teach it to someone. So the idea of the transfer of knowledge begins to emerge here. And so that's where we started our metrics approach.
So we came up with this idea of balanced metrics. There's lots of reading and research and ideas out there on balanced metrics already. We leaned into this idea of pairs. So if we're talking about flow metrics, if you only focus on throughput, what about variability of throughput as a pair? So now that shows you if you're predictable or stable in your production. So the balanced metrics system for observability, that would highlight where gaming was happening because if you overdrive on one, you'll see a negative effect in the other one, and that gives you an idea that things are out of balance.
We used the ODEM approach. So the ODEM approach was really interesting because it focused on new behaviors and conversations that we wanted to foster. And so what's the outcome we want to achieve? What is the decision that we need to make? What's the insight you hope that this metric gives? How are you measuring it? And then in the measurement category, the sampling frequency, how to collect it, how to calculate it, what are the data sources, how could it be displayed?
And so finally, we did land on a few metrics that might be usable, and I'm just going to give you a taste, a sampling of this. So the hunch metrics that we came up with that we thought would create observability of the work system would be flow metrics: WIP, throughput, variability of throughput, and among some others. But flow metrics is a general category, which is the flow of value through the work system.
Balanced with metrics centered on knowledge. How do we know we were learning? Where are the Brents? What is the optimal skills mix? How do we transfer knowledge? Skills liquidity. In finance, liquidity is how you can move your investments around. In this case, liquidity is the knowledge. So it's, what is our active strategy for moving knowledge around? We want to see how knowledge flows and to what extent that we can employ or that we can introduce the idea of intra-predictability.
Another thing that was really important in this is measurements and behavior. So in the ODEM approach, when we talked through the outcomes and the decisions and the insights and the metrics, a key concept that underlies that was the behaviors that we wanted to foster. So behavior changes depends on what's being measured. Management behaviors change when they have the information they didn't have before in their hunting ground, their hunting maps. And the desired behavior is more investment in learning, more investment in double-loop epistemic action, actions that create new knowledge.
One tweet I really liked was from Troy McGinnis that talked about three types of dashboards and why you need them: strategic, analytical, and operational, and some of the time concepts that were there. So we played with all these ideas to see how we could do this suite of metrics based on flow of value, flow of knowledge, how they would be visualized, how they might be visualized, and how they might be collected, and who would use them and who would benefit from them, who's using them to hunt, who's using them to make their way through the sociotechnical system.
And finally, I know this was just a sample and I went through the metrics part quickly and hopefully in a future talk, we will have an opportunity to explore that in more detail. This is work that we just got started on and it had some interruption and we could deep dive on those metrics, but we're looking to continue this work ultimately is what I was hoping to say and test this in practice and discover if in fact, these are universal patterns that expand the idea of transformation.
So thank you very much and catch me on Twitter if this work interests you or reach out in the Slack channel if you'd like to explore this more and any questions that you might have. Thank you so much.