Your Data Nerd Friends Need You
How the world of data analytics, science and insights is failing and how the principles from Agile, DevOps, and Lean are the way forward. #DataOps
Christopher Bergh is a CEO and Head Chef at DataKitchen. Chris has more than 25 years of research, engineering, analytics, and executive management experience. Chris has an M.S. from Columbia University and a B.S. from the University of Wisconsin-Madison.
Chris is a recognized expert on DataOps. He is the co-author of the "DataOps Cookbook" and the "DataOps Manifesto", and speaker on DataOps at many industry conferences. Chris began his career at the Massachusetts Institute of Technology's (MIT) Lincoln Laboratory and NASA Ames Research Center. There he created software and algorithms that provided aircraft arrival optimization assistance to Air Traffic Controllers at several major airports in the United States. Chris served as a Peace Corps Volunteer Math Teacher in Botswana, Africa.
Chris Bergh, CEO and Head Chef, DataKitchen
Chapters
Full transcript
The complete talk, organized by section.
Host Intro (Gene Kim)
I met the next speaker earlier this year, and I'm so grateful for how he helped me get up to speed on the vast data problem that almost every large organization faces: trying to get data from where it resides to where it needs to go.
It is so often stuck in systems of record, data warehouses, and fragile ETL processes, often requiring months to get data to where it needs to go, which is in the hands of the developers so that they can use it in their daily work.
During those conversations, it reminded me of a presentation that went back all the way to 2014, whether it was at Telstra, the Heather McMahon story at Target, or stories we heard here through Adidas, Optum, John Deere, National Bank of Canada, and so many others.
After several conversations with this gentleman, I got so excited that this actually became one of the central elements of "The Unicorn Project." He spent tens of hours with me on the phone, often on weekends, to review scenes with me so that I could hit my writing deadlines.
I'm very pleased with how it came out, and I'm so pleased that it's one of the central elements in the book.
In fact, just to share with you, one of my favorite new scenes is that in the middle of the catastrophic Phoenix release, Brent gets pulled into an urgent call because all the prices have disappeared from the e-commerce site and the mobile app. The reason is that someone in marketing uploaded a CSV file with a byte order mark, basically making sure that none of the fields matched.
Someone actually did that to me. That person is Dr. Nicole Forsgren. She passed me a CSV file with a byte order mark, and it took me a half day to figure out how to fix it.
Anyway, I asked Chris Bergh, CEO and Head Chef at DataKitchen, to share with you what he taught me and all of the terrible, horrible, unacceptable problems that exist in the data space and how this community can help. Please welcome Chris Bergh.
Chris Bergh
01Opening: why DevOps people should care about data analytics
Hi, everybody. I'm Chris Bergh.
Where did that Baby Shark meme come from? I remember my son, 15 years ago, doing Baby Shark at preschool. Now it's in Lebanon. Now we're doing it here. So how would you understand where the Baby Shark meme came from? You would do something called data analytics.
I'm sure all of you have heard of it, but the purpose of my talk is to say that what you know here, the ideas in DevOps, Agile, and Lean, actually apply to a whole different domain that's bigger, that has more potential, and has more people than the systems that you're working on.
The title of my talk is "Your Data Nerd Friends Need You." What am I going to talk about? First of all, it's a big problem. The people who are working in it, I'd like you to get some empathy for them. They are really, I think, suffering. They are a lot like you and I. I'm going to have an ask at the end: I'm going to ask you to help. So pay attention.
02Data is big, buzzy, and failing
You can't walk through an airport nowadays and not hear something. You can't watch a sports presentation without hearing analytics. Some people call data the new oil.
There's a lot of buzz around data: big data or small data, streaming data or batch data, structured data or unstructured data, and a whole bunch of techniques that people are applying to data. Some people call it machine learning, AI, data science, or data lakes. There's a lot of buzz and a lot of acronyms going on, and it's exciting.
It's a huge market, a $189 billion market from tools and technology. GitHub was bought by Microsoft for $7.5 billion. Recently, a company that does data visualization, in fact one of the 50 companies that does data visualization, was bought for $15 billion. The next day, another company that does data visualization was bought for $2 billion.
There are tens of millions of people whose job it is, full-time or part-time, to take data and wrest insight out of it in some way. It's more than software developers. Some people estimate it's one in 25 workers. If there are two or three billion workers in offices, you can do the math. It's a lot of people.
It's a big problem because a lot of this stuff is just not working. It's failing. Most data science projects fail. Only one in five, or one in ten, models goes from a data scientist's keyboard into production.
If you go to a Gartner conference that covers this space, they talk about wanting to be a digital company. In the data space, Gartner talks about being data-driven as the high-order bit that you should focus on. Yet most companies are saying that they're less likely to be data-driven than they were even a few years ago.
There are these numbers, 60 and 80. I've had a long career. I've been a teacher. I was a Peace Corps volunteer. I was a researcher at MIT and NASA. I did enterprise software development and was a leader, a CTO, ran teams, and wrote code. About 15 years ago, I started to focus on data and analytics full time. Whether you call it a data lake or a data warehouse, or whether you call the application of it machine learning or BI, most of these projects always fail.
Why is that? Why are they failing? That's the question.
I'd like you to walk down the hall when you get back to work and ask the people who do this in your company. I think you'll observe a few things: they have poor quality and a high error rate on what they do; minor changes to a model, some SQL code, or a visualization can take months; a lot of their work is hijacked by unplanned work; and they're beaten up, oversubscribed resources. There is a hair-shirt culture of, "Give it to me, I'll take it." Does that sound familiar to anyone?
03Data people took a different door
Let's talk about these people and why I'd like you to have empathy for them. They have lots of different roles. Some people are data scientists, some people are data engineers, some people do data visualization, some people are statisticians, some people are architects, some people are administrators, and some people are managers.
What's interesting is they're just like us. It's as if you got your BS in computer science and, upon graduation, you took a different door. You went to the data science and analytics door.
The work that they do is on an incredibly complex toolchain where every one of those roles has a tool that they use. Some of those tools, in our parlance, are low-code development tools, which are configuration. Some of those tools actually are code, like Python or R. People love their tools.
If you go into your organization, you're going to find three or four different tools that do data work, three or four different types of tools that do visualization work, and the prevalence of Excel doing all of it. So there's this complex world, and they're just like us.
They work in teams: data engineers, data scientists, people who do self-serve analytics, and people who do data governance. They have different personalities. Data engineers tend to be like back-end software engineers. They're a little grumpy. They carry a lunch pail. People who do data science are like the algorithm people on a team. They tend to be kind of mad scientists. People who do self-service or visualization see themselves as data artisans. Governance people are like governance people everywhere.
A typical day starts when a data engineer sources data from somewhere. There are different systems that you've created and different places in the organization. In a simple case, there's sales data. Then a data science team takes that and applies some algorithms to it. Maybe they segment it or cluster it, and create a high-value and low-value segment. They're adding into that data.
Then there's a self-service team visualizing that for an end customer. This self-service team may not only do the visualization; they may add more data into it, perhaps an owner, a west team, and an east team. They have these small data files they're adding in. They're actually doing a lot of data work themselves, in addition to the work the data science team is doing.
Then the data governance team comes along, trying to catalog it and asking, "Where did this data come from? Where does it go?" They work together. That's something that's changed in the last 10 or 15 years. We've moved out of the heroic age of data and analytics, where it was one person who could do everything, the full-stack data scientist. I speak at data and analytic conferences all the time, and there still are data scientists trying to be heroes, but it really is a team sport.
04Toolchains, bosses, and many-to-many Dev/Ops
There's this gigantic, massive, fragmented toolchain that people use. There are Tableau tools to do data visualization. There are self-service visualization tools where a business user can do it. There are tools like Cognos where they're more IT-focused. There are tools to do augmented analytics, storytelling tools, and tools to turn analytics into words. It's a lot of cool stuff.
There's been a blossoming of data science tools that help people do algorithms and democratize the work. There is a whole segment called AutoML, as well as stalwarts like SAS and open-source tools like Python.
There are data catalog tools. There are ETL tools, extract, transform, load. They take data, massage it, transform it, and change its form. That's another 50 companies. Then there are a variety of databases tuned to analytics. They tend to be clustered and parallel. You've heard of Hadoop and Spark and Snowflake and Redshift. They're more attuned to analytical needs because they can do joins really fast, and that's something you need to do in analytics.
They may work for the same boss. In the last five or 10 years, there's a role called a chief data officer or chief analytic officer that's kind of in parallel to a CIO. Their job is to help the organization become data-driven, and they play both offense and defense. They're trying to guard the organization to see if there will be data breaches. They're the one who gets hung if data leaves the organization. But they're also trying to get the organization to stop thinking about analytics as building a house, as doing reporting, and more as a continuous delivery of value, as a process of delivering insight over and over again. Conferences and organizations around CDOs have grown up exponentially.
They all may work for that same person. But as in a lot of organizations, your data and analytics people may not work for the same person. There may be people in IT, maybe people in line-of-business reporting, and maybe a data science team reporting to the CEO. They may all work on the same project.
In software, the relationship between dev and ops is often one-to-one. In data analytics, to use a database term, it's many-to-many. You might have a centralized development team and a centralized operations team. You might have decentralized development, where every line of business has someone doing Tableau, Looker, Python, or Excel work on top of a central data set. You might have decentralized development and decentralized production, where local teams can publish their own dashboards, while other parts of the pipeline still go through IT operations, data production teams, or manual review.
The end customer sees one dashboard, one number, one report. But behind that could be a data warehouse team, a data science team, a local self-service analyst, a governance team, and source systems owned by IT. When the CEO says, "Something's wrong in my dashboard," it is hard to figure out which team did it, where the calculation lives, and how to fix it quickly.
05Analytical systems are factories of insight
These teams source data from internal and external systems: CRM, ERP, supply chain, websites, financial systems, HR, open data, syndicated databases, APIs, and files. Many of those are DevOps-governed systems. The data teams take that data and run what I think of as a factory of insight.
There are assembly places in that factory: access, transform, model, visualize, report, and govern. Each station may be code running in a VM or a tool running in an engine for that tool, like an ETL tool. They need to pick pieces of that factory up, move them into development, create a dev environment, and create a sandbox.
For every software engineer, when you join a new group, you have some scripts and create your sandbox. If you're good, you can do it that first morning. But creating sandboxes in the data science and engineering world is surprisingly complicated. You need test data: small test data, big test data, clean test data without identifiable information. You want servers and software sort of like production.
Creating these development environments is hard, surprisingly hard, and they're often surprisingly out of date. Some people have fixed dev environments that were last updated six months ago, so they're not like production at all. Everyone here can see the problem with that.
These teams have diverse tools, diverse people, and diverse customers, and they have a process to deploy changes to that factory from dev into production. It's not great. Most companies take a long time to do this. We're not talking about continuous deployment or continuous integration. Most companies are taking months to do this.
These teams need to do both simultaneously: they need to run a factory that has low errors and lots of pipelines, but they also need to change the pipelines. They don't want to break production to change those pipelines.
06Your data nerd friends are suffering
They're people just like you. They took a different door. They have a different set of technologies. My thesis is that the ideas in the DevOps movement, the Lean movement, the ideas that came from Deming, and the book "Flow" all apply.
This group is really, I think, suffering. I'd like to talk to you about that suffering because it's a suffering that I experienced. When I talked to Gene about hero culture, fear culture, insanely high error rates, and the complete lack of automated testing, those ideas became part of "The Unicorn Project." There is a whole chapter in the book that talks about technology review boards, and I'm so happy Gene put it in.
To be particularly honest, I want a whole book devoted to the trials in data and analytics because I've been living it, it's a big area, and it has lots of problems. But I'm thankful to Gene for putting it in.
We did a survey with an independent consulting organization called Eckerson, which focuses on data and analytics, and asked a couple of questions. How many errors do you have per month? How often are you late, the data is wrong, or you miss your SLA? Eighty percent of the companies have an unacceptable error rate.
When I show this to people in the field, they think it's all bullshit. They think most people have no idea what their error rate is because they're not tracking it, because they're living in avoidance of error. They don't want to talk about it. Most of this data flowing through systems now isn't tested. It comes in, maybe they find out it works. Reports don't show up. The customers are testing it and telling if it's right.
In 1999, I worked on one of the first social websites. I still remember having a million people on this website and changing code on the production system. It was kind of cool that it was this living, breathing thing, and we were changing code because things were breaking left and right. But that's not a sustainable way to work. Having errors all the time certainly isn't.
In another part of the survey, we asked how long it takes to deploy changes from dev to production. It takes months. Ten lines of SQL can take months to move from a dev system into production. These are really good people, people out of major insurance companies, people like you and I. I like them, and it pains me to see how long it takes them to do this.
They're also really slow creating development environments.
07My story
About 2005, like I said, I'd been a CTO and I'd run development teams. I thought, "This data and analytics stuff, no problem." I joined a company that did analytics for the healthcare industry full-time. I was the COO, and I was the guy who made the trains run on time. I had data scientists, data engineers, and people who did data visualization. As the company grew, we had thousands and thousands of people who used our analytics.
I worked for a guy who knew a lot about healthcare. He went to Harvard Medical School, but he wasn't a technical guy. He would go off and talk to senior leaders in healthcare companies and come back with a great idea. I'd go into a room with a data scientist, a data engineer, and somebody who did data visualization. We'd whiteboard it up, and I'd come back to him and say, "David, this is going to take us two weeks to do."
He'd look at me over the top of his glasses like I had killed a patient on the table, saying, "Two weeks? Chris, I thought that should take two hours."
I'd go back to my office. You'd think people in healthcare companies are nice, but when you have the CEO get wrong data, or a whole bunch of salespeople have their incentive comp reports wrong, or even if you took a line chart and moved it to a different place on a dashboard, people would call me up and yell at me. I don't like being yelled at. I'm an introvert. I don't like having extroverted, socially powerful people call me up and ream me out.
We had hired all these smart people. They had master's degrees and PhDs. They wanted to do data science and visualization. There had been this blooming of open-source tools, and they wanted to try new stuff. They wanted to innovate.
How could I have a life where I could go fast and not break things, and let people try stuff out easily? It's actually fairly hard to do. That's my story.
08DataOps and the ask to DevOps leaders
How can you help? My company and my co-founders have been trying to get this idea of DataOps going. We didn't call it DataOps first. We called it Agile Analytic Operations. We called it Analytic Ops. We called it DevOps for data science, data engineering, and data visualization. We tried all these different names, and we settled on DataOps because it's short.
I know some of you don't like the idea of DataOps because the Ops term is getting overused, and I sympathize with that.
In 2017, we wrote a manifesto. We stole words from the Agile Manifesto, a DevOps manifesto, and a bunch of Lean ideas, and put it together into 18 points. Surprisingly enough, 6,000 people have signed it. DataOps got on Gartner's hype cycle, and we're seeing increased traffic and increased search on it.
What is DataOps? DataOps is the stuff that you know applied to the world of data science and engineering. Gene has a definition of DevOps, and I mutated it in my own way. DataOps is a set of technical practices and cultural norms and architecture that enable rapid experimentation and innovation for the fastest delivery of new insight to your customer.
That's the coin of the realm in data and analytics: insight. What you want to deliver is insight, and it's kind of a random walk to get there. In many ways, people are doing more experimentation and spike solutions in data science than in software, but they're still experimenting a lot. They want to do that with low errors, because people who have data that's wrong don't trust it. There is also collaboration across teams and locations and environments, and clear measurement and monitoring.
When you Google DataOps, we created a definition that acknowledges its intellectual heritage from Agile, DevOps, and Lean. But when you look at it, it's the same stuff. If you know those ideas, if you've come to this conference and heard the same ideas, take those same ideas and put them in a different reference. Put them in the data science and engineering world, and it's the same thing.
Teams now often have a change-fear mindset. Many organizations, because they've gotten yelled at by the head of sales, build a wall of process around the work. Deploying to dev, moving from dev to prod, checking for errors, and alerting are all manual. There is a lot of hope. Some of the biggest banks in the world bring data in from many sources and don't do automated testing to see if it's right. They hope it works, or they take code from dev to production and hope there isn't a regression. That's crazy.
As a result, there are heroes. "I want to do things. I'm going to work nights and weekends." I was at a conference where a guy talked about fixing a bug while sitting in the bathroom during his kid's birthday.
Between DevOps and DataOps, they're pretty much the same, with a couple of differences. My ask is that you join the DataOps movement. The ideas in Lean, Agile, and DevOps, applied to manufacturing and software development, need to be applied to the world of data science and engineering. It's a world that's broken. It's a world that needs your help, and you have a unique perspective and set of ideas.
When you go back, ask people in this role a bunch of impertinent questions. Are you using source control for your work? Do you have automated tests? Do you have regression, functional, and unit tests? How long does it take to deploy from dev to production? Is it automated? How up-to-date is your development environment? How often are your business users finding errors?
Ask these questions. You will be horrified at the answer. Just horrified.
If you want to learn more, Gene was kind enough to take chapter 16 of the book and allow us to syndicate it. We wrote a book on the ideas of DataOps. You can get it for free. You can sign the manifesto, or ask me for the slides. Thank you very much.