Scaling with LaunchDarkly: A feature toggle story!

I will tell the Tide feature toggle story from the beginning and the value that LaunchDarkly has added to it.
Chapters
Full transcript
The complete talk, organized by section.
Giorgos Ampavis
Hello, everyone. Thank you for joining me today. I hope everyone is well and safe.
My name is Giorgos, and today I'll be talking about Tide, our journey and experiences with feature toggles, and how our partnership with LaunchDarkly has helped us scale the product and the teams globally.
I'll start by introducing myself and Tide. Then I will tell you our story from the very beginning and the very early days of Tide, why eventually we moved to LaunchDarkly, how we're using it, and what is our roadmap for the near future. Finally, I will close this talk with the final thoughts.
So let's start with introductions. My full name is Giorgos Ampavis, and I'm identifying as a technologist and a product engineer. I love building cool products and solving problems with technology; a servant leader, my role is to enable my teams and people to deliver the best products; and a dad, as I'm raising a seven-year-old boy to be a respectful human being.
Here are a few stats about me. I'm 70% Athenian, of course, Greek too, as I was born and raised in Athens, Greece. However, I have lived in London, UK for 30% of my life, approximately, and so I feel like a Londoner too by now. I have been a software engineer and a technology leader for more than 16 years, and I have worked with a variety of companies of all shapes and sizes in different roles.
I'm currently the VP of Engineering at Tide, and I'm leading the product teams, but I'm also accountable for the mobile and web engineering domains. I have been with Tide a bit more than five years, and I joined Tide very early. I was the 25th employee, and at that time, Tide was occupying just a desk at a co-working space in London. We are now more than 1,000 people distributed globally in many countries.
I was also the ninth engineer to found the mobile and web teams. Actually, there was no web team, and I was leading only iOS at the time. And we are now more than 80 people across mobile and web, and more than 300 across engineering.
But what about Tide? Tide is a business financial services platform, a B2B, and it was founded in late 2015. And throughout 2016, we were occupying a few desks at a co-working space in Farringdon, London. We were working also with various remote vendors in Bulgaria for Android, Ukraine for iOS, and South Africa for backend.
We launched the first product in January 2017 at the Apple Store, and soon after, we followed up with the Android release at the Google Play Store. In 2017, we closed also the Series A the following summer, and we started building teams internally, slowly scaling them and establishing this way our first proper tech center in London, UK.
In 2018, we entered hypergrowth mode, and we started hiring permanent people also in Sofia, Bulgaria. In 2019, we opened our second tech center globally, in Sofia, Bulgaria, and soon after we closed Series B and we started scaling rapidly. In 2020, just before the pandemic, we opened our third tech center, this time in Hyderabad, India, and we started hiring our first permanent employees there. And throughout the pandemic, we actually managed to build and scale the team across the whole tech organization.
In 2021, we officially entered the India market, and we closed also Series C, and we continued rapid scaling, which literally never stopped since 2017. 2022 finds us continuing our rapid scaling journey, but also we are expanding the teams and the product across the globe.
And why do we do all this? Tide is targeting micro, small, and medium businesses, and we're on a mission to help our members save time and money so they can focus on the things they love. And our vision is to become the world's leading business financial platform. We are currently operating in UK, we recently entered India, and we are planning to open to a few more countries in the next couple of years. In other words, we are scaling the business and the products, and we are expanding globally.
So our journey with feature toggles starts really at the early days of Tide, and back then, we were only 10 to 12 engineers, less than 30 people in total. At that time, releases were inconsistent. They didn't happen often. We were releasing whenever we had confidence and whenever we were ready, and this translated to multiple weeks, six to eight on average.
Our quality was really low. We had no real quality controls. Every release was followed up by usually multiple hotfixes, and there were a lot of bugs that were reaching our members in production. The small team we had back then wasn't very happy when we had to release because that was a very stressful process for everyone.
And I remember very well two incidents that played a key part in us realizing that we need to change our ways of delivering software. And the first one actually happened the day that I joined Tide, and it was a not-so-sunny April day in 2017, and we had just released the new iOS version in the Apple Store, which had a critical bug. And this was crashing the app upon startup, and everyone, including the CEO and the CTO, were panicking. And it took us a couple of days to fix it, test it, release it, and actually go back to business.
The second incident happened approximately three or four months later, and this is when we wanted to release a new product. And this product was implemented by a sole engineer, and this engineer worked in his own branch, a long-lived branch. And unfortunately, he had never pulled from develop in these three, four months. So the merge conflicts we had were impossible to resolve, obviously.
And this is when one of our senior engineers threw the idea about using feature toggles, and this was my response: "Do you mean merging incomplete code back in develop? Are you serious?" And of course, he was dead serious.
Then we talked a bit more. He elaborated more on the benefits. We talked about the risks, but also how to mitigate them, and then he finally convinced me. And in no time, we had implemented the first local feature toggle. By local, I mean it was stored on device, not on a remote server. And this way, we could hide very easily a simple feature by switching a toggle in the debug menu.
And soon after, we started improving them. We took them to the server, and we started utilizing across the org and the other teams. Within a few years, we had implemented basic feature toggle functionality, like targeted releases, releasing to a small cohort of selected members; percentage rollout, releasing only to small percent of our members; circuit breakers, being able to switch back and forth between providers when needed, still manually; and kill switches to turn off misbehaving new or old products, again, manually.
All was good. Everything was working fine, but access was restricted to just backend engineers, and this was making us very slow. Every change, even the tiniest one, had to be done by a backend engineer, and actually someone from the team that put the toggle in place. As a result, we had huge bottlenecks in delivery, and we were reacting very slowly to everything.
In addition, we had to resort to other solutions for our experimentation and A/B testing needs. So we ended up using Firebase for that. Nevertheless, what we had built, our in-house solution, helped us scale the product and the teams really well and relatively fast.
Now, fast-forward to today, things have improved dramatically. We have scaled fast. We've been in hypergrowth mode since 2018. We currently have more than 300 engineers, multiple business areas, and multiple teams distributed globally. We are now releasing consistently value to our members.
Our chosen release cadence on mobile and web is once a week, and we have a weekly release train. For web, we are actually moving this quarter to continuous deployment, enabling multiple daily releases, something, of course, that is not possible in mobile due to the app review process. For backend, we have already moved to continuous deployment a few years back. Their releases just need to be synced with front-end releases.
So everything was going well the first few years. But we weren't able to improve our internal tools and process. And at the same time, there was the need for the business and the product to scale and expand globally. So we knew that what we had built, our current tooling, was inadequate, and it wouldn't help us scale.
So we wanted to automate, as an example, certain manual processes like changing a flag based on external conditions or new data, like a member becoming eligible for a product based on a business rule. We also wanted to be able to integrate with other tools, services like Datadog, like as an example, triggering automatically a kill switch if a certain service is underperforming.
We wanted to enable non-engineers to make changes and speed up our process, remove the bottlenecks. As an example, a product manager toggling a product on and off or adjusting the member cohorts without actually the need to change any code or database.
We also wanted the ability to go back in time and investigate for errors, any changes, and any misbehaviors. As an example, investigate why a flag is not performing as expected or seeing all the changes that were made in a flag and when.
We also wanted to have better insights on our flags, better insights on our tests and the ability to export all this data, and having statistics about flag evaluations and, as an example, exporting them to Segment. And we wanted more sophisticated experimentation tooling and better data insights.
And at the end, it was a buy versus build decision. We assessed building what was needed, or what we thought was needed, ourselves. But the truth is that internal capacity was never enough, and it is never enough, as we already had aggressive product roadmaps, and we're already fighting for capacity. And in addition, we would have to maintain it, sustain it, but also improve it.
So we started looking outside. We started testing various tools. LaunchDarkly was one of them, and as it's leading the race, it came out the top. The features, the platform, and the service, they score really high in our assessment, and they could offer us everything we needed and more, as we also found out about things that we could do that we wouldn't have thought ourselves, or it would have taken us some time to think about them.
But of course, there was pushback from some teams and some engineers. Many of them thought they could do a better job than the LaunchDarkly team or any team. And some probably could, right? But building a feature toggle management business didn't quite fit with our strategy. After several discussions, the decision was final. We could build everything ourselves, but we prefer spending our time and effort in building products for our members. We just needed to buy this toolkit for our toolkit.
So how are we using LaunchDarkly at Tide? I have grouped everything in four themes. The first one is called building, deploying, and releasing. First, we have decoupling deployments from releases, where now we're deploying to production multiple times per day, but we're also releasing on demand. And engineers, they commit to a trunk, but these changes are hidden behind a feature toggle.
Then we have moved to testing in production. We do targeted releases where we release new products or changes to selected members. We do canary releases to small cohorts of members in order to get early feedback on issues with changes. And we also do percentage rollouts, where we release to a small percent of actual members, which then increases slowly to 100%.
The second theme is called operations. First, we have kill switches. We started by implementing toggles that allow us to kill a feature or a product if there's an issue. And this, of course, has improved dramatically our confidence when releasing new products. Now we can switch off any product that is actually causing issues to our members.
Migrations: we're utilizing toggles when we do migrations, either with third-party providers, tools, services, databases, refactoring code, or adding new code. And this allows us to avoid big bangs, of course, and have smooth migrations across our tech stack or across our various integrations with third-party providers.
Another one is bidirectional changes. And now we have the ability to pull flag changes from LaunchDarkly but also push changes to the clients by the client-side or server-side, as the case.
We also have configurations. We're using toggles for configuring the system based on various factors like location, member data. As an example, this could be making a product available to certain locations or only to certain eligible members based on some criteria.
Lastly, we now have the ability to audit our systems for flag changes and monitor the performance of these flags, which, of course, has enabled us to identify and fix issues super fast.
The third theme is called experimenting. If you remember, I said for A/B testing, we're using Firebase. Now we have consolidated everything to LaunchDarkly, all the different tools we're using to LaunchDarkly, and now we're using it for A/B testing different variations of a new product or an improvement of an existing one.
The second one is beta testing here. We now beta test new products for defects and issues. We provide early access to selected and pre-registered members so we can get early feedback.
The last theme is called enabling, and it's all about enabling non-engineers to be able to create new flags, make changes, configure, investigate, and basically being able to use the system within their agreed guardrails and guidelines. With this, we're aiming to remove bottlenecks, speed up development, and react faster to market changes or member issues.
And what we have started doing is delegating responsibility for basic things like toggling flags on and off and adding members to cohorts.
Of course, we have just started. There are many, many more things that we want to implement and develop. And a few of the things that we have in immediate roadmap are approvals. We want to define better process around changing flags in production and investigate more complex workflows like feature workflows.
Adding automatic kill switches and circuit breakers. Integration with Datadog and flag triggers. Configurations based on business rules. We want the different domains of Tide to be configurable and that configuration to be captured in LaunchDarkly.
And any business decisions then can be configuration points. As an example, legal changing the terms and conditions without contacting any engineer; marketing pushing different campaigns in different markets and segments; product changing the price of a product or making a product available for a specific market, or changing validation rules, as an example, allowing business types per market based on risk readiness; production support deciding on appropriate providers in order to improve reliability metrics, as an example, switching from one provider to another if the first provider is unstable.
Our journey with feature toggles started later than others in the industry, but quite early for Tide anyway, as it was in the early days of Tide. In the beginning, progress was very slow, and in some cases was very painful.
By moving to LaunchDarkly, we're able to improve and modernize our CI/CD and other processes faster than before, but we're also able to scale the product and in a secure and safe way. We are able to scale the teams as well and make them also more efficient. More importantly, we were able to focus on what we love doing best, which is helping our members save time and money by saving our own time.
So finally, I'd like to pick the community's brain on these two topics. The first one is optimizing workflows for development and scale. We want to get some insights on how other companies are operating at scale with feature toggles or any solutions they have. And the second one is actually very specific to mobile, and it's about DORA metrics for mobile application development. We are very, very interested in learning how other companies, mobile-first like us, are defining and measuring the DORA metrics within their mobile development life cycle.
For anyone who wants to get in touch with me, this is how you can reach me. This is my email, my Tide email, and of course, my Twitter is, you can reach me on Twitter, but also on LinkedIn. I'll be happy to discuss this topic, but any other topic, really.
Yeah. And this is the end of the presentation. I hope I've added a bit of value to your day, and I hope you enjoy the rest of the conference. Thank you for joining me today, and have a great day. Thank you.