Targeting 100 Million Users: Continuous Improvement at TODO1
TODO1 is a leading mobile and online banking platform serving 10 Million users and over 2 Billion transactions per year. They began their DevOps journey in order to support the business’ goal to grow their user base 10x in three years.
Ironically, huge initial success with DevOps (like an 83% reduction in deployment times) revealed unexpected downstream challenges and barriers.
In this talk, TODO1’s Director of Technology Architecture will share some insights to their successes, and plans for continuing to optimize their release process, including:
- How agility gains in one place revealed bottlenecks elsewhere–and how they tackled them.
- The benefits of a “pipeline-first” rather than “stack-first” approach to automation (horizontal vs. vertical).
- How automation naturally forced adoption of a “more lean” deployment processes.
- Successful approaches to getting other teams on board.
Juan Felipe is the Director of Technology Architecture at TODO1 Services. He leads the Architecture Team and helps drive the company-wide DevOps Automation initiative. With more than 18 years in software development for the Financial Industry, he joined TODO1 in 2007, having several roles as Developer, Technical Lead, Solutions Architect, Product Design Manager, and agile sponsor.
With a strong experience in Digital Transformation in the Latin American Banking Industry, Juan Felipe worked in the design and implementation of Digital Banking, in both Retail and Corporate Banking, as well as Private Banking segments.
Juan Felipe plays the guitar (tries at least), reads sci-fi and loves super heroes.
Chapters
Full transcript
The complete talk, organized by section.
Juan Felipe Cardona
Welcome. My name is Juan Felipe Cardona. I'm going to share our journey with continuous improvement at TODO1.
First of all, I want to thank Electric Cloud for inviting me to do this presentation. I also want to thank my team back at TODO1 and John, our continuous automation engineer, who is here in the audience as well. Thank you for being here.
TODO1 is a company focused on humanizing the relationships between people and their financial institutions. We are not a fintech, since we don't build tools to sell to banks, but we create solutions as a service. That's our business. We offer solutions as a service to banks so they can complete their digital channel strategy, and we are focused on banks in Latin America.
We have more than 18 years of experience in the market. We work with world-class tools, platforms, and the best-of-breed platforms we can offer our clients. We're focused on the adoption of our solutions by our clients. We promise our clients the best user experience possible. We recently acquired a company that is expert in user experience and does all the development in that area, and all the user journey work for us to apply in our solutions.
We are experts in integration. We integrate with different back ends, the cores of our clients, and we also integrate different business applications to provide the best solutions for our clients.
Our business model is risk-based. That means we provide our users with our solutions and work very closely with them so their success becomes also our success. We don't just deliver the platforms or the tools for them to operate. We become business partners to our bank clients.
We're located in Miami and in Colombia, in South America. We have offices in Central America and Mexico, and we have clients in Colombia, Central America, and the Caribbean. We are now opening new markets in Peru and Argentina.
The type of applications that I'm going to speak about in this presentation are basically online and mobile banking, both for the retail and business segments. We also deliver PFM, personal finance management solutions, and BFM, business finance management solutions.
The architecture of our applications, and this is to give you more context of what we were trying to automate in our automation journey, is basically Java-based. Almost 95% of our applications are Java-based.
We have a UI layer for the mobile and web applications, in which we have our web installers. We deploy standard Java web artifacts. We also deploy in that layer, or generate the installer for, the mobile applications, both Android and iOS. We also deploy configuration files in that layer.
In the next layer, the back-end layer, where we have our APIs and business logic, we basically deploy standard .ear files, those are Java installers, and configuration files.
In the database layer, we work with DDL, data definition scripts, and data modeling scripts. In our integration layer, in which we have basically an SOA appliance, we deploy scripts for changing. In our deployments, we basically put configuration files.
The business goal given to us in IT was to target 100 million users in between three to five years. We're currently serving 15 million users among all of our clients in South America, Central America, and the Caribbean. So we still have a long way to go, but we are working on it.
Two additional business requirements given to us were basically to increase our release cadence by four times. In general, in Latin America, payrolls are usually paid on the 15th and by the end of the month. Those dates are frozen for us. We cannot deploy into production on those days, and we also cannot deploy a couple of days before and a couple of days after those payroll payment dates. That gives us only a couple of weeks to do our deployments in our production environment.
Additionally, the business was asking us to have better service quality. Of course, with our previous manual process, we were prone to human error, and the service availability and the credibility of our clients were being affected. So that is, of course, another business requirement that was asked of us.
To achieve those business requirements, we needed to get deployment automation in place. That's the only solution we had. Additionally, we had to adjust our processes. We had been doing continuous integration for more than 10 years. When we started, we were doing manual builds. We went forward to do continuous integration, and our teams have been doing that for all that time.
Now that we started our DevOps journey almost one year ago, our teams had to learn to do things again in a different way, and that's one of the challenges we had as well.
Our initial plan to have deployment automation: we decided to go with several applications, and we decided as well to focus on one of the layers of our architecture. We decided to go with the back-end layer, because that's the one in which most of the bugs were found. That's the most difficult one to change in terms of solving the bugs, because those are the API and logic layers.
Our initial plan then was to allow the developers to deploy into their integration environment and to allow the QA team to deploy into their QA environment without the intervention of our operations team. In a second phase, we decided to go into the production environment to pull this first plan of deploying that only layer in the production environment.
Then we started to see that, with that initial plan, which initially looked good to us, there were some smells. Some things were not good. But we started very shy because this was our first time doing DevOps, so we didn't want to mess with the production environment from the beginning.
This was our situation before doing automation. It was basically the development team doing all manual deployments in all of our layers and components, and the operations team was in charge of both the QA and production environments, also manually.
With that initial plan, we allowed the development team to automatically deploy into the back-end layer. We let the QA team deploy into the back-end layer. This part was deployed by the QA team, but the rest of the layers and the production environment were still being deployed by the operations team.
Even with that first approach, we had very good results, but we also realized some challenges, and we were not delivering what the business was asking from us.
For achieving this first part, we needed company-wide commitment. That means we had to break the silos. We had all our teams working together with our DevOps team to put that plan up. That was the first time we had the developers talking with the operations team and with the QA team, and we were working all together to achieve the same goal.
Additionally, management sponsorship was very important. These kinds of initiatives, because they are across the whole company, cannot be achieved by only one team pulling and doing all the effort there. Management sponsorship was very important for this initiative as well.
We had, of course, and we have seen in some of the other conferences, compliance. We still had to be compliant with our compliance, auditing, and security teams. So we relied on our automation tool, ElectricFlow, to provide us with the information we needed in terms of user auditing.
The reporting capabilities of the tool gave us the chance to know who did what and when, that's the user auditing, and to know what we changed between releases. The execution tracking also means that we are able to know exactly, or the auditors during the audits are able to know, what we changed, what artifacts we changed, what components were changed, and go exactly to the level they need to check for any change in our environments.
Additionally, we had to work with the infrastructure team to do some changes in our network segments. From the beginning, we have always had a separation between the development or integration network segment, the QA network segment, and the production network segment.
Now we needed to somehow send signals from the integration environment to the QA environment for the pipeline to continue its deployments. But we needed to be compliant with the condition that those segments were completely separated. So we worked with our infrastructure team and added an additional segment that is transverse to all other environments.
From that new network segment, we were able to connect and put all the security requirements that security and the compliance team were asking from us, to be able to push or handle the deployments in each of the separate environments.
We used a feature of the automation tool called entry gate rules, which allows us to put stops in the pipeline. When a developer starts a deployment, the tool stops, and only an authorized person from the next environment is able to do the review and approval for deploying in the next environment.
With that first phase, we got very good results. Initially, in our integration environment, we were deploying in 45 minutes, and we went from 45 to 12 minutes in two of our applications. With the new automation, we were able to have savings of 40 to 73 percent in time. With the number of deployments we have per month, that goes almost to 150 hours of savings that the developers were not spending doing deployments. Those hours are basically almost one developer working full-time during a month. So he should probably be creating more functionalities and adding more value to the business.
In our QA environment, we had even better results. We went from 120 minutes, two hours, to 20 minutes per deployment. That's an 83 percent savings of time that now the operations team is not spending doing manual processes. With the number of deployments that we were doing per month, or that we are doing per month, that's more than 400 hours that we are saving the operations team.
Q&A
Audience: Could you elaborate on that? Is that manual testing? Is that automated testing?
Juan Felipe Cardona: Say again, please?
Audience: Is that automated testing or manual testing?
Juan Felipe Cardona: No. I'm talking here only about the deployment. This is not testing.
Audience: Just to the QA environment?
Juan Felipe Cardona: Yes. This is deployment in the QA environments, not the testing. Those two hours also represent the queues that we were adding to the process for the operations team. Now that we took 100 minutes per deployment from the operations team, that's basically three operations team members who should be spending their time doing either more deployments or improving their monitoring in their environments.
Audience: Can you go back?
Juan Felipe Cardona: Sure.
Audience: You're telling us the what, but how did you achieve this?
Juan Felipe Cardona: For this, we had to work with processes. Let me go back to this one. We optimized our deployment process. We automated it, and we're only talking here about one of the layers of the architecture. So first part, automating. Second part, we made our processes leaner. We created less documentation. We relied on the capabilities of the tool to give the information that our manual processes were generating.
In our previous manual process, we were generating a lot of manual documentation that we had to deliver to the next team to operate. So we got rid of some of that documentation, and we started generating some of that documentation as evidence because the auditors were still requiring that information. Some of the reporting that we generated manually, for example, evidence of the deployments and evidence of restarting a server, is now being given automatically by the tool in the reporting capabilities. So basically two points there: automation of the deployment and leaner processes.
Juan Felipe Cardona
With these good results, we also generated a bottleneck in the QA environment, because now the developers were deploying faster in their environment and were delivering faster to QA. But the operations team, even with the savings that I showed, that should be almost three people doing other things, was not able to keep up with the new speed that the developers were having in their deployments and the speed that the QA team was having in that back-end layer. The QA team was still depending on the operations team for the deployment of the other layers.
So we went from automating a single layer in multiple applications to automating all layers in a single application. That's the change of strategy we had to do, and that's the change of plan. That's what we changed from our original plan. I'll go back again: we went from automating a single layer in multiple applications to automating all layers in a single application.
The application we chose was decided with the operations and business teams, and was one of the applications that the business was requiring us to deliver faster and to increase the cadence of the deployments in production.
The goal then was to have all the environments for that only application, all the layers, and to give the teams the ability to deploy with only one click. The developers were deploying in all their layers with only one click of the tool. The QA team is now deploying in all their layers with only one click.
In this first phase, we finished that part, and we are working and going toward pushing that into the production environment. In about two to three weeks, we're planning to do our first fully automated deployment in production. That's something that we're working on right now with our coworkers back in TODO1.
The second challenge we found, that we had to deal with, was the resistance of the teams with that change. Adjusting to this new process, to these new tools, required us to give them coaching, generate documentation for the new members of the development teams and QA teams that were coming, create demos, and create support.
The support increased because a lot of teams wanted to start doing automation once they saw the other teams deploying automatically. They started to come to us and ask us to include them. But we already had our plan, so the business was the one giving our priorities at that moment. The good thing is that once the teams started, they really loved it. It's the same phenomenon that happened 10 years ago when we started doing continuous integration.
Two important lessons we learned in our process: since we are not in the business of creating DevOps or automation platforms, we learned that we need a unified platform that we could build on and that provided us with enterprise-grade features like multi-tenancy and high availability.
The other important lesson is that we underestimated the number of nodes we needed to automate at the beginning. In our first plan, we only looked at a couple of applications and one layer. When we changed our strategy to focus on one application and all the layers, the number of nodes we needed to automate increased. So that was something we had to deal with in terms of budgeting as well. My advice for everyone is: plan and budget for the best from the beginning.
Our next steps in our process, in TODO1, in our continuous improvement, are to keep working toward the goal of serving 100 million users in three to five years. We want to implement deployment automation into all of our applications. We want to increase the integration with functional testing suites that we're currently working on.
We also want to go forward with building and delivering the mobile applications, in terms of delivering them to the testing markets or even delivering them to our final clients for them to deploy into their production environments in the markets by themselves.
Finally, in our next steps, we want to go forward with our database automation and the DML scripts, so we can have both automation of data manipulation and data definition. That's our other goal.
Thank you very much. That's all I have to share with you. If you have any questions, I'll be happy.
Q&A
Audience: Can you talk a little bit more about DDL automation?
Juan Felipe Cardona: Sure. In our current state, we are deploying only scripts for data manipulation. We learned that we still have a long journey to go there because we have mega-scripts. One of the things that we realized, and that we learned reading the documentation and reading other experiences, is that we have to have smaller pieces for our scripts. That way you can put those scripts in the automation pipeline and have them deployed and installed in the database without any problem. If you have smaller pieces, it's going to be easier even for doing rollbacks if you have to. Did I answer?
Audience: How do you validate DDL to ensure you're not going to destroy the environment before executing that DDL statement automatically in the environment?
Juan Felipe Cardona: How do we validate that information? The information that we are deploying there is basically configuration. It's not transaction information. It's not information about the transaction of our solutions. They are basically configuration files or configurations that work for our applications that are stored in the database. So that's the part that we are automating.
Audience: Do you have cases where you need to handle rollback if one of those layers fails?
Juan Felipe Cardona: Oh, sorry. We got the time out. We can speak here. Thank you very much.