Categorieën
Tech

From manually clicking our infrastructure together to infrastructure as code

Even with a small development team, we’ve always prioritised making our infrastructure scalable and secure. But as our product developed further, our infrastructure grew and became harder to maintain. Because we build all of our technology in-house, this complexity was compounded by the need to manage and scale multiple components simultaneously. We managed everything manually through the Azure portal, which required repetitive and time-consuming work. Setting up a new application meant deploying a variety of resources, and replicating these changes across multiple environments meant having to click through the same steps over and over.

We also had no way to track our infrastructure over time. If someone changed something, there was no way to track what had been modified or how it was configured before. We could only rely on our memory, which isn’t the most reliable after a few weeks, or even a few days.

We realised we needed a better solution, so we decided to migrate to Infrastructure as Code and chose Terraform for this.

Getting started

Because our development team at the start of this journey was so small, and none of us had any experience in this area, we decided to hire a consultancy company to help us with this migration. They used Microsoft’s Cloud Adoption Framework for our infrastructure setup, which is a comprehensive guide for setting up cloud infrastructure according to best practices. This approach provided clarity on areas we’d previously been uncertain about, particularly around network setup. They also created some initial Terraform modules to give us a head start.

Initially, we believed this collaboration would cover the majority of the work. In hindsight, the expertise of the consultancy company helped us adopt best practices early on, but as we got deeper into the migration, it became clear that the project was far more complex than we had anticipated. Our original time estimate for the project was one to two months—a timeline we now realise was overly optimistic, given the scope and intricacy of the migration.

Balancing short and long-term benefits

After the consultancy company finished the initial setup, it was time for our team to take over. Given our small size, we didn’t (and couldn’t) dedicate all of our development resources to this migration. Instead, one development team member took over expanding the infrastructure initially. We built and expanded a lot of modules, which helped us set up similar infrastructure more quickly and consistently later on.

As time went on, we struggled to balance the long-term benefits of this migration, such as scalability and easier management of our infrastructure, with the short-term demands of product development, such as customer satisfaction and feature delivery. We found ourselves going back and forth on whether to dedicate more time and effort to this migration or focus more on product features. The more resources we could dedicate to this migration, the quicker it would be completed, but that would also mean slowing down product development. This is a typical struggle for a product company in the early phase—finding the right balance between investing in scalable and secure solutions versus meeting immediate and innovative customer and product demands. For smaller companies, like ours, these trade-offs are even more pronounced, as resources are often limited.

Completing the migration

After months of this balancing act, we realised we needed to focus on getting this migration done. We set a deadline and decided to dedicate the majority of our development resources towards finishing the migration. First, we successfully set up the infrastructure for our test environment and wrote scripts to migrate our data. After completing the data migration for our test environment, we began running integration tests on the new infrastructure. Although we encountered some issues, by this point we were feeling confident in our Terraform skills and managed to resolve them.

With our test environment fully set up and validated, we could start setting up the infrastructure for our remaining environments. Having one environment fully up and running gave us the experience to set up the remaining environments much more smoothly.

After finishing the setup for our remaining environments and testing everything thoroughly, it was time for the final step: migrating our production environment. To minimise impact on our end users, we completed the migration at midnight. This was the first time the development team came together after hours. We started by migrating our data using the scripts we had written before, and once that was done, it was a matter of updating our DNS settings to point to our new infrastructure.

The migration was a success and went smoothly, with minimal impact on our end users. After a year, we finally had a new infrastructure in place!

What did we learn?

Now that the migration is complete and a few months have passed, we can look back objectively on what went well and what could’ve gone better. During the final sprint, we really came together as a team and realised just how much we can accomplish when we focus all of our efforts on a single goal.

Time required to migrate a full infrastructure to Infrastructure as Code

Looking back, we can admit that our initial time estimate for the migration was way off. We thought it would take about a month or two, but it ended up taking almost a year! The learning curve at the beginning was steep, and the process was a lot of work.

This was partly caused by our infrastructure already being quite large when we started this migration, which meant that we had to migrate a lot. It’s difficult to balance investing in projects like these and making sure they don’t come at the expense of innovation and creating the best possible product for the end user. But as your product develops, so will your infrastructure, and the earlier you start your migration, the less work it will ultimately be.

Terraform’s plan versus apply

We also learned that a successful plan, which shows a preview of the changes Terraform will make, doesn’t always guarantee a successful apply, which is when these changes are actually applied. Not all of Azure’s policies are validated in the plan stage, which can lead to unexpected errors when the changes are applied. This can be frustrating, and is something we still struggle with now.

Deploying a new version of an application

How Terraform handles updates, including small ones, also poses some challenges. When you update something as simple as a Docker image tag to deploy a new version of an application, Terraform still checks the entire infrastructure to determine which parts of the configuration have changed. This significantly slows down deploying applications, so we’re working on a solution to manage this outside of Terraform to address this, and deploy applications through CI/CD pipelines instead.

Best practices

Since completing the migration, more people in the development team have gained a better understanding of how to manage the infrastructure. Every engineer is end-to-end responsible for their own applications. With Terraform, everyone can now confidently deploy an application that follows best practices and meets security standard—something we think is really important.

Ultimately, migrating to Infrastructure as Code has made our infrastructure more consistent, scalable, secure, and easier to manage.