For the better part of the last two decades, the move towards utilizing public cloud infrastructure seemed like an inevitable, one-way tidal wave. Cost to end users would fall as providers continue to scale and an array of ever more fine-grained services would allow startups to stay lean and quickly adapt to surges or disruptions in demand.
During my three years working on the Stack Overflow podcast and blog, I've gotten the chance to talk with lots of interesting folks working with a wealth of microservices and containers. I’ve seen the push to build infrastructure as code and the appeal of going serverless. At the same time, I've chatted with lots of folks in the areas of observability and service meshes, who have found a new business supporting the sprawl of interconnections that exists in modern applications.

Recently, however, I've noticed a new trend growing in parallel. Yes, adoption of public cloud continues to grow, with many companies still in the process of deciding what to migrate off local servers. Tons of folks gathereveryday to share knowledge about AWS, Azure, and Google Cloud across our Stack Overflow Collectives.
At the same time, however, a growing number of organizations are also carving out space to repatriate work from public providers to on-prem private clouds and, for the growing world of edge computing and machine learning, going back to the future of actually owning and operating their physical hardware on-site.
A cloud to call your own
According to a 2022 report by Bessemer Ventures, there has been a significant uptick in the adoption of virtual private clouds. The report suggests that, “it is becoming easier to package SaaS products and deploy them inside a customer’s virtual private cloud (VPC). This is due in part to the standardization around Kubernetes as the operating system of the cloud. This makes it easier for SaaS companies to serve a wider range of customers that may prefer to keep certain sensitive data or applications in a VPC.”
A lot of our customers using Stack Overflow for Teams Enterprise edition opt for this approach. We build and upgrade the platform for knowledge sharing and collaboration, but it’s installed in a private or on-prem location and the conversations about their proprietary code clients are discussing stay privates on-prem.
Tom Limoncelli, a technical product manager on our site reliability team, has strong opinions on this trend. “Here's what's really happening,” he wrote to me:
a. Running your own datacenter encourages bad practices due to lack of governance and other reasons.
b. The cloud forces/encourages better practices, such as strict governance, infrastructure as code, and fully automated CD/CD pipelines.
c. People are building on-prem clouds, which emulate those best practices because they got a taste of in the cloud.
Another way to say that is... People aren't returning to the datacenter; they're returning to on-prem clouds because datacenters can be frustrating.
The way Limoncelli sees it, a fully DIY approach encourages bad practices. Owners tend to treat servers more like pets than cattle, adding their own customization. After decades of that, you get a datacenter that is just one big mess of bad configuration ideas, mismatched technologies, and political barriers that prevent any of that from being fixed. The smart way to move to private cloud is to be strict about standardization. In other words, no, you can't request a special machine with a weird ethernet connection because one person thinks it’s cool.
Cloud systems, on the other hand, provide an API to request new virtual machines in minutes, instead of a manual purchase process that took months. They install racks and racks of the same hardware configuration. They use standardized machine configs, guiding users away from bespoke configs. Costs are accounted for, which often doesn't happen in datacenters. On the software side, the move to the cloud is an opportunity to adopt automation like a CI/CD pipeline, weaning people off of manual deployments.
Users acquire resources via an API, not by a purchase order. Governance and automation are established from the start. The combination of accessing resources by a standardized API and automatically enforced governance results in a system that is more maintainable and enforces more modern practices.
Containers will set you free
As Nick Chase argues over at The New Stack, Kubernetes has been a powerful enabler for companies looking to gain more control over their usage of the cloud. It’s relatively agnostic about the real or virtual hardware you choose to employ because it can do a wealth of different things with the Linux kernel as its foundation.
Kubernetes was designed to make it simpler for a small group of people to manage a large constellation of applications by abstracting the underlying hardware. Barring a setup that confines your system to a scarce resource, it offers a level of resilience that customers ten years ago turned to public cloud providers for. Along with allowing users to update systems without going offline, it also offers capabilities for monitoring of services up and downstream, something that many microservice heavy organizations are now turning to third-party observability providers for.
“Kubernetes’ superpowers change the game in important ways,” writes Chase. “If you can deploy, scale and manage the lifecycle of Kubernetes, you can use it to pave over public and private cloud infrastructures, optimize costs and overheads aggressively, and treat everything underneath Kubernetes as a commodity.”
As my colleague Ryan Donovan pointed out during a recent conversation, “being able to abstract infrastructure has enabled a lot of cloud providers, but it's also allowed folks to have those containers located anywhere — within a public cloud, on a prem server in Lithuania, or a private cloud replicated across multiple locations.” Just because your infrastructure has moved to the cloud doesn’t mean you don’t care about proximity to users and the advantages that might provide you in terms of cost or latency.
Stack Overflow has taken advantage of some of these superpowers. Max Horstmann, formerly a Staff Software Engineer at Stack Overflow, now a principal software engineer on the Azure Kubernetes Service (AKS), wrote in depth about why Kubernetes might be a good choice and how we took advantage of it inside our organization. You read his article on it or listen to his podcast below.
“If you’re starting a new project from scratch — a new app, service, or website — your main concern usually isn’t how to operate it at web scale with high availability,” writes Horstmann. “Hence, when it comes to choosing the right set of technologies, Kubernetes — commonly associated with large, distributed systems — might not be on your radar right now. After all, it comes with a significant amount of overhead.”
Despite all this, he sees value in adopting it from the start. “When you're launching something new, your focus is typically to move fast and iterate quickly based on early feedback. Scaling is something for later. K8S is a tool that, in my view, allows you to do just that because it can accelerate your build/test/deploy loop, allows you to easily deploy and instrument different instances of your app, e.g. for split testing, customer demos etc."
If you’re lucky enough to find product market fit and start to see a surge in customer demand, Kubernetes proves valuable in this area as well. “The problems that come with scale — fault tolerance, load balancing, traffic shaping — are already handled," says Horstmann. "At no point will you hit that moment of being overwhelmed with success; you future-proofed your app without too much extra effort.”
This comment from a HashiCorp’s forum sums up the advantages well: “A Kubernetes cluster is a good example of an abstraction over compute resources: there are many hosted and self-managed implementations of it on different platforms, all of which offer a common API and common set of capabilities.”
A bridge between public and private clouds
The Bessemer report cites another emerging technology trend that pairs increased cloud adoption with on-prem data. “Emerging middleware platforms are making it easier to bring the power of the cloud to the data, wherever it may be. This has played out in industries like financial services, where a wave of modern fintech infrastructure helped build bridges between the cloud and legacy banking systems. We are seeing similar bridges being built in other large industries like supply chain, logistics, and healthcare to bring the power of the cloud to these on-premise data sources.”
It’s important to define what we mean by “middleware” here. As Red Hat points out, the term dates back to a 1968 NATO conference on software engineering, where it referred to code that sat between the assembler/compiler at the bottom of the pyramid and the application logic at the top. In the world of hybrid cloud, middleware refers to an evolved version of this same idea. As Asanka Abeysiinghe, Chief Tech Evangelist at WSO2 explains in a blog, this can look like, “mega clouds that provide infrastructure as a service (IaaS)-enabled middleware capabilities via APIs, which have become the new DLLs. So, for example, message queues, storage and security policies are open for developers to consume in applications running on the IaaS (Infrastructure-as-a-Service).”
Outside the big public cloud providers, Abeysinghe sees other alternatives catching on. “Kubernetes addresses the issue of cloud lock-in by bringing an open standard to the cloud-native world, and it enables basic middleware capabilities as components. In addition, the Cloud Native Computing Foundation (CNCF) brings a rich set of Kubernetes-centric middleware, and you can find them in the CNCF technology landscape. However, if the middleware capabilities provided by Kubernetes and the CNCF are not enough for your application development, you can add custom resources by defining them in a custom resource definition (CRD) because Kubernetes is built using open standards.”
When I spoke with Abeysinghe for this article, he was quick to point out that there was no data to indicate a trend of companies moving fully away from the cloud, far from it. There are still more folks migrating onto the public cloud than off it. He estimates that 80 percent of activity is still focused on the traditional shift from local to public cloud, with another 20 percent moving in the opposite direction. But that 20 percent is important, precisely because it flows against the prevailing tide we’ve seen over the last decade.
Abeysinghe believes that there is a realization, especially at organizations with a lot of legacy hardware infrastructure, that they now have a lot of machinery sitting idle. If you’re a big bank with decades of mainframes at your disposal, utilizing only five percent of that doesn’t make much sense. “Kubernetes lets you run a private cloud that better utilizes your existing on-prem hardware.” Cloud bursting technology lets you shift to third party resources when your local hardware is maxing out.
Not to be left out of the game, public cloud providers now offer physical server racks to clients who have jobs that are more efficient on-prem, or need to remain in-house for security and compliance reasons. Companies that once helped to migrate companies off local hardware now offer server-racks-as-a-service bundled with your public cloud offering, a truly full circle moment for the evolution of compute.
Bringing AI models in-house
One area where this trend seems particularly strong is among companies focused on artificial intelligence that work with large data sets and have created their own models. “Big cloud GPU compute is very expensive, whether it's for training or for inference,” says Dylan Fox, founder and CEO at Assembly AI, a startup that provides AI-as-a-service to companies that are seeking natural language capabilities in their offerings but don’t want to build the models or hire a team in-house.
“We do most of our training in on-prem instances. We have a couple hundred A100 NVIDIA cards, and we recently just purchased like a couple hundred more that we have for on-prem instances used to train." The crypto winter has been a blessing for this market, as a glut of GPUs has come onto the secondary market and prices for new and used hardware have fallen.
As David Linthcium wrote over at InfoWorld:
Companies are looking at other, more cost-effective options, including managed service providers and co-location providers (colos), or even moving those systems to the old server room down the hall. This last group is returning to “owned platforms” largely for two reasons.
First, the cost of traditional compute and storage equipment has fallen a great deal in the past five years or so. If you’ve never used anything but cloud-based systems, let me explain. We used to go into rooms called datacenters where we could physically touch our computing equipment — equipment that we had to purchase outright before we could use it. I’m only half kidding.
When it comes down to renting versus buying, many are finding that traditional approaches, including the burden of maintaining your own hardware and software, are actually much cheaper than the ever-increasing cloud bills.
Second, many are experiencing some latency with cloud. The slowdowns happen because most enterprises consume cloud-based systems over the open internet, and the multi-tenancy model means that you’re sharing processors and storage systems with many others at the same time. Occasional latency can translate into many thousands of dollars of lost revenue a year, depending on what you’re doing with your specific cloud-based AI/ML system in the cloud.
It’s not just small AI startups that want to crunch a lot of data at a low latency with homegrown models. Here’s an eye-opening quote from Protocol. “The on-prem trend is growing among big box and grocery retailers that need to feed product, distribution, and store-specific data into large machine learning models for inventory predictions, said Vijay Raghavendra, chief technology officer at SymphonyAI, which works with grocery chain Albertsons.”
Raghavendra left Walmart in 2020 after seven years with the company in senior engineering and merchant technology roles. “This happened after my time at Walmart. They went from having everything on-prem, to everything in the cloud when I was there. And now I think there's more of an equilibrium where they are now investing again in their hybrid infrastructure — on-prem infrastructure combined with the cloud,” Raghavendra told Protocol. “If you have the capability, it may make sense to stand up your own [co-location data center] and run those workloads in your own colo, because the costs of running it in the cloud does get quite expensive at certain scale.”
Chick-fil-A had a similar experience. In a blog written by Brian Chambers, the company’s head of Enterprise Architecture, he noted that, “In researching tools and components for the platform, we quickly discovered existing offerings were targeted towards cloud or data center deployments. Components were not designed to operate in resource constrained environments, without dependable internet connections, or to scale to thousands of active Kubernetes clusters. Even commercial tools that worked at scale did not have licensing models that worked beyond a few hundred clusters. As a result, we decided to build and host many of the components ourselves.”
Their solution allowed a DevOps Team and Smart Device Support to deploy, build, and update to thousands of restaurants.
Cloud, with control
Total spending on cloud computing is already enormous and still projected to grow over 20% this year, closing in on a half a trillion dollars. But it will be a far more varied and nuanced period of growth. “Cloud is the powerhouse that drives today’s digital organizations,” said Sid Nag, research vice president at Gartner. “CIOs are beyond the era of irrational exuberance of procuring cloud services and are being thoughtful in their choice of public cloud providers to drive specific, desired business and technology outcomes in their digital transformation journey.”
After a decade or more spent moving away from server racks, companies are finding there can be advantages to running local infrastructure for certain kinds of compute. There is also, perhaps, a generational shift at work. The engineers who cut their teeth building big public clouds inside large tech companies see now moving on to create startups or take senior roles at smaller companies that specialize in a subset of cloud offerings. What’s old is new again, but with a vast variety of new flavors and permutations to choose from.
TRANSCRIPT
Max Horstmann If you go back to Twitter and Hacker News over the last couple of years, you'll find a lot of people making like these snarky comments about like, you know, we know Kubernetes is hot, but why are you messing with it now? You don't need it right now, when you're at this early growth stage where really your focus should be on iterating and improving a product. You know, you can worry about the scale part later.
Ryan Donovan Yeah, I think they call it resume based development.
Yeah, exactly. Oh, that's like the same as blog posts, do things that don't scale. So that is absolutely true. I would argue, though, that setting up your cluster really isn't that hard anymore, because your cloud provider makes it easy. And you can even manage it as infrastructure as code. And it'll offer you another benefit next to the possibility to scale, it'll give you that ability to build test environments, and you know, iterate way more quickly.
[intro music]
Ben Popper Cockroach DB is the only book you'll ever love. Because it's the only one you don't have to worry about. As a low touch SQL database that automatically handles scale, operations, and uptime. Cockroach DB lets you focus on developing, get your free cluster and a free t-shirt at cockroachlabs.com/StackOverflow.
BP Hello everybody! Welcome to the Stack Overflow Podcast, a place to talk about software and technology. I am Ben Popper, Director of Content here at Stack Overflow. And today I'm joined by my colleague, Ryan Donovan, who is a content marketer on my team, running the blog and the newsletter. Hi, Ryan.
RD Hi, Ben, how you doing?
BP I'm pretty good. So one thing that we've written about a few times, and which I think we're going to discuss today, is the idea of using, you know, Kubernetes and containerization, to rethink how your company is architected. We've done a few like sponsored blog posts about this from big COEs and had pictures on it. But you know, it feels like overall, this is a big shift in the industry that's been playing out over the last few years.
RD Yeah, absolutely.
BP And, you know, a technology that is just exploding in popularity. What do you think is driving that?
RD Well, I think a lot of software folks don't want to manage their hardware. When a lot of people, a lot of organizations move to the cloud, it was just easier to handle resources. And now, right, you have Kubernetes, which you can treat the resource as code.
BP So this is part of the bigger sort of like infrastructure as code movement.
RD Yeah, I think so.
BP Well, we have a in house expert coming on the podcast today to discuss it. Max Horseman, who is a staff software engineer at Stack Overflow, and has recently been working on some really cool stuff in house related to exactly this topic. Welcome, Max.
MH Hey, Ben. Hey, Ryan. Thanks for having me.
BP Yeah, of course. Great to have you on. So for folks who don't know, you've been with Stack Overflow for I think, more than eight years you said?
MH Yeah, that's right. I joined back in 2012. So it's been eight and a half years now. Yeah.
BP Very cool. And you were at Microsoft before that?
MH That's right, spend a couple of years at Microsoft then took the year off for traveling, which was great. And then I joined this relatively young company back then called Stack Overflow.
BP We've all aged so gracefully together. So Max, you did a piece for us that I want you to talk about, maybe, you know, step back a minute and sort of tell people like, what your thesis is here and how that relates to some of the work you've been doing inside of Stack Overflow, building, essentially, what what is like new tooling that all of us get to use?
MH Absolutely. So the first couple of years here at stack, I worked on the Talent and Jobs team. So for those who don't know, if you're on Stack Overflow, there's also a job board. And if you're looking at a question or answer page on Stack, you'll, you'll sometimes see job listings advertised and you can go to the job board and find a job. And if you're an employer, and you're trying to hire developers, there's also a portal for that, and Stack Overflow Talent. So that's what I've done for many years. And when I when I joined the company, this was really at an early stage, and we needed to grow the product a lot and iterate a lot. So I mean, it's well established now that what you want to do at an early stage is you know, lean startup and all that sort of things, that you want to iterate and seek feedback early on, and then you know, make improvements based on that feedback.
RD Get that market fit.
MH Yeah. So getting to product market fit. And then based on that, grow the product and scale it and yeah, so at some point, we realized that really the tooling we have for that needs some improvement. And yeah, that's I'm here to talk about today.
BP Yeah, I feel like doesn't didn't one of our co founders have like a famous quote, somewhere along these lines. Maybe it was Jeff Atwood, from that sacrificial architecture piece. Right, Ryan? Am I thinking about this?
RD That performance is a feature. Basically.
BP Yes, performance is a feature.
RD You can build it in at the beginning, or you can add it in later. But it's, you should treat it as a feature.
BP So I thought that was kind of cool. Jeff would say performance as a feature. Many developers understood this to mean performance is the first thing to care about, but that's not quite right. I like the way that was put. So, Max, you were working on Talent. And I guess, you know, yeah, just to clarify, we should say, you know, the business of Stack Overflow has been evolving. We're kind of we talked about this publicly. It's no secret moving away from job slots, and we're gonna, you know, you started the scale and reach we have with developers to allow companies to do sort of awareness about their products or services or open roles they have for developers. Talk about their brand, and then focus on teams. So jobs, that's something you did in the past. Not anymore. But yeah, you have a different sort of take on Kubernetes, which often people talk about as being a little too complex or unwieldy or expensive for a startup to consider. It's something like we said people talk about is maybe a big re-architecture decision when you've scaled to a certain point. And you know, you feel like pain points or areas of friction or, or too much overhead is developing. But you feel like maybe that's no longer the case. And then I think you have kind of an interesting implementation of it in house, right?
MH Yeah, exactly. So. So to your first point, you're right. Kubernetes is often and still considered to be a relatively complex technology, right? Something that requires a lot of resources and overhead to create and to manage and running your own Kubernetes cluster is certainly something that is a complex task. But you know, in 2021, now, I think it's fair to say that with all these cloud offerings out there, that you know, take care of that for you, you know, offering Kubernetes, basically, as a utility or like a resource, I don't think that is any longer true that you would need to avoid Kubernetes, because of its complexity, in fact, it's gotten very, very straightforward to set up, it takes you a couple of minutes to set up your own cluster. If you you know, you mentioned infrastructures code earlier on, if you're a believer in infrastructure as code, you can use tools like terraform, to set up your cluster and basically manage it as code in your version control system.
BP So let me step back for a minute. And I'm sure Ryan will understand this better than I so I'll let him jump in. But we're talking about like, sort of the multiple layers of abstraction. This is just the natural evolution. So first, we say you know, you've got your own server, you've got your own hardware, then we say, okay, now you're going to use a Rackspace somewhere, no, you're just going to spin up, you know, like an AWS instance. And now people are saying, well, it's actually easier and better to use this, you know, containerization technology. But okay, even that's too tough. Like, well let some big company background spin that up, and we'll use their, you know, cluster. So you're like, three, four or five steps away now from setting up your own Raspberry Pi at home when it's time to do this stuff.
RD Yeah, I mean, it seems like from my understanding, Kubernetes is basically reducing it all to YAML files. Right? That just makes it another simpler thing. And if it's all attached to these cloud providers, like, why not, why not get it started early?
BP I mean, I guess, Max, like maybe not in layman's terms, but as you know, in sort of a summary way, like, what does it take? Yeah, what's involved in setting up and running a cluster?
MH Right, yeah. So just setting up a cluster can be done basically, in two ways. And we're talking about like a cloud hosted cluster, right, setting up your own cluster. And by the way, you can use your own Raspberry Pi hardware, or anything else at home, people have done that. There's stuff on Reddit, where people set up their own home Kubernetes cluster on Raspberry Pi's. But that's not we're talking about here today. Setting up your cluster in in a cloud hosted environment, you know, you can either just go to your favorite cloud providers portal and set it up through some sort of UI, or you can use like I mentioned to infrastructures code provider, like terraform. And write something, it's not quite YAML. That's quite similar, right? and define your cluster and then almost set it set up a deployment pipeline for, for your cluster and for infrastructure.
BP But I meant, step back one more level, like for people who don't know, and because we have lots of engineers, listen that but also lots of people who are software adjacent, like what do you get out of, you know, a Kubernetes cluster? What is involved in that? And what does it help you do?
MH Exactly, yeah. So I just want to give you an example, and talk about a project we've worked on here at Stack, because I think that's a really nice illustration, like, like people have said, you know, people think of Kubernetes very often as, like an infrastructure piece and something that allow you to scale and you know, that does some of the heavy lifting around fault tolerance, and all that sort of load balancing and that sort of stuff, right. But it can also be used for quite a number of different things. And one specific thing we did here at Stack was we're using Kubernetes, to host what we call PR environments, I should explain that a little bit. So this is something a lot of organizations and companies have been doing over the last couple of years. The idea is, when you're making a code change, you're typically using a version control system like Git, and or maybe GitHub and, and then typically, at some point, you're creating a pull request or PR, we all know that right? Where you're sharing your working progress with some of your peers maybe and seeking some comment on that. Now the problem is looking at the code will not tell you the whole story. And especially if you want to share that with someone maybe outside of the engineering team, let's say someone like you two in marketing, or maybe somebody in sales, or maybe in somebody and let's say the legal department, who knows, right? It would be very nice if you can just show them your code running. And ideally, you would just send them a link and say, Look, I have this work in progress here. Here's the link, can you just click on the link and then you'll see what I'm doing here and just try it out. Basically, this is like a completely isolated and custom test environment, just for those code changes you're working on.
RD I think this is a completely wild you know, advancement to the industry. Like the fact that you can see spin up this automatically, you don't have to download all the requirements to your local machine, you don't have to, you know, take over whatever the dedicated test environment is, you just have this one thing that's dedicated to the PR itself.
BP Ryan, you were mentioning previous jobs, you knew people who did this and half of their job was just going through that setup process.
RD Yeah, I mean, especially if you have a service oriented architecture, you know, you have to spin up each of those services, possibly put those in a Docker container. And it's just kind of a nightmare, you got to make sure other requirements are there. And if one of those requirements is broken, you're debugging a thing that shouldn't even be debugged.
BP Yeah, it's super interesting, you know, and especially right as, as more and more of the products lives, not, you know, on a floppy disk or CD ROM that you're going to ship people or you know, as an on prem delivery, but as something that is live on the web, you might very well want design, or marketing, if they're doing copy or product marketing, you know, to be able to look at these things and push the changes, you know, on a regular basis. So to me, it makes a lot of sense. And yeah, it's exciting for people like me, who are not software developers themselves, but have to work with you loons every day. I guess, tell us a little bit about sort of, yeah, like when you set out to build this, who was on the team internally? And yeah, what did it take to accomplish this task?
MH Yeah, absolutely. So the challenge in building something like environments is usually first you need to containerize your app, if you're lucky to start from scratch, but something like a new project now in 2021, right, chances are, you're just going to do that right away. And, and containerize your app right away from the beginning. But you know, maybe you're working on a 10, 20, 30 year old codebase. Who knows. So just getting your code to run on containers is, is probably the challenge that will depend heavily on your specific technical infrastructure, your tech stack and your environment. So in our case, for Stack Overflow, our code base is about 13 years old. It was written originally in 2008, in C sharp on dotnet framework, the classical dotnet framework, which as you know, runs on Windows, right. So it usually runs on Windows Server. And, you know, I'm in Microsoft at heart. And I love that technology when it's at its peak. But I would argue that and I hope my my friends at Microsoft will forgive me here, I would argue that when it comes to containers, and container technology, this is really more like a Linux driven ecosystem. So there is such a thing as Windows containers. And it's theoretically possible to run dotnet apps on Windows containers, even in Kubernetes. But this is just not really where the focus of the ecosystem is. And and dotnet framework itself is, at this point, officially considered legacy technology. Microsoft has like this new thing called dotnet Core and dotnet five. So really, one of the first things we had to do here was migrating the code base over to dotnet. Core. So thankfully, we had a team of very talented people across the company from SRE, to the product teams, I was not part of that team myself back then. But they they managed to migrate that old codebase over to dotnet core. And you know, that was kind of one of the basic building blocks for that.
BP And so yeah, how many people worked on this internally? How long did it take you? And yeah, did you along the way, were there any forking paths? Did you like think you might do something but then you end up changing it? How did how did your resource this internally?
MH Yeah, absolutely. So so once a code would run on dotnet core, the next thing is, you would want to build some sort of deployment pipeline for that, right. So you want this whole thing to be fully automated, like every time developer or designer, or really anyone creates a PR, you want to spin up this entire environment from scratch on a Kubernetes cluster. So the technology we've chosen, we've chosen was using GitHub Actions, which also required us to first migrate our entire code base over to GitHub hosted, so we could use the hosted version of GitHub, and we were using GitHub Enterprise on prem before. So that was like another migration project we had to do. And to your point, how many people involved almost the entire product team was involved here, because they were like countless dependencies that needed to be updated in our tooling, in our build system. And but once we were on the other side, there was like some time, early, middle to last year when we started it, and then we completed it in the fall last year. Once we were on on GitHub hosted now we had all the tools we needed, right? We had the code on GitHub Hosted, we had GitHub Actions as a workflow tool. And then we had the code base on dotnet. Core. So all we need to do is now containerize them and move it over two Kubernetes cluster. And like I said, we don't need to build our own cluster, we use one of the cloud providers, so using Azure, Azure Kubernetes services and Eks. So that piece was already there. We didn't have to build and maintain our own cluster.
RD When do you say containerize, does that require any changes to our code base? Or is that just building out the Kubernetes code?
MH Yeah, exactly. That's, that's an excellent question here. So containerizing mainly means you add, you know, you're using Docker as a tool for that and just writing a Docker file and getting your code to run in a container is usually the easy part. The hard part is usually to get your code to drop its assumptions about the infrastructure it's running on. So the code base we've been working with, had all these assumptions heavily baked in that it would run on servers on Windows Server in a data center specifically, right. So they were like changes all With a place around config around endpoints and hard coded things, and that's really worth the challenges. So maybe that's for our listeners as well, if you're, if you're thinking about containerizing something and then building maybe test environments and Kubernetes, just getting it to run in a container is not necessarily the hard part. But getting it to behave in a container is probably where, you know, we get to spending a lot of time.
BP I don't know why I've never heard that before. But something about saying the assumptions made really humanizes it for me, it makes the code sound like it's alive. I'm not sure why.
RD Well, it seems like this is part of the sort of modernization of the Stack Overflow code base we're doing with another post we worked on, talked about a lot of the sort of static stuff that was built in to make it a fast code base at the beginning. Yeah. And, and this seems like those worked at the time. But to run these PR environments, we have to modernize things.
BP And so Max, I know, you wanted to offer out sort of a hot take version of this, you know, I talked about it now. It's nuanced. But you know, like, the headline, obviously, so we can, so we can get attention is sort of like Kubernetes isn't too complicated, you know, for your startup or like, Don't overlook Kubernetes because you think it's too, you know, too complex?
RD Yeah. Containerize first.
BP Yeah, you know, maybe if you start building in this way, maybe the scale benefits won't be apparent immediately. But the cost of doing it now, like, as you said, just you know, picking it off menu from one of your cloud providers, means that you can have it in place early. And then obviously, you know, we've heard many people say, like, it's a real benefit as you grow and scale.
MH Yeah, yeah, absolutely. So I mean, if you go back to Twitter and Hacker News over the last couple of years, you'll find a lot of people making like these snarky comments about like, why, you know, we know Kubernetes is hot, but why are you messing with it now? You don't need it right now, when you're at this early growth stage, where really your focus should be on iterating and improving a product. You know, you can worry about the scale part later.
RD Yeah, I think they call it resume based development. [Max & Ben laugh]
MH Yeah, exactly. That's like the same as blog posts, do things that don't scale. So that is absolutely true. I would argue, though, that setting up your cluster really isn't that hard anymore, because your cloud provider makes it easy. And you can even manage it as infrastructure as code. And it'll offer you another benefit next to the possibility to scale, it'll offer you that ability to build test environments, and, you know, iterate way more quickly than usual, if you do that. So what you can do is, if you start using Kubernetes, from day one, you can build a similar system where for every change in every possible, you know, PR that you're going to be making to your system, you can easily create and spin up a test environment on your cluster. If you need many test environments, your cluster can scale up, if you need fewer, or if you're shutting down for the holidays, you know, you can just scale it down. So you don't have to pay all the time for like a permanent piece of infrastructure. But the point is doing that will allow you to iterate way quicker and seek feedback from people inside your organization. Even non technical stakeholders, maybe marketing sales. Yeah. And here's the other thing, you can even use that to maybe talk to customers or potential customers and easily pull off like a really customized demo for them and say, look, this is our product. In our case, it's Stack Overflow for Teams, our main product, right, and you can say, look, here's like a product. And we already customized it for you, you know, it has like your branding, and your styling and your logos in there. And this custom feature we said we might build for you, you can already see it here in your demo, right. And even if you'd like an early stage, something like that will allow you to iterate quicker and grow your product quicker.
BP I thought you were a staff engineer, you sound like a sales engineer now. [Max laughs[]
MH Yeah, you know, once in a while, we try to look outside of our engineering bubble and see what's going on the outside world.
BP Generous of you.
MH It's a wild, it's a wild ride out there.
BP Alright, so let me throw you a hypothetical just because I was reading a story about Just in Time Inventory, which is practice of working that was pioneered by Toyota. And you know, the idea is you're not holding, you know, lots of car doors and brakes and stuff in the factory, you know, costing you money, you know, everything is arriving, the factory just as is needed, and you put it on the assembly line, you get the car out. So this, you know, was a great innovation of Toyota, made their business far more valuable. And then it was picked up by companies from every, you know, industry, it didn't have to just be auto manufacturing. And it made them all a lot leaner, I'm sure their share prices went up and their, you know, their executives benefited from buybacks. But then when, you know, pandemic came along, and the supply chain, you know, got all messed up, it was difficult to recover, nobody was holding inventory. And now, lots of companies are stuck with demand that they can't meet. So I guess, you know, the parallel would be if from the beginning, your startup is building completely with infrastructure as code, and you know, everything is being spun up and spun down. And it's all dependent on a third party. You know, if some disaster comes along and knocks out every AWS cluster, or you know, if just something happens, that's more systemic to the internet, can you fall back and be you know, self reliant, essentially, like I remember when, you know, Hurricane Sandy came through New York, there's a great war story about people at Stack, you know, bailing out, you know, marching up and down stairs and bailing out water and keeping our servers running so that people can continue to use the service despite the fact that we had a local disaster. You know, as more and more of what you build is virtual and outsourced, do you run the risk at some point, yeah, of not being able to be self reliant, if such a, you know, systemic network effect should come down the line?
MH Great question. So here's what I'll say. So there's like the issue of vendor lock in or technology lock in which companies or like even startups want to avoid for basically what you just said. Let's say you're completely building your business on top of a single cloud provider. And you know, I mean, AWS is unlikely to go out of business. But let's say you're maybe partnering with a smaller one. And maybe that that provider has run into some sort of trouble. And then maybe you can no longer partner with them, what's like your plan B here. So you can basically consider Kubernetes as an abstraction layer, that will increase your independence and will make you even more technology agnostic. And that's just because for more or less, you know, you can switch Kubernetes providers. If you're using operator A today, to provide your Kubernetes service, the amount to switch over to a different provider is limited, there are some pieces,you'll have to change. You know, the way let's say your load balancer and your Ingress works will be different. If you're moving, let's say from Azure, to AWS, but really all the internal pieces like the internal networking, architecture of your cluster, your pods, your services, and all that sort of stuff. It'll be the same. So you can basically move over to a different provider more easily.
RD Yeah, and I mean, talking about the sort of disaster recovery stuff, all the virtualization stuff exists on top of real hardware, right? This lets you kind of let it exists on multiple data centers, in case one of them gets hit by a Godzilla, or something.
[music]
BP Alright, everybody, it is that time of the episode, I am going to shout out the winner of a lifeboat badge, that's somebody who came on Stack Overflow, and there was a question with a score of negative three or less, they gave it an answer, and it got up to a score of 20 or more. Today, we will shout out Mantas awarded May 26th: 'Determine if all the values in a PHP array are null'. So if you want to know how to determine it, we've got an answer for you. You can check it out in the show notes. I am Ben Popper, Director of Content here at Stack Overflow. You can always find me on Twitter @BenPopper and you can always email us with thoughts and suggestions podcast@stackoverflow.com. If you enjoyed the show, please do leave a rating and a review on whatever platform you're listening. It really helps.
RD I'm Ryan Donovan, you can find me on Twitter, I read my DMs or if you have a blog post idea, you can reach me at pitches@stackoverflow.com.
BP Max, who are you and where can people find you?
MH Awesome. I'm Max Horstmann. My last name is H O R S T M A N N. It's a German last name. You'll find me on the internet on the, in the usual places. I'm Max_Horstmann on Twitter, and you can also find me on maxhorstmann.net.
BP Yeah, and if you want to read more about what Max and some other folks at Stack built using Kubernetes, giving us all our little PRs, we can check out work in progress. We'll have a blog post up about that and we'll link it in the show notes.
[outro music]
