This week we chat with Max Horstmann, a software developer at Stack Overflow. He details our company's experience setting up a Kubernetes cluster, turning that into a managed service through cloud providers, and more recently, using Kubernetes to host what we call PR Environments. Based on this, Max argues that Kubernetes is no longer too challenging for startups to consider using early on.
Episode Notes
You can read Max's full article on Kubernetes on our blog here.
You can find Max on Twitter here and his personal website here.
Our lifeboat badge winner of the week is Mantas, who answered the question: Determine if all the values in a PHP array are null
TRANSCRIPT
Max Horstmann If you go back to Twitter and Hacker News over the last couple of years, you'll find a lot of people making like these snarky comments about like, you know, we know Kubernetes is hot, but why are you messing with it now? You don't need it right now, when you're at this early growth stage where really your focus should be on iterating and improving a product. You know, you can worry about the scale part later.
Ryan Donovan Yeah, I think they call it resume based development.
Yeah, exactly. Oh, that's like the same as blog posts, do things that don't scale. So that is absolutely true. I would argue, though, that setting up your cluster really isn't that hard anymore, because your cloud provider makes it easy. And you can even manage it as infrastructure as code. And it'll offer you another benefit next to the possibility to scale, it'll give you that ability to build test environments, and you know, iterate way more quickly.
[intro music]
Ben Popper Cockroach DB is the only book you'll ever love. Because it's the only one you don't have to worry about. As a low touch SQL database that automatically handles scale, operations, and uptime. Cockroach DB lets you focus on developing, get your free cluster and a free t-shirt at cockroachlabs.com/StackOverflow.
BP Hello everybody! Welcome to the Stack Overflow Podcast, a place to talk about software and technology. I am Ben Popper, Director of Content here at Stack Overflow. And today I'm joined by my colleague, Ryan Donovan, who is a content marketer on my team, running the blog and the newsletter. Hi, Ryan.
RD Hi, Ben, how you doing?
BP I'm pretty good. So one thing that we've written about a few times, and which I think we're going to discuss today, is the idea of using, you know, Kubernetes and containerization, to rethink how your company is architected. We've done a few like sponsored blog posts about this from big COEs and had pictures on it. But you know, it feels like overall, this is a big shift in the industry that's been playing out over the last few years.
RD Yeah, absolutely.
BP And, you know, a technology that is just exploding in popularity. What do you think is driving that?
RD Well, I think a lot of software folks don't want to manage their hardware. When a lot of people, a lot of organizations move to the cloud, it was just easier to handle resources. And now, right, you have Kubernetes, which you can treat the resource as code.
BP So this is part of the bigger sort of like infrastructure as code movement.
RD Yeah, I think so.
BP Well, we have a in house expert coming on the podcast today to discuss it. Max Horseman, who is a staff software engineer at Stack Overflow, and has recently been working on some really cool stuff in house related to exactly this topic. Welcome, Max.
MH Hey, Ben. Hey, Ryan. Thanks for having me.
BP Yeah, of course. Great to have you on. So for folks who don't know, you've been with Stack Overflow for I think, more than eight years you said?
MH Yeah, that's right. I joined back in 2012. So it's been eight and a half years now. Yeah.
BP Very cool. And you were at Microsoft before that?
MH That's right, spend a couple of years at Microsoft then took the year off for traveling, which was great. And then I joined this relatively young company back then called Stack Overflow.
BP We've all aged so gracefully together. So Max, you did a piece for us that I want you to talk about, maybe, you know, step back a minute and sort of tell people like, what your thesis is here and how that relates to some of the work you've been doing inside of Stack Overflow, building, essentially, what what is like new tooling that all of us get to use?
MH Absolutely. So the first couple of years here at stack, I worked on the Talent and Jobs team. So for those who don't know, if you're on Stack Overflow, there's also a job board. And if you're looking at a question or answer page on Stack, you'll, you'll sometimes see job listings advertised and you can go to the job board and find a job. And if you're an employer, and you're trying to hire developers, there's also a portal for that, and Stack Overflow Talent. So that's what I've done for many years. And when I when I joined the company, this was really at an early stage, and we needed to grow the product a lot and iterate a lot. So I mean, it's well established now that what you want to do at an early stage is you know, lean startup and all that sort of things, that you want to iterate and seek feedback early on, and then you know, make improvements based on that feedback.
RD Get that market fit.
MH Yeah. So getting to product market fit. And then based on that, grow the product and scale it and yeah, so at some point, we realized that really the tooling we have for that needs some improvement. And yeah, that's I'm here to talk about today.
BP Yeah, I feel like doesn't didn't one of our co founders have like a famous quote, somewhere along these lines. Maybe it was Jeff Atwood, from that sacrificial architecture piece. Right, Ryan? Am I thinking about this?
RD That performance is a feature. Basically.
BP Yes, performance is a feature.
RD You can build it in at the beginning, or you can add it in later. But it's, you should treat it as a feature.
BP So I thought that was kind of cool. Jeff would say performance as a feature. Many developers understood this to mean performance is the first thing to care about, but that's not quite right. I like the way that was put. So, Max, you were working on Talent. And I guess, you know, yeah, just to clarify, we should say, you know, the business of Stack Overflow has been evolving. We're kind of we talked about this publicly. It's no secret moving away from job slots, and we're gonna, you know, you started the scale and reach we have with developers to allow companies to do sort of awareness about their products or services or open roles they have for developers. Talk about their brand, and then focus on teams. So jobs, that's something you did in the past. Not anymore. But yeah, you have a different sort of take on Kubernetes, which often people talk about as being a little too complex or unwieldy or expensive for a startup to consider. It's something like we said people talk about is maybe a big re-architecture decision when you've scaled to a certain point. And you know, you feel like pain points or areas of friction or, or too much overhead is developing. But you feel like maybe that's no longer the case. And then I think you have kind of an interesting implementation of it in house, right?
MH Yeah, exactly. So. So to your first point, you're right. Kubernetes is often and still considered to be a relatively complex technology, right? Something that requires a lot of resources and overhead to create and to manage and running your own Kubernetes cluster is certainly something that is a complex task. But you know, in 2021, now, I think it's fair to say that with all these cloud offerings out there, that you know, take care of that for you, you know, offering Kubernetes, basically, as a utility or like a resource, I don't think that is any longer true that you would need to avoid Kubernetes, because of its complexity, in fact, it's gotten very, very straightforward to set up, it takes you a couple of minutes to set up your own cluster. If you you know, you mentioned infrastructures code earlier on, if you're a believer in infrastructure as code, you can use tools like terraform, to set up your cluster and basically manage it as code in your version control system.
BP So let me step back for a minute. And I'm sure Ryan will understand this better than I so I'll let him jump in. But we're talking about like, sort of the multiple layers of abstraction. This is just the natural evolution. So first, we say you know, you've got your own server, you've got your own hardware, then we say, okay, now you're going to use a Rackspace somewhere, no, you're just going to spin up, you know, like an AWS instance. And now people are saying, well, it's actually easier and better to use this, you know, containerization technology. But okay, even that's too tough. Like, well let some big company background spin that up, and we'll use their, you know, cluster. So you're like, three, four or five steps away now from setting up your own Raspberry Pi at home when it's time to do this stuff.
RD Yeah, I mean, it seems like from my understanding, Kubernetes is basically reducing it all to YAML files. Right? That just makes it another simpler thing. And if it's all attached to these cloud providers, like, why not, why not get it started early?
BP I mean, I guess, Max, like maybe not in layman's terms, but as you know, in sort of a summary way, like, what does it take? Yeah, what's involved in setting up and running a cluster?
MH Right, yeah. So just setting up a cluster can be done basically, in two ways. And we're talking about like a cloud hosted cluster, right, setting up your own cluster. And by the way, you can use your own Raspberry Pi hardware, or anything else at home, people have done that. There's stuff on Reddit, where people set up their own home Kubernetes cluster on Raspberry Pi's. But that's not we're talking about here today. Setting up your cluster in in a cloud hosted environment, you know, you can either just go to your favorite cloud providers portal and set it up through some sort of UI, or you can use like I mentioned to infrastructures code provider, like terraform. And write something, it's not quite YAML. That's quite similar, right? and define your cluster and then almost set it set up a deployment pipeline for, for your cluster and for infrastructure.
BP But I meant, step back one more level, like for people who don't know, and because we have lots of engineers, listen that but also lots of people who are software adjacent, like what do you get out of, you know, a Kubernetes cluster? What is involved in that? And what does it help you do?
MH Exactly, yeah. So I just want to give you an example, and talk about a project we've worked on here at Stack, because I think that's a really nice illustration, like, like people have said, you know, people think of Kubernetes very often as, like an infrastructure piece and something that allow you to scale and you know, that does some of the heavy lifting around fault tolerance, and all that sort of load balancing and that sort of stuff, right. But it can also be used for quite a number of different things. And one specific thing we did here at Stack was we're using Kubernetes, to host what we call PR environments, I should explain that a little bit. So this is something a lot of organizations and companies have been doing over the last couple of years. The idea is, when you're making a code change, you're typically using a version control system like Git, and or maybe GitHub and, and then typically, at some point, you're creating a pull request or PR, we all know that right? Where you're sharing your working progress with some of your peers maybe and seeking some comment on that. Now the problem is looking at the code will not tell you the whole story. And especially if you want to share that with someone maybe outside of the engineering team, let's say someone like you two in marketing, or maybe somebody in sales, or maybe in somebody and let's say the legal department, who knows, right? It would be very nice if you can just show them your code running. And ideally, you would just send them a link and say, Look, I have this work in progress here. Here's the link, can you just click on the link and then you'll see what I'm doing here and just try it out. Basically, this is like a completely isolated and custom test environment, just for those code changes you're working on.
RD I think this is a completely wild you know, advancement to the industry. Like the fact that you can see spin up this automatically, you don't have to download all the requirements to your local machine, you don't have to, you know, take over whatever the dedicated test environment is, you just have this one thing that's dedicated to the PR itself.
BP Ryan, you were mentioning previous jobs, you knew people who did this and half of their job was just going through that setup process.
RD Yeah, I mean, especially if you have a service oriented architecture, you know, you have to spin up each of those services, possibly put those in a Docker container. And it's just kind of a nightmare, you got to make sure other requirements are there. And if one of those requirements is broken, you're debugging a thing that shouldn't even be debugged.
BP Yeah, it's super interesting, you know, and especially right as, as more and more of the products lives, not, you know, on a floppy disk or CD ROM that you're going to ship people or you know, as an on prem delivery, but as something that is live on the web, you might very well want design, or marketing, if they're doing copy or product marketing, you know, to be able to look at these things and push the changes, you know, on a regular basis. So to me, it makes a lot of sense. And yeah, it's exciting for people like me, who are not software developers themselves, but have to work with you loons every day. I guess, tell us a little bit about sort of, yeah, like when you set out to build this, who was on the team internally? And yeah, what did it take to accomplish this task?
MH Yeah, absolutely. So the challenge in building something like environments is usually first you need to containerize your app, if you're lucky to start from scratch, but something like a new project now in 2021, right, chances are, you're just going to do that right away. And, and containerize your app right away from the beginning. But you know, maybe you're working on a 10, 20, 30 year old codebase. Who knows. So just getting your code to run on containers is, is probably the challenge that will depend heavily on your specific technical infrastructure, your tech stack and your environment. So in our case, for Stack Overflow, our code base is about 13 years old. It was written originally in 2008, in C sharp on dotnet framework, the classical dotnet framework, which as you know, runs on Windows, right. So it usually runs on Windows Server. And, you know, I'm in Microsoft at heart. And I love that technology when it's at its peak. But I would argue that and I hope my my friends at Microsoft will forgive me here, I would argue that when it comes to containers, and container technology, this is really more like a Linux driven ecosystem. So there is such a thing as Windows containers. And it's theoretically possible to run dotnet apps on Windows containers, even in Kubernetes. But this is just not really where the focus of the ecosystem is. And and dotnet framework itself is, at this point, officially considered legacy technology. Microsoft has like this new thing called dotnet Core and dotnet five. So really, one of the first things we had to do here was migrating the code base over to dotnet. Core. So thankfully, we had a team of very talented people across the company from SRE, to the product teams, I was not part of that team myself back then. But they they managed to migrate that old codebase over to dotnet core. And you know, that was kind of one of the basic building blocks for that.
BP And so yeah, how many people worked on this internally? How long did it take you? And yeah, did you along the way, were there any forking paths? Did you like think you might do something but then you end up changing it? How did how did your resource this internally?
MH Yeah, absolutely. So so once a code would run on dotnet core, the next thing is, you would want to build some sort of deployment pipeline for that, right. So you want this whole thing to be fully automated, like every time developer or designer, or really anyone creates a PR, you want to spin up this entire environment from scratch on a Kubernetes cluster. So the technology we've chosen, we've chosen was using GitHub Actions, which also required us to first migrate our entire code base over to GitHub hosted, so we could use the hosted version of GitHub, and we were using GitHub Enterprise on prem before. So that was like another migration project we had to do. And to your point, how many people involved almost the entire product team was involved here, because they were like countless dependencies that needed to be updated in our tooling, in our build system. And but once we were on the other side, there was like some time, early, middle to last year when we started it, and then we completed it in the fall last year. Once we were on on GitHub hosted now we had all the tools we needed, right? We had the code on GitHub Hosted, we had GitHub Actions as a workflow tool. And then we had the code base on dotnet. Core. So all we need to do is now containerize them and move it over two Kubernetes cluster. And like I said, we don't need to build our own cluster, we use one of the cloud providers, so using Azure, Azure Kubernetes services and Eks. So that piece was already there. We didn't have to build and maintain our own cluster.
RD When do you say containerize, does that require any changes to our code base? Or is that just building out the Kubernetes code?
MH Yeah, exactly. That's, that's an excellent question here. So containerizing mainly means you add, you know, you're using Docker as a tool for that and just writing a Docker file and getting your code to run in a container is usually the easy part. The hard part is usually to get your code to drop its assumptions about the infrastructure it's running on. So the code base we've been working with, had all these assumptions heavily baked in that it would run on servers on Windows Server in a data center specifically, right. So they were like changes all With a place around config around endpoints and hard coded things, and that's really worth the challenges. So maybe that's for our listeners as well, if you're, if you're thinking about containerizing something and then building maybe test environments and Kubernetes, just getting it to run in a container is not necessarily the hard part. But getting it to behave in a container is probably where, you know, we get to spending a lot of time.
BP I don't know why I've never heard that before. But something about saying the assumptions made really humanizes it for me, it makes the code sound like it's alive. I'm not sure why.
RD Well, it seems like this is part of the sort of modernization of the Stack Overflow code base we're doing with another post we worked on, talked about a lot of the sort of static stuff that was built in to make it a fast code base at the beginning. Yeah. And, and this seems like those worked at the time. But to run these PR environments, we have to modernize things.
BP And so Max, I know, you wanted to offer out sort of a hot take version of this, you know, I talked about it now. It's nuanced. But you know, like, the headline, obviously, so we can, so we can get attention is sort of like Kubernetes isn't too complicated, you know, for your startup or like, Don't overlook Kubernetes because you think it's too, you know, too complex?
RD Yeah. Containerize first.
BP Yeah, you know, maybe if you start building in this way, maybe the scale benefits won't be apparent immediately. But the cost of doing it now, like, as you said, just you know, picking it off menu from one of your cloud providers, means that you can have it in place early. And then obviously, you know, we've heard many people say, like, it's a real benefit as you grow and scale.
MH Yeah, yeah, absolutely. So I mean, if you go back to Twitter and Hacker News over the last couple of years, you'll find a lot of people making like these snarky comments about like, why, you know, we know Kubernetes is hot, but why are you messing with it now? You don't need it right now, when you're at this early growth stage, where really your focus should be on iterating and improving a product. You know, you can worry about the scale part later.
RD Yeah, I think they call it resume based development. [Max & Ben laugh]
MH Yeah, exactly. That's like the same as blog posts, do things that don't scale. So that is absolutely true. I would argue, though, that setting up your cluster really isn't that hard anymore, because your cloud provider makes it easy. And you can even manage it as infrastructure as code. And it'll offer you another benefit next to the possibility to scale, it'll offer you that ability to build test environments, and, you know, iterate way more quickly than usual, if you do that. So what you can do is, if you start using Kubernetes, from day one, you can build a similar system where for every change in every possible, you know, PR that you're going to be making to your system, you can easily create and spin up a test environment on your cluster. If you need many test environments, your cluster can scale up, if you need fewer, or if you're shutting down for the holidays, you know, you can just scale it down. So you don't have to pay all the time for like a permanent piece of infrastructure. But the point is doing that will allow you to iterate way quicker and seek feedback from people inside your organization. Even non technical stakeholders, maybe marketing sales. Yeah. And here's the other thing, you can even use that to maybe talk to customers or potential customers and easily pull off like a really customized demo for them and say, look, this is our product. In our case, it's Stack Overflow for Teams, our main product, right, and you can say, look, here's like a product. And we already customized it for you, you know, it has like your branding, and your styling and your logos in there. And this custom feature we said we might build for you, you can already see it here in your demo, right. And even if you'd like an early stage, something like that will allow you to iterate quicker and grow your product quicker.
BP I thought you were a staff engineer, you sound like a sales engineer now. [Max laughs[]
MH Yeah, you know, once in a while, we try to look outside of our engineering bubble and see what's going on the outside world.
BP Generous of you.
MH It's a wild, it's a wild ride out there.
BP Alright, so let me throw you a hypothetical just because I was reading a story about Just in Time Inventory, which is practice of working that was pioneered by Toyota. And you know, the idea is you're not holding, you know, lots of car doors and brakes and stuff in the factory, you know, costing you money, you know, everything is arriving, the factory just as is needed, and you put it on the assembly line, you get the car out. So this, you know, was a great innovation of Toyota, made their business far more valuable. And then it was picked up by companies from every, you know, industry, it didn't have to just be auto manufacturing. And it made them all a lot leaner, I'm sure their share prices went up and their, you know, their executives benefited from buybacks. But then when, you know, pandemic came along, and the supply chain, you know, got all messed up, it was difficult to recover, nobody was holding inventory. And now, lots of companies are stuck with demand that they can't meet. So I guess, you know, the parallel would be if from the beginning, your startup is building completely with infrastructure as code, and you know, everything is being spun up and spun down. And it's all dependent on a third party. You know, if some disaster comes along and knocks out every AWS cluster, or you know, if just something happens, that's more systemic to the internet, can you fall back and be you know, self reliant, essentially, like I remember when, you know, Hurricane Sandy came through New York, there's a great war story about people at Stack, you know, bailing out, you know, marching up and down stairs and bailing out water and keeping our servers running so that people can continue to use the service despite the fact that we had a local disaster. You know, as more and more of what you build is virtual and outsourced, do you run the risk at some point, yeah, of not being able to be self reliant, if such a, you know, systemic network effect should come down the line?
MH Great question. So here's what I'll say. So there's like the issue of vendor lock in or technology lock in which companies or like even startups want to avoid for basically what you just said. Let's say you're completely building your business on top of a single cloud provider. And you know, I mean, AWS is unlikely to go out of business. But let's say you're maybe partnering with a smaller one. And maybe that that provider has run into some sort of trouble. And then maybe you can no longer partner with them, what's like your plan B here. So you can basically consider Kubernetes as an abstraction layer, that will increase your independence and will make you even more technology agnostic. And that's just because for more or less, you know, you can switch Kubernetes providers. If you're using operator A today, to provide your Kubernetes service, the amount to switch over to a different provider is limited, there are some pieces,you'll have to change. You know, the way let's say your load balancer and your Ingress works will be different. If you're moving, let's say from Azure, to AWS, but really all the internal pieces like the internal networking, architecture of your cluster, your pods, your services, and all that sort of stuff. It'll be the same. So you can basically move over to a different provider more easily.
RD Yeah, and I mean, talking about the sort of disaster recovery stuff, all the virtualization stuff exists on top of real hardware, right? This lets you kind of let it exists on multiple data centers, in case one of them gets hit by a Godzilla, or something.
[music]
BP Alright, everybody, it is that time of the episode, I am going to shout out the winner of a lifeboat badge, that's somebody who came on Stack Overflow, and there was a question with a score of negative three or less, they gave it an answer, and it got up to a score of 20 or more. Today, we will shout out Mantas awarded May 26th: 'Determine if all the values in a PHP array are null'. So if you want to know how to determine it, we've got an answer for you. You can check it out in the show notes. I am Ben Popper, Director of Content here at Stack Overflow. You can always find me on Twitter @BenPopper and you can always email us with thoughts and suggestions podcast@stackoverflow.com. If you enjoyed the show, please do leave a rating and a review on whatever platform you're listening. It really helps.
RD I'm Ryan Donovan, you can find me on Twitter, I read my DMs or if you have a blog post idea, you can reach me at pitches@stackoverflow.com.
BP Max, who are you and where can people find you?
MH Awesome. I'm Max Horstmann. My last name is H O R S T M A N N. It's a German last name. You'll find me on the internet on the, in the usual places. I'm Max_Horstmann on Twitter, and you can also find me on maxhorstmann.net.
BP Yeah, and if you want to read more about what Max and some other folks at Stack built using Kubernetes, giving us all our little PRs, we can check out work in progress. We'll have a blog post up about that and we'll link it in the show notes.
[outro music]
