SPONSORED BY DOIT
In this episode, Ben and Ryan are joined by Joshua Fox, a senior cloud architect at DoiT, to discuss cloud cost optimization. They explore the importance of controlling and understanding cloud costs, the role of good architecture in cost optimization, and strategies for dealing with surprise costs.
Episode notes:
To learn more about the signs that indicate you may be paying more for your cloud computing than you should, check out DoiT’s seven red flags guide.
We’ve spoken with DoiT on the podcast before about LLM hallucinations and the security threats that LLMs open.
Partner with DoiT for cloud procurement, expertise or tooling in any combination, to drive value where you need it most.
Congrats to Lifeboat badge winner Sravan K Ghantasala for their answer to How to sort file lines in Bash?
Find Joshua at joshuafox.com.
Chapters
00:00 Introduction and Cloud Cost Control
01:08 Joshua Fox's Background
04:20 Understanding FinOps
06:17 The Importance of Good Architecture
08:18 Balancing Flexibility in Architecture
10:04 Surprise Costs and Dealing with Them
13:19 Bracing for Unexpected Cloud Costs
25:41 The Future of Cloud Cost Optimization
27:09 Closing Remarks
TRANSCRIPT
[intro music plays]
Ben Popper Hello, everybody. Welcome back to the Stack Overflow Podcast, a place to talk all things software and technology. I am Ben Popper, Director of Content here, joined as I often am by my colleague, Ryan Donovan. Ryan, how’re you doing today?
Ryan Donovan I'm good. How’re you doing, Ben?
BP I am pretty good. So I think one thing you and I have talked about many times over the years with many different people from companies big and small is cloud costs, controlling them, understanding them, expecting them, and how easy it is sometimes to lose sight of that or get smacked with a bill that you weren't expecting.
RD That's right.
BP So today's episode is sponsored by the fine folks at DoiT. It's not the first time we've had them on or they've written for the blog– we’ve done a lot of great stuff with them that you can check out– but today we are lucky to have Joshua Fox with us, who is a Senior Cloud Architect over at DoiT. We're going to be chit-chatting a lot about how you manage and optimize what you're doing in cloud and how DoiT can sort of help you in that arena. So Josh, welcome to the podcast.
Joshua Fox Glad to be here.
BP So first things first, talk to us a little bit about your background. How'd you get into the world of software and technology and what led you to the role you're at with DoiT today?
JF Well, I've been in software for 24 years. I spent many years in product development as an architect developing and coding a lot of software, and then about four years ago I moved over to cloud consulting. I'd been working with Doit, they'd been advising my company, and now I'm advising other companies. That's what I do all day.
BP You were a satisfied customer before you were an employee.
JF I was, yeah. I had a great working relationship with them, and I gave a talk at a multi-cloud meetup that DoiT sponsored and that's how we got to know each other.
RD Nice. So we're talking about cloud cost optimization, and I remember at my last role they had a service oriented architecture. They did a sort of audit of everything that was in their cloud and they found ways to save five/six figures a month. How do people go about doing that themselves– finding where those excess expenses are and then optimizing their expenses?
JF Okay, so the first thing to remember in any optimization is to find the bottlenecks. It's the Pareto rule, the 80/20 rule, which is to say that most of the problems are going to be in a few small areas. And it really has to be that way, because if you're spending 10% of your cloud spend on some service and you cut that in half, well, great, but you've only reduced it from 10 to 5%. And furthermore, if your cloud spread is spread out across 20-30 services pretty much equally, then you're actually doing great because there is no one area with obvious over the top waste. Because that's where the real expense comes from. It's from the extreme bugs, extreme mistakes, and that's what you need to work to reduce. Now, you asked me how to find this, and you need to dice and slice and cut up your expenses. We at DoiT have a cloud analytics tool which I think is great. And there are tools within the Google Cloud Platform or AWS, you can use those, but ours lets you attribute the cost to different teams. So let's say you've got three teams and you want to give a budget to each one so you can have different ways of assigning, let's say, per project or by labels, and now you can focus in on one team's expenses. And now you cut it up mostly by service, that's what I always do first. So let's say there's a Cloud SQL database or a Kinesis Stream, and the report will just make it very clear that one of these areas is way over the top, and that's what I do first. And then I happen to know some areas where there is usually waste just from experience, we could talk about those, and that's where I go next. But that's my answer.
BP So one thing I'm curious about is how this sort of gets out of control in the first place. And when we chatted with you, you mentioned cloud optimization. You also mentioned something called FinOps which I had never heard of before. So can you describe FinOps to me? I know that's sort of a proactive way to control cost in advance. And it does seem like a lot of times what we're talking about is not, “How do I set this up so that this doesn't happen to me?” but rather the inverse of, “Oh, we finally did the audit and figured out we were paying X too much and now we're saving those costs.”
JF So as you say, FinOps just means planning in advance, and that's a good idea as with anything. So you should think carefully about where your expenses are going and maybe sketch out that a certain percent is going to go to virtual machines and another percentage to a database. Obviously it's good to plan ahead. I would caution that you never know the future so you cannot, if you're a fast growing startup, say, “Oh, we're going to spend this much on databases.” Well, this month, yes, but in a year you're going to grow, and then not only will that database cost grow, but it might grow disproportionately to, let's say, your compute costs. So as always with planning, it's essential, but plans never last. And when you think ahead about your FinOps planning, you can most effectively do that if you're not a fast growing startup, let's say if you're a bank with 30,000 employees with tens of thousands of machines, old fashioned machines, hardware sitting around that you're migrating to the cloud.
BP Mainframes, yeah.
JF That's right. If you're one of those, then you probably can do FitOps because you're not changing so much year by year. That's great, but when I work at DoiT, we're almost always working with fast growing companies, and therefore the FinOps is just a plan. But there is one thing that you absolutely can and must plan ahead of time, and that is a good architecture. Now this talk is not about everything to do with architecture. It's a great topic, it's what I do professionally. I can touch on a few points. But good architecture is defined as flexible. And when a year or two has gone by and everything has grown tremendously and you're happy but you see that there's now a bottleneck of costs and one specific part is spending 90%– let's say one of my customers was using a Google Kubernetes environment and they had a lot of memory on each virtual machine, or what's called a node in Kubernetes. And at first they said, “Well, we need memory because what we're doing is a complex algorithm against what is effectively an in-memory database.” Can't argue with that, okay. Remember that so-called in-memory database is going to grow tremendously. And once they had a lot of memory there, it was very hard to back out because all the code assumed that the data was immediately accessible in memory. So that's an example of something that's hard to change. Conversely, if they could have thought, I'm not saying this is easy, but they could have thought about a stateless architecture in which data comes into the compute, in this case Kubernetes, it is calculated, transformed, and goes out and all, the state, as much as possible, stays in, let's say, a database. In that case, the nodes, the compute, could be horizontally scaled out. Moreover, it's not just that instead of a very large memory node you have ten small memory nodes, it's that you can correctly take the different types of resources. Because when you buy a huge memory node, you're often forced to buy too much CPU that you don't really need. So when you keep it small you can control the types of resources that you're using. That's an example of how planning an architecture does not give you an immediate answer but will give you a solution when things start to change.
RD So we ran an article a while back about how much future-proofing is too much. So I'm wondering if there's a point where their architecture can be too flexible, where there's things you have to watch out for to make sure it's not overly flexible architecture.
JF Our job as architects is to predict what will change and what will not. And as I like to say, that's why we get those big salaries, because you just don't know what will change. Well, unless you have years of expertise, wisdom, and understanding the business requirements will change, too, the business landscape might change. Your carefully balanced resources, let's say one company was using AWS Lambda against DynamoDB, and they'd correctly chosen the allocation of resources. But I was able to tell them that because the Lambda had a slow startup time and actually a slow execution time as well, that disproportionately there would be a burden on the Lambda. That was something I could know in advance, and yes, that is the key to architecture. I said good architecture is defined as flexible– flexible to change. So it has to be flexible to the change that you know is going to come.
BP It's interesting, I would love to chat a little bit about surprise costs and some of the fun stuff you've seen. I know you don't want to talk about Gen AI, maybe we'll chat a little bit of MLOps, but in my mind as you were saying that before, how can you plan in advance but you can't plan everything? If I was a startup two or three years ago and I was growing fast and now all of a sudden somebody is coming to me and asking about vector databases or GPUs, those are things that I just may not have thought about in the context of my company until all of a sudden that was what everybody was talking about. So tell us from your perspective, what are some of the interesting things you've seen that surprised folks and then how they dealt with them?
JF Okay, so we can talk about MLOps. And I'm not an expert in Gen AI, although many of my colleagues are, but in the area of MLOps, I've been exploring a few very interesting approaches. So let's talk about training, first of all. Training and AI is expensive. It is slow, it takes a long time, it uses GPUs which are in a great shortage nowadays, in addition to the fact that they're expensive. So every time you train your AI, you're going to spend however many dollars. Now, one of my customers asked me, “How much does it cost?” And I said that the rumor is that ChatGPT costs $150 million for one training session. I said, “Okay, you're probably not going to spend that much, but there is no limit.” So what that means is that the very unpredictable training run may have to be redone because you may realize that you got some of the hyperparameters wrong. You set it up a little bit wrong and you redo. Hey, that's life, we always redo things. But now you're going to spend more money and more time on that rerun. This is called hyper-optimization. So first of all, make sure you do it on a platform such as Google Vertex or AWS SageMaker. They will make sure that the correct resources are used. You have to pay for them, but if you do it yourself, you might have your GPUs running 24/7 when you're not training 24/7, or whatever. They have thought, the people at Google or Amazon have thought about what it means to give the resources needed when needed. Secondly, if you have to rerun your training, do it methodically. And one approach is automatic hyperparameter tuners, such as SageMaker Autopilot or Google AutoML. Another approach is something called Vertex Vizier from Google, and that lets you run the training sessions over and over trying to make them better and better, but each time it gives you a suggestion. You as a human could try to tweak the parameters and rerun. This tool will suggest better parameters. That gives you full control over the expensive training process. You can do your best to lower the price as opposed to handing it off to AutoML or SageMaker Autopilot. Those are a few tips about training. When we're talking about inference, which is at the end of the process after the model is trained and you're asking questions, in this case I again advise using a service. Inference endpoints from Google, from Amazon, again, that will give the amount of resources needed, no less, no more, and if necessary, scaling up, and again, scaling down. This is generally true about the cloud. You should always use the more and more managed service, unless you have a very special need, which you probably don't. And that is true in the case of the training, the inference, and everything in between them which is tied together with tools called pipelines.
RD So how about surprise cloud costs? There was a story out recently about somebody getting a $100,000 monthly bill because of a DDoS attack. Are there ways to sort of brace yourself for surprises like that?
JF Yes. So I unfortunately, too, have advised customers who have seen tremendous losses. One of the most common is Bitcoin mining when a key leaks out. I should say one case, but in actually several cases, the hackers would spin up instances in every region all over the world to avoid quotas and limits in every given region, and every given type of virtual machine costing over half a million dollars in a few days. Usually the cloud providers will refund that once. Don't mess around and let your keys leak a second time because they're not going to be so friendly the second time. In another case, I saw an SSH server had been left with the default password, admin admin or whatever it was, and then hackers had transmitted nine terabytes to China using this instance over a weekend. So very simple, so how do you deal with these things? One is to think hard about security, and we at DoiT advise on that. I advise on that. So get your security down right. One example is to use organizational policies to block certain regions, certain instance types. If your entire system is in AWS Europe West 1, then there's no reason hacker scripts should be allowed to spin up virtual machines in other areas. And another thing is that, after some of these hacks had occurred, I advised the customers in setting up alerts. Now, we at DoiT have anomaly detection which will send out an alert when the costs go up unexpectedly. It's not as simple as it sounds. You can set up budgets that say over $1,000 is too much. You can do that in AWS, in Google, in our tool, but the thing is that what if you have a viral event? Good, we're happy. The cost just went up, that's wonderful. So you can't just say, “If the costs go up.” And our anomaly detection uses machine learning to look at patterns, time of year, and give you an alert if there's an anomaly in costs. You can use ordinary budgets if you want. One warning about this sort of thing is that cloud data takes hours up to a day, sometimes more, to arrive. That's true of every tool you use, whether you use the built-in Google or Amazon, or whether you use our cloud console, simply that's how the cloud providers gather the data. I'm not an architect at Google or Amazon, but I understand that they have some sort of batch process that runs periodically. So these anomaly alerts will only kick in after a certain number of hours. If you want to go beyond that, then you should have good security practices and alerts, which is not our topic today, but alerts for potential security leaks.
BP I don't know if you know the answer to this, but that's a double whammy. Not only are you getting hacked, but you're paying the freight costs for them to take that data somewhere else. I don't know if people can get insurance on that or not, but it sounds rough.
JF I've never heard of insurance. I would not sell an insurance policy on that because the people who can most control it are the customers themselves, and unfortunately some of my customers have better practices than others. But as I say, the cloud providers themselves can be flexible within limits, which is true of insurance companies as well.
BP All right, so you were just talking about best practices. What are some of your top tips on cost control?
JF First of all, a good architecture, which is flexible to change, as I said earlier. That's the most important thing. Often startups like to move fast and break things, and I agree, that's what I do too, but they should realize that the cost will come later in many ways. And the most important flexibility is to create new features that your customers want to use. Every company needs to do that. But the other type of flexibility is to adjust when you identify new security threats or when you see that costs are running out of control, so you need to make your architecture modular and test each module. It might mean microservices, but it doesn't have to. If you're doing more of a monolith, then break it into Java modules or Python modules and make sure that you have well-defined interfaces between them. Again, this is all a lesson in Architecture 101. That will be the next podcast. But a specific thing to think about within the context of architecture is statelessness. I mentioned earlier one type of statelessness and why it's important– not letting your systems scale vertically. Even if in theory, one small virtual machine which is one tenth the size of a large one is a tenth the cost. That's actually true. But you can most optimally avoid wastage if you keep your state, your memory in this case, in small blocks. But memory is not the only kind of state. In one case that I actually touched on earlier, the AWS Lambda was taking a long time to reply to every request. It was accessing the database more than once, and they can say, “Okay, I don't mind if it's taking half a second to reply,” that's a type of state. Because as the request is executing, the Lambda has to keep whatever it is in memory and that drives up the cost for the same reason that having a large virtual machine drives up the cost. But with Lambda it's far worse. People ask, “Are serverless systems more or less expensive?” And the answer is that it's more expensive per unit resource, but if you keep it tight and stateless, it can be negligible cost for thousands, millions of, let's say, Lambda requests. So in the case of this customer, they had to figure out that they could compress multiple database requests into one and that accelerated the Lambda and meant that instead of taking half a second, it was taking a few milliseconds, and they radically reduced the cost by eliminating that state. Another example of state is in data processing pipelines. So one of my customers was using Dataflow, and that was taking input data, transforming it, outputting it, and it would take quite a long time to do that. Well, that's not a problem if it crashes so long as that pipeline can be restarted. It ran for, let's say, two minutes, it crashed, it started again. As long as everything is stateless and the job can be rerun, then that could be okay. Once you have state, once the processing pipeline is holding a lot of memory internally, or is taking far too long, it's taking in one case 20 hours, then that restart, which previously was negligible, is now a very expensive challenge. So this is true of data processing pipelines as well. State is your enemy. Keep state in the database and keep your processing units small.
RD All right. The enemy of the state, I guess, huh? You mentioned earlier that when you're looking for waste, you have some typical areas where you're like, “Oh, I know these are the places where there's going to be waste in a cloud.” Can you share a couple of them?
JF Sure. Okay, so I mentioned earlier that my process is to slice it down to attribution groups and look at the most expensive services. Let's imagine that in one case I see it's EC2 or as Google Compute Engine, and then I look at the most common reasons for wastage there. And again, cloud cost comes from waste, not from choice of service. People ask me, “Should I be using EC2 or should I be using EKS, ECS?” I say that there are pluses and minuses, we can talk about it, and one might be more expensive, but that's not really where the costs come from. The cost comes from blatant mistakes. So let's talk about compute. First of all, keep your instance types, disks or CPUs, up to date. So if Amazon moves from M4 to M5, they provide something new, Amazon is going to pressure you to move to the new one, and pressure in this case means cost, or to look at it a little bit more cheerfully, the new one is cheaper. Good. So migrate to the newest thing, of course, once you've determined that it's stable. Use spot instances in both Google or Amazon. Spot instances are virtual machines that the cloud provider can kill with very little warning at any time, and that sounds scary, but they're way cheaper. They can be 90% cheaper. And in order to make that possible, you need statelessness. Yes, I said it before. You need a situation where you don't really care if a virtual machine is killed and replaced by another because it has no identity. It has no state. That's completely possible. When I talk to my customers and some of them say, “Oh, no problem using spot instances,” I say, “Besides the fact you're saving 90%, I know you have a good architecture.” Vice versa the ones who say that they're scared because the cloud provider might kill their instance. I say, “Well, maybe you have other issues in architecture that you should look at.” Another area of cost saving is commitment. Let's say you pay for three years in advance, or you promise to pay for three years in advance, one year in advance, you can save a lot of money. Now we're all a little cautious about commitment in everything in life, but in this too. I've seen companies that were sure they were going to grow, and actually they did grow, but that department unfortunately had to be cut back. So that commitment maybe was a waste. So one approach is Flexsave, which we at DoiT provide Flexsave within AWS, and that means that we take on the commitment, the risk, and you get your virtual machines from us as if you'd committed to a year but with no commitment. So those are a set of easy ways to save money in virtual machines. There's also a hard way which I don't usually discuss with my customers, and that's making sure your application code internally is well built. If your Java code, your Python code is creating data structures that use up way too much memory, or if you're doing exponential complexity algorithm, then that is going to also increase cloud costs. I'm happy to discuss it, but usually with my customers we're talking about the cloud infrastructure level itself.
RD Yeah, the regular code optimization stuff like spotting memory leaks and The Big(O) Notation.
JF Very important.
BP Josh, it was great to have this conversation with you. You have so much experience in the field and now it seems working with customers across all of the new things that are coming down the pipe. As you look out to the future, are there things you're particularly excited about or challenges that you think folks who are listening should be especially keen to?
JF Yes, I think the future of cloud cost optimization is that it's going to happen more than it did in the past, because in the past, customers were excited by the flexibility of just spinning up new compute resources whenever needed and not waiting for one to be shipped from the factory. But now they realize that they have to cut back on that tremendous waste that sometimes happens, and we at DoiT help them with that. We advise them, we have tools, and so the customer is successfully reducing their cost and is going to see more of that. But on the other hand, costs are going to grow as new services consume massive resources. Machine learning I mentioned earlier, it’s insane how expensive it can get, and that's just going to happen. Costs are going to go up and new optimization techniques will be found, but the customers are going to have to deal with the greatly increased costs. There's a rule in computing that whatever looked like a waste in the last generation looks completely normal. I remember in ‘96 when I saw a one-inch video going through the university network. And I saw this video and I said, “What a waste of bandwidth. It's going to crash the network.” Well, nowadays, of course we stream hundreds of hours of YouTube across the world, and that's completely normal. So this is going to happen with the cloud. There will be a lot more usage of resources and we're going to always have to think about optimizing the cost.
BP Definitely.
[music plays]
BP All right, everybody. Thanks so much, as always, for listening. As we do at the end of every show, I want to shout out someone who came on Stack Overflow and helped to share a little knowledge. A Lifeboat Badge was awarded to Sravan K: “How to sort file lines in Bash.” Sravan has the answer for you and has helped 14,000 people, so we appreciate it and congrats on your Lifeboat Badge. As always, I am Ben Popper, Director of Content here at Stack Overflow. You can find me on X @BenPopper. If you want to come on the podcast or chat with us, hit us up, podcast@stackoverflow.com. And if you enjoyed today's discussion, leave us a rating and a review.
RD I'm Ryan Donovan. I edit the blog here at Stack Overflow. You can find it at stackoverflow.blog. And if you want to reach out to me on X/Twitter, you can find me @RThorDonovan.
JF And this is Joshua. I've enjoyed this conversation, and you can find me at joshuafox.com and be in touch.
BP Thanks so much for listening, and we will talk to you soon.
[outro music plays]
