1 00:00:09,446 --> 00:00:13,046 Lee: Hello and welcome to Software Architecture Insights, your go-to 2 00:00:13,046 --> 00:00:17,186 resource for empowering software architects and aspiring professionals 3 00:00:17,456 --> 00:00:21,686 with the knowledge and tools they require to navigate the complex 4 00:00:21,686 --> 00:00:24,476 landscape of modern software design. 5 00:00:26,550 --> 00:00:29,929 Last month, I wrote about a 3:00 AM page. 6 00:00:30,619 --> 00:00:34,399 A payment service was failing, and the engineer who answered 7 00:00:34,409 --> 00:00:36,399 the page never touched it. 8 00:00:37,379 --> 00:00:40,959 The next hour was spent finding someone who actually knew how to fix it. 9 00:00:42,180 --> 00:00:46,160 That article was called "When Everything is Critical, Nothing Is." 10 00:00:46,980 --> 00:00:50,449 Today, I want to go further than the article had room to discuss. 11 00:00:50,820 --> 00:00:55,310 The article told you what's missing: two decisions, which 12 00:00:55,310 --> 00:00:58,020 services matter and who owns them. 13 00:00:59,050 --> 00:01:03,330 What it didn't tell you is how to make those decisions, and that's 14 00:01:03,330 --> 00:01:05,490 where most teams tend to get stuck. 15 00:01:06,590 --> 00:01:11,530 So that's this episode, how you add a service tier to your services, 16 00:01:11,760 --> 00:01:16,420 and how you make ownership stick. Let's start with a quick recap 17 00:01:16,420 --> 00:01:17,810 in case you missed the article. 18 00:01:18,210 --> 00:01:21,470 By the way, if you'd like to read the original article, there's a link to 19 00:01:21,470 --> 00:01:27,180 it in the show notes. When an outage runs long, we usually blame the 20 00:01:27,180 --> 00:01:30,690 technology or we blame communication. 21 00:01:31,840 --> 00:01:34,380 But look at where the time actually goes. 22 00:01:34,700 --> 00:01:38,320 In a story from the article, the actual problem took 11 minutes to 23 00:01:38,320 --> 00:01:43,630 find, and the other 62 minutes went to finding a person responsible. 24 00:01:44,690 --> 00:01:47,940 No technical fix would have saved those 62 minutes. 25 00:01:48,580 --> 00:01:51,780 What was missing were two key decisions. 26 00:01:52,620 --> 00:01:54,620 The first is criticality. 27 00:01:55,770 --> 00:01:58,440 Which of your services matter the most? 28 00:01:59,399 --> 00:02:01,100 The second decision is ownership. 29 00:02:01,760 --> 00:02:07,300 For each service, which one team, one single team, is accountable for 30 00:02:07,309 --> 00:02:08,940 keeping that service up and running? 31 00:02:10,160 --> 00:02:15,730 Now, most organizations have made neither decision, or even worse, they 32 00:02:15,739 --> 00:02:20,330 think they've made the decision, and the answer on paper doesn't really match what 33 00:02:20,330 --> 00:02:22,400 really happens at 3:00 in the morning. 34 00:02:23,400 --> 00:02:25,640 So let's start with service tiers. 35 00:02:26,770 --> 00:02:30,760 I've used a four-tier model for a long time now. 36 00:02:31,539 --> 00:02:34,160 I wrote about it in my book, "Architecting for Scale." 37 00:02:34,980 --> 00:02:37,019 Here's the short version of that. 38 00:02:38,310 --> 00:02:42,480 A Tier 1 service is a service where if it goes down, the 39 00:02:42,480 --> 00:02:44,310 business is hurt right away. 40 00:02:45,110 --> 00:02:50,990 Customers can't buy, they can't log in, and money stops flowing into the business. 41 00:02:51,949 --> 00:02:56,269 A Tier 1 service outage means the building is on fire. 42 00:02:57,269 --> 00:03:01,530 A Tier 2 service is a service where failure hurts, but 43 00:03:01,530 --> 00:03:03,510 the business can keep going. 44 00:03:04,810 --> 00:03:06,799 Customers will notice something is wrong. 45 00:03:06,809 --> 00:03:11,289 You know, maybe search is slow or not working, or recommendations are 46 00:03:11,290 --> 00:03:16,759 missing from products, or something like that, but things generally still work. 47 00:03:17,319 --> 00:03:20,689 And most importantly, the business can keep moving forward. 48 00:03:20,779 --> 00:03:22,259 Money is still being made. 49 00:03:23,159 --> 00:03:28,240 Commitments are still being kept. A Tier 3 service is a service 50 00:03:28,470 --> 00:03:31,940 whose failure most customers wouldn't even notice right away. 51 00:03:32,150 --> 00:03:34,110 Maybe eventually, but not right away. 52 00:03:35,300 --> 00:03:39,999 Something in the background, like an email digest or a report that runs 53 00:03:40,000 --> 00:03:41,769 overnight, or something like that. 54 00:03:41,830 --> 00:03:45,109 Those are examples of Tier 3 services. 55 00:03:46,109 --> 00:03:51,959 A Tier 4 service is internal and doesn't touch customers at all. 56 00:03:52,910 --> 00:03:58,400 If a Tier 4 service goes down, somebody on your team might get annoyed, but that's 57 00:03:58,400 --> 00:04:00,409 really the full extent of the impact. 58 00:04:01,790 --> 00:04:04,249 Now, the definitions are the easy part. 59 00:04:04,799 --> 00:04:08,109 Most teams can agree on those definitions in about 10 minutes. 60 00:04:08,259 --> 00:04:12,199 The hard part, though, is applying them to your application. 61 00:04:13,579 --> 00:04:18,079 In this article, I described an exercise where 53 services got sorted, 62 00:04:18,559 --> 00:04:21,289 and 48 of them came back as Tier 1. 63 00:04:21,980 --> 00:04:25,459 I've seen versions of that sort of problem occurring more than once. 64 00:04:26,709 --> 00:04:29,570 And I want to be fair to the people that are doing it because 65 00:04:29,669 --> 00:04:31,459 nobody's being dishonest here. 66 00:04:32,360 --> 00:04:35,689 Each team gets asked, "Is your service important?" 67 00:04:36,059 --> 00:04:39,999 And every team says, "Yes, of course," because it is 68 00:04:40,000 --> 00:04:42,529 important, at least it is to them. 69 00:04:43,529 --> 00:04:45,589 The problem is the question. 70 00:04:46,429 --> 00:04:47,389 Is it important? 71 00:04:47,409 --> 00:04:49,529 We'll always get a yes answer. 72 00:04:49,709 --> 00:04:54,539 After all, if a service isn't important, well, why does it exist at all? 73 00:04:55,509 --> 00:04:56,969 So change the question. 74 00:04:57,809 --> 00:05:01,299 Here are three questions that I recommend asking instead of 75 00:05:01,309 --> 00:05:03,129 the is it important question. 76 00:05:04,259 --> 00:05:09,319 Start with this one: What happens to a customer in the first 77 00:05:09,319 --> 00:05:11,389 hour this service goes down? 78 00:05:12,189 --> 00:05:14,849 Not the first day, the first hour. 79 00:05:15,889 --> 00:05:20,139 If the answer is they can't complete a purchase, then this 80 00:05:20,139 --> 00:05:21,889 service is a Tier 1 service. 81 00:05:22,929 --> 00:05:28,179 But if the answer is something more like a nightly report is late, then 82 00:05:28,179 --> 00:05:33,470 this most definitely is not a Tier 1 service. Then ask this question. 83 00:05:34,470 --> 00:05:38,340 If this service and the checkout service both fail at the exact same 84 00:05:38,340 --> 00:05:41,680 time, which one do you fix first? 85 00:05:42,680 --> 00:05:44,270 Everyone knows the answer. 86 00:05:44,580 --> 00:05:45,630 The checkout service. 87 00:05:45,730 --> 00:05:49,289 The checkout service is probably one of your most critical services, period, if 88 00:05:49,289 --> 00:05:51,710 you're a e-commerce application at least. 89 00:05:52,619 --> 00:05:59,539 Because without the checkout service, the business itself is dead. But the moment 90 00:05:59,569 --> 00:06:04,589 you ask the question out loud, the moment you ask how does your service compare to 91 00:06:04,589 --> 00:06:11,429 the checkout service, you've ranked the two services against each other. And the 92 00:06:11,429 --> 00:06:17,949 last question: would you pay for a second cloud region for this service in order 93 00:06:17,949 --> 00:06:22,609 to improve availability, or would you put an engineer on call for it overnight? 94 00:06:23,749 --> 00:06:27,439 Making a service Tier 1 comes with a price tag. 95 00:06:27,569 --> 00:06:32,959 It costs money to duplicate a service for availability, and on-call engineers are 96 00:06:32,959 --> 00:06:35,089 expensive to your other projects as well. 97 00:06:36,239 --> 00:06:40,349 If nobody's willing to pay for those things, well, then the service just 98 00:06:40,469 --> 00:06:43,059 plain can't be a Tier 1 service. 99 00:06:44,059 --> 00:06:49,359 That last question changes the conversation completely because every tier 100 00:06:49,360 --> 00:06:55,020 assignment is now a budget decision. One more thing that helps is putting a limit 101 00:06:55,120 --> 00:06:58,890 on the number of Tier 1 services you're allowed to have in your application. 102 00:06:59,300 --> 00:06:59,970 Pick a number. 103 00:07:00,270 --> 00:07:02,789 Maybe it's five, maybe it's eight, maybe it's fifteen. 104 00:07:03,490 --> 00:07:06,790 The right number absolutely depends on your application. 105 00:07:07,429 --> 00:07:09,409 But pick one and write it down. 106 00:07:10,569 --> 00:07:16,479 Then, if a team wants to add a new service to the Tier 1 list and it, 107 00:07:16,519 --> 00:07:20,309 the list is already full, something else has to come off the list. 108 00:07:21,589 --> 00:07:26,270 Or they have to make the case in front of everyone why the 109 00:07:26,280 --> 00:07:27,780 limit needs to be increased. 110 00:07:29,140 --> 00:07:34,670 Now, this may sound bureaucratic, but in practice, it moves the ranking 111 00:07:34,670 --> 00:07:39,579 conversation into a meeting on a Tuesday afternoon, which is a whole 112 00:07:39,579 --> 00:07:43,769 lot better than a bridge call at three o'clock in the morning during a crisis. 113 00:07:44,769 --> 00:07:47,339 You're going to rank your services either way. 114 00:07:47,769 --> 00:07:52,609 You can do it calmly ahead of time, or you can let whoever answers the 115 00:07:52,609 --> 00:07:54,499 page do it in the middle of the night. 116 00:07:55,039 --> 00:07:56,009 It's your choice. 117 00:07:57,009 --> 00:08:01,320 Now, something this article didn't get into, and it trips up almost 118 00:08:01,399 --> 00:08:06,529 every team that tiers for the first time, and that is dependencies. 119 00:08:07,979 --> 00:08:10,850 Say the checkout service is a Tier 1 service. 120 00:08:11,740 --> 00:08:15,849 Now, checkout calls a tax calculation service. 121 00:08:16,849 --> 00:08:21,729 Somebody ranked the tax calculation service as a Tier 3 service 122 00:08:22,249 --> 00:08:25,419 because it's small and nobody really thinks that much about it. 123 00:08:25,599 --> 00:08:27,639 It was a poor ranking, but that's what they came up with. 124 00:08:28,639 --> 00:08:31,619 But what happens when the tax service goes down? 125 00:08:32,459 --> 00:08:35,259 Well, in most cases, checkout goes down with it. 126 00:08:35,389 --> 00:08:40,029 Your Tier 1 service is only as reliable as the least reliable 127 00:08:40,039 --> 00:08:41,919 thing that it can't live without. 128 00:08:42,919 --> 00:08:45,009 So here's the rule I use. 129 00:08:45,669 --> 00:08:51,049 When a Tier 1 service, such as the checkout service, depends on a lower 130 00:08:51,079 --> 00:08:54,019 tier service, you have two choices. 131 00:08:54,899 --> 00:09:00,039 You raise the dependent service to be a Tier 1 status service on its own right, 132 00:09:00,769 --> 00:09:04,149 with everything that comes with that and all the costs that are associated with 133 00:09:04,149 --> 00:09:11,459 that, or you make the Tier 1 service, like the checkout service, able to survive 134 00:09:11,529 --> 00:09:18,189 without the dependency. Maybe checkout can use a cached tax rate for a few minutes. 135 00:09:18,829 --> 00:09:21,869 Maybe it lets the order through and calculates the tax later. 136 00:09:22,449 --> 00:09:25,989 Maybe it turns off one feature instead of failing the whole page. 137 00:09:26,989 --> 00:09:32,869 Either raise the tier of the dependency or make the dependency itself optional. 138 00:09:33,589 --> 00:09:34,769 Either choice works. 139 00:09:35,229 --> 00:09:40,549 The one that fails is a Tier 1 service quietly depending on a Tier 3 service 140 00:09:40,549 --> 00:09:45,589 as a necessity with nobody noticing until the night it really matters. 141 00:09:46,759 --> 00:09:50,459 When you do your first tiering pass, walk the dependencies 142 00:09:50,459 --> 00:09:51,939 of every Tier 1 service. 143 00:09:52,439 --> 00:09:56,019 That's usually where some of the surprises are going to show up. 144 00:09:57,419 --> 00:10:01,929 Okay, let's talk about ownership, the second question. 145 00:10:03,129 --> 00:10:06,879 Every service has exactly one owning team. 146 00:10:07,639 --> 00:10:10,849 The team is named, and the team is current. 147 00:10:11,849 --> 00:10:15,469 That's part of a principle that I call STOSA, Single Team 148 00:10:15,469 --> 00:10:17,029 Oriented Service Architecture. 149 00:10:17,609 --> 00:10:19,849 One service, one owning team. 150 00:10:20,749 --> 00:10:25,059 The rule is simple, but living it is actually quite a bit harder. 151 00:10:26,509 --> 00:10:31,879 Owner is one of those words people use and nobody really defines. 152 00:10:32,279 --> 00:10:35,179 So let's try coming up with a definition of ownership. 153 00:10:36,179 --> 00:10:40,929 An owning team is the team that gets paged when a service breaks. 154 00:10:41,859 --> 00:10:45,329 It's the team that has the service on its roadmap, so 155 00:10:45,399 --> 00:10:47,869 upgrades and fixes get scheduled. 156 00:10:48,889 --> 00:10:52,739 And it's the team that can change the service and deploy it without 157 00:10:52,780 --> 00:10:54,419 asking anybody's permission. 158 00:10:55,789 --> 00:10:58,999 Paged, planned, permitted. 159 00:11:00,219 --> 00:11:04,180 If a team has all three of those attributes, it owns the service. 160 00:11:04,660 --> 00:11:07,079 If it's missing any one of them, it doesn't. 161 00:11:08,359 --> 00:11:09,649 So here's a test you can run. 162 00:11:10,270 --> 00:11:11,699 Pick a service, any service. 163 00:11:12,590 --> 00:11:16,349 Ask the team that's listed as the owner three questions. 164 00:11:17,679 --> 00:11:21,220 First, when this breaks at 3:00 AM, does your phone ring? 165 00:11:22,599 --> 00:11:26,179 Is there work for this service in your plan for this quarter? 166 00:11:27,179 --> 00:11:30,819 Could you deploy a change to it today without asking another team? 167 00:11:32,160 --> 00:11:34,919 If the answer to all three of those questions is yes, 168 00:11:35,100 --> 00:11:37,170 great, you have ownership. 169 00:11:38,170 --> 00:11:41,380 But what you'll often find is two yeses and a no. 170 00:11:42,380 --> 00:11:47,360 The team gets paged, but they can't deploy without also calling in the platform team. 171 00:11:48,240 --> 00:11:52,340 Or they can deploy it, but it's never on their roadmap, so nothing 172 00:11:52,350 --> 00:11:55,170 gets upgraded until it breaks. 173 00:11:56,170 --> 00:12:00,239 That no is where your next long outage is coming from. 174 00:12:01,239 --> 00:12:05,239 The hardest ownership cases are inherited services. 175 00:12:06,269 --> 00:12:11,169 Something written by a team that no longer exists, or by one person who left. 176 00:12:12,409 --> 00:12:15,839 Nobody wants those, and I understand why. 177 00:12:17,049 --> 00:12:21,179 Taking ownership means taking the 3:00 AM pages for code that you didn't 178 00:12:21,179 --> 00:12:23,709 write and don't fully understand. 179 00:12:24,959 --> 00:12:26,869 So don't hand it over and walk away. 180 00:12:27,569 --> 00:12:30,439 If you're asking a team to own an inherited service, 181 00:12:31,109 --> 00:12:32,499 give them time to learn it. 182 00:12:33,259 --> 00:12:34,579 Put that time on the roadmap. 183 00:12:35,269 --> 00:12:38,719 Let them fix the worst of the problems that occur in a service 184 00:12:38,969 --> 00:12:40,929 before they're on the hook overnight. 185 00:12:42,029 --> 00:12:46,019 And if a service isn't worth that level of investment, well, 186 00:12:46,039 --> 00:12:47,569 that tells you something as well. 187 00:12:48,569 --> 00:12:53,339 Maybe it belongs in Tier 4 instead of Tier 3 or two. 188 00:12:54,319 --> 00:12:56,519 Maybe the service should be retired. 189 00:12:57,629 --> 00:13:03,389 And that's where the two decisions start working together. The article made this 190 00:13:03,409 --> 00:13:07,139 point, but I'll make it again because it's the heart of the whole matter. 191 00:13:07,869 --> 00:13:12,299 A tier without an owner is a promise nobody made. 192 00:13:13,129 --> 00:13:18,399 You can call checkout a Tier 1 service, and if no single team is accountable 193 00:13:18,399 --> 00:13:24,009 for that, nobody is going to defend it when a deadline shows up. An owner 194 00:13:24,019 --> 00:13:30,489 without a tier is a team defending everything at equal priority, which means 195 00:13:30,519 --> 00:13:32,449 really defending nothing in particular. 196 00:13:33,269 --> 00:13:37,569 They'll spend their effort on whatever broke most recently, not 197 00:13:37,589 --> 00:13:42,989 what is most critical. Put those two decisions together and each one 198 00:13:42,999 --> 00:13:44,639 makes the other one work harder. 199 00:13:45,209 --> 00:13:49,909 The tier tells the team how much reliability you can buy, and 200 00:13:50,029 --> 00:13:56,059 ownership gives them the authority and the obligation to buy it. So 201 00:13:56,309 --> 00:13:58,379 here's what I'd do this week. 202 00:13:59,379 --> 00:14:01,719 Take your 10 busiest services. 203 00:14:02,329 --> 00:14:07,099 For each one, write down the tier and the owning team from memory, 204 00:14:07,389 --> 00:14:09,029 not from the service catalog. 205 00:14:10,339 --> 00:14:11,879 Then check the catalog. 206 00:14:12,969 --> 00:14:20,019 Then ask the owning teams the three questions: paged, planned, permitted. 207 00:14:21,019 --> 00:14:24,409 For your Tier 1 services, walk their dependencies. 208 00:14:25,089 --> 00:14:29,269 Find the lower tier service your Tier 1 service just can't live without. 209 00:14:30,269 --> 00:14:33,999 You don't have to fix everything you find, but you'll know where your 210 00:14:34,019 --> 00:14:38,649 next 3:00 AM page is probably going to come from, and you'll know who's 211 00:14:38,649 --> 00:14:42,879 going to answer it. The original article is linked in the show notes. 212 00:14:43,449 --> 00:14:46,439 So is the ownership gap diagnostic that I mentioned there. 213 00:14:47,229 --> 00:14:50,689 It's a spreadsheet, a simple spreadsheet that you can download that walks you 214 00:14:50,689 --> 00:14:53,939 through all of this for your own services. 215 00:15:12,493 --> 00:15:15,403 Lee: Thank you for joining us on Software Architecture Insights. 216 00:15:16,003 --> 00:15:20,113 If you found this episode interesting, please tell your friends and colleagues. 217 00:15:20,533 --> 00:15:23,143 You can listen to Software Architecture Insights on all 218 00:15:23,143 --> 00:15:25,063 of the major podcast platforms. 219 00:15:26,083 --> 00:15:30,003 And if you want more from me, take a look at some of my many articles 220 00:15:30,003 --> 00:15:32,863 at softwarearchitectureinsights.com. 221 00:15:32,863 --> 00:15:38,763 And while you're there, join the 2000 people who have subscribed to my 222 00:15:38,763 --> 00:15:43,623 newsletter, so you always get my latest content as soon as it's available. 223 00:15:48,830 --> 00:15:51,680 Thank you for listening to Software Architecture Insights.