Presentation: The Incident Lifecycle: How a Culture of Resilience Can Help You Accomplish Your Goals

MMS Founder
MMS Vanessa Huerta Granda

Transcript

Granda: Incidents prevent us from meeting our goals. Your goal could be to sell all the tickets to the Taylor Swift concert. Shouldn’t be that hard, everyone wants to go. Or to get everyone home for the holidays without a major scandal. Again, shouldn’t be that hard, people are traveling during this time. Maybe you want to get goods shipped across the globe, or have everyone watch the last season of the really hot show that everyone’s been talking about. Then incidents happen, and they prevent us from meeting these goals even though it seemed like all the stars were aligned for this to happen. Incidents never happen in a vacuum. For every single one of these high-profile incidents, we can highlight a number of similar events that happened before. We can highlight a number of items that led to them happening the way that they did. For example, for the Southwest scandal, things go back to decisions made decades ago when the airlines were first being deregulated. In order to meet our goals to achieve them, organizations need to make investments. They need to make investments in a culture of resilience.

I’m going to highlight this with an example from my life. These are my kids, they’re almost 2 now. When they were first born, I’m very lucky, I didn’t have any major complications, even though there were two of them, and I’m 5′ 1″. I had a scheduled C-section at one of the top hospitals in the nation. Everything went great. The nurses, they took care of everything. I was even able to watch The Bachelor one night. Then we went home. If anyone here has had kids, you may have had a similar experience of chaos. I ran major incidents for a long time. I can tell you that this was the toughest incident I have ever experienced, because they just wouldn’t sleep. If one slept, the other one was crying. It’s not that I had kids thinking that they wouldn’t cry, and they would sleep perfectly, but it was just chaos. We tried troubleshooting them. It was like that incident that fixed itself. At 9 p.m., they just started crying, and then at 5 a.m., they were just feeling better, and they wanted to snuggle. We didn’t know what was up.

If we think of our lives, again, we didn’t have kids so that we could just never sleep again. We wanted to enjoy our babies at some point, and eventually, I was supposed to go back to work after my leave was over. Those first few weeks, I wasn’t doing any of that. I was lucky if I was able to shower. We were sleep deprived, and we were doing the best that we could, and it just kept happening. Until, and here I’m acknowledging my privilege, we started being able to get on top of it, because we started investing our time, our expertise, and our money into it. First, we had to get better at the actual response, those crazy hours from 9 p.m. to 5 a.m. The problem is that we were sleep deprived, so we tried different things. We tried some shifts. We tried some formula, so my husband could actually let me sleep. We invested in a night nanny. With this, we were able to get some time to just start getting better.

We could have stayed here but I don’t have unlimited money and night nannies are expensive. With the extra energy, we started trying to understand what was going on. We talked to our friends. We read books. We just got better at being parents. Every day, we would just hold a post-mortem, we would talk about what had happened and what could go better, and trying to think how we could apply those learnings into our lives. We fixed some processes around swaddling. We figured that Costco actually delivers formula so we didn’t have to grab our two car seats in the middle of winter in Chicago and go to Costco every time we needed formula. We actually realized how much formula and diapers we were going through. Infant twins go through a lot. We were able to figure out our capacity better. That gave us more energy, more bandwidth. We started looking at the trends.

All incidents are different. All babies are different. We realized what worked and what didn’t work. We figured out that there were times that one of us could handle both babies, but there were times when it was all hands on deck. I later came to learn that that happens quite a bit for infants. We were doing our cross-incident analysis and we started coming up with some ideas. Maybe they were a little out there compared to other people, but we had some data points that made us believe that we were going in the right direction, we might as well try them. We were like, let’s just drop the swaddle, she clearly hates it, she wants to be free. Maybe we can move them to their own room, like that could be fine. Eventually, things got better by investing time and effort into the process, then having our post-mortems. Then that cross-incident analysis. It gave us the bandwidth that we needed to meet our goals. By the time that my maternity leave was over, I was able to go back to work. I was able to enjoy being back in the world. I was able to enjoy being a mother. This was a couple of months ago. I was able to cook something more than just cereal. That was a very long example. That’s what we’re going to talk about here, the lifecycle of incidents, and how we can learn from them, and how we can get better as a system so that we can accomplish all of the goals that we have.

Career Journey

I’m Vanessa. I’m an Engineering Manager at Enova, where I manage the resiliency team. When I’m not wrangling adorable twins, I spend a lot of time in incidents. I have spent the last decade in technology and site reliability engineering, focusing on incidents. When I say incidents, I mean the whole lifecycle of it. I’ve done many roles. I’ve been the only incident commander on-call for many years. I’ve owned the escalation process. I’ve trained others. I’ve ran retrospective programs. Most importantly, I have scaled them. Scaling is actually the hardest thing when it comes to doing this work. While I do this for fun, and because I like it, I also know that the reason that I get paid to do what I do, is because having a handle on incidents allows companies to do the work that they need to do. It allows businesses to sell those Taylor Swift tickets, to get people to String Succession. All of those things. Previous to my current job, for the past couple of years, I had the chance to work at Jeli.io. I work with many other companies, helping them make the most out of their incidents, helping them establish and improve their incident programs. I have seen firsthand and from a consultative side, how engineering teams are able to focus on resilience and achieve their goals.

Incident Lifecycle

Let’s go back to this picture. We will never be at zero incidents. We’ll never be at zero incidents, because we will continue doing work. We continue releasing code. We continue making changes. We continue coming up with new products. That’s a good thing. I believe that incidents are a lifecycle and they’re part of how a system gets work done, how the system achieves their goals. First, you have a system, and then something happens. Maybe you announce a new product, maybe you announce a new album, a new TV show, a new concert, and tons of people try to access that system, and then the system fails. Now you’re in an incident. You work on it. It’s resolved, and life is back to normal. Ideally, you do something after the incident, you do something related to learning. Even if you don’t think that you’re doing any traditional learning, your company doesn’t do post-mortems, your company doesn’t do write-ups, folks naturally like to talk and debrief. The learning activity, it can be a chat with your coworkers while getting drinks after work. It can be like, “I’m slacking my best year.” We’re going to be like, that sucked. It can be your post-mortem meeting, your incident report, whatever it is that you want it to be. Those learnings are then applied back into the system. Maybe we decide that we need to make a change to our on-call schedule. Maybe I’m just gossiping with my boss, and we’re like, I’m just never going to trust that person again. Maybe we change the way that we enqueue our customers. In the Taylor Swift example, maybe we have a congressional investigation around antitrust legislation. Even if we don’t make any actual physical changes, just the fact that we have the experience of the incident in our brains means that we have changed the system, because the system is not just your codebase. Your system is also that sociotechnical system, it’s also us, the people who are working with the technologies. That is the new system now, and that’s why it’s the incident lifecycle.

Why do we care about this? Why am I here on this talk? It’s because these things are expensive for companies. Let’s think about the cost of an incident. You have the incident itself, “We lost money because people couldn’t access our site for 20 minutes.” There’s the reputational damage, “I’m never going to trust this website again, because I wasn’t able to rent a car from this website.” There’s the workload interruption. The incident that lasted an hour, 2 hours, you had 10 engineers involved. It’s not like after the incident these engineers are going to go back to their laptop and their keyboards and focus on their OKRs. Incidents are a huge interruption and takes engineers a lot of time to go back to what they were doing. Then it has the impact that it has on our goals and our planning. If engineers are fighting fires all day, they’re not working on your features. When I was fighting the fires of my children, I wasn’t cooking, I wasn’t working out, I wasn’t really doing anything else. What we see is that unless we do something about it, people end up getting caught in this cycle of incidents where there’s no breathing room to get ahead. The impact is not just the incident, and the impact is not just the engineers that are working on it. An incident impacts the whole team, the whole organization.

There is some good news about incidents. Your outages mean that somebody cares about the work that you’re doing. That it matters, whether you’re up and down. That it matters that you’re doing the things that your organization is there to do. That’s the thing about working in technology. Our users don’t care about the languages we use, the methods we use. They don’t care if we do retrospectives, or if we use Jira, or Slack, or Notion, or whatever it is that you’re talking about. They just don’t care. They care that they’re able to rent cars. They care that they’re able to get home for the holidays. They care about being able to make a bank transfer, get tickets to a concert. Those are the things that matter to them. We owe it to our users that they’re able to do those things. Here’s where resilience helps. Resilience is about having the capacity to withstand or to recover quickly from difficulties, to recover from your outages, from your incidents. Resilience can help us turn those incidents into opportunities.

You’re here for a reason. You care about this stuff. We care about improvement. That’s not how things usually work. I don’t know if you’ve had the experience to work at an organization where things don’t work that way but I certainly have. Usually, what happens at some organizations is that you have an incident, you’re down. You resolve it, and then you move on. You have like my friend here, Chandler Bing. This is how folks who are stuck in a cycle of incident work, not because they want to. I have tons of engineers in my life. I’m one myself, my dad, my brother. We just like fixing things. We don’t like seeing something that’s broken. Sometimes we just don’t have the bandwidth to do anything else other than move on.

Resilience in Incident Analysis

There’s a better way. I’ve been lucky enough to work at and work with a number of organizations that are leading the way in doing this work, so that’s what we’re going to get into. This better way includes focusing on three different things. One, focusing on the incident response process. Focusing on learning from your individual incidents, that’s your post-mortems, your incident reports. Then, focusing on macro insights. You don’t have to do it all. You can if you want to. It’s often hard to find an organization that’s going to give you the bandwidth to do it all at once. If your organization doesn’t even have a process for incidents, going up to the CTO and saying like, “I want to do cross-incident analysis, give me 100k.” That’s going to be hard. Usually, you have to prove your way along the way. We’re going to go through these three ways to improve resilience in the incident lifecycle. There are some caveats. There are some things that can go wrong. That’s because a lot of this requires skills that are just not traditionally engineering skills. This work is difficult. It takes time. It takes investment. It requires selling. It requires telling other people that this is the right thing to do, and making them feel good about it. Selling isn’t really your best skill as technologists. It’s not my best skill as a technologist. Instead of calling it selling it, we should think of it as the scientific method. We should think that we’re trying things out, we’re presenting our data at work, and we’re making our case. The good news is that I have seen this work, and we can get many small wins along the way.

A Focus on Incident Response

What does it look like to focus on incident response? It makes sense to start here at incident response. That’s the thing that we’re already paying for. We’re already spending time in incidents, so we might as well focus on those things. We might as well focus on coordination, collaboration, and communication, which is usually what makes up an incident. On the coordination front, I recommend people write up their current workflow, even if the workflow is just call that person that has been here for 10 years. Ideally, you have something more defined than that. Then ask yourself, how do folks come together to solve a problem? Look for any gaps, write them down. What can be done to make those gaps even just a little bit smaller? Maybe people are just not communicating their insights, so maybe a virtual whiteboard, or maybe getting them on Zoom, anything like that.

On the collaboration front, how do you get people in the room? How do you know who to call? I had an organization that I worked at that actually, whenever I would have incidents if I was at home, I would literally have to open up my laptop. Get on the VPN, get on the wiki, look up a team, look up that person, and get their phone numbers, and then grab my iPhone 4. That’s a lot of work. That’s not great. There’s improvements that can be made there. We can ask ourselves, are we always calling the same people? Are we burning them out? Take a look at your on-call rotations and your expectations. What little things can we do to make their lives a little bit easier? Then the communication front, how do you communicate to stakeholders what is happening? How do you tell your customers what is happening? There’s a difference between those two groups. There’s a difference between the responders as well. What’s hard about it? What are they asking for? Write up some loose guidelines to help manage those expectations from folks both inside and outside the immediate team that’s responsible for responding.

Why aren’t people doing this? They’re not doing this because it requires work. You need to train people on this process. This process is different than the way we do traditional engineering work. When you’re doing normal, everyday feature work, everyone has different ways of working. During an outage, we need to get services back up ASAP. There are some certain procedures that make sense here. You want to teach folks the skills that help move incidents forward. You’re not focusing on hierarchy. You’re not focusing on roles. You’re not blaming people or anything that happened. You want them to know about the roles that make a difference. Communication, having somebody in charge of coordination. Understanding who’s a stakeholder versus who is a responder. Don’t waste your time trying to explain to a stakeholder, things that a responder needs to know. Oftentimes, these coordination tasks are what leads to the highest cognitive load. Oftentimes, focusing on the response is more work than actually fixing something. These things can be automated, or they can be achieved through tools. Sometimes the people working on these teams don’t have the right tools for it, or they want to create these tools, they want to build it themselves, but they just don’t have the bandwidth to develop them. Again, they’re overloaded with incidents so they’re stuck fighting fires.

How have people been able to do this in the real world? A lot of them have focused on those two things: the process, and they’ve leveraged automation and bots. I can highlight the work of my friend, Fred Hebert at Honeycomb. He wrote a great blog post on their incident response process. He says that as they were growing as a company, they were running into some issues with their incident response process. He wrote that while they’re trying to ramp up, they’re trying to balance two key issues. They’re trying to balance avoiding the incident framework being more demanding to the operator than solving the issue. That’s what I talked about earlier. They don’t want the process to be more difficult than the engineering work itself. They’re trying to provide support and tools that are effective to the people who are newer to being on-call for whom those guidelines are useful. That’s that training thing. You want to make it easier for people onboarding. What they did, and I worked with them on this process, was to automate some tasks using the Jeli incident bot. They were automating the creation of incident channels. They were communicating status updates to stakeholders. Doing that left the engineers the time to focus on the actual engineering part.

The idea is that you get quick wins even a week into the improvement process. You can restructure on-call rotations. You can automate the response task. All of that allows us to spend less time on the tasks that demand our attention during engineers, but aren’t necessarily engineering tasks, because we never ever want to automate the creativity or the engineering skills that come with responding. We do want to reduce that cognitive load during a high stress time. Doing these little things can help build momentum and get buy-in for the larger changes that you’re going to make. If I show to my stakeholders that incidents are just easier, because we’ve done some automations and because we’ve set some process, then they’re more willing to help me out with the next things I want to do. A focus on the incident response part is going to help improve the process itself. It’s going to lead to a better experience for your users and your engineers that are going to focus less time on the repetitive task. It’s going to lead to easier onboarding for on-call, more efficient process, and lower impact to your customers.

Anti-Pattern: MTTX

Here is where we’re going to talk about an anti-pattern when it comes to incident response, and it’s this idea of MTTX. MTTX is mean time to discovery, mean time to recovery, to resolution, whatever it is that you want to call it. I am on the record on saying that those numbers don’t mean anything. That doesn’t mean that I think that we should just not care at all about how long our incidents last. I just think that a single number should not be the goal. One time I had an issue where our data center caught fire, and the fire marshal wasn’t going to let us go back on. That’s just going to take longer than reverting code. We want to make the experiences for our users and our engineers better, because that’s actually what’s going to help us get ahead. It’s what’s actually going to help people get online and get home for the holidays, and then buy their tickets and all of that. Yes, there are things that we can do to make incident response easier, so that our engineers are better equipped to resolve them. Just like we’re never going to be at zero incidents, we are not in control of everything.

A Focus on Incident Analysis

The next way that we apply resilience is when we focus on incident analysis. After you have an incident, you want to learn about it. I believe that the best way to do this is through a narrative-based approach, where you get to highlight what happened, how folks experienced the incident from their different points of view. After all, if I’m telling you the story of an incident from my point of view, as the subject matter expert, that experience is going to be different from the point of view of someone who is new to the organization, or from the point of view of somebody in customer support, people who are still interacting with our systems and are still impacted by the incident. What we’ve seen is that template filling sessions, root cause analysis, Five Whys of the past, while they were helpful in the past, today, they are not as much. They can actually sometimes cause harm, because they give us this false sense of security that we’re doing something. I’ve seen this many times. You have a root cause analysis session and we say that the root cause is human error, so let’s just not do it again. Maybe we put on an alert for that specific thing, because if that very exact same thing happens again, we’ll be prepared to handle it. That’s not a very good action item to come out of a meeting. If we have a blame-aware learning focused review, we can highlight that the on-call engineer didn’t have access to some of the dashboards to understand what was going on and that’s why it took longer to resolve. Or maybe when we were releasing this code, the QA process didn’t account for some new requirements, and that shows a disconnect between teams that maybe engineering and marketing weren’t on the same page and didn’t realize that there was a capacity need. Those are things that we can learn from, and that can actually help move the needle when it comes to resilience.

Why aren’t folks doing this right now? It’s not that folks aren’t doing this work, it’s just that they could be doing it better. They’re not doing this because creating a narrative timeline can take a lot of work. There’s a lot of copying and pasting. I don’t know if you’ve ever created a timeline, but sometimes you have three screens, and you got Jira. You’re summarizing conversations from Slack, and you’re just copying and pasting. It’s hard to do that. Then, if you’re not doing that, then you’re going into your post-mortem, and that meeting then becomes the meeting where you create that timeline. You’re spending 50-minute meetings with 10 engineers, just creating a timeline, and maybe you leave 5 minutes to discuss interesting tidbits. That’s not a very good return on your investment, and it ends up not being worth it. If you want to learn about how to do incident analysis that actually moves the needle, there’s a number of sources out there including the Etsy Debriefing Guide and the Jeli Howie guide. It’s a free guide. It helps you understand how an incident happened the way that it did. It helps you understand who should be the person in charge of leading the review. What data you can collect. You have your Git, Jira, Slack, whatever you want to call it, how to lead an incident review, how to finalize, how to share your learnings. It’s a lot of work.

I’m just going to give you my short version here. I recommend that you make it your own. When you’re making it your own, there’s four key steps. First, you want to identify and focus on your data. You want to focus on who was involved in the incident. Where in the world were they? Does the incident take place in Slack or Zoom? What do you have going on? Based off of this, you can decide, we should have a meeting to review this, or, let’s just create a document and all collaborate asynchronously. Whatever you wanted to do, you make it your own. You still have to prepare for it. That means that you take the data that you have, you create your narrative timeline. You jot down any questions that you have, and that you’re inviting the right folks. So often, we have these retrospective meetings where it’s like you and the person who was the incident commander, and then maybe one other person that released the code. That’s not leading to collaboration. Invite people in marketing. Invite people in product management. Invite people in customer support. You will learn so much from inviting people in customer support. When you meet with them, you want to make sure that it’s a collaborative conversation. If I’m leading a review meeting, it should never be the Vanessa show. Others should actually be speaking more than me, because we want to hear how they experience it. That’s how we know that the on-call engineer didn’t have access to the things. That’s how we know that the person in customer support was getting inundated with calls and was dropping things. Then you can finalize your findings in a format that other people can understand. Maybe your findings can be a really long report, or maybe they just are an executive summary and highlight action items. That’s ok. You have to think of your audience.

Because I talked so much about scaling this work, we want to make it as easy as possible for people to do this. Doing this narrative-based post-mortems, it can take a lot of time, but the more you do it, the easier it becomes. It becomes more natural. The idea is that you take all your Zoom, your Slack transcripts, your PRs, your dashboards, and start telling the story of your incident. You don’t tell the story in the way of like, we release something, we reverted it back to normal. You tell the story of like, maybe three years ago, we acquired this other company and that’s part of the reason why we have some legacy systems. It’s really important to give people the ability to create these artifacts, so they can share them with others and allow folks to tell the story from their own points of view. The narrative can then inform action items for improvements. These items reflect what we learned about an incident’s contributing factors, as well as their impact.

We’ve seen this happen in the real world, where folks did this narrative-based post-mortem. They talked to people involved, and they realized that oftentimes, some responders don’t have access to all the tools or they have outdated information. That’s actual things that you can improve that can then help future responders in a future incident, so they can more quickly understand how to solve an issue because now they actually have access to the observability tools. In my current organization, actually, after an analysis that we did, we were able to understand just how vendor relations impact the way that an incident is handled, in order to improve the process. There was nothing we can do about the vendors, but even knowing how to contact them, even knowing what part of the process was them versus us, helps us get better at these incidents and just resolve them faster. A focus on incident analysis can help identify the areas of work that are going to lead to engineers working effectively to resolve issues. Then the next time that they encounter a similar incident, they will be better enabled to handle them, again, leading to lower customer impact and fewer interruptions. It’s all about making people’s lives better.

Anti-Pattern: Action Items Factory

When I talk about resilience, I want to make sure that I’m staying away from the anti-pattern of being an action items factory. The idea that you have an incident and a post-mortem, and all we do is play Whack a Mole, and you’re creating alerts for that specific issue. Or we have action items that are just fixes that never get prioritized, and they live in this Google Doc, forever to be forgotten. That’s not good, because it’s going to erode trust in the process. I’ve seen it firsthand. If the goal of an incident review is to come up with a fix, and I’m not here to argue that it is or it isn’t, but if it is, and then the fixes never actually get completed, then your engineers and your stakeholders are not going to take them seriously. Instead, they’re going to think that it’s a waste of time for a group that’s already inundated in outages. We’ve all been there. If I’m already inundated with work, the last thing I want to do is waste time in a meeting that’s never going to have any clear outcomes. Instead, I think we should look at the resilience coming out of incidents as part of different categories. Because action items today are flawed, so some of the problems are that they don’t get completed, they don’t move the needle, and they’re not applicable to large parts of the organization. Sometimes you encounter what I like to call the Goldilocks problem, where there are too small action items and you have too large action items and you just want the ones that are just right.

When it comes to too small action items, these are usually items that have already been taken care of, either during the incident or right after: a quick bug fix, a quick updating the database, sending an email letting people know. These items shouldn’t wait until the post-mortem in order to be addressed, because they’re just not worth spending too much time on them during the incident reviews. Folks should still get credit for them. People really like getting credit for the things that they do and you should still give them credit. They should still be part of your artifact. You should still include them in your incident report, we just shouldn’t focus too much time on them. I do think we should put them in the incident report, because if this happens in the future, you want to know what you did. These are just easy things that should be done with, don’t worry about them.

Then you have your too big action items. This is the thing that bothers a lot of people. They’re important, you want to include them in your incident reviews. There’s a problem with having them in the same category as the other things that I mentioned, because sometimes they take too long to complete. They can take years to complete. If I’m like, let’s switch out of AWS. That’s not an action item that just one person can take. If you’re tracking completion, which a lot of organizations, in order to close an incident report, you have to close all the action items, that’s going to skew your numbers, and nobody’s going to be happy. Why is this? Why are these items not going to be resolved? It’s often because the folks in the incident review meeting who come up with these items are not the ones who do the work or who get to decide that this work gets done. They are usually larger initiatives that need to be accounted for during planning cycles, so you need managers or even principal level folks to agree to do this and to actually take on it. There’s still important takeaways from post-mortems. There are still important things that I think we should include as part of the review, and should be highlighted. What I actually recommend is that you don’t call them action items, you call them initiatives, and you have them as a separate thing. Or, if you want to keep them as action items, change the wording of it. The wording can be to start a conversation or to drive the initiative. Other things that I think make a difference with getting these completed are having a due date and having an owner. If you’re at the review meeting, and you realize that the owner isn’t in the room, or people aren’t agreeing on what it is, or the due date, maybe that highlights that we should rethink of this as an action item.

Really, the perfect action item in my perspective is something that needs to get prioritized. Something that can be completed within the next month or two. Something that can be decided by the people in the room. Finally, something that moves the needle in relation to the incident. I really like this chart, because it shows the difference between all of these things that I’ve just talked about. It shows the difference between what a quick fix is, what a true action item should be, what a larger initiative is, and then, finally, what a true transformation can come out of these incidents and learning from them. I’m going to go through an example actually of how we can come up with a true action item. Here you can see that we’re creating the narrative of an incident, and we’re asking questions of the things that are happening. Here I’m asking a question about service dependencies and how they relate to engineers understanding information from the service map vendor. I’ve been in incidents like this, where you have a vendor for your service map, and they have a definition for something that is different from my definition, and then things just don’t add up. Here, you can see how that question that I asked during the post-mortem, made it into a takeaway from the incident. There’s a takeaway around what service those touches and what depends on them. Then, finally, when you’re writing your report, you can see that there’s action items around the understanding of dependency mapping. You can see here that you have action items that meet the requirements I talked about earlier, that are doable, have ownership, that have due dates, and then move the needle. You go from creating your timeline, to asking questions, to a takeaway, to an action item.

A Focus on Cross-Incident Insights

Finally, the last way in which we use resilience to accomplish our goals is to focus on cross-incident insights. Cross-incident analysis is the key to getting the most out of your incidents. You’ll get better high-quality cross-incident insights the more you focus on your individual high-quality reviews. Cross-incident analysis is a great evolution from the process that we just discussed. The more retrospectives that you do, the more data you have that you can then use to make recommendations for large scale changes. Those transformations or those initiatives that I showed in that chart. Doing this should not be one person’s job. We’ve talked a lot about the value of collaboration, about hearing different points of views. Usually when I do this, it’s in collaboration with leadership as well as multiple engineering teams, multiple product teams, marketing, customer support, principals. I want to engage as many people as possible. Here’s your chance to drive cultural transformations. Here’s your chance to say, based on all of these incidents, I think that we need more headcount for this team, or I think we need to make this large change to our platform, because you now have this breadth of data with context to provide to leadership for any future focused decisions.

Why aren’t people doing this now? People aren’t doing this now because they don’t have the data. Like I mentioned, high-quality insights come from individual high-quality incident reviews, and people still aren’t doing that. It’s also because whatever home you have for your incident reviews are usually not very friendly for data analysis. Oftentimes, what we see is people have their incident reports living in a Google Doc or a Confluence, and so you don’t have a good way to query those numbers out, is basically what I’m saying. The other reason is that engineers are not analysts. When we’re trying to get somebody to do this work, to drive it, you need to think like an analyst. You need to think about how to read the data, how to present the data in a friendly format, how to understand how to tell the story. This is all of this work that I’ve done at different organizations. It’s actually one of the hardest things.

Here’s an example of how cross-incident insights can come with context. Instead of simply saying, we had 50 incidents. We can take look at the technologies that are involved. We can say that, “This past year, Vault was involved in a majority of our incidents. How about we take a look at those teams?” Maybe you give them some love. Maybe you help them in restructuring some of their processes. Maybe you focus on their on-call experiences. We’re not shaming them, we’re helping them. I’ve done this with a number of teams, and they’ve been able to look at a number of their incidents in a given timeframe. They’ve been able to make organizational-wide decisions. Because focusing on cross-incident analysis can help identify large initiatives that allow our engineers to thrive and help our businesses achieve their goals. We’ve seen people use these insights to decide on things like feature flag solutions, or decide to switch vendors, maybe even make organizational changes. All of this creates an environment where engineers are able to do their best work.

Anti-Pattern: MTTX (Again)

Here’s this anti-pattern again, MTTX. Here’s the part where I will never tell an analyst or an engineer to tell their CTO that the metrics that they’re requesting are silly, and they’re just not going to do it. Mostly because they don’t get paid enough, and you shouldn’t fight that fight. I’m never going to tell an analyst to do that because of that. Also, because when you have leadership asking for these metrics, you should actually take this as an opportunity to show your work. Take this as an opportunity to show the awesome insights that you’ve gathered through your individual incident analysis and through your cross-incident analysis. I call this the veggies in the pasta sauce method, mostly because I grew up not eating many vegetables. The idea here is that you’re taking something that maybe you think doesn’t have much nutritional value, which is that single MTTR number, like mean time to resolution was 45 minutes this month. You’re adding context to make it much richer, to bring a lot more value to the organization. Like, incidents impacting these technologies lasted longer than the incidents impacting these other technologies. It can help you make strategic recommendations from the data that you have towards the direction that you want to go. A lot of the times they say that people doing this work shouldn’t have an agenda. I have never not had an agenda in my life. There’s one thing that I wanted to get done, and I find my way to get it done.

Anti-Pattern: Not Communicating the Insights

Which leads me to my last anti-pattern. It’s something that we’re not, again, super-duper good as engineers, which is communicating our insights. Because if we have insights and nobody sees them, then what is the point? In the Howie guide when discussing sharing incident findings, we explain that your work shouldn’t be completed to be filed, it should be completed so that it can be read and shared across the business, even after we’ve learned, even after we’ve done our action items. This is true for individual incidents. It’s even more true about cross-incident analysis. I say this because I’ve often been in a review meeting with engineers, and they’ve said, “We all know what we need to do, we just never do it.” I’m sitting there and I’m like, do we all know because nobody has told me? Have we actually communicated this? Often, when people don’t interact with our learnings, it’s because the format is unfriendly for them. The truth is that different audiences need to learn different things from what you’re sharing. If you’re sharing insights to folks in the C-suite, maybe your CTO is going to understand your technical terms, but then you’re going to lose most of your audience. When you’re sharing these insights, you should use language that your audience is going to understand. You should focus on your goal. Don’t spend too much time going through details that aren’t important. If you want to get approval to rearchitect a system, explain what you think is worth doing and what the problem is based on the data that you have, and then who should be involved, and leave it at that.

Focus on your format, whether that includes sharing insights, numbers, telling a story, sometimes technical details. What I like to do usually, is I like to think of my top insight. I will say something like, incidents went up last quarter, but usage also went up. SREs are the only on-call incident commanders for 100% of the incidents, so as we grow, this is proving to be unsustainable, I think we should change the process and include some other product teams in the rotation. This is how you actually get transformations happening from your individual incident analysis to your cross-incident insights, to actually communicating them and getting things done. Another example of this is if you say something like, x proportion of incidents are tied to an antiquated CI/CD pipeline, I said tied, not root cause. This impacts all of our teams across all of our products, and it’s making it hard to onboard new team members. Based on feedback from stakeholders and from experts, we recommend we focus next quarter on rearchitecting this. We’re not suggesting a 4-month long project just because you feel like it. You’re suggesting it because you have the data to back that up and you’re showing your work.

Going back to this chart. Instead of thinking of action items being the only output of your incident work, let’s think of this entire chart as a product of treating incidents as opportunities. Let’s think of this entire chart as the product of what we can accomplish when we put effort, when we put resources into the resilience culture of our organizations. Because doing this work is how we get ahead. We will never be at zero incidents. Technology evolves and new challenges come up, but we focus on incident response and incident analysis and cross-incident insights while looking out for those anti-patterns, we can lower their cost. We can be better prepared to handle incidents, which is going to lead to a better experience for our users. It’s going to lead to a culture where engineers are engaged, and they have the bandwidth not only to complete their work, but also to creatively solve problems to help us achieve the goals of our organizations.

Questions and Answers

Participant 1: [inaudible 00:42:25].

Granda: You’re asking about how to deal with people who are just bummers about the whole process?

Participant 1: [inaudible 00:42:54].

Granda: I work with those things. I think it’s a couple of things. One, we all know that when you’re trying to deprecate something, it’s probably going to take longer than you think it’s going to take, so just highlighting that. Then, I would not even focus sometimes so much on the technical parts of it, but maybe tell them that we’re going to focus on how we work as a team, on the socio part of that sociotechnical system. I definitely think it’s still worth fighting that good fight.

Participant 2: You spoke a little bit about the challenges with cross-incident analysis, which actually resonated with a lot of people here and other challenges we’ve already overcome. Are there any additional thoughts on some resources or strategies that you’ve seen or have been effective in practically overcoming some of those challenges?

Granda: I have seen people do this at a smaller scale. It’s hard to do cross-incident analysis at scale. I think the biggest issue on that, is just getting good data. What I presented was around training folks, and not expecting an engineer to do everything, especially something that they just have never done before. Focusing on maybe hiring an analyst for your team, or giving people the chance to work on those skills, giving some formalized training, is what I’ve seen as probably the most helpful. I know that there are some vendors out there that are helping with the repository of your data. That’s hard for a number of reasons, but I haven’t found the key to it yet.

See more presentations with transcripts

Subscribe for MMS Newsletter

By signing up, you will receive updates about our latest information.

  • This field is for validation purposes and should be left unchanged.

Presentation: Two Years of Incidents at 6 Different Companies: How a Culture of Resilience Can Help You Accomplish Your Goals

MMS Founder
MMS Vanessa Huerta Granda

Transcript

Granda: We’re going to talk about some incidents. Incidents prevent us from meeting our goals. You could have any goal that you want, maybe your goal is to get all of the Taylor Swift fans to make it to her concert. Maybe get folks home for the holiday seamlessly, or get goods shipped across the world, or just get people to watch Succession, or Game of Thrones, or whatever it is that HBO is doing nowadays. Sometimes things happen, and that prevents us from meeting our goals. Incidents never happen in a vacuum. For all of these high-profile incidents, we can usually highlight a number of similar experiences, or other incidents. We can also highlight a number of items that led to them happening the way that they did. For the Southwest outage, it goes back to decisions made decades ago when the airline industry was first being deregulated. My first job out of college was actually at an airline. All of these things are dear and near to my heart. In order for organizations to improve and to achieve their goals in spite of these incidents, there needs to be investment made on their end, investment in a culture of resilience.

Twins Analogy to Incidents

I have 17-month-old twins. They were born a bit early, like twins usually are. If you’ve had children, you probably know that while you’re at the hospital, everything is just like perfect. We gave birth in Chicago and we sent them to the nursery. I got to rest. I even got to watch The Bachelor actually. Then we took them home. I have been in incidents for many years. I can tell you that this was like the hardest major incident of my life. Every night for the first few weeks, they just would not sleep. If one slept, the other one was crying. It’s like, they were just playing tag team on me. We tried troubleshooting them. It was like that incident that just fixes itself at some point. At 9 p.m., they start crying, and then at 5 a.m., they were like, ok, now we’re cool. Basically, if we think of our lives, if we think of our goals, at that time in my life, my goal was to just enjoy my children, maybe keep a home, maybe get to eat, and then eventually go back to work. These first few weeks, I was lucky if I even got to shower. We were sleep deprived. We were doing our best, but it just kept happening until, and I’m really acknowledging my privilege here, we started being able to get on top of it. We were able to invest our time, our expertise, and really our money.

First, we needed to get better at the response, like at those actual nights when things were hitting the fan. We tried a few things, even though we were sleep deprived. We tried some shifts. We tried formula. We invested in a night nanny, and that was the big one. With this, we were able to start getting some time for ourselves. Being able to sleep actually cleared our heads for us to understand what was going on. We could have stopped here, but I didn’t have unlimited money for a night nanny. With some of that extra energy, I started trying to understand what was going on. I talked to some friends. I read some books. Every day, we would hold a post-mortem, just while eating yogurt or whatever it is that we could get our hands on. We would start applying those learnings into our lives. We fixed some of our processes around swaddling. We realized that Costco actually delivers formula and diapers, and so that’s a big-time savings not having to get two car seats into your car in the winter in Chicago, and then drive to Costco. That gave us a bit more energy, more bandwidth. We could start looking at trends. Like all incidents are different, all babies are different. We realized there were some times maybe when one parent could do it on their own, and the other one gets to clean up or cook. There were times of the day when we both needed to be there. We were doing our own cross-incident analysis, and we started coming up with some ideas. Some of them were a little out there. Like, let’s just move them to their own bedroom. Let’s just drop the swaddle because she clearly hates the swaddle.

Investing time and effort into the process, then into the post-mortems, and then doing that cross-incident analysis gave us the bandwidth to meet our goal. By the end of my maternity leave, I was able to go back to work. I was able to cook and shower. Really, I was able to enjoy my life and enjoy my new family. Live with them is always going to bring challenges. By living on this culture of, let’s just keep learning, let’s just keep trying, let’s be resilient, we know that we can pivot and do the things that we enjoy. That’s basically what we’re going to talk about. This is the lifecycle of incidents and how we can become resilient, how we can learn from them. How we can get better as a system so that we can accomplish our goals.

Background

I’m Vanessa. I work at jeli.io. When I’m not wrangling adorable twins, I spend a lot of time on incidents. I’ve spent the last decade working in technology and site reliability engineering, focusing on incidents. My background is on FinTech actually, in the airline industry. I focus on the entire lifecycle of incidents. I have been the only incident commander on-call for many years. I have owned the escalation process. I have trained others. I have ran retrospective programs. Most importantly, I’ve scaled them. That’s actually the most difficult but also the most important thing to make sure that it doesn’t just live with you, with one person. While I do this for fun, because I enjoy it, I also know that the reason why I get paid to do this is because having a handle on your incidents allows your business to achieve their goals. In the past couple of years, I’ve had the chance to work with other companies, helping them establish their programs. I’ve seen how engineer teams are able to accomplish their goals by focusing on resilience.

The Mirage of Zero-Incidents

Let’s go back to this picture. We will never have zero incidents. If we do, that means we’re not doing anything, or nobody’s using our products. I believe that incidents are a lifecycle. They’re part of how a system gets work done. First, we have a system, say Ticketmaster, and then something happens. Maybe Taylor Swift announced her world tour. Tons of people try to access the system, the system fails, and now you’re in an incident. People work on it, and it’s resolved, and the site is back up. Life is back to normal. As you can see, though, we’re not even halfway through the cycle here. Ideally, after the incident is done, you do something that is focused on learning. Even if you don’t think you’re doing any traditional learning, people naturally like to talk to each other, people like to debrief. Your learning might be just like getting together, grabbing some coffee with your coworkers after the incident. Or it can be like an actual post-mortem meeting, retrospective, whatever you want to call it, your 5 why’s. Or it can be like your incident report, whatever you want it to be. Those learnings are then applied back into the system. Maybe we make a change into our on-call schedule, or the way that we enqueue customers, or our antitrust legislation. That is the new system. That’s what we’re working with.

The Negative Effects of Incidents and Outages

Why do we care about this? We have this whole track here, resilience, talking about outages. We care about this because incidents are expensive to companies. If we think of the cost of the incident, there’s the incident itself. Like, no, we lost money, because the site was down for like 20 minutes. There’s the reputational damage. Like, no, I wasn’t able to get these tickets. I wasn’t able to fly home. I’m just never going to fly this airline again. There’s the workload interruption. If you have an incident that lasts an hour, and you have 10 engineers, that’s time that they’re spending on this incident. Then when the incident is over, they’re not going to go back to their laptops and like, I’m going to work on my OKRs. That just takes a big part of your brain, that takes some time to get back to what you’re doing. It has an impact on our goals and our plans. If our engineers are fighting fires, they’re not working on new features. What we see is that unless we do something about it, people end up getting caught in a cycle of incident where there’s just no breathing room to get ahead.

Incidents are part of doing work, but that doesn’t mean that we can’t get better at them. If engineers are constantly fighting fires, it’s going to impact the way that they deliver and the speed at which they deliver it. If our customers are constantly seeing outages, it’s going to impact how they interact with us. It’s going to impact our goals. As the company, you got to make money. If your customers aren’t using you, then that’s a problem. There’s good news about incidents, I mentioned this earlier. Having the means that somebody cares about your work, that it matters to them whether you’re up or down or that you’re doing the thing that you say you’re doing. I think that’s the thing about tech. Our users don’t care about the language we use or the methods we use, they care about being able to rent cars or get their money or get tickets. We owe it to our users that they’re able to do these things.

The Importance of Resilience in Incidents

Here’s where resilience helps. It’s the capacity to withstand or to recover quickly from difficulties, from your outages, from your errors, from your incidents. Resilience can help us turn them into opportunities. Most of the time, people don’t care about resilience. What usually happens is that you have an incident, you resolve them, and then you move on. We have Chandler Bing here saying, “My work here is done.” Not because they want to. I’ve met tons of engineers. My dad’s an engineer. I’m an engineer. We’re always trying to fix things. Sometimes we just don’t have the bandwidth to do anything else other than move on. That’s how we get stuck in this cycle of fighting incidents. There’s a better way. I have been lucky enough to work with and for a number of organizations that are leading the way in improving resilience in the tech world. The better way includes a focus on three things, a focus on the incident response process, a focus on learning from individual incidents, that’s like your post-mortems, your incident reports, just chatting with people outside the war room. Then, macro insights. You don’t have to do it all. You can if you want to. Often, it’s hard to find an organization that’s going to say like, yes, go ahead, spend all your time doing all of this. There are ways that you can start doing this, one by one. I will go through all of them throughout this talk. First, some caveats. Being successful at resilience is not easy. It’s not easy. It’s not fast, and it’s not cheap. A lot of this requires just like selling this new way of working with resilience as a focus. Selling isn’t our best skill as technologists. Maybe instead of selling, we should call it, presenting our data and our work and making a case for it. The good news is that I have seen this work, and we can definitely get many small wins along the way.

A Focus on Incident Response

Let’s do our first area of focus. What does it look like to focus on incident response? When you focus on incident response, you’re focusing on these three things. You’re focusing on coordination, collaboration, and communication. It makes sense to start here. This is the thing that we are already paying for. We’re already spending time in outages, so we might as well focus on them. When it comes to coordination, write up your current response workflow. How do folks come together to solve a problem? Look for any gaps. Where are things breaking down? What can be done to make those gaps even just a tiny little bit smaller? When it comes to collaboration, how do you get people in the same room? Is that room Zoom? Is that room Slack? Is that room Google Meet? How do you know who to call? Is it the same people every time? Maybe it’s that one expert that wrote the thing 10 years ago. You should examine your on-call rotations and your expectations. Again, what are the little things that we can do to make the lives of those on-callogists a little bit better, just a little bit easier? Then finally, communication. How do you communicate to your stakeholders what is happening? How do you tell your customers? What is hard about it? What are they asking for? Do your engineers know what your users care about, or are they going like, Kafka. When really people are caring about like, can I get approved for my loan or not? Write up some loose guidelines to help manage expectations for folks both inside and outside the immediate teams that are responsible for responding to incidents.

While it makes sense to start an incident response, why aren’t we perfect at it yet? Why aren’t we putting in the effort into getting better at this? There are a few reasons for this, including that we need to train folks on the process. I think, oftentimes, we get a new engineer, we give them the pager, and then they say good luck. Every organization has different needs and different ways of working. When we need to get our services back ASAP, there are certain procedures that make sense. You want to teach folks the skills that help the incident move forward, that urgency. That like, don’t come at me with like, “If only we had done this, if only we had done that?” That idea of hierarchy versus roles. You don’t want people during an incident to focus on hierarchy, even though that’s how they’re used to working. Thinking about the roles that makes sense for incident response, like who is the person in charge of this incident? Who is the person that’s in charge of communicating? Who are your stakeholders versus who are your responders? Do they need to know different information? Do they need to act differently? Additionally, many of the changes in this area of focus will require specific tools and specific automations. Sometimes you have these tools that lead to the highest cognitive load, and these things can be automated, but maybe teams don’t have the right tools for this, or they just don’t have the bandwidth to create these tools to develop them if they’re stuck fighting fires.

Here’s where I quote my friend, Fred Hebert. He actually wrote a great blog post for The New Stack about Honeycomb’s incident response process. They basically are growing as a company, like a lot of our organizations are, and they were running into some issues with their incident response process. He said, “While we ramp up and reclaim expertise, and expand it by discovering scaling limits that weren’t necessary to know about before, we have to admit that we are operating in a state where some elements are not fully in our control.” Fred was trying to balance two key issues. They’re trying to avoid an incident framework that’s more demanding to the operator than solving the issue. We’ve all been there, where you have these runbooks that they’re so long, and it’s like, this is taking me away from actually resolving the thing. You don’t want your runbook to just take so much time and so much effort that it is keeping you from resolving your issue. You also want to provide support and tools to people who have less on-call experience, and for whom these clear rules and guidelines actually do make sense. I’ve seen this throughout my career. If I’m an experienced responder, I don’t need to read all of those things. I don’t need to follow all these procedures. If I’m trying to onboard somebody, like they’re not going to be able to read my brain. We worked together on a process that automated certain tasks, used our incident bot. Some tasks that can be automated, like creating an incident channel or communicating status updates to stakeholders. Doing that leaves the engineers the time to focus on the actual engineering part.

At Honeycomb, as well as other organizations, the idea is to get quick wins, like just weeks into the process of trying to fix your incident response. You can do this by restructuring on-call rotations, making sure that the page is going to the right responder. Automating some response tasks, like I mentioned, creating Slack channels. Automating how you update some folks. This allows us to spend less time on the task that demand our attention during incidents, but aren’t necessarily engineering tasks. Because you never ever want to automate the creativity and the engineering skills of responding. There are things that we can do to just reduce that cognitive load during the high stress times. Improving little things, it’s going to help you build momentum and get buy-in around making larger changes, both from leaderships as well as the folks holding the pagers. If you think back to my example of my children, getting those small changes of like, let’s switch to formula, allowed me to get some sleep so that then I could work on the bigger changes. A focus on incident response can help improve the process, leading to a better experience to your users and engineers. Leads to spending fewer time on repetitive tasks, easier onboarding for on-call. Just a more efficient process and lower customer impact, which is obviously what we care about.

Anti-Pattern: MTTX

Which leads me to my first anti-pattern that I will discuss, my version of MTTX. You’ll hear me talk about MTTX a lot. By that I mean your mean time to discovery, mean time to recovery, mean time to resolution. I’m on the record of saying that they don’t mean anything. Because if I say, we had 50 incidents in Q1, and they average 51 minutes. What is that telling us? It’s not really telling us anything. That doesn’t mean that like, I think we should just not care about how long our incidents last. I wanted to get those Taylor Swift concert tickets, I’m like, I wanted to get them now. I just think that a single number should not be our goal. What I do believe in is that we want to make the experiences for our users and our engineers better, because that is actually what’s going to help us get ahead. There are things that we can do to make the incident response easier and faster so that our engineers are better equipped to resolve them. Just like we will never be at like incident zero, we’re not in control of everything. That single timing net metric, just should not be the goal.

A Focus on Incident Analysis

The next way that we apply resilience is when we focus on incident analysis, that’s like you’re learning from incidents. After an incident, you want to learn from it, or about it. I believe that the best way to do this is this narrative-based approach that can highlight what happened, and can highlight how folks experience the incident from their different points of view. How I experience an incident as a responder is going to be different how my user experiences an incident, how my customers or folks experience an incident, how my stakeholders experience an incident. What we have seen is that those template filling sessions, those 5 why’s of root cause analysis, they can be helpful. Sometimes they’re not helpful. Sometimes it can actually cause harm, because they make you feel like you’re doing something. They give us this false sense of security, when in reality they’re not doing much. There’s more that you can do. I’ve done this myself. I’ve seen it many times. You have this root cause analysis session, and we say, the root cause is human error, so let’s just never do that again. The action item is for that person to just never do it again. That’s not much to get out of 1 hour with 15 engineers. If you have this blame-aware, learning-focused review, you can highlight different things. Maybe you can highlight that the new on-call engineer didn’t have access to some of the dashboards that can help out with the response. Maybe you can highlight that you have a QA process that just doesn’t account for certain requirements. Or my favorite, when you realize that engineering and marketing aren’t talking to each other, and there’s a new feature that they’re announcing, and we just don’t have enough capacity to handle it. Those are things that we can learn from and that can actually move the needle when it comes to resilience.

Why aren’t folks doing this right now? Like I said, people think that they’re doing this. People think that they’re doing 5 why’s of root cause analysis. That is not the same as a narrative approach. I think part of it is that when we think of a narrative approach, we usually default to timelines. Creating timelines is a lot of work. I’ve done this for many years. You have two screens, you have three screens, and you’re copying and pasting. You got Slack here. You got a Google Doc here. Then you have to open GitHub and PagerDuty. You’re switching from different data sources, and you’re summarizing conversations, and sometimes stuff just gets along the way. Then if you’re like, I don’t want to do this. I’m not going to do this prep. You have your post-mortem. Like I said, you have like 15 engineers for an hour. Then you’re spending that time just building that timeline. That’s not really a good use of your time. Your people aren’t actually talking and collaborating. Then this leads to folks just not trusting the post-mortem, the learning process. If you want to know more about how to do incident analysis, we actually put out a free guide for doing this. This is called the Howie guide for how we got here. Dr. Laura Maguire was one of the co-authors, as well as myself. Here we outline a 12-step process to help you best understand how you got to where you ended up. I laugh at the 12-step process, because it’s long. You assign. You accept an investigation. You identify your data. You prepare for interviews. You write up this calibration document. You help wrap up your investigation. You lead a learning review. You do an incident report. You do more findings, and you share it. You do all of this. John Allspaw is amazing at this. I love it. I think it’s great. This is a lot. You can do this or you can actually take from this and take the spirit of what this is, which really is just like a collaborative narrative approach.

If you break it down to the basics, you want to identify your data sources. You want to try and understand who wasn’t actually involved in the incident, not just the person who responded to the incident, but like, who were the stakeholders? Are there PMs involved? Are there marketing folks involved? Are there your customer support folks involved? Where in the world were they? What were they dealing with? Like if I’m in Chicago, and I have somebody in Sydney, we are interacting differently because we had different parts of the day. Where did it take place? Were people in-person? Were they in Slack? Were they looking at each other face-to-face? Then you want to prepare for your meeting. You want to create your timeline. Usually this is where I go in Slack, I create a narrative timeline. I jot down any questions that I have for people. I can have an interview, or I can have those questions be at the review meeting in front of everyone. You can cheat. You can look at the top moments, but really, the key moments are the ones where people are not knowing what they’re doing, not knowing what to do. Then you have your meeting. When I’m leading a meeting, or when I’m doing an interview, it should never be the Vanessa show. It shouldn’t be like this, where I’m here and I’m lecturing you all. It should be a collaborative conversation. People should be able to tell me what they experience from their own point of view. Like, “I’m a customer support person and I was getting inundated with requests from our users.” That’s impacting how I experience the incident. Then you finalize your findings. You finalize them in a format that you can share with others, that others are going to understand. Depending on your audience, you will want to share different information with them.

We want to make it as easy as possible for people to do this. Doing a narrative-based approach can take weeks, or it can take 20 minutes. Here’s a narrative builder, which we have in Jeli. The idea is that you take your different data sources, and you create this narrative. Here, you can see that an incident is never like, we released this bug, and we reverted the bug, and now we’re done. You can understand that maybe the reason why the incident lasted as long as it did was because some people realized halfway through the incident, that they didn’t have the full impact of what was going on, because they were looking at different dashboards. Knowing that can lead to actual learnings and actual change in how we do our work, which has direct impact into engineers meeting their goals. Engineers not looking at the right dashboards has nothing to do with the bug, but it has a lot to do with how we respond to incidents. It’s really important to give people the ability to create these documents, to create these artifacts that they can then share with others, to give them the ability to share how the incident happened from their own points of view.

Then we talk about action items, because, historically, we all think of retrospectives as a source of action items. I really believe that the action items should reflect the insights that we gained from the narrative and from the incident reviews. That it should reflect what we learned about the contributing factors and the impact. The action items shouldn’t be just like, revert this bug and be done with it. We’ve seen this in the wild. We’ve seen some examples of people that I’ve worked with. At Zendesk, they had an incident that highlighted the need to just rethink their documentation and the information that’s in their runbooks. They did this narrative builder exercise, and they realized that the responders just don’t have access to the right things, or they’re working with outdated information. I’ve been there as a responder as well. I’ve been there where I pull up a wiki page, and they have just things that worked two years ago. You’re understanding what things we can do in the future to make things better for the engineers themselves that are solving incidents today.

At Chime, another organization that we work with, they did an incident review, and they realized how some of the vendor relations impact the way that an incident is handled, in order to improve the process. Because at the end of the day, it’s important to know how to get a hold of a vendor. Sometimes the people who are on-call are not the ones that either know how to find that person or have the access to do that. Even knowing what parts of the process belongs to vendors versus us makes a huge difference next time that we have an incident. Again, we’re never going to be at incident zero, so these action items are helping us solve incidents in the future. A focus on incident analysis can identify the areas of work that’s going to lead to engineers working effectively to resolve issues, so that the next time that they encounter another incident, similar incident or otherwise, there’ll be better positioned to handle it, leading to just lower customer impact and fewer interruptions.

Anti-Pattern: Action Items Factory

Then my next anti-pattern is this idea of the action items factory. When I talk about resilience, I just want to make sure to stay away from the anti-pattern of being an action items factory. The idea that we have an incident and a post-mortem, and all we do is just play Whack a Mole, and like, let’s have an alert for this, an alert for that, an alert for that. We’re going to put them all in this Google Doc and this Jira thing that’s never going to get prioritized, and it’s going to be there forever. Because that’s not good. It erodes trust in the process. I remember when I started at a previous organization, I’d go in there. We have 100 incident reports open because they all have 10 action items that are still open from 5 years before. I was still in college back then. If the goal of the incident review is to come up with fixes, but the fixes never get done, then engineers and stakeholders aren’t going to take them seriously. Instead, they’re going to feel like the retrospectives are just a waste of time for a group of people that are already really busy. Instead, we should look at resilience that’s coming out of incidents as part of different categories.

What’s the problem with the action items today? What we hear often is that action items just don’t get completed. If they do, they don’t move the needle. That’s like, let’s just play Whack a Mole and create more alerts. Another problem is that they’re not applicable for the whole organization. Maybe it makes sense for one person, but it’s not going to get done. It’s not going to actually move the needle. It’s almost as if finding the right action item is like a bit of a Goldilocks situation. When we think of action items that are too small, I think they are things that shouldn’t really be action items in the first place, because they’re already being taken care of. Engineers love solving issues. I was trying to troubleshoot my newborns. They shouldn’t wait for the post-mortem to address these issues. They shouldn’t wait for the post-mortem to do a cleanup, or fix a bug, or something like that. We shouldn’t spend our precious post-mortem time on these. I do think that folks should still get credit for them. I think they should still be included in the artifact, because when we have an incident that is similar to that, again, we’re going to want to look back and try to see what we did.

Then we have the two big action items. The problem here is that they’re going to take just too long to complete. You have action items that are going to take whole years to complete, and by then we are not even having that product anymore. If you’re tracking completion, which a lot of companies do, that’s going to skew your numbers. We can talk more about numbers and stuff like that later. The other problem is that the people who are in the meeting, the people who are in the post-mortem aren’t the ones who get to do the work, aren’t the ones that get to decide. Maybe it’s a cross-team initiative, or maybe I say I’m going to do this, but my manager is never going to give me the time to do it. These large action items are still important takeaways from the post-mortems. In that case, I actually recommend that you keep them and you call them initiatives, and you just have them in a separate space in your artifact. Or, rethink the scope. Instead of the action item being like, let’s rearchitect this thing. It’s like, ok, let’s have a conversation about what it would take to rearchitect this thing. Or like, let’s have somebody drive the initiative. Another tip that I have is to always have an owner for your action item. This actually really helps with accountability. I actually recommend having a direct owner as well as an indirect owner. This is a hot tip that I did at a previous organization where like, if I had the action item assigned to my name, my director also gets it assigned to their name. Then eventually, they feel embarrassed for having so many action items open, and they’re like, “Ok, Vanessa, how about you get the time to do them.”

What does the perfect action item look like? It’s usually something that needs to get prioritized, and that will get prioritized. It’s something that can be completed within the next month or two. You don’t want it to last forever. Again, that’s an initiative. It’s something that can be decided by the folks in the room. It’s also a good flag to invite a diverse group of people to these reviews: you want your managers, you want your product managers. It’s something that moves the needle in relation to the incident at hand. I really like this chart, mostly because I created it. It shows the difference between all of these things that I mentioned. It shows the difference between a quick fix, like your cleanups, and an action item. Redesigning a process, updating documentation. Your larger initiatives, switching vendors, rearchitecting your systems. Then a transformation that can actually take place out of an incident, like org changes, headcount changes, things like that.

Let’s go into an example of how to get there. Here, you can see that we are creating a narrative. We’re looking at a specific line in the Slack transcript. We’re asking a question, in this case, about service dependencies, and how they relate to engineers understanding the information that’s in their service map vendor. I’ve been in incidents like this, where you realize that you have the service map vendor, and they have an understanding of how things work, and we have a different understanding, and things just don’t match up. Here you can see how that question made it into a takeaway around what the service touches and what it depends on. Finally, in the report section, you can see that there are action items around the understanding of the dependency mapping. You can see that these action items meet the requirements that I said earlier. They have ownership, they have due dates, and they move the needle in how we respond to incidents in the future.

A Focus on Cross-Incident Insights

Finally, we have a focus on cross-incident insights. Cross-incident insights is the key to getting the most out of your incident. You will get better quality cross-incident insights the more you focus on your individual high-quality reviews. It’s really a great evolution from that process. The more retrospectives, post-mortems you have, the more data you can then use to make recommendations for larger scale changes. Doing this shouldn’t be one person’s job. We’re going to talk a lot about collaboration. We’ve talked a lot about it the past few days. This is done, I believe, in collaboration between leadership, multiple engineering teams, product teams. It’s your chance to drive those cultural transformations that I mentioned earlier. Maybe you’re deciding that you want to DevOps part of your process, or you want to rearchitect the system, because you have the breadth of data, and you have the context to provide leadership for future focused decisions. We spoke earlier about having goals around metrics. Doing cross-incident analysis is actually allowing you to provide context around those timing metrics and make recommendations for how to improve the experience for our users and our engineers.

Why aren’t folks doing this? A lot of people want to do this, and they actually think that they’re doing this, but they aren’t. That’s because they just don’t have the data. High quality insights come from high quality individual incident reviews. It’s very hard to get good cross-incident analysis if you’re not doing actual incident analysis. Also, the place where our incident reports live, are not friendly for data analysis. If you have something living in a Google Doc, like Google Docs are great for narrative storytelling. They’re not very easily searchable, or queryable, or anything like that. Most importantly, and this is something that is close to my heart, is that a lot of people think that they can be analysts, but engineers are not necessarily analysts. That’s ok. If you want to do this work, you need training and presenting data in friendly formats. You need to know how to use the data to tell a story. All of this is hard work.

Here are some examples of some cross-incident insights that are coming with context. This is coming from Jeli’s Learning Center. Instead of simply saying we had 300 incidents this year, we can take a look at the technologies that are involved, and we can make recommendations based off of that. I can say that something in the past year related to console was involved in most of the incidents, let’s take a look at those teams, and maybe give them some love. Maybe help them restructure their processes, focus on their on-call experience.

Make those recommendations to make their lives better, if we want to lower that number. Again, this stuff works in the wild. We worked with Zendesk. They were able to look at their incidents, and they were able to make recommendations around the rotations of certain teams and technologies. Had they not had the data from individual incident analysis, making this case would have been a lot more difficult. We have our friends at Honeycomb as well. They use this information to drive continuous improvement of the software. We’ve seen people use these insights to decide on a number of things, feature flag solutions, switching vendors, even making organizational changes from cross-incident analysis. I actually was part of a reorg. Thanks to that, where I was like, our SRE team is like, we have so much stuff going on, let’s try to split things up. Let’s try just experimenting something. These are the larger initiatives that we discussed. All of this creates an environment where engineers are able to do their best work and achieve the goals of their organizations.

Now let’s go back to MTTX. Here’s the part that I will tell people, “I will never tell an analyst to tell their CTO that the metric that they’re requesting is silly, and that they’re just not going to do it.” Because I’ve been that engineer, I’ve been that analyst, and I did not get paid enough to have that conversation. When you have leadership asking for these metrics, take this as an opportunity to show your work. I call this my veggies in the pasta sauce approach. Now that I have children, I should call it like the veggies and the smoothie’s approach. It’s the idea that you’re taking something that maybe you don’t think has much nutritional value, which is that single MTTR number, like 51 minutes. You’re adding context to make it much richer, to bring a lot more value to your teams and to your organizations. It can help you make strategic recommendations about the data that you have. It really helps move the needle in that direction that you want. You can show that maybe incidents lasted longer this quarter, but the teams that maybe had the shorter incidents, they have some really cool tools that they’ve been experimenting with. Maybe we should change that and give other teams access to that, spread them across the organization.

Anti-Pattern: Not Communicating the Insights

That leads me to the last anti-pattern, which is, again, something we’re not very good at as engineers, and it’s communicating our insights. Because if we have insights but nobody sees them, what’s the point? In Jeli’s Howie guide to incident analysis, when discussing sharing incident findings, we explain that your work shouldn’t be completed to be filed, it should be completed so it can be read and shared across the business, even after the learning review has taken place and the corrective actions have been taken. This is true about individual incident analysis, but even more so about cross-incident analysis. I’ve been an engineer my whole life, and I know how we are. We’re like, we all know what we need to do, we just are not doing it. Senior leadership would never do this. We’re like, yes, that’s true, but do we all know that? Have we told other people this? Often, when people don’t interact with our learnings, it’s because the format just isn’t meant for them. Different audiences learn from different things in different ways. If you’re sharing insights with your C-suite folks, maybe your CTO is going to understand the extremely technical language that you’re using, but everybody else is going to be looking at their phones. You need to use language that your audience is going to understand. You should focus on your goal. Why are you sharing this? Think of the impact. Don’t spend too much time going through details that just aren’t important to your goal. If you want to get approval to rearchitect a system, explain what you think it’s worth doing, what problem it’s going to solve, and who should be involved.

Focus on your format. Whether that includes sharing insights, sharing numbers, telling a story, and sometimes those technical details. I usually like to think of my top insights. For example, I will go up to CTO and say, “Incidents went up this quarter but usage also went up.” Or the example that I mentioned earlier, SREs were the only on-call incident commanders for 100% of our incidents. There’s your number. As we grow, this is proving to not be sustainable, so I think we should change this process and get product teams involved. That’s how you move away from like, we all know what’s wrong, and nobody’s doing anything, to actual changes. I’ve had this experience before. The SREs in getting product teams involved made a change. I’ve rearchitected actual CI/CD pipelines by going up to people and saying, “So many of our incidents, the root cause isn’t how we release things, but it is impacting how long it takes us to resolve them. After talking to all of these subject matter experts, I think we need to have a project this year where we rearchitect this whole thing. I know it’s going to take some time. I know it’s going to take some money, but it’s necessary if we want to move forward.” Because you’re not suggesting this 4-month project because you feel like it, you’re suggesting it because you have the data to back it up. If you’re asking me for these MTTR numbers, I’m going to give them to you but I’m also going to use that as an opportunity to give you the context, and to give you the recommendation of the thing we all know we need to do.

Going back to this chart, so instead of thinking of action items, or just recovering from downtime, being the only output of your resilience work, let’s think of this as the product of treating incidents as opportunities. Doing this work is how organizations get ahead: how we make org changes, and headcount changes, and transformations. Because we will never be at zero incidents. Technology evolves and new challenges come up. By focusing on resilience, on incident response, incident analysis, cross-incident insights, we can lower the cost of them. We can be better prepared to handle incidents, leading to just a better experience to our users. A culture of where engineers are engaged, and they have the bandwidth to not only complete their work, but to creatively improve things. That’s why we’re hiring engineers and paying them a lot of money. To just achieve and surpass our goals, which means that people can make more money.

See more presentations with transcripts

Subscribe for MMS Newsletter

By signing up, you will receive updates about our latest information.

  • This field is for validation purposes and should be left unchanged.

Podcast: Resilience and Incident Management with Vanessa Huerta Granda

MMS Founder
MMS Vanessa Huerta Granda

Subscribe on:






Transcript

Good day, folks. This is Shane Hastie for the InfoQ Engineering Culture podcast. Today I’m sitting down with Vanessa Huerta Granda. Welcome. Thanks for taking the time to talk to us today.

Vanessa Huerta Granda: Thank you for having me.

Shane Hastie: So Vanessa, you were the track host for the Resilience Engineering, and I love the second half of that, Culture As a System Requirement track at QCon San Francisco. Let’s start delving into the track. What was the message you were trying to convey when putting that track together?

Resiliency in sociotechnical systems [00:52]

Vanessa Huerta Granda: I, and a lot of folks from the Learning from Incidents community, when we’re thinking about our tech systems, we don’t just think of them as technological systems. We like to think about the sociotechnical system. And so socio, that people part is part of our systems. Everything that we do, it depends on people, it’s running because people are making decisions at some point or another. Maybe when something is getting first developed, when something’s first getting architected, but also when you’re maintaining something, when you’re handling issues, incidents, when you’re learning from them or anything like that. So the idea here is that culture is part of the system, is part of that sociotechnical system. And resiliency, making sure that your organization is a resilient culture is a huge part of that.

Shane Hastie: Taking a step backwards, what brought you to focusing on this? What’s your background?

Introducing Vanessa [01:43]

Vanessa Huerta Granda: I am an industrial engineer, which weirdly, actually fits in perfectly with resiliency work, with understanding how it is that people work, how it is that people create the software, write it and maintain it. So I have worked in operations as a Site Reliability Engineer and leader for the past decade. I am currently the Manager of Resiliency Engineering at Enova, and previously, I was at a startup called jeli.io, focusing on products to help people handle their incidents, learn from them, the entire lifecycle.

Honestly, I got into this kind of work because I just really enjoy solving problems and I really enjoy talking to people. And at some point, a boss of mine realized that was actually a good fit for this kind of work, and so I’ve been doing it ever since.

Shane Hastie: So going right down to first principles, what do you mean by resilience?

Defining resilience [02:34]

Vanessa Huerta Granda: A mentor of mine once said that resilience is something your system does rather than what your system is. Resiliency is our ability to sustain challenges, to sustain fractures, failures, whatever it is. I think we often think that having a resilient system means that our system is never going to break. That’s just not true at all. It means that your system, including the sociotechnical system, sometimes something bad is going to happen, something is going to break, and how do we recover from that?

Shane Hastie: And what does a culture of resilience look like and feel like?

Vanessa Huerta Granda: Well, I can tell you I’ve experienced that all week. So a culture of resilience really means when the folks that are working at your organization are able to handle whatever is happening, whatever is thrown at them. So we have this idea that we have these plans for the quarter, we’re going to get so many projects done because we have so many T-shirt size things that we’re going to do. But at some point, something is going to happen, someone’s going to go on vacation, someone’s going to get sick, or maybe the code that we understood actually doesn’t work the way that it does.

And so that’s when you have your socio part of the system figuring out a way to make it work. And that can be through automatically having failovers in your system or just having a process where people talk to each other and figure out like, “Oh, hey, can you help me with this? Can you help me with that? Let’s prioritize this. Let’s prioritize that.”

Shane Hastie: Incidents and resilience, of course go hand in hand, but incident management in my experience is something that many organizations do haphazardly at best and often very badly.

Vanessa Huerta Granda: I hate that.

Shane Hastie: So what does good incident management look like?

Good incident management [04:11]

Vanessa Huerta Granda: Oh my gosh, how much time do you have? When I think about incident management, I think about allowing your engineers to do their best work. So as an incident manager, I am not here to yell at anyone, and I hate this idea. We often think of the person leading an incident, we call them incident commander. I can tell you that when I first started this job, I was 25, the only woman in the room, the only Latina woman in the room, and I definitely did not feel like a commander. What I did feel like was somebody who could talk to people and get them to actually discuss what was happening and figure out the best way forward. Incident management that actually works is understanding that we’re all working for the same team, we’re all working for the same goal, which is to just get our systems back to normal and that we need to work together to do that.

The other side of the coin is that when you’re in an incident, if your company’s making money, that means that during the incident you’re not making money or something bad is happening. And clearly people are going to care, clearly your stakeholders, your leaders, they need to be aware of what’s happening. And so I like to tell the responders to, “Worry about responding, worry about doing the engineering things. Use your brain towards that. I’m going to focus on the communication. I’m going to focus on the coordination and the collaboration, and I’m going to make sure that you’re not answering a million things from your CTO, that you’re not worried that your CMO is going to be upset because the website is down. I’m going to take that away so we can work towards resolution.”

Shane Hastie: And then what happens afterwards?

Vanessa Huerta Granda: Oh, my favorite part, you gossip about it. So we have an incident that’s over, and I work from home a lot more nowadays than I did back then, but you’re outside of the office and you’re outside of the war room, whatever you want to call it. And you’re talking about it, right? You’re discussing, “This is something that happened, that is something happened.” Maybe you go to lunch, maybe you go to happy hour. People are always going to talk about it. In cultures where there’s not that culture of resiliency, you have a postmortem, and that postmortem is usually some sort of document, that incident report that tells you, “This incident started at 5:00 AM and it was over by 7:00 AM and it was Shane’s fault, and Shane really sucks.”

Shane Hastie: Yes, the who-can-we-blame session?

Vanessa Huerta Granda: Right?

Shane Hastie: That’s a really, really important part.

Vanessa Huerta Granda: Oh my gosh. And that to me is really, really sad because an incident is just this big red arrow pointing to you towards something that is happening at your organization that you can learn from. And so, I like to think of incidents as learning opportunities. So a retrospective can be that people are talking about your incidents either way, people are talking after the incident, I’m certainly slapping my work bestie to be like, “Hey, can you believe that happened?” Or during lunch afterwards we’re going to talk about our incident and so we might as well talk about it together and learn from it.

So usually what I like to do during a retrospective is make sure that people are sharing what happened from their own point of view because at my previous role, I was doing a lot of consultant work. I was not an engineer, and for the first time when I was in an incident, I wasn’t seeing what was happening in the code. I was seeing what our customer was experiencing. And so that is a different point of view than the engineer, and that can certainly make a difference in how we move forward. The idea is that you have a learning review, a postmortem, retrospective, whatever it is that you want to call it. You’re learning from that. And then you’re coming up with action items that are helping move the needle in the future.

And it can be as easy as like, “You know what? Maybe we need to have better post-release checks.” Or maybe it’s something like, “You know what? This process that we’ve been working, it probably doesn’t make sense anymore. It made sense back a year or two years ago. It doesn’t make sense anymore. Maybe we need to do some more training,” et cetera, et cetera. There’s many things that you can learn from an incident.

Shane Hastie: How do we avoid it being that blamestorming activity?

Avoiding blamestorming – turn postmortems into learning opportunities [07:59]

Vanessa Huerta Granda: Well, that’s the part of the culture, right? That’s where you have to as an engineering leader and as individual contributors, make sure that when you’re leading a retrospective, when you’re leading a postmortem, that you are not just filling out a document, that you are speaking up and you’re letting other voices heard. So there’s a lot of good information out there. There’s the Etsy’s Debriefing Guide. There is the Howie Guide from jeli.io. I co-authored it a few years ago, and it helps people understand how to best position themselves so they can turn their postmortems into learning opportunities.

From my standpoint, I can give you the best advice that I can give you is to… If you’re leading a retrospective, never be the only person that’s speaking. Let other people speak up, let other people share from their points of view what happened and be forceful of that, “This is not a blaming game. We’re all on the same team.”

Shane Hastie: And communicating the outcomes. What I want to explore there is getting off the hamster wheel of just incident, incident, incident response and breaking what feels to me at times like a never-ending spiral for folks.

Vanessa Huerta Granda: I think I’ve mentioned this earlier, I like to think of incidents as an incident lifecycle where you have your incident, something breaks, then you have your retrospective where you’re learning, out of the retrospective, there’s action items. And so that’s feeding into the process and there’s things that you can do throughout that entire lifecycle to make things better. This is the part where I can give you my example that actually, let me understand… Not let me understand, but really highlighted how the incident lifecycle can be applied to anything.

I currently have 2-year-old twins and they’re my only children. When they were first born and we took them home from the hospital, it was the craziest incidents I had ever had in my life. And I had had a lot of incidents, but it was kind of bananas. I was like, “Shoot, I can’t sleep at all.” And my husband couldn’t sleep either because one wasn’t crying, the other one was. And so we felt like we were in this hamster wheel fighting this incident over and over again, and we just did not have the brain power to do anything about it. And so what we did was like, “Let’s try to make this process easier for us.” And that’s what I recommend a lot of organizations start doing. Make the process easier for your responders. A lot of organizations start introducing tools, introducing Slack bots or team spots, whatever sort of chat platform you’re using, start communicating, automating some of the things.

And what we did is we hired an night nanny. So once you have the bandwidth, you’re out of constantly, constantly fighting incidents, then you’re able to start putting those productive retrospectives, productive postmortems and start figuring out what it is that you can do at a higher level. Right? At that point, that’s when my husband and I were like, “You know what? We don’t have to drag our two infants in their car seats to Costco. We can just have the diapers delivered.” So we’re thinking of a change to the process that’s going to make the entire lifecycle easier, not just one specific incident. And I think I gave that example earlier, right? After an incident, maybe once we are able to have a retrospective, we realize, “You know what?We need better controls for this processor. We need to have a testing suite that makes more sense or blue-green deployments,” whatever it’s that you want to call it. And so those action items make it easier, lead to fewer incidents, give you more bandwidth.

And then the last thing that we like to do is cross-incident analysis where you’ve made those easy changes, you’ve addressed the low hanging fruit and then taking a holistic look at your incidents. Maybe you’re taking a look at all the incidents that you had in the last quarter, and you’re able to say like, “Okay, you know what? This team has a lot of incidents. Let’s try to maybe give them a little bit more headcount.” Or, “It seems like these two teams are working on something similar or it seems like a lot of these incidents are related to this antiquated pipeline. Let’s maybe give more resources to do all of this.” And so you go from making small changes to the incident process itself to making changes out of those incidents. And then you’re making larger transformations.

And actually, during my time, we were able to make a case for changing, for going into more of a DevOps organization, moving away from an ops team that everything was funneled through them, through having the SRE team. And that’s the funnel that not everything goes through them because we were able to look at the incidents and because we were able to try to find patterns there. And I did the same thing with my children and now I love being a mom.

Shane Hastie: Two-year-old twins, your life is a hurricane.

Vanessa Huerta Granda: It’s fun. I do incidents for a living.

Shane Hastie: One of the things that we touched on earlier, but I would like to dig in deeper is stress and burnout. How do we help folks reduce the stress and avoid burnout? Because we certainly know that burnout is a significant issue in our industry at the moment.

Reducing stress and avoiding burnout [12:51]

Vanessa Huerta Granda: Absolutely. We’ve seen it everywhere, right? Burnout. I don’t see people getting burnt out because they’re working so much, as much as they get burnt out because they’re working hard, but it never stops. Nothing ever changes. And that’s why I am so passionate about the learning part of things and not just the resolving problems, but if you’re in the hamster wheel, I want to hear what you are seeing and I want to see where it is that I can help. And that’s why I like to make those… You have those short-term action items that can help maybe things a little bit, but then when we’re making those higher level recommendations, we include things like, “Let’s add more headcount because things are hitting a fan.” Or you always hear engineers saying, “The problem is that we have this outdated architecture that made sense a while ago, but no one’s going to put the resources.”

Well, if I’m the person that has all of the data around incidents and I’m able to go up to your leadership team and tell them, “All of these incidents that you’re having, maybe you should put in some resources into doing something different.” That’s going to allow people to see that their hard work isn’t just for nothing. I’m also a manager of my specific team, and I take that very seriously and I make sure to have personal connections with them and make sure to give them the time that they need to rest, to listen to them, and to be proactive about that, right? Like, “If you were up all night working on an incident, please don’t come in today. Please take some time to sleep.” And the same goes with your personal life, right? Like, “If you were handling your three-month-old baby overnight, there are more important things out there.”

Shane Hastie: You are a manager, you’re a leader of a team. A lot of our audience are stepping into that role, often for the first time. What advice would you have for them?

Advice for new leaders [14:40]

Vanessa Huerta Granda: When you’re an individual contributor, it’s sometimes hard to understand the constraints that management is working with. I think becoming a first-time manager, I wish I had given myself a little bit more grace and realized that I can’t change everything. One of my mentors actually, her mantra was, “Grit and grace.” Yes, try to work through things with grit, but also give yourself grace. Give other people grace. No one’s out to get you. And I feel like it’s taken me a little bit to realize that, especially when you’re working with incidents, when you’re trying to work with people from different functions, they’re all working with their own constraints. And so remember that you’re on the same team, I think makes a lot of difference.

And then when you’re managing your own team, I mentioned that I take that very seriously. These are people’s livelihoods that I have on my hands, right? I’m their manager, and so they spend a lot of time working, and I just want to make sure that I’m listening to them, that I’m understanding where they’re coming from, not making assumptions, giving them grace as well.

Shane Hastie: Grit and grace. I like the combination, grit and grace. Thank you very much. Vanessa, if people want to continue the conversation, where will they find you?

Vanessa Huerta Granda: I guess on X is now what it’s called, I am the v_hue_g. You can also find me on LinkedIn. My name, Vanessa Huerta Granda. And yeah, I talk about incidents all the time and a little bit about reality TV, mostly about incidents.

Shane Hastie: Thank you so much for taking the time to talk to us today.

Vanessa Huerta Granda: Thank you, Shane.

Mentioned:

About the Author

.
From this page you also have access to our recorded show notes. They all have clickable links that will take you directly to that part of the audio.

Subscribe for MMS Newsletter

By signing up, you will receive updates about our latest information.

  • This field is for validation purposes and should be left unchanged.