<?xml version="1.0" encoding="UTF-8" ?>
<?xml-stylesheet href="https://rss.buzzsprout.com/styles.xsl" type="text/xsl"?>
<rss version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:podcast="https://podcastindex.org/namespace/1.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:psc="http://podlove.org/simple-chapters" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <atom:link href="https://rss.buzzsprout.com/1793843.rss" rel="self" type="application/rss+xml" />
  <atom:link href="https://pubsubhubbub.appspot.com/" rel="hub" xmlns="http://www.w3.org/2005/Atom" />
  <title>The VOID</title>

  <lastBuildDate>Wed, 13 May 2026 04:54:44 -0400</lastBuildDate>
  <link>https://podcast.thevoid.community</link>
  <language>en-us</language>
  <copyright>© 2026 The VOID</copyright>
  <podcast:locked>yes</podcast:locked>
    <podcast:guid>7efb3953-e4d8-5ee3-a9a0-bd018d08d65f</podcast:guid>
  <itunes:author>Courtney Nash</itunes:author>
  <itunes:type>episodic</itunes:type>
  <itunes:explicit>false</itunes:explicit>
  <description><![CDATA[<p>The VOID makes public software-related incident reports available to everyone, raising awareness and increasing understanding of software-based failures in order to make the internet a more resilient and safe place. This podcast is an insider's look at software-related incident reports. Each episode, we pull an incident report from the VOID (https://www.thevoid.community/), and invite the author(s) on to discuss their experience both with the incident itself, and the also the process of analyzing and writing it up for others to lean from.&nbsp;</p>]]></description>
  <generator>Buzzsprout (https://www.buzzsprout.com)</generator>
  <itunes:owner>
    <itunes:name>Courtney Nash</itunes:name>
  </itunes:owner>
  <image>
     <url>https://storage.buzzsprout.com/wj52muavdz6s86dg6hur3livmgco?.jpg</url>
     <title>The VOID</title>
     <link></link>
  </image>
  <itunes:image href="https://storage.buzzsprout.com/wj52muavdz6s86dg6hur3livmgco?.jpg" />
  <itunes:category text="Technology" />
  <item>
    <itunes:title>Uptime Labs and the Multi-Party Dilemma (Part II)</itunes:title>
    <title>Uptime Labs and the Multi-Party Dilemma (Part II)</title>
    <itunes:summary><![CDATA[Watch on YouTube In Part II of the Multi-Party Dilemma (MPD) drill retrospective, we reconvene to dig deeper into the implications and nuances of the simulated incident exercise hosted on the Uptime Labs platform. Eric Dobbs (incident analyst), Alex Elman (deputy IC), and Sarah Butt (incident commander) continue their debrief with Courtney, reflecting on how team behavior evolved under stress, the importance of expertise in managing non-technical aspects of an incident like saturation, and ho...]]></itunes:summary>
    <description><![CDATA[<p><a href='https://www.youtube.com/watch?v=OAX7_4KRWGI'>Watch on YouTube</a></p><p>In Part II of the Multi-Party Dilemma (MPD) drill retrospective, we reconvene to dig deeper into the implications and nuances of the simulated incident exercise hosted on the Uptime Labs platform. Eric Dobbs (incident analyst), Alex Elman (deputy IC), and Sarah Butt (incident commander) continue their debrief with Courtney, reflecting on how team behavior evolved under stress, the importance of expertise in managing non-technical aspects of an incident like saturation, and how deeply held assumptions often go unspoken until tested under pressure.</p><p>This episode emphasizes the complex social and cognitive dimensions of incident response, such as how people coordinate, communicate, and construct shared understanding. It highlights the value of analyzing drills not for failure points, but for what they reveal about real work, adaptation, and human coordination.</p><p>Key Highlights</p><ul><li><b>Incident Analysis as a Practice</b>:<ul><li>Eric Dobbs emphasized <b>understanding how people make sense of unfolding events</b>, rather than judging decisions in hindsight.</li><li>The goal is to study the “why it made sense at the time,” not what was “right” or “wrong.”</li></ul></li><li><b>Drills Expose Hidden Assumptions</b>:<ul><li>Even experienced responders bring unspoken mental models into incidents.</li><li>The drill revealed assumptions about communication flows, authority boundaries, and vendor interactions that were not made explicit in planning.</li></ul></li><li><b>The Value of Human Expertise</b>:<ul><li>Everyone involved in this incident brought an unparalleled level of expertise to the work. </li><li>Often this kind of expertise goes unnoticed or is taken for granted, however this kind of knowledge is precisely what makes for smoother, better coordinated (and sometimes), faster incident response.</li></ul></li><li><b>Importance of Framing</b>:<ul><li>The way questions are asked in retrospectives can shape what is revealed—e.g., “What made that hard?” is more productive than “What did you miss?”</li><li>Reframing incidents around constraints and tradeoffs leads to deeper insight.</li></ul></li><li><b>Team Learning and Culture</b>:<ul><li>Safe, high-trust environments enable better learning during drills.</li><li>Psychological safety allows team members to admit confusion or raise alternate interpretations during real incidents.</li></ul></li></ul><p><b>Resources and References</b></p><ul><li><a href='https://podcast.thevoid.community/1793843/episodes/17564236-uptime-labs-and-the-multi-party-dilemma-part-i'>Episode I</a></li><li><a href='https://www.researchgate.net/publication/327427067_The_Theory_of_Graceful_Extensibility_Basic_rules_that_govern_adaptive_systems'>Model of Overload/Saturation as part of the Theory of Graceful Extensibility</a></li><li><a href='https://surfingcomplexity.blog/2017/06/24/a-conjecture-on-why-reliable-systems-fail/'>Lorin&apos;s Law</a></li></ul>]]></description>
    <content:encoded><![CDATA[<p><a href='https://www.youtube.com/watch?v=OAX7_4KRWGI'>Watch on YouTube</a></p><p>In Part II of the Multi-Party Dilemma (MPD) drill retrospective, we reconvene to dig deeper into the implications and nuances of the simulated incident exercise hosted on the Uptime Labs platform. Eric Dobbs (incident analyst), Alex Elman (deputy IC), and Sarah Butt (incident commander) continue their debrief with Courtney, reflecting on how team behavior evolved under stress, the importance of expertise in managing non-technical aspects of an incident like saturation, and how deeply held assumptions often go unspoken until tested under pressure.</p><p>This episode emphasizes the complex social and cognitive dimensions of incident response, such as how people coordinate, communicate, and construct shared understanding. It highlights the value of analyzing drills not for failure points, but for what they reveal about real work, adaptation, and human coordination.</p><p>Key Highlights</p><ul><li><b>Incident Analysis as a Practice</b>:<ul><li>Eric Dobbs emphasized <b>understanding how people make sense of unfolding events</b>, rather than judging decisions in hindsight.</li><li>The goal is to study the “why it made sense at the time,” not what was “right” or “wrong.”</li></ul></li><li><b>Drills Expose Hidden Assumptions</b>:<ul><li>Even experienced responders bring unspoken mental models into incidents.</li><li>The drill revealed assumptions about communication flows, authority boundaries, and vendor interactions that were not made explicit in planning.</li></ul></li><li><b>The Value of Human Expertise</b>:<ul><li>Everyone involved in this incident brought an unparalleled level of expertise to the work. </li><li>Often this kind of expertise goes unnoticed or is taken for granted, however this kind of knowledge is precisely what makes for smoother, better coordinated (and sometimes), faster incident response.</li></ul></li><li><b>Importance of Framing</b>:<ul><li>The way questions are asked in retrospectives can shape what is revealed—e.g., “What made that hard?” is more productive than “What did you miss?”</li><li>Reframing incidents around constraints and tradeoffs leads to deeper insight.</li></ul></li><li><b>Team Learning and Culture</b>:<ul><li>Safe, high-trust environments enable better learning during drills.</li><li>Psychological safety allows team members to admit confusion or raise alternate interpretations during real incidents.</li></ul></li></ul><p><b>Resources and References</b></p><ul><li><a href='https://podcast.thevoid.community/1793843/episodes/17564236-uptime-labs-and-the-multi-party-dilemma-part-i'>Episode I</a></li><li><a href='https://www.researchgate.net/publication/327427067_The_Theory_of_Graceful_Extensibility_Basic_rules_that_govern_adaptive_systems'>Model of Overload/Saturation as part of the Theory of Graceful Extensibility</a></li><li><a href='https://surfingcomplexity.blog/2017/06/24/a-conjecture-on-why-reliable-systems-fail/'>Lorin&apos;s Law</a></li></ul>]]></content:encoded>
    <enclosure url="https://www.buzzsprout.com/1793843/episodes/17564423-uptime-labs-and-the-multi-party-dilemma-part-ii.mp3" length="40529347" type="audio/mpeg" />
    <itunes:image href="https://storage.buzzsprout.com/am65mjc32seqfyacew14gqum840l?.jpg" />
    <itunes:author>Courtney Nash</itunes:author>
    <guid isPermaLink="false">Buzzsprout-17564423</guid>
    <pubDate>Wed, 06 Aug 2025 06:00:00 -0700</pubDate>
    <podcast:transcript url="https://www.buzzsprout.com/1793843/17564423/transcript" type="text/html" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/17564423/transcript.json" type="application/json" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/17564423/transcript.srt" type="application/x-subrip" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/17564423/transcript.vtt" type="text/vtt" />
    <itunes:duration>3373</itunes:duration>
    <itunes:keywords></itunes:keywords>
    <itunes:season>2</itunes:season>
    <itunes:episode>3</itunes:episode>
    <itunes:episodeType>full</itunes:episodeType>
    <itunes:explicit>false</itunes:explicit>
  </item>
  <item>
    <itunes:title>Uptime Labs and the Multi-Party Dilemma (Part I)</itunes:title>
    <title>Uptime Labs and the Multi-Party Dilemma (Part I)</title>
    <itunes:summary><![CDATA[Watch on YouTube In this episode I'm joined by a group of seasoned incident response professionals to discuss a simulated incident drill conducted on the Uptime Labs platform. The conversation centers around the Multi-party Dilemma—the challenge of coordinating incident response across teams or organizations with different missions, contexts, or incentives. Eric Dobbs, our incident analyst, joins to break down the drill and provide deep insights into the incident dynamics, team interactions, ...]]></itunes:summary>
    <description><![CDATA[<p><a href='https://www.youtube.com/watch?v=dxZbvThQGDA'>Watch on YouTube</a></p><p>In this episode I&apos;m joined by a group of seasoned incident response professionals to discuss a simulated incident drill conducted on the Uptime Labs platform. The conversation centers around the <b>Multi-party Dilemma</b>—the challenge of coordinating incident response across teams or organizations with different missions, contexts, or incentives.</p><p>Eric Dobbs, our incident analyst, joins to break down the drill and provide deep insights into the incident dynamics, team interactions, and what true incident analysis looks like when it&apos;s done well. Participants Alex Elman and Sarah Butt, who served as deputy and lead incident commanders respectively during the drill, recount their roles and experiences, highlighting realistic stress responses, decision-making, and coordination failures and successes. Hamed Silatani, CEO of Uptime Labs, provides context and insights into the behind-the-scenes work he and his team provide as the other &quot;characters&quot; driving the narrative of the drill.</p><p>The episode uniquely showcases the <b>value of structured incident analysis</b> and the benefits of using drills to expose hidden assumptions and improve resilience in complex systems.</p><p>A few key highlights include:</p><ul><li>How detailed incident analysis leads to an understanding of the context and rationale behind responders&apos; actions, rather than identifying errors or assigning blame.</li><li>The real goal is to learn how the system and people <em>actually</em> function, not just fix a broken component.</li><li>Themes revealed by the analysis and subsequent discussion<ul><li><b>Saturation</b> and the value of trust in delegation (especially between Sarah and Alex).</li><li>The role of <b>deep expertise</b> and how it often makes work appear effortless.</li><li>Importance of recognizing the <b><em>real work</em></b><b> done during incidents</b>—often messy and improvisational.</li></ul></li></ul><p><b>References/Resources </b></p><ul><li><a href='https://uptimelabs.io/what-experts-see-that-the-rest-of-us-miss-during-incidents/'>What Experts See That the Rest of Us Miss During Incidents</a></li><li><a href='https://uptimelabs.io/virtual-festival-2025/'>Incident Fest</a> (Uptime Labs event)</li><li><a href='https://www.researchgate.net/publication/220628177_Beyond_Simon&apos;s_Slice_Five_Fundamental_Trade-Offs_that_Bound_the_Performance_of_Macrocognitive_Work_Systems'>Law of Fluency</a> </li><li><a href='https://www.researchgate.net/publication/376354074_Handling_the_Multi-Party_Dilemma'>Handling the Multi-Party Dilemma</a> (Sarah &amp; Alex paper)</li><li><a href='https://www.youtube.com/watch?v=CbSiKAtO7Fk'>Embracing the Multi-Party Dilemma</a> (Sarah &amp; Alex conference talk)</li></ul><p><br/></p><p><br/></p><p><br/></p>]]></description>
    <content:encoded><![CDATA[<p><a href='https://www.youtube.com/watch?v=dxZbvThQGDA'>Watch on YouTube</a></p><p>In this episode I&apos;m joined by a group of seasoned incident response professionals to discuss a simulated incident drill conducted on the Uptime Labs platform. The conversation centers around the <b>Multi-party Dilemma</b>—the challenge of coordinating incident response across teams or organizations with different missions, contexts, or incentives.</p><p>Eric Dobbs, our incident analyst, joins to break down the drill and provide deep insights into the incident dynamics, team interactions, and what true incident analysis looks like when it&apos;s done well. Participants Alex Elman and Sarah Butt, who served as deputy and lead incident commanders respectively during the drill, recount their roles and experiences, highlighting realistic stress responses, decision-making, and coordination failures and successes. Hamed Silatani, CEO of Uptime Labs, provides context and insights into the behind-the-scenes work he and his team provide as the other &quot;characters&quot; driving the narrative of the drill.</p><p>The episode uniquely showcases the <b>value of structured incident analysis</b> and the benefits of using drills to expose hidden assumptions and improve resilience in complex systems.</p><p>A few key highlights include:</p><ul><li>How detailed incident analysis leads to an understanding of the context and rationale behind responders&apos; actions, rather than identifying errors or assigning blame.</li><li>The real goal is to learn how the system and people <em>actually</em> function, not just fix a broken component.</li><li>Themes revealed by the analysis and subsequent discussion<ul><li><b>Saturation</b> and the value of trust in delegation (especially between Sarah and Alex).</li><li>The role of <b>deep expertise</b> and how it often makes work appear effortless.</li><li>Importance of recognizing the <b><em>real work</em></b><b> done during incidents</b>—often messy and improvisational.</li></ul></li></ul><p><b>References/Resources </b></p><ul><li><a href='https://uptimelabs.io/what-experts-see-that-the-rest-of-us-miss-during-incidents/'>What Experts See That the Rest of Us Miss During Incidents</a></li><li><a href='https://uptimelabs.io/virtual-festival-2025/'>Incident Fest</a> (Uptime Labs event)</li><li><a href='https://www.researchgate.net/publication/220628177_Beyond_Simon&apos;s_Slice_Five_Fundamental_Trade-Offs_that_Bound_the_Performance_of_Macrocognitive_Work_Systems'>Law of Fluency</a> </li><li><a href='https://www.researchgate.net/publication/376354074_Handling_the_Multi-Party_Dilemma'>Handling the Multi-Party Dilemma</a> (Sarah &amp; Alex paper)</li><li><a href='https://www.youtube.com/watch?v=CbSiKAtO7Fk'>Embracing the Multi-Party Dilemma</a> (Sarah &amp; Alex conference talk)</li></ul><p><br/></p><p><br/></p><p><br/></p>]]></content:encoded>
    <enclosure url="https://www.buzzsprout.com/1793843/episodes/17564236-uptime-labs-and-the-multi-party-dilemma-part-i.mp3" length="34251440" type="audio/mpeg" />
    <itunes:image href="https://storage.buzzsprout.com/vrvw3joxeys430unz4ak3i9oq1o3?.jpg" />
    <itunes:author>Courtney Nash</itunes:author>
    <guid isPermaLink="false">Buzzsprout-17564236</guid>
    <pubDate>Tue, 29 Jul 2025 06:00:00 -0700</pubDate>
    <podcast:transcript url="https://www.buzzsprout.com/1793843/17564236/transcript" type="text/html" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/17564236/transcript.json" type="application/json" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/17564236/transcript.srt" type="application/x-subrip" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/17564236/transcript.vtt" type="text/vtt" />
    <itunes:duration>2850</itunes:duration>
    <itunes:keywords></itunes:keywords>
    <itunes:season>2</itunes:season>
    <itunes:episode>2</itunes:episode>
    <itunes:episodeType>full</itunes:episodeType>
    <itunes:explicit>false</itunes:explicit>
  </item>
  <item>
    <itunes:title>Canva and the Thundering Herd</itunes:title>
    <title>Canva and the Thundering Herd</title>
    <itunes:summary><![CDATA[Greetings fellow incident nerds, and welcome to Season 2 of The VOID podcast. The main new thing for this new season is we’re now available in video—so if you’re listening to this and prefer watching me make odd faces and nod a lot, you can find us here on YouTube.  The other new thing is we now have sponsors! These folks help make this podcast possible, but they don’t have any say over who joins us or what we talk about, so fear not.  This episode’s sponsor is Uptime Labs. Uptime L...]]></itunes:summary>
    <description><![CDATA[<p>Greetings fellow incident nerds, and welcome to Season 2 of The VOID podcast. The main new thing for this new season is we’re now available in video—so if you’re listening to this and prefer watching me make odd faces and nod a lot, you can <a href='https://www.youtube.com/watch?v=3lSslOY4rPs'>find us here on YouTube</a>. </p><p>The other new thing is we now have sponsors! These folks help make this podcast possible, but they don’t have any say over who joins us or what we talk about, so fear not. </p><p>This episode’s sponsor is Uptime Labs. Uptime Labs is a pioneering platform specializing in immersive incident response training. Their solution helps technical teams build confidence and expertise through realistic simulations that mirror real-world outages and security incidents. When most of investment these days in the incident space goes to technology and process, Uptime Labs focuses on sharpening the human element of incident response.</p><p>In this episode, we talk to Simon Newton, Head of Platforms at Canva, about their first public incident report. It’s not their first incident by any means, but it’s the first time they chose as a company to invest in sharing the details of an incident with the rest of us, which of course we’re big fans of here at the VOID. </p><p>We discuss:</p><ul><li>What led to Canva finally deciding to publish a public incident report</li><li>What the size and nature of their incident response looks like (this incident involved around 20 different people!)</li><li>Their progression from a handful of engineers handling incidents to having a dedicated Incident Command (IC) role</li><li>Avoiding blame when a known performance fix was ready to be deployed but hadn&apos;t yet, which contributed to the incident getting worse as it progressed</li><li>The various ways the people involved in the incident collaborated and improvised to resolve it</li></ul><p><br/></p>]]></description>
    <content:encoded><![CDATA[<p>Greetings fellow incident nerds, and welcome to Season 2 of The VOID podcast. The main new thing for this new season is we’re now available in video—so if you’re listening to this and prefer watching me make odd faces and nod a lot, you can <a href='https://www.youtube.com/watch?v=3lSslOY4rPs'>find us here on YouTube</a>. </p><p>The other new thing is we now have sponsors! These folks help make this podcast possible, but they don’t have any say over who joins us or what we talk about, so fear not. </p><p>This episode’s sponsor is Uptime Labs. Uptime Labs is a pioneering platform specializing in immersive incident response training. Their solution helps technical teams build confidence and expertise through realistic simulations that mirror real-world outages and security incidents. When most of investment these days in the incident space goes to technology and process, Uptime Labs focuses on sharpening the human element of incident response.</p><p>In this episode, we talk to Simon Newton, Head of Platforms at Canva, about their first public incident report. It’s not their first incident by any means, but it’s the first time they chose as a company to invest in sharing the details of an incident with the rest of us, which of course we’re big fans of here at the VOID. </p><p>We discuss:</p><ul><li>What led to Canva finally deciding to publish a public incident report</li><li>What the size and nature of their incident response looks like (this incident involved around 20 different people!)</li><li>Their progression from a handful of engineers handling incidents to having a dedicated Incident Command (IC) role</li><li>Avoiding blame when a known performance fix was ready to be deployed but hadn&apos;t yet, which contributed to the incident getting worse as it progressed</li><li>The various ways the people involved in the incident collaborated and improvised to resolve it</li></ul><p><br/></p>]]></content:encoded>
    <enclosure url="https://www.buzzsprout.com/1793843/episodes/17162771-canva-and-the-thundering-herd.mp3" length="26968178" type="audio/mpeg" />
    <itunes:image href="https://storage.buzzsprout.com/v488gorjvo5ax0ythqe0fqi5ts5k?.jpg" />
    <itunes:author>Courtney Nash</itunes:author>
    <guid isPermaLink="false">Buzzsprout-17162771</guid>
    <pubDate>Wed, 14 May 2025 11:00:00 -0700</pubDate>
    <podcast:transcript url="https://www.buzzsprout.com/1793843/17162771/transcript" type="text/html" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/17162771/transcript.json" type="application/json" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/17162771/transcript.srt" type="application/x-subrip" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/17162771/transcript.vtt" type="text/vtt" />
    <itunes:duration>2244</itunes:duration>
    <itunes:keywords></itunes:keywords>
    <itunes:season>2</itunes:season>
    <itunes:episode>1</itunes:episode>
    <itunes:episodeType>full</itunes:episodeType>
    <itunes:explicit>false</itunes:explicit>
  </item>
  <item>
    <itunes:title>Episode 8: A Tale of A Near Miss</itunes:title>
    <title>Episode 8: A Tale of A Near Miss</title>
    <itunes:summary><![CDATA[On this episode of the VOID podcast, I’m joined by Nick Travaglini, who is a Technical Customer Success Manager at Honeycomb. Nick wrote up a near miss that his team tackled towards the end of 2023, and I’ve been really wanting to discuss a near miss incident report for a very long time. What’s a Near Miss you might ask, or how is that an incident, or is it? What IS an incident? Keep listening, because we’re going to get into those questions, along with discussing whether or not it’s a good i...]]></itunes:summary>
    <description><![CDATA[<p>On this episode of the VOID podcast, I’m joined by Nick Travaglini, who is a Technical Customer Success Manager at Honeycomb. Nick wrote up a near miss that his team tackled towards the end of 2023, and I’ve been really wanting to discuss a near miss incident report for a very long time. What’s a Near Miss you might ask, or how is that an incident, or is it? What IS an incident? Keep listening, because we’re going to get into those questions, along with discussing whether or not it’s a good idea to say nasty things about other companies in your incident reports. </p><p><b>Related Resources</b></p><ul><li><a href='https://www.honeycomb.io/blog/preempting-problems-sociotechnical-system'>Preempting Problems in a Sociotechnical System</a> (the incident report)</li><li><a href='https://humanisticsystems.com/2016/12/05/the-varieties-of-human-work/'>Work as Imagined vs Work as Done</a></li><li><a href='https://resilienceinsoftware.org/'>Resilience in Software Foundation</a></li><li><a href='https://www.upress.umn.edu/9781517904876/on-the-mode-of-existence-of-technical-objects/'>On the Mode of Existence of Technical Objects</a></li><li><a href='https://www.dukeupress.edu/hitting-the-brakes'>Hitting the Brakes</a></li><li><a href='https://www.thevoid.community/report-2024'>2024 VOID Report</a><br/><br/> </li></ul>]]></description>
    <content:encoded><![CDATA[<p>On this episode of the VOID podcast, I’m joined by Nick Travaglini, who is a Technical Customer Success Manager at Honeycomb. Nick wrote up a near miss that his team tackled towards the end of 2023, and I’ve been really wanting to discuss a near miss incident report for a very long time. What’s a Near Miss you might ask, or how is that an incident, or is it? What IS an incident? Keep listening, because we’re going to get into those questions, along with discussing whether or not it’s a good idea to say nasty things about other companies in your incident reports. </p><p><b>Related Resources</b></p><ul><li><a href='https://www.honeycomb.io/blog/preempting-problems-sociotechnical-system'>Preempting Problems in a Sociotechnical System</a> (the incident report)</li><li><a href='https://humanisticsystems.com/2016/12/05/the-varieties-of-human-work/'>Work as Imagined vs Work as Done</a></li><li><a href='https://resilienceinsoftware.org/'>Resilience in Software Foundation</a></li><li><a href='https://www.upress.umn.edu/9781517904876/on-the-mode-of-existence-of-technical-objects/'>On the Mode of Existence of Technical Objects</a></li><li><a href='https://www.dukeupress.edu/hitting-the-brakes'>Hitting the Brakes</a></li><li><a href='https://www.thevoid.community/report-2024'>2024 VOID Report</a><br/><br/> </li></ul>]]></content:encoded>
    <enclosure url="https://www.buzzsprout.com/1793843/episodes/16700096-episode-8-a-tale-of-a-near-miss.mp3" length="25676372" type="audio/mpeg" />
    <itunes:image href="https://storage.buzzsprout.com/j9v1fog6lboiirnlwnwg1ypm4fe5?.jpg" />
    <itunes:author>Courtney Nash</itunes:author>
    <guid isPermaLink="false">Buzzsprout-16700096</guid>
    <pubDate>Fri, 28 Feb 2025 08:00:00 -0800</pubDate>
    <podcast:transcript url="https://www.buzzsprout.com/1793843/16700096/transcript" type="text/html" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/16700096/transcript.json" type="application/json" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/16700096/transcript.srt" type="application/x-subrip" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/16700096/transcript.vtt" type="text/vtt" />
    <podcast:soundbite startTime="31.0" duration="38.5" />
    <itunes:duration>2136</itunes:duration>
    <itunes:keywords></itunes:keywords>
    <itunes:episode>8</itunes:episode>
    <itunes:episodeType>full</itunes:episodeType>
    <itunes:explicit>false</itunes:explicit>
  </item>
  <item>
    <itunes:title>Episode 7: When Uptime Met Downtime</itunes:title>
    <title>Episode 7: When Uptime Met Downtime</title>
    <itunes:summary><![CDATA[We took a bit of a hiatus from recording last year, but we're back with an episode that I think everyone is really going to enjoy. Late last year, John Allspaw told me about this new company called Uptime Labs. They simulate software incidents, giving people a safe and constructive environment in which to experience incidents, practice what response is like, and bring what they learn back to their own organizations.  For the record, this is not a sponsored podcast. I legitimately just love wh...]]></itunes:summary>
    <description><![CDATA[<p>We took a bit of a hiatus from recording last year, but we&apos;re back with an episode that I think everyone is really going to enjoy. Late last year, John Allspaw told me about this new company called Uptime Labs. They simulate software incidents, giving people a safe and constructive environment in which to experience incidents, practice what response is like, and bring what they learn back to their own organizations.<br/><br/>For the record, this is not a sponsored podcast. I legitimately just love what they do. And I had the sincere privilege to meet Uptime&apos;s cofounder and CEO, Hamed Silatani at SRECon EMEA in November, where he gave a fantastic talk about some of the things they&apos;ve learned about incident response for running hundreds of simulations for their customers.<br/><br/>They recently had their first serious outage of their own platform. And so Hamed is joined by Joe McEvitt, cofounder and director of engineering at Uptime to discuss with me the one time that Uptime met downtime.</p>]]></description>
    <content:encoded><![CDATA[<p>We took a bit of a hiatus from recording last year, but we&apos;re back with an episode that I think everyone is really going to enjoy. Late last year, John Allspaw told me about this new company called Uptime Labs. They simulate software incidents, giving people a safe and constructive environment in which to experience incidents, practice what response is like, and bring what they learn back to their own organizations.<br/><br/>For the record, this is not a sponsored podcast. I legitimately just love what they do. And I had the sincere privilege to meet Uptime&apos;s cofounder and CEO, Hamed Silatani at SRECon EMEA in November, where he gave a fantastic talk about some of the things they&apos;ve learned about incident response for running hundreds of simulations for their customers.<br/><br/>They recently had their first serious outage of their own platform. And so Hamed is joined by Joe McEvitt, cofounder and director of engineering at Uptime to discuss with me the one time that Uptime met downtime.</p>]]></content:encoded>
    <enclosure url="https://www.buzzsprout.com/1793843/episodes/16528135-episode-7-when-uptime-met-downtime.mp3" length="37342091" type="audio/mpeg" />
    <itunes:image href="https://storage.buzzsprout.com/toy71ux2w4cnwk26icov33df01lj?.jpg" />
    <itunes:author>Courtney Nash</itunes:author>
    <guid isPermaLink="false">Buzzsprout-16528135</guid>
    <pubDate>Thu, 30 Jan 2025 11:00:00 -0800</pubDate>
    <podcast:transcript url="https://www.buzzsprout.com/1793843/16528135/transcript" type="text/html" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/16528135/transcript.json" type="application/json" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/16528135/transcript.srt" type="application/x-subrip" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/16528135/transcript.vtt" type="text/vtt" />
    <podcast:soundbite startTime="1281.0" duration="33.0" />
    <itunes:duration>3108</itunes:duration>
    <itunes:keywords></itunes:keywords>
    <itunes:episode>7</itunes:episode>
    <itunes:episodeType>full</itunes:episodeType>
    <itunes:explicit>false</itunes:explicit>
  </item>
  <item>
    <itunes:title>Episode 6: Laura Nolan and Control Pain</itunes:title>
    <title>Episode 6: Laura Nolan and Control Pain</title>
    <itunes:summary><![CDATA[In the second episode of the VOID podcast, Courtney Wang, an SRE at Reddit, said that he was inspired to start writing more in-depth narrative incident reports after reading the write-up of the Slack January 4th, 2021 outage. That incident report, along with many other excellent ones, was penned by Laura Nolan and I've been trying to get her on this podcast since I started it.   So, this is a very exciting episode for me. And for you all, it's going to be a bit different because instead of ju...]]></itunes:summary>
    <description><![CDATA[<p>In the <a href='https://podcast.thevoid.community/1793843/9633162-episode-2-reddit-and-the-gamestop-shenanigans'>second episode</a> of the VOID podcast, Courtney Wang, an SRE at Reddit, said that he was inspired to start writing more in-depth narrative incident reports after reading the write-up of the <a href='https://slack.engineering/slacks-outage-on-january-4th-2021/#:~:text=On%20January%204th%2C%20one%20of,to%20scale%20transparently%20to%20us.'>Slack January 4th, 2021 outage</a>. That incident report, along with many other excellent ones, was penned by <a href='https://twitter.com/lauralifts'>Laura Nolan</a> and I&apos;ve been trying to get her on this podcast since I started it. <br/><br/>So, this is a very exciting episode for me. And for you all, it&apos;s going to be a bit different because instead of just discussing a single incident that Laura has written about, we get to lean on and learn from her accumulated knowledge doing this for quite a few organizations. And she&apos;s come with <em>opinions</em>.  <br/><br/>A fun fact about this episode, I was going to title it &quot;Laura Nolan and Control Plane Incidents,&quot; but the automated transcription service that I use, which is typically pretty spot on (thanks, Descript!), kept changing &quot;plane&quot; to &quot;pain&quot; and well, you&apos;re about to find out just how ironic that actually is... <br/><br/>We discussed: </p><ul><li>A set of incidents she&apos;s been involved with that featured some form of control plane or automation as a contributing factor to the incident.</li><li>What we can learn from fields of study like Resilience Engineering, such as the notion of Joint Cognitive Systems</li><li>Other notable incidents that have similar factors</li><li>Ways that we can better factor in human-computer collaboration in tooling to help  make our lives easier when it comes to handling incidents</li></ul><p>References:<br/><a href='https://slack.engineering/slacks-outage-on-january-4th-2021/#:~:text=On%20January%204th%2C%20one%20of,to%20scale%20transparently%20to%20us.'>Slack&apos;s Outage on Jan 4th 2021</a><br/><a href='https://slack.engineering/a-terrible-horrible-no-good-very-bad-day-at-slack/'>A Terrible, Horrible, No-Good, Very Bad Day at Slack</a><br/><a href='https://sre.google/sre-book/automation-at-google/#xref_automation_diskerase-sidebar'>Google&apos;s &quot;satpocalypse&quot;</a><br/><a href='https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/'>Meta (Facebook) outage</a><br/><a href='https://www.reddit.com/r/RedditEng/comments/11xx5o0/you_broke_reddit_the_piday_outage/'>Reddit Pi-day outage</a><br/><a href='https://www.complexcognition.co.uk/2021/06/ironies-of-automation.html'>Ironies of Automation </a>(Lissane Bainbridge)</p><p><br/></p>]]></description>
    <content:encoded><![CDATA[<p>In the <a href='https://podcast.thevoid.community/1793843/9633162-episode-2-reddit-and-the-gamestop-shenanigans'>second episode</a> of the VOID podcast, Courtney Wang, an SRE at Reddit, said that he was inspired to start writing more in-depth narrative incident reports after reading the write-up of the <a href='https://slack.engineering/slacks-outage-on-january-4th-2021/#:~:text=On%20January%204th%2C%20one%20of,to%20scale%20transparently%20to%20us.'>Slack January 4th, 2021 outage</a>. That incident report, along with many other excellent ones, was penned by <a href='https://twitter.com/lauralifts'>Laura Nolan</a> and I&apos;ve been trying to get her on this podcast since I started it. <br/><br/>So, this is a very exciting episode for me. And for you all, it&apos;s going to be a bit different because instead of just discussing a single incident that Laura has written about, we get to lean on and learn from her accumulated knowledge doing this for quite a few organizations. And she&apos;s come with <em>opinions</em>.  <br/><br/>A fun fact about this episode, I was going to title it &quot;Laura Nolan and Control Plane Incidents,&quot; but the automated transcription service that I use, which is typically pretty spot on (thanks, Descript!), kept changing &quot;plane&quot; to &quot;pain&quot; and well, you&apos;re about to find out just how ironic that actually is... <br/><br/>We discussed: </p><ul><li>A set of incidents she&apos;s been involved with that featured some form of control plane or automation as a contributing factor to the incident.</li><li>What we can learn from fields of study like Resilience Engineering, such as the notion of Joint Cognitive Systems</li><li>Other notable incidents that have similar factors</li><li>Ways that we can better factor in human-computer collaboration in tooling to help  make our lives easier when it comes to handling incidents</li></ul><p>References:<br/><a href='https://slack.engineering/slacks-outage-on-january-4th-2021/#:~:text=On%20January%204th%2C%20one%20of,to%20scale%20transparently%20to%20us.'>Slack&apos;s Outage on Jan 4th 2021</a><br/><a href='https://slack.engineering/a-terrible-horrible-no-good-very-bad-day-at-slack/'>A Terrible, Horrible, No-Good, Very Bad Day at Slack</a><br/><a href='https://sre.google/sre-book/automation-at-google/#xref_automation_diskerase-sidebar'>Google&apos;s &quot;satpocalypse&quot;</a><br/><a href='https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/'>Meta (Facebook) outage</a><br/><a href='https://www.reddit.com/r/RedditEng/comments/11xx5o0/you_broke_reddit_the_piday_outage/'>Reddit Pi-day outage</a><br/><a href='https://www.complexcognition.co.uk/2021/06/ironies-of-automation.html'>Ironies of Automation </a>(Lissane Bainbridge)</p><p><br/></p>]]></content:encoded>
    <enclosure url="https://www.buzzsprout.com/1793843/episodes/12683592-episode-6-laura-nolan-and-control-pain.mp3" length="20363962" type="audio/mpeg" />
    <itunes:image href="https://storage.buzzsprout.com/tqr5qfte8a2uhw5dhhu0jv7wvfyo?.jpg" />
    <itunes:author>Courtney Nash</itunes:author>
    <guid isPermaLink="false">Buzzsprout-12683592</guid>
    <pubDate>Tue, 25 Apr 2023 09:00:00 -0700</pubDate>
    <podcast:transcript url="https://www.buzzsprout.com/1793843/12683592/transcript" type="text/html" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/12683592/transcript.json" type="application/json" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/12683592/transcript.srt" type="application/x-subrip" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/12683592/transcript.vtt" type="text/vtt" />
    <podcast:soundbite startTime="1323.0" duration="29.5" />
    <itunes:duration>1694</itunes:duration>
    <itunes:keywords></itunes:keywords>
    <itunes:episode>6</itunes:episode>
    <itunes:episodeType>full</itunes:episodeType>
    <itunes:explicit>false</itunes:explicit>
  </item>
  <item>
    <itunes:title>Episode 5: Incident.io and The First Big Incident</itunes:title>
    <title>Episode 5: Incident.io and The First Big Incident</title>
    <itunes:summary><![CDATA[What happens when you use your own incident management software to manage your own incidents but said incident takes out your own incident management product? Tune in to find out...  We chat with engineer Lawrence Jones about: How their product is designed, and how that both contributed to, but also helped them quickly resolve, the incidentThe role that organizational scaling (hiring lots of folks quickly) can play in making incident response challengingWhat happens when reality doesn't line ...]]></itunes:summary>
    <description><![CDATA[<p>What happens when you use your own incident management software to manage your own incidents but said incident takes out your own incident management product? Tune in to find out...<br/><br/>We chat with engineer Lawrence Jones about:</p><ul><li>How their product is designed, and how that both contributed to, but also helped them quickly resolve, the incident</li><li>The role that organizational scaling (hiring lots of folks quickly) can play in making incident response challenging</li><li>What happens when reality doesn&apos;t line up with your assumptions about how your system(s) works</li><li>The importance of taking a step back and making sure the team is taking care of each other when you can get a break in the urgency of an incident</li></ul>]]></description>
    <content:encoded><![CDATA[<p>What happens when you use your own incident management software to manage your own incidents but said incident takes out your own incident management product? Tune in to find out...<br/><br/>We chat with engineer Lawrence Jones about:</p><ul><li>How their product is designed, and how that both contributed to, but also helped them quickly resolve, the incident</li><li>The role that organizational scaling (hiring lots of folks quickly) can play in making incident response challenging</li><li>What happens when reality doesn&apos;t line up with your assumptions about how your system(s) works</li><li>The importance of taking a step back and making sure the team is taking care of each other when you can get a break in the urgency of an incident</li></ul>]]></content:encoded>
    <enclosure url="https://www.buzzsprout.com/1793843/episodes/12252408-episode-5-incident-io-and-the-first-big-incident.mp3" length="22852994" type="audio/mpeg" />
    <itunes:image href="https://storage.buzzsprout.com/dshjjgyaciopx5w873q10sk8bdno?.jpg" />
    <itunes:author>Courtney Nash</itunes:author>
    <guid isPermaLink="false">Buzzsprout-12252408</guid>
    <pubDate>Tue, 14 Feb 2023 10:00:00 -0800</pubDate>
    <podcast:transcript url="https://www.buzzsprout.com/1793843/12252408/transcript" type="text/html" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/12252408/transcript.json" type="application/json" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/12252408/transcript.srt" type="application/x-subrip" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/12252408/transcript.vtt" type="text/vtt" />
    <itunes:duration>1902</itunes:duration>
    <itunes:keywords></itunes:keywords>
    <itunes:episode>5</itunes:episode>
    <itunes:episodeType>full</itunes:episodeType>
    <itunes:explicit>false</itunes:explicit>
  </item>
  <item>
    <itunes:title>Episode 4: Emily Ruppe and The Inaugural LFI Conference</itunes:title>
    <title>Episode 4: Emily Ruppe and The Inaugural LFI Conference</title>
    <itunes:summary><![CDATA[In this episode we take a delightful detour from our usual VOID programming to have Emily Ruppe, a Solutions Engineer at Jeli.io and member of the Learning From Incidents (LFI) community, on the program to discuss the upcoming LFI Conference happening in Denver in February. Find out more about the goals and some of the featured speakers for the event, and we hope to see you there!  Discussed in this episode: Jeli.io Learning From Incidents The LFI Conference (Feb 15-16, 2023 in Denver, CO) ]]></itunes:summary>
    <description><![CDATA[<p>In this episode we take a delightful detour from our usual VOID programming to have Emily Ruppe, a Solutions Engineer at Jeli.io and member of the Learning From Incidents (LFI) community, on the program to discuss the upcoming LFI Conference happening in Denver in February. Find out more about the goals and some of the featured speakers for the event, and we hope to see you there!<br/><br/>Discussed in this episode:<br/><a href='https://www.jeli.io/'>Jeli.io</a><br/><a href='https://www.learningfromincidents.io/'>Learning From Incidents</a><br/><a href='https://www.learningfromincidents.io/learning-from-incidents-conference-2023'>The LFI Conference</a> (Feb 15-16, 2023 in Denver, CO)</p>]]></description>
    <content:encoded><![CDATA[<p>In this episode we take a delightful detour from our usual VOID programming to have Emily Ruppe, a Solutions Engineer at Jeli.io and member of the Learning From Incidents (LFI) community, on the program to discuss the upcoming LFI Conference happening in Denver in February. Find out more about the goals and some of the featured speakers for the event, and we hope to see you there!<br/><br/>Discussed in this episode:<br/><a href='https://www.jeli.io/'>Jeli.io</a><br/><a href='https://www.learningfromincidents.io/'>Learning From Incidents</a><br/><a href='https://www.learningfromincidents.io/learning-from-incidents-conference-2023'>The LFI Conference</a> (Feb 15-16, 2023 in Denver, CO)</p>]]></content:encoded>
    <enclosure url="https://www.buzzsprout.com/1793843/episodes/12033276-episode-4-emily-ruppe-and-the-inaugural-lfi-conference.mp3" length="8714232" type="audio/mpeg" />
    <itunes:image href="https://storage.buzzsprout.com/xgupaixhxrwblmjro4obm52mhbe7?.jpg" />
    <itunes:author>Courtney Nash</itunes:author>
    <guid isPermaLink="false">Buzzsprout-12033276</guid>
    <pubDate>Thu, 12 Jan 2023 14:00:00 -0800</pubDate>
    <podcast:transcript url="https://www.buzzsprout.com/1793843/12033276/transcript" type="text/html" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/12033276/transcript.json" type="application/json" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/12033276/transcript.srt" type="application/x-subrip" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/12033276/transcript.vtt" type="text/vtt" />
    <itunes:duration>724</itunes:duration>
    <itunes:keywords></itunes:keywords>
    <itunes:episode>4</itunes:episode>
    <itunes:episodeType>full</itunes:episodeType>
    <itunes:explicit>false</itunes:explicit>
  </item>
  <item>
    <itunes:title>Episode 3: Spotify and A Year of Incidents</itunes:title>
    <title>Episode 3: Spotify and A Year of Incidents</title>
    <itunes:summary><![CDATA[If you or anyone you know  has listened to Spotify, you're likely familiar with their year end Wrapped tradition. You get a viral, shareable little summary of your favorite songs, albums and artists from the year. In this episode, I chat with Clint Byrum, an engineer whose team helps keep Spotify for Artists running, which in turn keeps well, Spotify running.   Each year, the team looks back at the incidents they've had in their own form of Wrapped. They tested hypotheses with incid...]]></itunes:summary>
    <description><![CDATA[<p>If you or anyone you know  has listened to Spotify, you&apos;re likely familiar with their year end Wrapped tradition. You get a viral, shareable little summary of your favorite songs, albums and artists from the year. In this episode, I chat with Clint Byrum, an engineer whose team helps keep Spotify for Artists running, which in turn keeps well, Spotify running. <br/><br/>Each year, the team looks back at the incidents they&apos;ve had in their own form of Wrapped. They tested hypotheses with incident data that they&apos;ve collected, found some interesting results and patterns, and helped push their team and larger organization to better understand what they can learn from incidents and how they can make their systems better support artists on their platform.<br/><br/>We discussed:</p><ul><li>Metrics, both good and bad</li><li>Moving away from MTTR after they found it to be unreliable</li><li>How incident analysis is akin to archeology</li><li>Getting managers/executives interested in incident reviews</li><li>The value of studying near misses along with actual incidents</li></ul>]]></description>
    <content:encoded><![CDATA[<p>If you or anyone you know  has listened to Spotify, you&apos;re likely familiar with their year end Wrapped tradition. You get a viral, shareable little summary of your favorite songs, albums and artists from the year. In this episode, I chat with Clint Byrum, an engineer whose team helps keep Spotify for Artists running, which in turn keeps well, Spotify running. <br/><br/>Each year, the team looks back at the incidents they&apos;ve had in their own form of Wrapped. They tested hypotheses with incident data that they&apos;ve collected, found some interesting results and patterns, and helped push their team and larger organization to better understand what they can learn from incidents and how they can make their systems better support artists on their platform.<br/><br/>We discussed:</p><ul><li>Metrics, both good and bad</li><li>Moving away from MTTR after they found it to be unreliable</li><li>How incident analysis is akin to archeology</li><li>Getting managers/executives interested in incident reviews</li><li>The value of studying near misses along with actual incidents</li></ul>]]></content:encoded>
    <enclosure url="https://www.buzzsprout.com/1793843/episodes/11537920-episode-3-spotify-and-a-year-of-incidents.mp3" length="22891871" type="audio/mpeg" />
    <itunes:image href="https://storage.buzzsprout.com/be5pxhtoifh2uali7dd1or27vqvi?.jpg" />
    <itunes:author>Courtney Nash</itunes:author>
    <guid isPermaLink="false">Buzzsprout-11537920</guid>
    <pubDate>Thu, 20 Oct 2022 11:00:00 -0700</pubDate>
    <podcast:transcript url="https://www.buzzsprout.com/1793843/11537920/transcript" type="text/html" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/11537920/transcript.json" type="application/json" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/11537920/transcript.srt" type="application/x-subrip" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/11537920/transcript.vtt" type="text/vtt" />
    <podcast:soundbite startTime="387.0" duration="27.0" />
    <itunes:duration>1905</itunes:duration>
    <itunes:keywords></itunes:keywords>
    <itunes:season>1</itunes:season>
    <itunes:episode>3</itunes:episode>
    <itunes:episodeType>full</itunes:episodeType>
    <itunes:explicit>false</itunes:explicit>
  </item>
  <item>
    <itunes:title>Episode 2: Reddit and the Gamestop Shenanigans</itunes:title>
    <title>Episode 2: Reddit and the Gamestop Shenanigans</title>
    <itunes:summary><![CDATA[ At the end of January, 2021, a group of Reddit users organized what's called a "short squeeze." They  intended to wreak havoc on hedge funds that were shorting the stock of a struggling brick and mortar game retailer called GameStop. They were coordinating to buy more stock in the company and drive its price further up.  In large part, they were successful—at least for a little while. One hedge fund lost somewhere around $2 billion and one Reddit user purportedly made off with around $1...]]></itunes:summary>
    <description><![CDATA[<p><br/>At the end of January, 2021, a group of <a href='https://www.reddit.com/'>Reddit</a> users organized what&apos;s called a &quot;short squeeze.&quot; They  intended to wreak havoc on hedge funds that were shorting the stock of a struggling brick and mortar game retailer called GameStop. They were coordinating to buy more stock in the company and drive its price further up.<br/><br/>In large part, they were successful—at least for a little while. One hedge fund lost somewhere around $2 billion and one Reddit user purportedly made off with around $13 million. Things managed to get even weirder from there, when online trading company Robinhood restricted trading for GameStop shares and sent its values plummeting losing three fourths of its value in just over an hour. But that&apos;s less relevant to this episode. <br/><br/>What matters is that while all this was happening, traffic to a very specific page on Reddit, called a subreddit, <a href='https://www.reddit.com/r/wallstreetbets/'>r/wallstreetbets</a> went to the moon. Long after the dust had settled, and the team had a chance to recover and reflect, some of the engineers wrote up an anthology of reports based on the numerous incidents they had that week. We talk to <a href='https://twitter.com/CKWang'>Courtney Wang</a>, <a href='https://twitter.com/garrettleeh'>Garrett Hoffman</a>, and <a href='https://www.linkedin.com/in/frang/'>Fran Garcia</a> about those incidents, and their write-ups, in this episode.<br/><br/>A few of the things we discussed include:</p><ul><li>The precarious dynamic where business successes (traffic surges based on cultural whims) are hard to predict, and can hit their systems in wild and surprising ways.</li><li>How incidents like these have multiple contributing factors, not all of which are purely technical</li><li>How much they learned about their company&apos;s processes, assumptions, organizational boundaries, and other &quot;non-technical&quot; factors</li><li>How people are the source of resilience in these complex sociotechnical systems</li><li>Creating psychologically safe environments for people who respond to incidents</li><li>Their motivation for investing so much time and energy into analyzing, writing, and publishing these incident reviews</li><li>What studying near misses illuminated for them about how their systems work</li></ul><p><br/>Resources mentioned in this episode include:</p><ul><li>Reddit&apos;s <a href='https://www.reddit.com/r/RedditEng/comments/o4y1yv/the_rwallstreetbets_incident_anthology/'>r/wallstreetsbets incident anthology</a>, which links to all the reports we discuss.</li><li>&quot;<a href='https://www.youtube.com/watch?v=qNk_UfXcq6k'>Work as imagined and work as done</a>&quot; by Steven Shorrock (video)</li></ul><p><br/><br/></p>]]></description>
    <content:encoded><![CDATA[<p><br/>At the end of January, 2021, a group of <a href='https://www.reddit.com/'>Reddit</a> users organized what&apos;s called a &quot;short squeeze.&quot; They  intended to wreak havoc on hedge funds that were shorting the stock of a struggling brick and mortar game retailer called GameStop. They were coordinating to buy more stock in the company and drive its price further up.<br/><br/>In large part, they were successful—at least for a little while. One hedge fund lost somewhere around $2 billion and one Reddit user purportedly made off with around $13 million. Things managed to get even weirder from there, when online trading company Robinhood restricted trading for GameStop shares and sent its values plummeting losing three fourths of its value in just over an hour. But that&apos;s less relevant to this episode. <br/><br/>What matters is that while all this was happening, traffic to a very specific page on Reddit, called a subreddit, <a href='https://www.reddit.com/r/wallstreetbets/'>r/wallstreetbets</a> went to the moon. Long after the dust had settled, and the team had a chance to recover and reflect, some of the engineers wrote up an anthology of reports based on the numerous incidents they had that week. We talk to <a href='https://twitter.com/CKWang'>Courtney Wang</a>, <a href='https://twitter.com/garrettleeh'>Garrett Hoffman</a>, and <a href='https://www.linkedin.com/in/frang/'>Fran Garcia</a> about those incidents, and their write-ups, in this episode.<br/><br/>A few of the things we discussed include:</p><ul><li>The precarious dynamic where business successes (traffic surges based on cultural whims) are hard to predict, and can hit their systems in wild and surprising ways.</li><li>How incidents like these have multiple contributing factors, not all of which are purely technical</li><li>How much they learned about their company&apos;s processes, assumptions, organizational boundaries, and other &quot;non-technical&quot; factors</li><li>How people are the source of resilience in these complex sociotechnical systems</li><li>Creating psychologically safe environments for people who respond to incidents</li><li>Their motivation for investing so much time and energy into analyzing, writing, and publishing these incident reviews</li><li>What studying near misses illuminated for them about how their systems work</li></ul><p><br/>Resources mentioned in this episode include:</p><ul><li>Reddit&apos;s <a href='https://www.reddit.com/r/RedditEng/comments/o4y1yv/the_rwallstreetbets_incident_anthology/'>r/wallstreetsbets incident anthology</a>, which links to all the reports we discuss.</li><li>&quot;<a href='https://www.youtube.com/watch?v=qNk_UfXcq6k'>Work as imagined and work as done</a>&quot; by Steven Shorrock (video)</li></ul><p><br/><br/></p>]]></content:encoded>
    <enclosure url="https://www.buzzsprout.com/1793843/episodes/9633162-episode-2-reddit-and-the-gamestop-shenanigans.mp3" length="31742478" type="audio/mpeg" />
    <itunes:image href="https://storage.buzzsprout.com/g8fq3x42xrqjzox9fpgbpy2dhz05?.jpg" />
    <itunes:author>Courtney Nash</itunes:author>
    <guid isPermaLink="false">Buzzsprout-9633162</guid>
    <pubDate>Wed, 01 Dec 2021 08:00:00 -0800</pubDate>
    <podcast:transcript url="https://www.buzzsprout.com/1793843/9633162/transcript" type="text/html" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/9633162/transcript.json" type="application/json" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/9633162/transcript.srt" type="application/x-subrip" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/9633162/transcript.vtt" type="text/vtt" />
    <podcast:soundbite startTime="474.0" duration="24.0" />
    <itunes:duration>2642</itunes:duration>
    <itunes:keywords></itunes:keywords>
    <itunes:season>1</itunes:season>
    <itunes:episode>2</itunes:episode>
    <itunes:episodeType>full</itunes:episodeType>
    <itunes:explicit>false</itunes:explicit>
  </item>
  <item>
    <itunes:title>Episode 1: Honeycomb and the Kafka Migration</itunes:title>
    <title>Episode 1: Honeycomb and the Kafka Migration</title>
    <itunes:summary><![CDATA["We no longer felt confident about what the exact operational boundaries of our cluster were supposed to be."  In early 2021, observability company Honeycomb dealt with a series of outages related to their Kafka architectural migration, culminating in a 12-hour incident, which is an extremely long outage for the company. In this episode, we chat with two engineers involved in these incidents, Liz Fong-Jones and Fred Hebert, about the backstory that is summarized in this meta-analysis they pub...]]></itunes:summary>
    <description><![CDATA[<p>&quot;We no longer felt confident about what the exact operational boundaries of our cluster were supposed to be.&quot;<br/><br/>In early 2021, observability company <a href='https://www.honeycomb.io/'>Honeycomb</a> dealt with a series of outages related to their Kafka architectural migration, culminating in a 12-hour incident, which is an extremely long outage for the company. In this episode, we chat with two engineers involved in these incidents, <a href='https://twitter.com/lizthegrey'>Liz Fong-Jones</a> and <a href='https://twitter.com/mononcqc'>Fred Hebert</a>, about the backstory that is summarized in <a href='https://www.honeycomb.io/blog/kafka-migration-lessons-learned/'>this meta-analysis</a> they published in May. <br/><br/>We cover a wide range of topics beyond the specific technical details of the incident (which we also discuss), including:</p><ul><li>Complex socio-technical systems and the kinds of failures that can happen in them (they&apos;re always surprises)</li><li>Transparency and the benefits of companies sharing these outage reports</li><li>Safety margins, performance envelopes, and the role of expertise in developing a sense for them</li><li>Honeycomb&apos;s incident response philosophy and process</li><li>The cognitive costs of responding to incidents</li><li>What we can (and can&apos;t) learn from incident reports</li></ul><p>Resources mentioned in the episode:</p><ul><li><a href='https://www.honeycomb.io/blog/kafka-migration-lessons-learned/'>Kafka Migration and Lessons Learned</a> by  Honeycomb</li><li><a href='https://queue.acm.org/detail.cfm?id=3380779'>Managing the Hidden Costs of Coordination</a> by Laura McGuire</li><li><a href='https://queue.acm.org/detail.cfm?id=3380777'>Above the Line, Below the Line</a> by Richard Cook</li><li><a href='https://www.researchgate.net/publication/220579378_Those_found_responsible_have_been_sacked_Some_observations_on_the_usefulness_of_error'>&quot;Those found responsible have been sacked&quot;: Some observations on the usefulness of error</a> by Richard Cook and Christopher P. Nemeth</li></ul><p><br/><em>Published in partnership with </em><a href='https://www.indeed.com'><em>Indeed</em></a><em>.</em></p>]]></description>
    <content:encoded><![CDATA[<p>&quot;We no longer felt confident about what the exact operational boundaries of our cluster were supposed to be.&quot;<br/><br/>In early 2021, observability company <a href='https://www.honeycomb.io/'>Honeycomb</a> dealt with a series of outages related to their Kafka architectural migration, culminating in a 12-hour incident, which is an extremely long outage for the company. In this episode, we chat with two engineers involved in these incidents, <a href='https://twitter.com/lizthegrey'>Liz Fong-Jones</a> and <a href='https://twitter.com/mononcqc'>Fred Hebert</a>, about the backstory that is summarized in <a href='https://www.honeycomb.io/blog/kafka-migration-lessons-learned/'>this meta-analysis</a> they published in May. <br/><br/>We cover a wide range of topics beyond the specific technical details of the incident (which we also discuss), including:</p><ul><li>Complex socio-technical systems and the kinds of failures that can happen in them (they&apos;re always surprises)</li><li>Transparency and the benefits of companies sharing these outage reports</li><li>Safety margins, performance envelopes, and the role of expertise in developing a sense for them</li><li>Honeycomb&apos;s incident response philosophy and process</li><li>The cognitive costs of responding to incidents</li><li>What we can (and can&apos;t) learn from incident reports</li></ul><p>Resources mentioned in the episode:</p><ul><li><a href='https://www.honeycomb.io/blog/kafka-migration-lessons-learned/'>Kafka Migration and Lessons Learned</a> by  Honeycomb</li><li><a href='https://queue.acm.org/detail.cfm?id=3380779'>Managing the Hidden Costs of Coordination</a> by Laura McGuire</li><li><a href='https://queue.acm.org/detail.cfm?id=3380777'>Above the Line, Below the Line</a> by Richard Cook</li><li><a href='https://www.researchgate.net/publication/220579378_Those_found_responsible_have_been_sacked_Some_observations_on_the_usefulness_of_error'>&quot;Those found responsible have been sacked&quot;: Some observations on the usefulness of error</a> by Richard Cook and Christopher P. Nemeth</li></ul><p><br/><em>Published in partnership with </em><a href='https://www.indeed.com'><em>Indeed</em></a><em>.</em></p>]]></content:encoded>
    <enclosure url="https://www.buzzsprout.com/1793843/episodes/9471681-episode-1-honeycomb-and-the-kafka-migration.mp3" length="22705921" type="audio/mpeg" />
    <itunes:author>Courtney Nash</itunes:author>
    <guid isPermaLink="false">Buzzsprout-9471681</guid>
    <pubDate>Mon, 01 Nov 2021 12:00:00 -0700</pubDate>
    <podcast:transcript url="https://www.buzzsprout.com/1793843/9471681/transcript" type="text/html" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/9471681/transcript.json" type="application/json" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/9471681/transcript.srt" type="application/x-subrip" />
    <podcast:transcript url="https://www.buzzsprout.com/1793843/9471681/transcript.vtt" type="text/vtt" />
    <podcast:soundbite startTime="689.0" duration="23.5" />
    <itunes:duration>1891</itunes:duration>
    <itunes:keywords></itunes:keywords>
    <itunes:episodeType>full</itunes:episodeType>
    <itunes:explicit>false</itunes:explicit>
  </item>
</channel>
</rss>
