<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>Site Reliability Engineering on Engineering Leadership in AI &amp; Software</title>
		<link>https://engineering-leadership-preview.hinshelwood.com/tags/site-reliability-engineering/</link>
		<description>Recent content in Site Reliability Engineering on Engineering Leadership in AI &amp; Software</description>
		<generator>Hugo</generator>
		<language>en</language>
		
		
		
		
			<lastBuildDate>Tue, 16 Jun 2026 17:40:35 +0000</lastBuildDate>
		
			<atom:link href="https://engineering-leadership-preview.hinshelwood.com/tags/site-reliability-engineering/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>Resilience is Part of the Product, Not an Afterthought</title>
				<link>https://engineering-leadership-preview.hinshelwood.com/articles/resilience-is-part-of-the-product-not-an-afterthought/</link>
				<pubDate>Mon, 09 Jun 2025 09:00:00 +0000</pubDate>
				<guid>https://engineering-leadership-preview.hinshelwood.com/articles/resilience-is-part-of-the-product-not-an-afterthought/</guid>
				<description>Resilience must be designed into your product from the start, not added later or left to individual heroics. Building resilience means engineering for failure containment, rapid recovery, and continuous improvement, using tools like telemetry, feature flags, and safe deployment practices. Make resilience a core part of your development process and culture, treating it as a critical feature to avoid costly outages and business risks.</description>
			</item>
			<item>
				<title>How to Build for Business Resilience and Continuity</title>
				<link>https://engineering-leadership-preview.hinshelwood.com/articles/how-to-build-for-business-resilience-and-continuity/</link>
				<pubDate>Mon, 26 May 2025 09:00:00 +0000</pubDate>
				<guid>https://engineering-leadership-preview.hinshelwood.com/articles/how-to-build-for-business-resilience-and-continuity/</guid>
				<description>Building business resilience requires intentional design, strong observability, and aggressive decoupling so failures do not cascade across systems. Empower teams to act quickly, treat deployments as routine, and design for fast recovery using practices like chaos engineering and circuit breakers. Make resilience a core part of your culture and operations, not a one-time project, and use real metrics to guide continuous improvement.</description>
			</item>
			<item>
				<title>Fragile by Design: The Cost of Pretending to Be Resilient</title>
				<link>https://engineering-leadership-preview.hinshelwood.com/articles/fragile-by-design-the-cost-of-pretending-to-be-resilient/</link>
				<pubDate>Mon, 12 May 2025 09:00:00 +0000</pubDate>
				<guid>https://engineering-leadership-preview.hinshelwood.com/articles/fragile-by-design-the-cost-of-pretending-to-be-resilient/</guid>
				<description>Most systems fail under real pressure because resilience is often treated as a checkbox or afterthought rather than a core product capability that is engineered, tested, and verified in real-world conditions. High-profile outages at Spain’s grid, Oracle, and Heathrow show that bad engineering, shallow product thinking, and leadership denial lead to systemic fragility. To avoid costly failures, development managers must make resilience a disciplined, ongoing practice by designing for failure, testing recovery regularly, and confronting weaknesses head-on.</description>
			</item>
			<item>
				<title>Mastering Site Reliability: Insights from Azure DevOps on Building a Resilient Live Site Culture</title>
				<link>https://engineering-leadership-preview.hinshelwood.com/videos/mastering-site-reliability-insights-from-azure-devops-on-building-a-resilient-live-site-culture/</link>
				<pubDate>Thu, 04 Jun 2020 02:05:28 +0000</pubDate>
				<guid>https://engineering-leadership-preview.hinshelwood.com/videos/mastering-site-reliability-insights-from-azure-devops-on-building-a-resilient-live-site-culture/</guid>
				<description>The Azure DevOps team at Microsoft has built a resilient live site culture by prioritising transparency with customers, investing in comprehensive telemetry, and automating deployment processes. Cross-functional teams and structured incident response drive continuous improvement and reliability at scale. Development managers should focus on these practices to boost both agility and system robustness while maintaining customer trust.</description>
			</item>
			<item>
				<title>Building a Resilient Token Server: Engineering for Flow, Fault Tolerance, and Speed</title>
				<link>https://engineering-leadership-preview.hinshelwood.com/engineering-notes/building-a-resilient-token-server-engineering-for-flow-fault-tolerance-and-speed/</link>
				<pubDate>Thu, 08 May 2025 09:00:00 +0000</pubDate>
				<guid>https://engineering-leadership-preview.hinshelwood.com/engineering-notes/building-a-resilient-token-server-engineering-for-flow-fault-tolerance-and-speed/</guid>
				<description>Aiming for a resilient, fast, and fault-tolerant token counting system, the author replaced fragile server restarts with a batch-wide server lifecycle, added retry logic for transient failures, and implemented a local fallback to ensure uninterrupted processing. These changes improved reliability, reduced downtime, and provided clear logs for troubleshooting. Development managers should focus on building systems that handle real-world failures gracefully, prioritize flow, and include observability and fallback mechanisms from the start.</description>
			</item>
			<item>
				<title>Site Reliability Engineering</title>
				<link>https://engineering-leadership-preview.hinshelwood.com/tags/site-reliability-engineering/</link>
				<pubDate>Mon, 05 May 2025 10:17:24 +0000</pubDate>
				<guid>https://engineering-leadership-preview.hinshelwood.com/tags/site-reliability-engineering/</guid>
				<description>Site Reliability Engineering (SRE) is a discipline that utilises software engineering principles to develop scalable and reliable systems, effectively bridging the gap between development and operations. Originating from the need to embed reliability within the software development lifecycle, SRE ensures that systems maintain functionality and resilience under diverse conditions. This methodology prioritises automation, monitoring, and incident response, which allows teams to deliver consistent value in a sustainable manner. SRE teams establish service level objectives (SLOs) and service level indicators (SLIs) to create clear performance and reliability metrics, fostering a culture of accountability and continuous improvement. This proactive approach to problem-solving and engineering solutions to operational challenges enhances overall system performance, distinguishing SRE from traditional operations roles. By promoting shared responsibility for reliability across teams, SRE encourages collaboration and knowledge sharing, which not only improves user satisfaction but also drives positive business outcomes. Ultimately, the integration of reliability into the development process supports organisational strategic goals and enhances competitive advantage, making SRE a vital component in agile, DevOps, and product development frameworks.</description>
			</item>
	</channel>
</rss>
