In dealing with monitoring and observability strategy on large scales, two differing approaches manifest, what I am calling inherited and developed strategies. To effectively manage large-scale organizations, it is essential to balance the two strategies. This balance provides observability and monitoring tools, best practices, and solutions while fostering a sense of ownership among engineering teams, encouraging their "buy-in."

A reliability strategy that skews far on the inherited side means a team or small cohort that centrally manages and owns all of the reliability solutions in the organization. For example, the automatic distribution of metric or logging collectors to compute en masse. It could include the auto generation of alerts and alerting strategies for development teams. An inherited strategy prioritizes providing solutions for other teams with little input or customization for those teams. It attempts to cover a large swath of use cases through one or multiple general solutions applied for others.

Conversely, a developed strategy drops the central management and instead relies on individual teams' solutions to fill reliability requirement gaps. While this approach can encourage innovation and ownership within teams, it also presents challenges such as a lack of alignment across the organization and difficulties in scaling consistent practices. On the other hand, it allows teams to tailor solutions to their specific needs, potentially leading to more effective and context-sensitive reliability measures. One team may opt for one approach, another a different approach. Or, through trial and error, several approaches may be attempted until one is decided upon at a higher level by the engineering community within the organization. This strategy prioritizes individual team ownership but leaves behind the benefits of utilizing a centrally managed reliability team or department. This could lead to problems of alignment and strategy if too many teams are doing too many different things.

An inherited strategy, where engineering teams may have no reliability strategy or site reliability engineer of their own in house, is prone to a sense of "someone else handles reliability" forming. A "throw it over the wall" mentality, as many refer to it. With that comes a growing unfamiliarity with the centrally managed tooling provided to teams not included in the reliability team or department. This unfamiliarity can undermine the effectiveness of the organization’s reliability efforts, as teams may struggle to fully leverage the tools or integrate them into their workflows. Additionally, as organizations become more complex it may become more difficult for the central team to understand the requirements of their internal users as their numbers grow.

However, a strategy lacking some level of inherited principles loses the ability to align teams across large organizations to a cohesive strategy on how to properly measure what does matter to their customers. Without a cohesive strategy, it becomes increasingly difficult to have productive conversations across teams, or more broadly, vertical business department leadership. For instance, if one team measures reliability solely through uptime while another prioritizes error rates, aligning on a unified metric can become contentious. This lack of alignment often leads to misunderstandings or inefficiencies during cross-department discussions about overall reliability priorities. With no alignment on how metrics that matter are being collected and surfaced, can a team in department A realistically be asked to understand the individual strategies for reliability in department B, C, and D to have a conversation?

To strike a balance, engineering teams need clear guidance, not shipped solutions, on how to develop a reliability strategy and what approaches to use for different use cases. For example, this guidance could include detailed playbooks for setting reliability targets, frameworks for selecting appropriate observability tools, or templates for defining Service Level Objectives (SLOs) tailored to specific team needs. Additionally, they must develop a clear understanding of how to uncover the signals that reflect their customer experience, with support from an experienced reliability team when needed. This involves identifying key metrics, setting appropriate thresholds, and interpreting data to inform decision-making. This combined approach is the most optimal path. A mixture of inherited and developed strategies allows for broad alignment while ensuring engineering teams have enough skin in the game in regard to their reliability strategy to stay familiar with it. To check in on it, update it. To understand what is coming through to their pager. It allows for the reliability teams to not have to own everything, but to play an orchestrator role. Assisting where needed in both the day to day improvement of reliability by collaborating hands on with other teams while also developing and maintaining the broader strategy.