All Perspectives

Your Operations Center Is the Most Valuable Team in the Organization.

On This Page
  1. The Operations Center Has a Reputation
  2. The Operations Center Should Operate
  3. Alerts Need More Than a Recipient
  4. Autonomy Has to Be Earned
  5. A Great OC Is a Talent Pipeline
  6. Escalation Should Follow Expertise
  7. You Get the OC You Build

The Operations Center Has a Reputation

Have you ever worked with a truly great Operations Center?

Calling the Operations Center the most valuable team in an organization probably sounds like a hot take, and maybe it is, but there is a reason I make the argument. Operations Centers have developed a reputation in many engineering organizations as teams that acknowledge alerts, open tickets, and find someone else who can actually fix the problem. There is some truth behind that reputation. Plenty of Operations Centers function exactly that way, but that does not mean they have to. In many cases, it is simply the role the organization has built for them.

A great Operations Center can look very different. It can independently resolve much of the routine work required to keep a platform running and handle incidents that fall within its capabilities. When something moves beyond those capabilities, it should know when to bring in deeper technical expertise.

Throughout this Perspective, I will refer to the Operations Center as the OC. I will also use escalation team as a general term for the technical teams the OC may need to involve when a problem moves beyond its capabilities. Depending on the organization and the problem, that might mean SRE, Engineering, Network Operations, etc. The names and organizational structures change, but the relationship is largely the same.

That is where I think the real value of a great OC is found. It is not measured by the number of alerts it acknowledges or tickets it closes, but by how much engineering capacity it gives back to everyone around it.

The Operations Center Should Operate

The name should give away much of the job. An Operations Center should be able to operate.

That means doing more than receiving an alert, gathering a little context, and finding the team that owns the affected system. A capable OC should be able to investigate what happened, perform remediation within its area of responsibility, and validate that the problem has been resolved. It should also know when the problem requires expertise beyond its own.

For known and recurring problems, that entire process should often happen within the OC. The same can be true for incidents. Not every incident requires an SRE, engineer, or other specialized team simply because it has been labeled an incident. A mature OC can handle many incidents from detection through resolution, including the operational communication and status updates that accompany them.

There will always be a boundary. Some failures are unfamiliar or sufficiently severe that they require expertise the OC should not be expected to have. A great OC understands where that boundary is and provides a meaningful escalation when it has been crossed.

When that relationship works well, the engineering teams behind the OC remain accountable for the systems they build without needing to be the first responder to every problem those systems encounter. The OC handles what it is capable of handling, allowing the engineering teams behind it to remain focused until a problem actually requires their expertise.

Alerts Need More Than a Recipient

A capable OC needs more than access and good monitoring. It needs enough operational knowledge to understand what the platform is telling it and what should happen next. Every actionable alert the OC receives should have a runbook.

I will take that a step further. Every actionable alert should have a runbook, regardless of who is expected to respond to it.

An alert that requires human attention should have a documented expected human response. The runbook should provide enough context to understand why the alert fired, what should be investigated, and when the problem needs to move beyond the person currently responding to it. Without that context, alerting someone is little more than telling them that something is wrong and leaving them to figure out the rest.

The OC should have a meaningful role in developing and maintaining the runbooks it uses. The people repeatedly responding to an alert develop an understanding of how the failure actually presents in production, and that experience should make its way back into the documentation. Over time, the runbooks capture the operational knowledge the team has built rather than remaining a set of instructions handed down from somewhere else in the organization. As those runbooks improve, knowledge gained by experienced members of the OC becomes part of how the next person learns the platform rather than remaining knowledge they have to rediscover for themselves.

None of this means a runbook should remove judgment from the response. Production has a habit of finding new ways to fail. A good runbook provides a starting point and an expected response, while a capable OC still needs to recognize when the problem no longer fits the documented path.

Autonomy Has to Be Earned

A highly capable OC cannot be created simply by expanding production permissions. The access required to operate a platform safely has to develop alongside the capability and judgment of the team using it.

This takes time. An OC may begin with relatively narrow responsibilities and expand its footprint as the team gains experience and demonstrates that it can operate safely within its boundaries. Trust develops through that experience, and greater autonomy becomes a consequence of demonstrated capability rather than something granted in anticipation of it.

The relationship between the OC and its escalation teams is crucial throughout that process. The OC needs somewhere to take unfamiliar problems and gain a deeper understanding of the systems it supports. The escalation teams, in turn, need confidence that the OC understands its boundaries and will involve them when their expertise is genuinely necessary. Knowing when to ask for help is part of operating responsibly, not an indication that someone is incapable, and demonstrating that judgment is part of what allows the OC's responsibility to expand over time.

A Great OC Is a Talent Pipeline

A mature OC can be an unusually effective environment for developing talent because it provides continuous exposure to real production systems and the problems that occur within them.

There are often people on an OC team who are eager to understand more than what an alert or runbook tells them. They want to understand why something failed and why the remediation works rather than simply knowing which steps to follow. When the OC has enough responsibility to actually investigate and resolve problems, there is room for that curiosity to turn into experience.

Over time, repeatedly working with production systems builds a kind of knowledge that is difficult to reproduce through training alone. Following a runbook can turn into understanding why the runbook works, and eventually into recognizing when it does not. Troubleshooting becomes less about finding the next documented step and more about understanding the system well enough to determine what should happen next.

Not everyone in an OC will develop in the same direction, nor should that be the expectation. Some may gravitate toward deeper technical work or leadership, while others may build their careers within the OC itself. None of those paths is inherently better than another.

A strong OC therefore does more than protect the engineering capacity that already exists. It can also become a place where talented people build careers, whether those careers remain within the OC or eventually take them somewhere else.

Escalation Should Follow Expertise

Not every escalation is the same. Sometimes a problem has moved beyond what the OC can reasonably handle and requires someone with deeper knowledge of the system. Other times, the OC understands the problem and its remediation but lacks the access or permissions required to act.

The first is exactly what an escalation path is for. The second deserves more scrutiny. There are legitimate reasons to restrict production access, but when routine remediation consistently requires escalation solely because someone else has the necessary permissions, the escalation has not added expertise to the problem. It has added permission.

The cost of that kind of escalation is easy to underestimate because the remediation itself may only take a few minutes. During the workday, someone has to stop what they are doing, gather enough context to act, and then find their way back to whatever they were working on. During an on-call rotation, that same interruption may mean waking someone to perform work that never required their expertise in the first place, just their account permissions.

Platforms change, and the security and operating models around them often change as well. A new platform may require much stricter access controls than the one it replaces, and those restrictions may be completely justified. The impact on the OC still needs to be considered as part of that change. If the access model changes what the OC is capable of handling, the operational consequences need to be understood as deliberately as the security benefits.

That does not necessarily mean restoring broader direct access. The mechanism may need to change with the platform, but the OC still needs a safe way to perform the work that belongs within its boundaries.

This is where the force multiplier works in both directions. A capable OC gives engineering teams back time and focus by absorbing the operational work it can safely handle. Without that capability, more of the same work moves toward specialized teams regardless of whether their expertise is actually necessary.

Those engineering teams should still remain accountable for the systems they build and understand how they behave in production, but that accountability does not require specialized engineers to personally perform every routine operational task associated with them.

You Get the OC You Build

The ticket-pushing OC is not inevitable. What an OC becomes is heavily influenced by the environment the organization creates around it.

An OC cannot develop meaningful operational capability without the opportunity to actually operate. At the same time, that capability cannot simply be handed over through broader access or a larger set of responsibilities in one fell swoop. It develops as the team gains experience, demonstrates good judgment, and gradually earns greater responsibility.

This is why the reputation of the ticket-pushing OC has always felt incomplete. There are capable people sitting in OCs who want to understand more, troubleshoot harder problems, and take greater ownership. Whether that potential develops into a stronger OC depends heavily on the environment the organization creates around them.

That brings me back to the hot take. The OC is not the most valuable team because its work is somehow more important than the teams around it. Its value comes from the effect it can have on all of them. A great OC gives those teams more time to do the work they were actually built to do while ensuring that the continuous work of keeping production running has a capable owner.

A great Operations Center is an engineering force multiplier.