On This Page
Good Decisions in Isolation
The way observability platforms evolve sometimes reminds me of driving at night. You can see what's directly in front of you. You can avoid the obstacle that's right ahead. What you can't see is where you'll be a mile down the road. Every decision makes sense in the moment, but without a destination, it's easy to arrive somewhere you never intended.
They usually begin because somebody needs monitoring. A few dashboards get built, some alerts are added, and metrics start flowing. As new teams and problems emerge, more gets added to the platform. None of those decisions are wrong. In fact, they're usually exactly the right decision for the problem being solved at that moment.
That's what makes this interesting. The people building the platform are usually making good engineering decisions. They're solving real problems for real teams. Five years later though, the observability platform has become the sum of thousands of those individual decisions. Everything in that platform probably exists for a reason. The problem isn't that engineers made bad decisions. It's that they were making good decisions in isolation.
Most observability platforms have an owner on paper. They rarely have someone whose primary responsibility is making them better. In my experience, observability is usually assigned to a team like Platform Engineering, SRE, or DevOps. There's nothing inherently wrong with that. The problem is that observability is often just one responsibility among dozens. The platform continues to evolve one decision at a time, but very few organizations step back and ask what the observability platform should become over the next few years.
Without that vision, the platform slowly becomes a collection of excellent tools that never quite feels like one cohesive platform. The reason becomes obvious once you look at how different teams actually use it.
Every Team Uses a Different Product
One of the things I've come to appreciate over the years is that every team naturally uses an observability platform differently.
They're all using the same underlying telemetry, but they're trying to answer different questions because their responsibilities are different. Management is usually looking for confidence. Engineering is looking for answers. Operations lives somewhere between those two worlds. They need enough information to understand the health of the environment and gather meaningful context before escalating an issue, but not so much that they become overwhelmed by unnecessary detail. If they can quickly identify what is affected and gather enough context to understand the nature of the problem, they've already made the next engineer's job much easier.
None of those experiences are more important than the others. They're all valid, and they're all consumers of the same platform. That's one of the reasons I think observability should be thought of as a product. Products have different types of users, each with different needs and different ways of interacting with them. An observability platform is no different.
The natural tendency is for every team to shape the platform around the problems they're trying to solve. That's exactly what they should do. The problem begins when those decisions happen without a shared vision for the platform as a whole.
Without a shared vision, those experiences slowly drift apart. Over time, every team ends up with a slightly different observability platform built on the same underlying technology. The tooling is shared, but the experience isn't. Maintaining and improving the platform becomes increasingly fragmented because each part has been optimized independently instead of intentionally designed as one cohesive product.
Meet Bob
Eventually, someone becomes Bob.
Sometimes Bob inherited the observability platform because somebody had to own it. It isn't the part of his job he's most passionate about. He's keeping it healthy because it's important, but his interests are somewhere else.
Other times, Bob genuinely enjoys observability. He wants to improve the platform and spend time understanding what its users actually need. The problem is that observability is only one part of his job, competing with the rest of the infrastructure he owns and whatever production issue became today's highest priority.
Those are two very different engineers, but they often end up in exactly the same place.
The platform runs. Things get fixed when they break and routine maintenance gets done. From the outside, everything appears healthy. That's often where the story ends. The platform "works," so it quietly falls behind everything else competing for Bob's attention.
The difference between maintaining a platform and improving one isn't desire. It's time.
Improving a platform requires making time for work that rarely feels urgent. It means understanding where people struggle and having enough space to think beyond the next immediate problem. Production issues will almost always feel more pressing.
When you're responsible for enough systems, something else is almost always competing for that time. The observability platform continues doing its job, so improving it slips another week, then another month. That work may not produce immediate business value next week, but it can produce significantly better outcomes six months from now.
Eventually the platform becomes something that simply exists. It still provides value, but it stops getting noticeably better. That isn't Bob's fault. Whether Bob loves observability or simply inherited it doesn't really matter. The outcome is often the same because the organization has asked someone to own a product without giving them the time to think like a product owner.
Thinking Like a Product Team
This is why I've started thinking about observability platforms as products rather than infrastructure.
The biggest difference isn't the technology. It's the questions that get asked.
Infrastructure is often measured by whether it's available. Products are measured by whether they're improving. That shift changes the conversation from whether the monitoring stack is running to whether the experience of using it is getting better. Engineers should be finding answers faster. Incidents should begin with better context. The platform should become easier to understand and operate over time.
Those aren't infrastructure questions. They're product questions.
Products have owners. They have roadmaps. They have documentation. They evolve based on feedback from the people using them. They aren't considered finished simply because nothing has broken recently. The same should be true for an observability platform.
I actually like when an observability platform has a name. It sounds like a small detail, but I think it changes how people think about it. Documentation has somewhere to live. Roadmaps have an identity. Conversations become clearer. Instead of talking about "the Grafana dashboards" or "our monitoring system," people start talking about the platform itself. It becomes something the organization recognizes as an important part of the engineering ecosystem rather than another collection of infrastructure components quietly running in the background.
Product thinking also changes how improvements are prioritized. Work that improves the experience of using the platform stops being something someone might eventually get around to. It becomes an intentional investment in the platform itself. Products aren't successful because they exist. They're successful because they're continually improved.
Company size obviously matters. Not every organization needs a dedicated observability team. Some companies genuinely only need one engineer spending part of their week on it. The important part isn't how many people work on the platform. It's that someone is empowered to think beyond keeping the lights on. Whether that's one engineer or an entire team, someone should be thinking about where the platform should be a year from now, not just whether it made it through this week.
Building With Intention
The best observability platforms weren't defined by the tools they used. Some ran commercial software. Others were entirely open source. Some were built and maintained by dedicated teams. Others were largely the responsibility of a single engineer. The technology mattered, but it was never the defining characteristic.
What they all had in common was intention. Someone was continually thinking about how the platform should evolve, who it served, and how all of its pieces fit together. Improvements weren't driven solely by the latest production issue or the next team's immediate need. They were guided by a broader vision of what the platform should become over time.
That's the difference between an observability platform that simply works and one that becomes an asset to the engineering organization. The first fulfills a technical need. The second becomes part of how the organization operates.
An observability platform doesn't become that because the right collection of tools was assembled. It becomes that because someone continually invests in making the platform better.
That's why an observability platform is a product, not a collection of tools.
