Who Tests Observability
Software products almost always have people intentionally testing them. Whether it's a dedicated QA team, beta users, or internal testing, somebody is constantly validating the experience before it reaches production. What they find becomes feedback, and over time the product gets better because people are continually using it and questioning the experience.
Observability platforms don't get that luxury.
The people consuming the platform every day become the closest thing it has to a QA team, whether they realize it or not. Every frustrating experience becomes feedback about the platform itself. The challenge is that very little of that feedback ever finds its way back into improving the platform.
3000 Pages
One of the experiences that shaped how I think about observability happened while I was working in an Operations Center.
A major outage starts. I'm the only engineer on shift.
Within a few seconds the alerts begin firing and my phone starts going off almost immediately.
At the same time, I have to understand what's actually broken. Is this a single service or something much larger? Which engineering team owns it? Does this warrant declaring an incident? All of the coordination and communication that comes with a major outage still needs to happen while I'm trying to understand what I'm looking at.
Meanwhile my phone just keeps going.
Ding....Ding....Ding....Ding.
I can already see nearly 3,000 pages queued behind the ones I'm receiving. I'm already actively investigating the outage. Every new notification is competing for the same attention I need to understand what's actually happening.
Eventually I turn my phone off.
Not because the alerts are wrong, but because they're no longer helping. At that point they're simply pulling my attention away from the work that actually needs to get done.
After incidents like that I'd usually go back to the engineers responsible for the observability platform. We'd talk through what happened, what helped, and what made the investigation harder. Sometimes improvements were made and sometimes nothing changed.
After you've had that experience enough times, it's surprisingly easy for frustration with the platform to become resentment. To be honest, I think some of that resentment is probably warranted depending on how feedback is handled.
The platform owners weren't ignoring the feedback. Observability was only one responsibility among many, with plenty of other systems and production problems competing for their attention. Improving the observability platform wasn't the only thing on their plate. Many weeks it probably wasn't even on the radar.
That's something I understand much better now that I've been on both sides of that conversation.
Your Users Are Your QA Team
One of the easiest traps for an observability platform owner to fall into is becoming disconnected from the people actually using the platform every day. When you're responsible for keeping the platform healthy, it's natural to experience it differently than everyone else. You're thinking about everything required to keep it operational. The engineers using the platform experience something entirely different. They experience whether the dashboard answered the question they were trying to ask, whether the alert gave them enough context to act, and whether the platform made an already stressful outage easier or harder to navigate.
That's exactly why their feedback is so valuable.
They're stress testing the platform under real production conditions every single day. They encounter the kinds of friction that are nearly impossible to find from the platform owner's perspective. The people relying on the platform every day become the feedback loop.
I don't think the cadence really matters. Weekly syncs, monthly syncs, whatever works for the organization. What matters is creating a reliable way for those engineers to influence how the platform evolves. Those conversations rarely need much structure because the people using the platform almost always arrive with something they want to improve. Those aren't complaints. They're product feedback.
Sometimes the people providing that feedback become some of the platform's main contributors. That's actually how I got my hands dirty with observability. It didn't start with ownership of the platform. It started by using it every day and becoming frustrated by parts of the experience. That turned into suggesting improvements and eventually helping implement them. Looking back, I think that's one of the healthiest feedback loops an observability platform can have because the people improving it are often the same people depending on it during production incidents.
Continuous Feedback
One of the biggest shifts in my thinking happened when I stopped viewing the people using an observability platform as consumers and started viewing them as partners in improving it.
Your observability platform may not have a dedicated QA team. It has engineers using it under real production conditions every single day, and their experience is telling you where the platform needs to improve.
The platform is already being tested every day. The organizations that improve the fastest are usually the ones that learn to treat that feedback as one of the platform's most valuable inputs.
