← All work Next projectBrewhaus
Sentinel
Infrastructure monitoring that flags anomalies across 3,000 services before they become outages.
The challenge
An engineering org with 3,000 services only learned about incidents from customers, and alert noise meant real problems were missed.
What we built
- 1
Unified metrics and traces with anomaly detection tuned per service.
- 2
Alert grouping that cut noise and routed issues to the right team.
- 3
Incident timelines generated automatically for post-mortems.
- −68%
- alert noise
- −45%
- time to detect
- 3,000
- services covered
Stack
- Go
- Python
- OpenTelemetry
- Kubernetes
Sample case study. Swap in real screenshots, numbers and a client quote before launch.
