How Enterprise SRE Observability Works | Metrics, Logs, Traces & OpenTelemetry
DevOpsForAll | Cloud Architecture Deep Dives
0:00 / 0:00
How Enterprise SRE Observability Works | Metrics, Logs, Traces & OpenTelemetry
120 просмотров · 5 дн. назад
DevOpsForAll | Cloud Architecture Deep Dives
86 подписчиков
120 просмотров · 5 дн. назад
How do modern engineering teams build an observability platform that connects metrics, logs and distributed traces across applications running on Kubernetes?
In this SRE Observability Deep Dive, we demonstrate an end-to-end observability architecture using OpenTelemetry, Prometheus, Loki, Tempo, Grafana, Grafana k6, SigNoz and OpenObserve.
The session focuses on a practical SRE problem: collecting telemetry from applications, routing it through OpenTelemetry, visualizing system health, correlating metrics, logs and traces, and building alerting workflows.
This is a hands-on technical demonstration built around an AWS EKS / Kubernetes environment.
━━━━━━━━━━━━━━━━━━━━━━
🏗️ OBSERVABILITY ARCHITECTURE
Applications
↓
OpenTelemetry / OTEL
↓
OTEL Collector
↓
┌────────────┬────────────┬────────────┐
│ METRICS │ LOGS │ TRACES │
│ Prometheus │ Loki │ Tempo │
└────────────┴────────────┴────────────┘
↓
Grafana
↓
Dashboards & Alerting
↓
Slack
━━━━━━━━━━━━━━━━━━━━━━
🚀 WHAT WE DEMONSTRATE
• Application telemetry
• OpenTelemetry instrumentation
• OpenTelemetry Collector
• Metrics collection with Prometheus
• Prometheus Exemplars
• Log aggregation with Loki
• Distributed tracing with Tempo
• Grafana dashboards
• Metrics / Logs / Traces correlation
• Alerting to Slack
• Load testing with Grafana k6
• Observability workflows
• SRE investigation patterns
• AWS EKS / Kubernetes observability
• SigNoz
• OpenObserve
━━━━━━━━━━━━━━━━━━━━━━
🛠️ TECHNOLOGY STACK
OpenTelemetry
OTel Collector
Prometheus
Prometheus Exemplars
Grafana
Grafana k6
Loki
Tempo
SigNoz
OpenObserve
AWS EKS
Kubernetes
Slack Alerting
━━━━━━━━━━━━━━━━━━━━━━
🔎 METRICS + LOGS + TRACES
A modern observability workflow should not stop at collecting data.
When an application problem occurs:
Metrics help identify WHAT is changing.
Logs help understand WHAT the application reported.
Traces help identify WHERE the request spent time.
Correlation helps connect these signals into one investigation workflow.
In this session we demonstrate how these signals can work together.
━━━━━━━━━━━━━━━━━━━━━━
🎯 WHO IS THIS FOR?
• SRE Engineers
• DevOps Engineers
• Platform Engineers
• Kubernetes Engineers
• Cloud Engineers
• Observability Engineers
• Site Reliability Engineers
• DevOps Architects
• Cloud Architects
• Engineering Managers
• Technology Leaders
━━━━━━━━━━━━━━━━━━━━━━
☁️ ENVIRONMENT
AWS EKS
Kubernetes
Cloud-native applications
OpenTelemetry-based telemetry pipeline
━━━━━━━━━━━━━━━━━━━━━━
🌐 TECHNOLOGY & OPEN-SOURCE ECOSYSTEM
This session demonstrates technologies from the broader cloud-native and observability ecosystem, including:
• Kubernetes — CNCF
• OpenTelemetry — CNCF
• Prometheus — CNCF
• Grafana / Loki / Tempo / k6 — Grafana
• SigNoz — Open-source observability platform
• OpenObserve — Open-source observability platform
• AWS EKS — Amazon Web Services
━━━━━━━━━━━━━━━━━━━━━━
💡 DEVOPSFORALL
DevOpsForAll focuses on practical Enterprise DevOps, Platform Engineering, Kubernetes, GitOps, Cloud Engineering, SRE, Observability and production-oriented engineering practices.
We don't just learn individual tools.
We understand how the pieces work together.
Practical Architecture.
Production Thinking.
Subscribe to DevOpsForAll for more deep-dive engineering sessions.
#SRE #Observability #OpenTelemetry #Prometheus #Grafana