Перейти к содержимому

How Enterprise SRE Observability Works | Metrics, Logs, Traces & OpenTelemetry

DevOpsForAll | Cloud Architecture Deep Dives

0:00 / 0:00

How Enterprise SRE Observability Works | Metrics, Logs, Traces & OpenTelemetry

120 просмотров · 5 дн. назад
DevOpsForAll | Cloud Architecture Deep Dives
86 подписчиков
120 просмотров · 5 дн. назад
How do modern engineering teams build an observability platform that connects metrics, logs and distributed traces across applications running on Kubernetes? In this SRE Observability Deep Dive, we demonstrate an end-to-end observability architecture using OpenTelemetry, Prometheus, Loki, Tempo, Grafana, Grafana k6, SigNoz and OpenObserve. The session focuses on a practical SRE problem: collecting telemetry from applications, routing it through OpenTelemetry, visualizing system health, correlating metrics, logs and traces, and building alerting workflows. This is a hands-on technical demonstration built around an AWS EKS / Kubernetes environment. ━━━━━━━━━━━━━━━━━━━━━━ 🏗️ OBSERVABILITY ARCHITECTURE Applications ↓ OpenTelemetry / OTEL ↓ OTEL Collector ↓ ┌────────────┬────────────┬────────────┐ │ METRICS │ LOGS │ TRACES │ │ Prometheus │ Loki │ Tempo │ └────────────┴────────────┴────────────┘ ↓ Grafana ↓ Dashboards & Alerting ↓ Slack ━━━━━━━━━━━━━━━━━━━━━━ 🚀 WHAT WE DEMONSTRATE • Application telemetry • OpenTelemetry instrumentation • OpenTelemetry Collector • Metrics collection with Prometheus • Prometheus Exemplars • Log aggregation with Loki • Distributed tracing with Tempo • Grafana dashboards • Metrics / Logs / Traces correlation • Alerting to Slack • Load testing with Grafana k6 • Observability workflows • SRE investigation patterns • AWS EKS / Kubernetes observability • SigNoz • OpenObserve ━━━━━━━━━━━━━━━━━━━━━━ 🛠️ TECHNOLOGY STACK OpenTelemetry OTel Collector Prometheus Prometheus Exemplars Grafana Grafana k6 Loki Tempo SigNoz OpenObserve AWS EKS Kubernetes Slack Alerting ━━━━━━━━━━━━━━━━━━━━━━ 🔎 METRICS + LOGS + TRACES A modern observability workflow should not stop at collecting data. When an application problem occurs: Metrics help identify WHAT is changing. Logs help understand WHAT the application reported. Traces help identify WHERE the request spent time. Correlation helps connect these signals into one investigation workflow. In this session we demonstrate how these signals can work together. ━━━━━━━━━━━━━━━━━━━━━━ 🎯 WHO IS THIS FOR? • SRE Engineers • DevOps Engineers • Platform Engineers • Kubernetes Engineers • Cloud Engineers • Observability Engineers • Site Reliability Engineers • DevOps Architects • Cloud Architects • Engineering Managers • Technology Leaders ━━━━━━━━━━━━━━━━━━━━━━ ☁️ ENVIRONMENT AWS EKS Kubernetes Cloud-native applications OpenTelemetry-based telemetry pipeline ━━━━━━━━━━━━━━━━━━━━━━ 🌐 TECHNOLOGY & OPEN-SOURCE ECOSYSTEM This session demonstrates technologies from the broader cloud-native and observability ecosystem, including: • Kubernetes — CNCF • OpenTelemetry — CNCF • Prometheus — CNCF • Grafana / Loki / Tempo / k6 — Grafana • SigNoz — Open-source observability platform • OpenObserve — Open-source observability platform • AWS EKS — Amazon Web Services ━━━━━━━━━━━━━━━━━━━━━━ 💡 DEVOPSFORALL DevOpsForAll focuses on practical Enterprise DevOps, Platform Engineering, Kubernetes, GitOps, Cloud Engineering, SRE, Observability and production-oriented engineering practices. We don't just learn individual tools. We understand how the pieces work together. Practical Architecture. Production Thinking. Subscribe to DevOpsForAll for more deep-dive engineering sessions. #SRE #Observability #OpenTelemetry #Prometheus #Grafana