






The Grafana & Prometheus for Metrics-Driven Observability course is designed to help learners build the practical skills required to collect, query, visualize, and act on metrics for effective observability. Tailored for DevOps engineers, cloud engineers, platform engineers, SREs, Kubernetes administrators, and developers, this course provides a comprehensive introduction to metrics-driven observability using Prometheus and Grafana. Through conceptual lessons, guided demonstrations, and hands-on exercises, you'll learn how to instrument applications, collect and query metrics, build meaningful dashboards, and create effective alerts for modern cloud-native environments.
Throughout the course, you'll gain practical experience working with Prometheus, Grafana, PromQL, Kubernetes metrics, OpenTelemetry, and the kube-prometheus-stack. You'll explore the fundamentals of metrics, the Four Golden Signals of Monitoring, and service-level concepts such as SLIs, SLOs, and SLAs. You'll learn how Prometheus collects and stores metrics, work with metric types including counters, gauges, histograms, and summaries, and query data using PromQL.
As you progress, you'll learn how to instrument applications and expose meaningful metrics through the /metrics endpoint. You'll explore both in-code instrumentation and auto-instrumentation concepts with OpenTelemetry, while also learning how to design scalable metric systems and avoid common issues such as high-cardinality metrics.
You'll then put your skills into practice by building Grafana dashboards that transform raw metrics into actionable insights. You'll create time series visualizations, Golden Signal dashboards, histogram and heatmap visualizations, KPI panels, state timelines, and saturation panels to effectively monitor application and infrastructure health.
The course also covers alerting philosophy and Grafana Alerting, helping you understand how to create alerts that provide meaningful signals without generating unnecessary noise. You'll additionally explore metrics observability for AI inference workloads, giving you insight into how observability principles can be applied to modern AI-powered systems.
By the end of this course, you'll have the practical knowledge needed to design and implement a metrics-driven observability solution using Prometheus and Grafana. You'll be able to instrument applications, collect and query metrics, build effective dashboards, identify performance and reliability issues, and create actionable alerts for cloud-native and Kubernetes environments.
Build a strong foundation in metrics-driven observability. Learn what metrics are, how Grafana supports observability, and how to use the Four Golden Signals to monitor latency, traffic, errors, and saturation. You'll also explore SLIs, SLOs, and SLAs and understand how they help measure and manage service reliability.
Learn why Prometheus and Grafana are widely used for cloud-native observability and gain hands-on experience deploying them in Kubernetes. Explore the /metrics endpoint, understand counters, gauges, histograms, and summaries, learn PromQL fundamentals, and use the kube-prometheus-stack Helm chart to collect and query Kubernetes metrics.
Learn how to expose and collect meaningful application metrics. Explore in-code metric instrumentation, Prometheus metric instrumentation, and the fundamentals of auto-instrumentation with OpenTelemetry. You'll understand how to design metrics that provide useful operational insights.
Understand how metric cardinality affects the scalability and performance of Prometheus. Learn how to identify and diagnose high-cardinality metrics, avoid common metric design mistakes, and work with Prometheus native histograms for efficient metric collection and analysis.
Transform metrics into actionable insights by building effective Grafana dashboards. Create time series panels, Golden Signal visualizations, histograms, heatmaps, KPI panels, state timelines, and saturation panels to monitor application performance, reliability, and resource utilization.
Learn the principles of effective alerting and understand how to create meaningful alerts using Grafana Alerting. Explore alerting philosophy and learn how to focus on actionable signals while reducing unnecessary alert noise.
Explore how metrics-driven observability applies to AI inference workloads. Learn about the key metrics and observability considerations involved in monitoring the performance and behavior of modern AI-powered applications.
Review the key concepts covered throughout the course and reinforce your understanding of metrics-driven observability using Prometheus and Grafana.
Build the practical skills needed to instrument applications, collect and query metrics, create meaningful Grafana dashboards, and implement effective alerting using Prometheus and Grafana for metrics-driven observability.

Reece Iriye is a Software Engineer and Site Reliability Engineer with over 2 years of experience scaling high-demand systems, most recently at Verizon, where he designed the core DNS routing framework for the company's 5G core infrastructure, supporting billions requests per hour nationwide. He holds 2 patents for automated 5G network failover solutions in on-prem Kubernetes environments and has led incident response for nationwide outage war rooms, prioritizing rapid restoration of critical services. With a background in mathematics and machine learning, Reece brings a strong analytical foundation to his engineering work. As an instructor, he's passionate about teaching complex topics and helping learners build hands-on skills to tackle production-grade infrastructure with confidence.