Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps Using Open Source Tools
- Key concepts and strategic benefits of AIOps
- The role of Prometheus and Grafana in the modern observability stack
- The place of ML in AIOps: contrasting predictive and reactive analytics
Setting Up Prometheus and Grafana
- Installation and configuration of Prometheus for time series data collection
- Building dynamic dashboards in Grafana utilizing real-time metrics
- Investigating exporters, relabeling rules, and service discovery mechanisms
Data Preprocessing for Machine Learning
- Techniques for extracting and transforming Prometheus metrics
- Preparing robust datasets for anomaly detection and forecasting tasks
- Implementing data transformations via Grafana features or Python pipelines
Applying Machine Learning for Anomaly Detection
- Introduction to basic ML models for outlier detection (e.g., Isolation Forest, One-Class SVM)
- Strategies for training and evaluating models on time series data
- Visualizing detected anomalies within Grafana dashboards
Forecasting Metrics with Machine Learning
- Developing simple forecasting models (introduction to ARIMA, Prophet, and LSTM)
- Predicting system load and resource utilization trends
- Leveraging predictions to inform early alerting and scaling decisions
Integrating Machine Learning with Alerting and Automation
- Defining alert rules based on ML outputs or dynamic thresholds
- Configuring Alertmanager and managing notification routing
- Triggering scripts or automation workflows upon anomaly detection
Scaling and Operationalizing AIOps
- Integrating with external observability tools (e.g., ELK stack, Moogsoft, Dynatrace)
- Operationalizing ML models within observability pipelines
- Best practices for implementing AIOps at enterprise scale
Summary and Next Steps
Requirements
- A solid understanding of core system monitoring and observability concepts
- Prior experience utilizing Grafana or Prometheus
- Familiarity with Python programming and fundamental machine learning principles
Target Audience
- Observability engineers
- Infrastructure and DevOps teams
- Monitoring platform architects and Site Reliability Engineers (SREs)