Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI for IT Operations
- The evolution of IT automation: moving from static runbooks to reasoning agents
- Understanding agent anatomy: reasoning loops, tool utilization, memory, and planning
- Determining when to automate versus when to maintain human oversight
Agent Frameworks and Architectural Patterns
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent architectures: supervisor, hierarchical, and swarm patterns
- Comparing frameworks: LangGraph, CrewAI, AutoGen, and custom agent solutions
- Building your first operational agent: querying monitoring systems, diagnosing issues, and proposing solutions
Tool Integration for IT Operations
- Connecting agents to APIs for Prometheus, Grafana, Datadog, and PagerDuty
- Agent-based log querying: integrating with Elasticsearch, Loki, and Splunk
- Leveraging infrastructure tools: using kubectl, Terraform, and Ansible via agent actions
- Designing secure tool interfaces with parameter validation and idempotency
Incident Response Automation
- Automated incident triage: severity classification and routing
- Generating root cause hypotheses and gathering supporting evidence
- Automated remediation strategies: restarting, scaling, rolling back, and failing over
- Creating an incident runbook agent with progressive levels of autonomy
Safety, Guardrails, and Human-in-the-Loop Strategies
- Classifying actions: read-only, low-risk, high-risk, and destructive
- Establishing approval gates and escalation policies for critical operations
- Implementing guardrail patterns: action allowlists, blast radius limits, and rollback guarantees
- Maintaining audit trails and decision provenance for compliance purposes
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialist agents: triage, diagnosis, and remediation roles
- Managing inter-agent communication and shared context
- Resolving conflicts when agents propose contradictory actions
- Simulating end-to-end major incidents with multi-agent responses
Observability and Evaluation
- Tracing agent reasoning chains for debugging and audit purposes
- Evaluating agent decision quality: precision, recall, and time-to-resolution
- Establishing feedback loops: learning from operator overrides and outcomes
- Tracking costs and analyzing token economics for operational agents
Production Deployment and Operations
- Deploying agents as services: utilizing APIs, webhooks, and scheduled jobs
- Rolling out gradual autonomy: moving from shadow mode to full auto-remediation
- Creating runbooks for agent failures: addressing scenarios where the agent itself malfunctions
- Building a business case and measuring ROI for autonomous operations
Requirements
- Practical experience in IT operations, DevOps, or SRE practices.
- Proficiency in Python scripting and REST APIs.
- A foundational understanding of LLM capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation strategies.
- Platform engineers focused on building self-healing infrastructure.
- IT operations leaders evaluating agentic AI solutions for incident management.
14 Hours