Course Outline
Foundations of Cloud Operations on AWS
- Defining operational roles and responsibilities within a cloud environment.
- Understanding AWS account structures, organizations, and multi-account strategies.
- Exploring core operational services, including CloudWatch, CloudTrail, and AWS Config.
Infrastructure as Code and Provisioning
- Core principles of IaC and the concept of immutable infrastructure.
- Executing provisioning tasks using Terraform and AWS CloudFormation.
- Managing state, modules, and the promotion of environments.
CI/CD and Deployment Strategies
- Designing CI/CD pipelines optimized for cloud-native applications.
- Implementing Blue/green, canary, and rolling deployment models.
- Automating rollback procedures, health checks, and release validation processes.
Monitoring, Observability, and Alerting
- Handling metrics, logs, and traces: collection, storage, and analysis.
- Utilizing CloudWatch, X-Ray, and third-party observability tools.
- Establishing SLOs/SLIs, alerting policies, and on-call operational practices.
Security Operations and Identity Management
- Applying IAM best practices, least privilege principles, and cross-account access controls.
- Managing secrets, utilizing KMS, and securing parameter stores.
- Operational security measures, including patching strategies, vulnerability scanning, and audit trails.
Resilience, Backup, and Disaster Recovery
- Architecting for fault tolerance and high availability.
- Developing backup strategies, automating snapshots, and defining restore procedures.
- Planning disaster recovery initiatives and creating comprehensive runbooks.
Cost Optimization and Governance
- Enhancing cost visibility through billing analysis, tagging, and cost allocation strategies.
- Optimizing resources via rightsizing, reserved instances, savings plans, and budget controls.
- Enforcing governance through policies, guardrails, and compliance automation.
Containers, Serverless, and Runtime Operations
- Operational considerations for managing ECS, EKS, and Lambda.
- Managing service discovery, autoscaling, and resource limits.
- Logging, tracing, and debugging techniques for containerized workloads.
Incident Response, Playbooks, and Chaos Engineering
- Conducting runbook-driven incident response and postmortem analysis.
- Automating remediation actions and implementing self-healing patterns.
- Introduction to chaos engineering experiments for validating system resilience.
Hands-on Workshop: Operating a Sample Workload
- Deploying a sample application using IaC and a CI/CD pipeline.
- Implementing monitoring, alerts, and automated remediation scripts.
- Simulating incidents to practice runbook-based response procedures.
Summary and Next Steps
Requirements
- Foundational knowledge of cloud computing concepts and networking principles.
- Proficiency with the Linux command line and scripting languages.
- Practical experience with source control systems (such as Git) and an understanding of basic CI/CD workflows.
Intended Audience
- Cloud operations engineers.
- Site Reliability Engineers (SREs) and platform engineers.
- DevOps engineers and technical team leads.
Custom Corporate Training
Training solutions designed exclusively for businesses.
- Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
- Flexible Schedule: Dates and times adapted to your team's agenda.
- Format: Online (live), In-company (at your offices), or Hybrid.
Price per private group, online live training, starting from 3900 € + VAT*
Contact us for an exact quote and to hear our latest promotions
Testimonials (1)
I've find out new interesting things about Lambda and Serverless