Get in Touch

Course Outline

Foundations of Cloud Operations on AWS

  • Defining operational roles and responsibilities within a cloud environment.
  • Understanding AWS account structures, organizations, and multi-account strategies.
  • Exploring core operational services, including CloudWatch, CloudTrail, and AWS Config.

Infrastructure as Code and Provisioning

  • Core principles of IaC and the concept of immutable infrastructure.
  • Executing provisioning tasks using Terraform and AWS CloudFormation.
  • Managing state, modules, and the promotion of environments.

CI/CD and Deployment Strategies

  • Designing CI/CD pipelines optimized for cloud-native applications.
  • Implementing Blue/green, canary, and rolling deployment models.
  • Automating rollback procedures, health checks, and release validation processes.

Monitoring, Observability, and Alerting

  • Handling metrics, logs, and traces: collection, storage, and analysis.
  • Utilizing CloudWatch, X-Ray, and third-party observability tools.
  • Establishing SLOs/SLIs, alerting policies, and on-call operational practices.

Security Operations and Identity Management

  • Applying IAM best practices, least privilege principles, and cross-account access controls.
  • Managing secrets, utilizing KMS, and securing parameter stores.
  • Operational security measures, including patching strategies, vulnerability scanning, and audit trails.

Resilience, Backup, and Disaster Recovery

  • Architecting for fault tolerance and high availability.
  • Developing backup strategies, automating snapshots, and defining restore procedures.
  • Planning disaster recovery initiatives and creating comprehensive runbooks.

Cost Optimization and Governance

  • Enhancing cost visibility through billing analysis, tagging, and cost allocation strategies.
  • Optimizing resources via rightsizing, reserved instances, savings plans, and budget controls.
  • Enforcing governance through policies, guardrails, and compliance automation.

Containers, Serverless, and Runtime Operations

  • Operational considerations for managing ECS, EKS, and Lambda.
  • Managing service discovery, autoscaling, and resource limits.
  • Logging, tracing, and debugging techniques for containerized workloads.

Incident Response, Playbooks, and Chaos Engineering

  • Conducting runbook-driven incident response and postmortem analysis.
  • Automating remediation actions and implementing self-healing patterns.
  • Introduction to chaos engineering experiments for validating system resilience.

Hands-on Workshop: Operating a Sample Workload

  • Deploying a sample application using IaC and a CI/CD pipeline.
  • Implementing monitoring, alerts, and automated remediation scripts.
  • Simulating incidents to practice runbook-based response procedures.

Summary and Next Steps

Requirements

  • Foundational knowledge of cloud computing concepts and networking principles.
  • Proficiency with the Linux command line and scripting languages.
  • Practical experience with source control systems (such as Git) and an understanding of basic CI/CD workflows.

Intended Audience

  • Cloud operations engineers.
  • Site Reliability Engineers (SREs) and platform engineers.
  • DevOps engineers and technical team leads.
 21 Hours

Custom Corporate Training

Training solutions designed exclusively for businesses.

  • Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
  • Flexible Schedule: Dates and times adapted to your team's agenda.
  • Format: Online (live), In-company (at your offices), or Hybrid.
Investment

Price per private group, online live training, starting from 3900 € + VAT*

Contact us for an exact quote and to hear our latest promotions

Testimonials (1)

Provisional Upcoming Courses (Contact Us For More Information)

Related Categories