Get in Touch

Course Outline

Foundations of Agentic Systems in Production

  • Agentic architectures: loops, tools, memory, and orchestration layers.
  • The agent lifecycle: from development and deployment to continuous operation.
  • Key challenges in managing agents at production scale.

Infrastructure and Deployment Models

  • Deploying agents across containerized and cloud environments.
  • Scaling strategies: comparing horizontal and vertical scaling, along with concurrency and throttling mechanisms.
  • Orchestration of multi-agent systems and workload balancing.

Monitoring and Observability

  • Essential metrics: latency, success rates, memory usage, and agent call depth.
  • Tracing agent activities and mapping call graphs.
  • Implementing observability instrumentation with Prometheus, OpenTelemetry, and Grafana.

Logging, Auditing, and Compliance

  • Establishing centralized logging and structured event collection.
  • Ensuring compliance and auditability within agentic workflows.
  • Designing audit trails and replay mechanisms to facilitate debugging.

Performance Tuning and Resource Optimization

  • Minimizing inference overhead and refining agent orchestration cycles.
  • Leveraging model caching and lightweight embeddings for accelerated retrieval.
  • Conducting load testing and stress scenarios for AI pipelines.

Cost Control and Governance

  • Analyzing agent cost drivers: API calls, memory consumption, compute resources, and external integrations.
  • Tracking costs at the agent level and implementing chargeback models.
  • Enforcing automation policies to curb agent sprawl and prevent idle resource usage.

CI/CD and Rollout Strategies for Agents

  • Embedding agent pipelines into CI/CD systems.
  • Developing testing, versioning, and rollback strategies for iterative agent updates.
  • Executing progressive rollouts and safe deployment mechanisms.

Failure Recovery and Reliability Engineering

  • Architecting for fault tolerance and graceful degradation.
  • Applying retry, timeout, and circuit breaker patterns to enhance agent reliability.
  • Implementing incident response and post-mortem frameworks for AI operations.

Capstone Project

  • Developing and deploying an agentic AI system equipped with comprehensive monitoring and cost tracking.
  • Simulating loads, measuring performance, and optimizing resource utilization.
  • Presenting the final architecture and monitoring dashboard to peers.

Summary and Next Steps

Requirements

  • A solid grasp of MLOps and production-grade machine learning systems.
  • Practical experience with containerized deployments, such as Docker and Kubernetes.
  • Working knowledge of cloud cost optimization and observability tools.

Target Audience

  • MLOps Engineers
  • Site Reliability Engineers (SREs)
  • Engineering Managers responsible for AI infrastructure
 21 Hours

Custom Corporate Training

Training solutions designed exclusively for businesses.

  • Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
  • Flexible Schedule: Dates and times adapted to your team's agenda.
  • Format: Online (live), In-company (at your offices), or Hybrid.
Investment

Price per private group, online live training, starting from 3900 € + VAT*

Contact us for an exact quote and to hear our latest promotions

Testimonials (3)

Provisional Upcoming Courses (Contact Us For More Information)

Related Categories