Get in Touch
 Duration 21 hours

Course Outline

NiFi Fundamentals and Data Flow Concepts

  • Data in motion versus data at rest: underlying concepts and challenges
  • NiFi architecture: cores, flow controller, provenance, and bulletin boards
  • Essential components: processors, connections, controllers, and provenance

Big Data Context and Integration

  • The role of NiFi within Big Data ecosystems (Hadoop, Kafka, cloud storage)
  • Overview of HDFS, MapReduce, and contemporary alternatives
  • Use cases: stream ingestion, log shipping, and event pipelines

Installation, Configuration & Cluster Setup

  • Installing NiFi on a single node and in cluster mode
  • Cluster configuration: node roles, ZooKeeper integration, and load balancing
  • Orchestrating NiFi deployments using Ansible, Docker, or Helm

Designing and Managing Dataflows

  • Routing, filtering, splitting, and merging flows
  • Configuring processors (e.g., InvokeHTTP, QueryRecord, PutDatabaseRecord)
  • Managing schemas, enrichment, and transformation operations
  • Handling errors, retry mechanisms, and backpressure

Integration Scenarios

  • Connecting to databases, messaging systems, and REST APIs
  • Streaming data to analytics systems such as Kafka, Elasticsearch, or cloud storage
  • Integrating with Splunk, Prometheus, or logging pipelines

Monitoring, Recovery & Provenance

  • Utilising the NiFi UI, metrics, and provenance visualiser
  • Designing autonomous recovery and graceful failure handling
  • Backup strategies, flow versioning, and change management

Performance Tuning & Optimisation

  • Tuning JVM, heap, thread pools, and clustering parameters
  • Refining flow design to minimise bottlenecks
  • Resource isolation, flow prioritisation, and throughput control

Best Practices & Governance

  • Flow documentation, naming conventions, and modular design
  • Security measures: TLS, authentication, access control, and data encryption
  • Change control, versioning, role-based access, and audit trails

Troubleshooting & Incident Response

  • Addressing common issues: deadlocks, memory leaks, and processor errors
  • Log analysis, error diagnostics, and root cause investigation
  • Recovery strategies and flow rollback procedures

Hands-on Lab: Implementing a Realistic Data Pipeline

  • Constructing an end-to-end flow covering ingestion, transformation, and delivery
  • Implementing error handling, backpressure management, and scaling strategies
  • Performance testing and pipeline tuning

Summary and Next Steps

Requirements

  • Proficiency with the Linux command line
  • Basic knowledge of networking and data systems
  • Familiarity with data streaming or ETL concepts

Target Audience

  • System administrators
  • Data engineers
  • Developers
  • DevOps professionals

Custom Corporate Training

Training solutions designed exclusively for businesses.

  • Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
  • Flexible Schedule: Dates and times adapted to your team's agenda.
  • Format: Online (live), In-company (at your offices), or Hybrid.
Investment

Price per private group, online live training, starting from 3900 € + VAT*

Contact us for an exact quote and to hear our latest promotions

Testimonials (7)

Provisional Upcoming Courses (Contact Us For More Information)

Related Categories