Get in Touch
 Duration 14 hours (2 days)

Course Outline

Introduction to Speech Recognition Technologies

  • Historical context and the evolution of speech recognition
  • The roles of acoustic models, language models, and decoding processes
  • Contemporary architectures: RNNs, transformers, and Whisper

Audio Preprocessing and Transcription Fundamentals

  • Managing audio formats and sample rates
  • Audio cleaning, trimming, and segmentation techniques
  • Converting audio to text: distinguishing between real-time and batch processing

Practical Application with Whisper and External APIs

  • Installation and utilisation of OpenAI Whisper
  • Leveraging cloud-based APIs (such as Google and Azure) for transcription
  • Comparative analysis of performance, latency, and associated costs

Language Diversity, Accents, and Domain-Specific Adaptation

  • Processing multiple languages and diverse accents
  • Implementing custom vocabularies and enhancing noise tolerance
  • Managing specialised terminology in legal, medical, or technical contexts

Output Formatting and System Integration

  • Incorporating timestamps, punctuation, and speaker identification labels
  • Exporting results into text, SRT, or JSON formats
  • Integrating transcription outputs into applications or databases

Scenario-Based Implementation Laboratories

  • Transcribing content from meetings, interviews, or podcasts
  • Developing voice-to-text command systems
  • Generating real-time captions for video and audio streams

Performance Evaluation, Limitations, and Ethical Considerations

  • Defining accuracy metrics and conducting model benchmarking
  • Addressing bias and fairness within speech models
  • Navigating privacy concerns and compliance requirements

Recap and Future Directions

Requirements

  • A foundational understanding of general AI and machine learning principles
  • Familiarity with standard audio or media file formats and associated tools

Intended Audience

  • Data scientists and AI engineers specialising in voice data processing
  • Software developers creating applications reliant on transcription technologies
  • Organisations investigating speech recognition for process automation

Custom Corporate Training

Training solutions designed exclusively for businesses.

  • Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
  • Flexible Schedule: Dates and times adapted to your team's agenda.
  • Format: Online (live), In-company (at your offices), or Hybrid.
Investment

Price per private group, online live training, starting from 2600 € + VAT*

Contact us for an exact quote and to hear our latest promotions

Provisional Upcoming Courses (Contact Us For More Information)

Related Categories