Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours (2 days)
Course Outline
Introduction to Speech Recognition Technologies
- Historical context and the evolution of speech recognition
- The roles of acoustic models, language models, and decoding processes
- Contemporary architectures: RNNs, transformers, and Whisper
Audio Preprocessing and Transcription Fundamentals
- Managing audio formats and sample rates
- Audio cleaning, trimming, and segmentation techniques
- Converting audio to text: distinguishing between real-time and batch processing
Practical Application with Whisper and External APIs
- Installation and utilisation of OpenAI Whisper
- Leveraging cloud-based APIs (such as Google and Azure) for transcription
- Comparative analysis of performance, latency, and associated costs
Language Diversity, Accents, and Domain-Specific Adaptation
- Processing multiple languages and diverse accents
- Implementing custom vocabularies and enhancing noise tolerance
- Managing specialised terminology in legal, medical, or technical contexts
Output Formatting and System Integration
- Incorporating timestamps, punctuation, and speaker identification labels
- Exporting results into text, SRT, or JSON formats
- Integrating transcription outputs into applications or databases
Scenario-Based Implementation Laboratories
- Transcribing content from meetings, interviews, or podcasts
- Developing voice-to-text command systems
- Generating real-time captions for video and audio streams
Performance Evaluation, Limitations, and Ethical Considerations
- Defining accuracy metrics and conducting model benchmarking
- Addressing bias and fairness within speech models
- Navigating privacy concerns and compliance requirements
Recap and Future Directions
Requirements
- A foundational understanding of general AI and machine learning principles
- Familiarity with standard audio or media file formats and associated tools
Intended Audience
- Data scientists and AI engineers specialising in voice data processing
- Software developers creating applications reliant on transcription technologies
- Organisations investigating speech recognition for process automation
Custom Corporate Training
Training solutions designed exclusively for businesses.
- Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
- Flexible Schedule: Dates and times adapted to your team's agenda.
- Format: Online (live), In-company (at your offices), or Hybrid.
Price per private group, online live training, starting from 2600 € + VAT*
Contact us for an exact quote and to hear our latest promotions