Multimodal Generative AI: Vision, Speech, and Assistants
About this course
This four-week course provides a hands-on deep dive into the full spectrum of modern AI capabilities. You will master Image-to-Text (Vision), Text-to-Speech (TTS), and Speech-to-Text (Whisper), before culminating in the development of sophisticated AI Assistants. By the end of the course, you’ll be able to build intelligent, multi-modal applications that can see, hear, speak, and solve complex problems.
Price shown by edX — confirm on their site.
Enroll on edXYou'll be redirected to edX to complete enrollment.
- Listed & compared by CourseAsk
- English · All Levels
More courses like this
Coursera
Multi-Agent Systems Design: AI Customer Support with n8n
Coursera · MOOC / Non-credit
Using AI to Code: Your Step-by-Step Guide
Udemy · MOOC / Non-credit
ChatGPT and AI Master Course: Prompt Engineering and AI Tool
Udemy · MOOC / Non-credit
Modern IT Service Management 4 Foundation || UPDATED ||
Udemy · Certificate
More courses from edX
edX
فن التواصل في العصر الرقمي - The Art of Communication in Digital Age
LEORON Institute · MOOC / Non-credit
edX
Emotional Resilience and the Ecological Self
The University of Wisconsin-Madison · Certificate
edX
Environmental Impact Assessment
Adelaide University · MOOC / Non-credit
edX
المصادر الإستراتيجية وتقنيات المشتريات المتقدمة Strategic Sourcing and Advanced Procurement Techniques
LEORON Institute · MOOC / Non-credit