Agent Skills: multimodal-ai
Patterns for building multimodal AI applications that combine text, images, audio, and video. Covers vision APIs, audio transcription, and unified pipelines. Use when "multimodal AI, vision API, image understanding, GPT-4V, Claude vision, audio transcription, Whisper, document extraction, image to text, " mentioned.
UncategorizedID: omer-metin/skills-for-antigravity/multimodal-ai
15022
Install this agent skill to your local
Skill Files
Browse the full folder contents for multimodal-ai.
Loading file tree…
Select a file to preview its contents.