Agent Skills: Multimodal Fusion for Speaker Diarization

Combine visual features (face detection, lip movement analysis) with audio features to improve speaker diarization accuracy in video files. Use OpenCV for face detection and lip movement tracking, then fuse visual cues with audio-based speaker embeddings. Essential when processing video files with multiple visible speakers or when audio-only diarization needs visual validation.

UncategorizedID: benchflow-ai/skillsbench/Multimodal Fusion for Speaker Diarization

Install this agent skill to your local

pnpm dlx add-skill https://github.com/benchflow-ai/skillsbench/Multimodal Fusion for Speaker Diarization

Skill Files

Browse the full folder contents for Multimodal Fusion for Speaker Diarization.

Download Skill

Loading file tree…

Select a file to preview its contents.