Generative AIComputer VisionNLP
Speech, frames, and a language model, run in parallel to summarise a video.
Working Colab prototype with recorded output.
Whisper smallViT-GPT2 captioningPhi-3 mini 128kHugging Face TransformersmoviepyOpenCVPython threading
Overview
A multimodal pipeline that turns a video into a written summary. Audio is transcribed with Whisper, frames are sampled once per second and captioned, and Phi-3 writes the summary from both streams.
Problem
Neither the transcript nor the frames tell the whole story on their own, and running three large models one after another is slow.
Approach
Split the work into three concurrent stages on threads: extract and transcribe audio in 30-second chunks, caption one frame per second, then prompt Phi-3 mini with the captions and transcript together.
Architecture
How it fits together
Select a component to see what it does. Blue packets show the direction data moves.
- Video to Audio track
- Video to Frames @1 fps
- Audio track to Whisper small
- Frames @1 fps to ViT-GPT2
- Whisper small to Phi-3 mini (transcript)
- ViT-GPT2 to Phi-3 mini (captions)
- Phi-3 mini to Summary
What it does
- Audio extraction and chunked transcription with openai/whisper-small.
- One-frame-per-second sampling and captioning with nlpconnect/vit-gpt2-image-captioning.
- Summary generation with microsoft/Phi-3-mini-128k-instruct.
- Parallel execution of the stages with Python threads.
Recorded results
- Output
- Qualitative
- Recorded summary of a wildlife clip; no automatic metric
- Source: code.ipynb, cell 10
Limitations
- Runs on a Colab GPU; the input path is fixed to /content/video.mp4.
- No quantitative evaluation of summary quality.