- Genial
- Tutoriais de IA e automação
- Stop Wasting Hours: Analyze Any Video with AI
Stop Wasting Hours: Analyze Any Video with AI
You will create a Gemini API key, give it and a video file to an AI agent like Codex or Claude Code, and get back a description of what happens in the video, with timestamps, even when the clip has no sound.
O que você precisa
A video file to analyze (I use a silent 16-second clip)
A Google account for Google AI Studio
An AI agent that can run tasks: Codex, Claude Code or ChatGPT Work
Passo a passo
Pick the right level of video analysis
Decide what you need from the video. Level 1 samples snapshots every couple of seconds but loses the sound. Level 2 transcribes the audio but loses the visuals. Level 3 is native video understanding: the model watches the images and hears the audio together. Only Gemini does this, through the API.
Use free transcription if speech is all you need
If the words hold the information (I'd say that is true for about 90% of videos), run a free speech-to-text model on your Mac or PC, such as OpenAI Whisper or NVIDIA Parakeet. Skip to Gemini when what happens on screen matters.
OpenAI Whisperhttps://github.com/openai/whisper
NVIDIA Parakeethttps://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
Prepare a test video
Record a short clip to test with. I film 16 seconds with no sound: counting with my fingers, moving my face and showing an equation on my phone. A silent clip proves Gemini is reading the images, not a transcript.
Create a Gemini API key in Google AI Studio
Open Google AI Studio and click Create an API key. Name it (I use "test" for this video), then click Create key. Copy the key.
Google AI Studiohttps://aistudio.google.com
Open your agent and drop in the video
Open Codex (or Claude Code, or ChatGPT Work). Drag and drop the video file into the chat so the agent knows which file to analyze.
Ask for a native analysis and give it the key
Paste the prompt below, then paste your Gemini API key. Any of these agents works as long as it has the API key it needs.
PromptI'm going to give you an API key for Gemini. I want you to analyze this video natively. Gemini should be able to tell me exactly what's happening inside the video. API key: YOUR_GEMINI_API_KEY
Review the results and check the timestamps
Read the timeline Gemini returns. In my run, Gemini 3.1 Pro picked up each finger count, the hand lowering and the integral shown on the phone, with no sound at all. Check the timestamps against the video before you cut anything with them.
Switch to Flash-Lite to cut costs
For a cheaper option, ask your agent to use Gemini 3.1 Flash-Lite instead of Pro. I use it for all of my video editing.
Fique atento a
Snapshot sampling (level 1) drops everything in the sound, and a transcript (level 2) drops everything visual. Use Gemini when you need both.
Gemini's timestamps are not fully reliable yet. I find they still need refinement, so check them before cutting clips.
The agent can only call Gemini if you give it the API key. Without it, Codex, Claude Code or ChatGPT Work cannot run the analysis.
Create a separate test key for experiments, as I do, instead of reusing a key you depend on.




