目录 / Contendeo
Contendeo
Multimodal video analysis MCP. Extracts transcripts, OCR text, and visual chart data from YouTube, Vimeo, and Instagram. 10 free credits to start. Transcripts capture half the signal. When a speaker says "look at this chart" and points at a number on screen, a transcript-only tool loses the data. Contendeo returns it. Contendeo is a true multimodal video analysis MCP server. We process the actual video frames alongside the audio to give your LLM complete context. Our pipeline (yt-dlp → ffmpeg → Groq Whisper → Tesseract OCR → Claude Vision) extracts timestamped transcripts, keyframe descriptions, and hard OCR data into one structured output. Works with YouTube, Instagram Reels, Vimeo, Twitter/X, TikTok, and direct video URLs. **Four tools:** - `quick_transcribe` — fast audio transcription with speaker labels (1 credit) - `deep_analyze` — full multimodal extraction: transcript + visual keyframe analysis + OCR + chart/diagram data (5 credits) - `clip_context` — analyze a specific timestamp range without processing the full video (1–3 credits) - `batch_analyze` — parallel processing of up to 10 videos with cross-video synthesis **Pricing & Free Tier:** Pay only for processing. Cache hits are free, and failures are refunded automatically. Get 10 free credits on signup, no card required. Create your account at [contendeo.app](https://contendeo.app).
接入信息
- 传输形态
- http
- 鉴权方式
- 鉴权未知
- 端点
https://contendeo--highcryptoclub.run.tools
{
"mcpServers": {
"Contendeo": {
"url": "https://contendeo--highcryptoclub.run.tools"
}
}
}
能力清单
| 工具 | 说明 |
|---|---|
| quick_transcribe | TRANSCRIPTION ONLY (no visual analysis). For vision, OCR, charts, or keyframe analysis, use deep_analyze instead. Returns timestamped transcript with speaker labels. Supports YouTube, Instagram Reels, Vimeo, Twitter/X, TikTok, and direct video URLs. Costs 1 credit. |
| deep_analyze | PRIMARY tool for video understanding. Full multimodal pipeline: transcript + visual keyframe analysis + OCR + chart/diagram extraction. Returns unified analysis with Summary, Key Claims, Visual Assets, Data Extracted, Entities Mentioned. Costs 5 credits. |
| clip_context | Analyze a specific segment of a video by timestamp range. Default is full multimodal analysis (transcript + vision + OCR, 3 credits). Pass mode='quick' for transcript-only (1 credit). Use when you only need a section, not the full video. |
| batch_analyze | Process multiple videos and get cross-video synthesis. Max 10 URLs. Returns individual results plus common themes, entity overlap, and contradictions. Credit cost is per-video rate with 10% discount on 5+ videos. |
提交举报 / 纠错
侵权举报经核验成立后,我们会即时下线该条目并删除已存的内容副本。