Ai Audio Generation
AI audio generation for agents through Image Skill's zero-setup hosted creative runtime
共 69338 条已上架资产
AI audio generation for agents through Image Skill's zero-setup hosted creative runtime
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language
Professional audio production for music, podcasts, and sound design
Convert, compress, merge, split, clip, inspect, and extract audio locally with FFmpeg, plus decode supported music-cache…
Configure and use Gladia audio intelligence features: speaker diarization, translation, sentiment analysis, named entity…
Convert, compress, merge, split, clip, inspect, and extract audio locally with FFmpeg, plus decode supported music-cache…
游戏音频原则
Generate MP3 speech with Fish Audio through RunAPI
使用 fal
兼顾速度与精度的文字识别
Extract text from images using Tesseract
调用 OCR
Extract text from images using Tesseract
截图文字识别
OCR document extraction - extract text from scanned documents, photos, and images using OCR
Extract text from images using Tesseract
Professional-grade OCR for PDFs and images using MinerU
Multi-model OCR benchmark and comparison tool
[macOS only] Use this skill when the user requests OCR (Optical Character Recognition), image/PDF text extraction
Extract text from images using Tesseract