What does this skill do?

The AI Skill Media Studio allows you to generate high-quality images, video, and audio using fal.ai models through its MCP integration. It covers text-to-image, text-to-video, text-to-speech, and editing of existing images (inpainting, outpainting, style transfer), with cost estimates provided before each generation for total budget control.

Text to Image
Generate images from text descriptions using models such as Nano Banana 2, controlling the size, number of images, and seed.
Text to Video
Create dynamic, cinematic videos from text or still images using templates such as Seedance 1.0 Pro.
Text-to-Speech
Create natural, conversational-sounding speech with CSM-1B, ideal for narration, demos, and synchronized sound effects.
Image Editing
Edit existing images using inpainting, outpainting, or style transfer by first uploading the source image.

Usage examples

🖼️ Generate image
Generate an image of a futuristic urban landscape at sunset, in the cyberpunk style, in 16:9 aspect ratio.
🎬 Create a video
Create a 5-second, cinematic video of a drone flight over a mountain lake during the golden hour.
🔊 Natural voice
Generate a voice that says: "Hello, welcome to the demo. I'll show you how this works."
✏️ Edit image
Upload this image and regenerate it in a watercolor style while keeping the same composition.

Features

Unified Media Creation Images, video, and audio in a single skill using fal.ai MCP models.
Preliminary Cost Estimate Use `estimate_cost` to check the cost of each operation before running resource-intensive tasks such as video processing.
Model Discovery Find the right model for each task using the search tool (text-to-image, text-to-video, voice, etc.).
Upload of base resources Upload images or videos in advance to use them as a basis for editing or image-to-image generation.
Asynchronous Monitoring Check the status of jobs in progress and retrieve the final file with the results when it is ready.

Frequently asked questions

You need to configure the fal.ai MCP server in ~/.claude.json with your API key (FAL_KEY). Without this configuration, the skill will not be able to generate media.
Yes. The skill includes the `estimate_cost` tool, which returns a breakdown of the estimated cost per unit of generation before the task is executed.
You can generate images (text-to-image), videos (text-to-video or image-to-video), natural-sounding speech (text-to-speech), and edit existing images using inpainting, outpainting, or style transfer.
Not always. Some generations, especially video ones, are asynchronous. You'll receive a job ID and can monitor its status using the status tool until it's complete.
AI Media Studio — Unified Creation of Images, Video, and Audio with Claude AI

¿Prefieres escuchar el contenido? Genera la narración de audio con un clic.