fal-ai/controlfoley
Optional text prompt describing the desired audio. When combined with the video it provides text-controlled video-to-audio (TC-V2A) generation; leave empty for pure video-to-audio (V2A).
Negative text prompt — describe audio characteristics to avoid.
URL of the video to generate synchronized audio for.
Target audio duration in seconds. Truncated to source video length when shorter.
Run the model to see the result here.