Audio-driven talking-avatar videos from a single image. Upload a reference image and a speech clip, describe the scene, and get a ~5 second lip-synced video (model, code).
Longer, more descriptive prompts (appearance, actions, scene) give better results.