
Whisper Large v3 Audio Encoder Model.
Used to extract speech features, perform lip-syncing, and analyze emotional rhythm, enhancing lip accuracy and audio alignment in digital human video broadcasting.
Suitable for LongCat Avatar audio-driven workflows.
暂无作品
