Skip to main content
Audio models are selected on the audio-generating nodes. They cover three jobs: turning text into speech, composing music from a description, and adding sound effects to a video.

Families

Settings

Each audio node exposes the settings relevant to its model. Availability depends on the model you pick.

Speech

Music

| lyrics | MiniMax Music 3 only: the words the singer sings, with tags like [verse] or [chorus] on their own lines. Leave it empty for an instrumental track. force_instrumental wins if both are set. | See Model settings for the full reference.

Generate Speech

Convert text to spoken audio with an ElevenLabs voice.

Generate Music

Compose a music track from a text description.

Add Sound Effects

Generate and attach sound effects to a video.