Generate and edit audio

Generate speech, music, and sound effects on the canvas, then preview, adjust, download, or reuse the result.

Create an audio node to generate speech, music, or sound effects from a description. The result stays in the node, where you can preview it, download it, save it to the Library, or connect it to a video node that accepts audio input.

1. Create an audio node

  1. Select + in the left canvas toolbar;
  2. Choose Audio;
  3. A new audio node appears on the canvas. Select it to open the audio input panel below the node;
  4. Enter what you want to generate.

If you already have an audio file, choose + > Upload, or drag it onto the canvas. The upload becomes a playable audio node. With Seed audio 1.0, you can add that node as reference audio; you can also connect it to a later node that accepts audio input. See Upload files to the canvas for supported formats and file sizes.

2. Choose a model and use case

The model selector is at the lower left of the audio input panel. Some models also show a use-case selector for Text to Speech, Music, or Sound Effect.

What you want to doModels to start withWhat to enter
Generate speech from text or reference audioSeed audio 1.0The text to read and how it should be delivered; you can also add reference audio
Generate narration or dialogueElevenLabs V3The text to read, plus any voice or delivery options shown on the page
Generate a song or background musicMiniMax Music 2.6, Mureka V8, Mureka O2, or Sonilo MusicStyle, mood, tempo, use case, and any lyrics or duration requested by the model
Generate ambience or sound effectsElevenLabs V3 or Sonilo MusicThe sound source, distance, rhythm, duration, and whether it should loop

The models and tools available can vary by account. Use the options currently shown on your page.

3. Use Seed audio 1.0

Seed audio 1.0 can generate speech from text alone or use audio already on the canvas as a reference. Use it when you want new speech to sound similar to a reference clip, or when you need to adjust speed, pitch, volume, format, or subtitles.

Describe the audio

Describe the content, type of sound, pacing, environment, and ending in the input field. A Seed audio 1.0 prompt can contain up to 3,000 characters.

Generate a calm female voice saying, “Meet me by the water at eight tonight.” Use a slightly slow pace and close microphone position. Keep the background quiet and end the sentence naturally.

Add reference audio

  1. Place the audio you want to reference on the same canvas;
  2. Select the target audio node and choose Seed audio 1.0;
  3. Select + in the reference row above the input field;
  4. Select the audio nodes you want to reference. You can select up to three; when finished, select Exit in the top toolbar or press Esc.

One Seed audio 1.0 node can use up to three audio references. Images are not currently supported as references for this model.

The selected audio appears above the input field. In the prompt, say what to preserve and what to change:

Use the voice and speaking rhythm from the connected reference audio to say, “The new product launches this Friday.” Keep a similar vocal tone and microphone distance, add a little more anticipation, and do not add background music.

Set the output

Select the current format beside the model (for example, WAV) to set subtitles and output format. Select the gear to open advanced settings.

SettingAvailable valuesWhen to change it
SubtitlesOn / OffTurn this on when you also need editable text
Output formatWAV, MP3, OGG_OPUSWAV suits further production; MP3 is convenient for previews and sharing
Sample rate8k, 16k, 24k, 32k, 44.1k, 48kKeep the default 24k unless your delivery has a specific requirement
Speech rate0.5–2.0Make speech slower or faster
Pitch-12–+12Make the result lower or higher
Volume0.5–2.0Change the output loudness

Keep the defaults for your first generation. After listening, change one setting at a time so you can tell what affected the result.

4. Generate and preview

  1. Review the prompt, references, and settings;
  2. Select Generate at the lower right of the input panel;
  3. The audio node shows a progress state while the model is running;
  4. When it finishes, the node shows a waveform and playback controls.

Every model call uses Tapies. The Generate button shows an estimate. Seed audio 1.0 is ultimately charged by the model's output duration, so the estimate can differ from the final usage.

Drag the waveform to move to a specific point, or use the 10-second rewind and fast-forward buttons. Pausing playback or moving the playhead does not modify the audio file.

If subtitles are enabled for Seed audio 1.0 and the model returns subtitle text, the canvas creates a text node named Audio subtitles and connects it to the audio node.

5. Regenerate, download, or connect the result

If the result is not right, change the prompt, references, or settings, then select Generate again. The audio node shows the new result. To keep the current version, download it or save it to the Library before generating again.

ActionWhat happens
DownloadSaves the current audio file to your computer
Save to LibraryOpens the save panel and stores the audio in your personal or team Library
Connect the right handle to a video nodeUses the audio as input for a video node that supports audio references

The waveform and playback controls are for previewing audio; they do not trim the file. To change the content, pacing, or sound, update the input or settings and generate again.

Next: Create text & 3D