Generate and edit audio
Generate speech, music, and sound effects on the canvas, then preview, adjust, download, or reuse the result.
Create an audio node to generate speech, music, or sound effects from a description. The result stays in the node, where you can preview it, download it, save it to the Library, or connect it to a video node that accepts audio input.
1. Create an audio node
- Select
+in the left canvas toolbar; - Choose Audio;
- A new audio node appears on the canvas. Select it to open the audio input panel below the node;
- Enter what you want to generate.
If you already have an audio file, choose + > Upload, or drag it onto the canvas. The upload becomes a playable audio node. With Seed audio 1.0, you can add that node as reference audio; you can also connect it to a later node that accepts audio input. See Upload files to the canvas for supported formats and file sizes.
2. Choose a model and use case
The model selector is at the lower left of the audio input panel. Some models also show a use-case selector for Text to Speech, Music, or Sound Effect.
| What you want to do | Models to start with | What to enter |
|---|---|---|
| Generate speech from text or reference audio | Seed audio 1.0 | The text to read and how it should be delivered; you can also add reference audio |
| Generate narration or dialogue | ElevenLabs V3 | The text to read, plus any voice or delivery options shown on the page |
| Generate a song or background music | MiniMax Music 2.6, Mureka V8, Mureka O2, or Sonilo Music | Style, mood, tempo, use case, and any lyrics or duration requested by the model |
| Generate ambience or sound effects | ElevenLabs V3 or Sonilo Music | The sound source, distance, rhythm, duration, and whether it should loop |
The models and tools available can vary by account. Use the options currently shown on your page.
3. Use Seed audio 1.0
Seed audio 1.0 can generate speech from text alone or use audio already on the canvas as a reference. Use it when you want new speech to sound similar to a reference clip, or when you need to adjust speed, pitch, volume, format, or subtitles.
Describe the audio
Describe the content, type of sound, pacing, environment, and ending in the input field. A Seed audio 1.0 prompt can contain up to 3,000 characters.
Generate a calm female voice saying, “Meet me by the water at eight tonight.” Use a slightly slow pace and close microphone position. Keep the background quiet and end the sentence naturally.
Add reference audio
- Place the audio you want to reference on the same canvas;
- Select the target audio node and choose Seed audio 1.0;
- Select
+in the reference row above the input field; - Select the audio nodes you want to reference. You can select up to three; when finished, select Exit in the top toolbar or press
Esc.
One Seed audio 1.0 node can use up to three audio references. Images are not currently supported as references for this model.
The selected audio appears above the input field. In the prompt, say what to preserve and what to change:
Use the voice and speaking rhythm from the connected reference audio to say, “The new product launches this Friday.” Keep a similar vocal tone and microphone distance, add a little more anticipation, and do not add background music.
Set the output
Select the current format beside the model (for example, WAV) to set subtitles and output format. Select the gear to open advanced settings.
| Setting | Available values | When to change it |
|---|---|---|
| Subtitles | On / Off | Turn this on when you also need editable text |
| Output format | WAV, MP3, OGG_OPUS | WAV suits further production; MP3 is convenient for previews and sharing |
| Sample rate | 8k, 16k, 24k, 32k, 44.1k, 48k | Keep the default 24k unless your delivery has a specific requirement |
| Speech rate | 0.5–2.0 | Make speech slower or faster |
| Pitch | -12–+12 | Make the result lower or higher |
| Volume | 0.5–2.0 | Change the output loudness |
Keep the defaults for your first generation. After listening, change one setting at a time so you can tell what affected the result.
4. Generate and preview
- Review the prompt, references, and settings;
- Select Generate at the lower right of the input panel;
- The audio node shows a progress state while the model is running;
- When it finishes, the node shows a waveform and playback controls.
Every model call uses Tapies. The Generate button shows an estimate. Seed audio 1.0 is ultimately charged by the model's output duration, so the estimate can differ from the final usage.
Drag the waveform to move to a specific point, or use the 10-second rewind and fast-forward buttons. Pausing playback or moving the playhead does not modify the audio file.
If subtitles are enabled for Seed audio 1.0 and the model returns subtitle text, the canvas creates a text node named Audio subtitles and connects it to the audio node.
5. Regenerate, download, or connect the result
If the result is not right, change the prompt, references, or settings, then select Generate again. The audio node shows the new result. To keep the current version, download it or save it to the Library before generating again.
| Action | What happens |
|---|---|
| Download | Saves the current audio file to your computer |
| Save to Library | Opens the save panel and stores the audio in your personal or team Library |
| Connect the right handle to a video node | Uses the audio as input for a video node that supports audio references |
The waveform and playback controls are for previewing audio; they do not trim the file. To change the content, pacing, or sound, update the input or settings and generate again.
Next: Create text & 3D