Skip to content
Sign in

AI Studio: images, music, songs & speech

An illustrated guide to generating images, background music, songs and voiceovers, with example prompts, settings and result management.

Open AI Studio

Choose a generator

Open Dashboard > AI Studio and select a tab. Prepare a description for images or music, a theme and lyrics for a song, or the text for a voiceover. Check the signed-in account and any displayed credit estimate before submitting.

Choose a generator
TabUse it forPrepare
ImagesProduct visuals, illustrations and image editsA prompt; optional reference images
Music5–30 seconds of instrumental background musicStyle, mood, instruments and duration
Songs10–240 seconds of vocals or instrumental musicA musical idea, prompt, lyrics and vocal language
Text-to-speechSpoken narration and voiceoversUp to 200 characters; optional reference voice

Screenshots show the actual generator forms with example text and demo model data, before submission. They are not generated results. Models, available voices and credit estimates may vary by account and configuration.

Generate images

Describe the image you want, or add reference images to guide an edit. The model menu adapts to the number of references.

Open this generator
Images: prompt, reference attachments, model, resolution, quality and estimated credits.
Images: prompt, reference attachments, model, resolution, quality and estimated credits. Open full-size screenshot

Steps

  1. Select Images and enter the subject, scene, composition and visual style in the prompt box.
  2. Optionally attach or drag reference images into the composer. Remove any references you no longer need before choosing a model.
  3. Choose a model, resolution and quality. Wait for the model list and credit estimate to finish loading, then select Generate Image.
  4. Keep the page open until the image appears. Open the result to inspect it, or use Reuse & edit to bring its prompt and settings back into the composer.

Settings

References
The composer displays its attachment limit; each model may accept fewer images. Supported uploads include JPG, PNG, GIF and WebP.
Resolution & quality
Choose 1K, 2K or 4K and low, medium or high quality. The credit estimate refreshes when these settings or references change.
Further edits
Reuse & edit restores the prompt and settings and adds the image as a reference. Use as reference adds only that image to the current composition.

Example input

A white ceramic coffee cup on a light gray tabletop, soft morning light from the left, a little steam, clean product photography, square composition, no text or logo.

The progress steps describe the waiting process; they are not an exact percentage or a promised completion time.

Generate background music

Use Music for short instrumental background tracks. For lyrics, vocals or longer compositions, use Songs.

Open this generator
Music: the background-music description and 5–30 second duration slider.
Music: the background-music description and 5–30 second duration slider. Open full-size screenshot

Steps

  1. Select Music and describe the genre, mood, instruments, tempo and intended use.
  2. Move the duration slider to a value from 5 to 30 seconds in 5-second steps; the initial value is 15 seconds.
  3. Select Generate Music. Follow the task in the generation queue below the form.
  4. When ready, listen in the latest result or recent generations and open the audio file to save it.

Settings

Music prompt
Up to 2,000 characters. Be specific about instruments and mood, and say instrumental or no vocals when that is your intention.
Duration
5–30 seconds. Match it to the video segment that needs background music.

Example input

Instrumental background music for a coffee product video: light jazz piano, soft brushed drums, relaxed morning mood, steady gentle rhythm, no vocals.

Submitting a task starts generation. Wait for the audio result before evaluating it or submitting another version.

Create songs

Create a song from a musical idea or write the prompt and lyrics yourself. Generating a prompt and lyrics creates a draft; generating the song creates the audio.

Open this generator
Songs: idea, editable prompt and lyrics, instrumental mode, vocal language, key, tempo and duration.
Songs: idea, editable prompt and lyrics, instrumental mode, vocal language, key, tempo and duration. Open full-size screenshot

Steps

  1. Select Songs and enter a theme in Song idea. Choose Generate / regenerate song prompt and lyrics if you want a draft; you may also fill in the prompt and lyrics directly.
  2. Review the song prompt and lyrics. Describe the genre, mood and instruments; organize lyrics with markers such as [Verse] and [Chorus]. You can edit or regenerate the draft.
  3. Confirm the vocal language, key signature, tempo and duration. Enable Instrumental to disable the lyrics and vocal-language controls.
  4. Select Generate Song, wait for the queued task to finish, then listen to the resulting audio.

Settings

Vocal language
The initial choice follows the song idea, not the website language. Review it manually, especially after generating a new draft.
Key & tempo
Choose a key from the menu and a tempo from 30 to 240 BPM.
Duration
10–240 seconds. Draft generation can update the duration, tempo, language and instrumental mode, so review them before submitting.
Instrumental
Ignores lyrics and creates music without the vocal-lyrics condition.

Example input

A warm acoustic pop song about starting a new day, gentle guitar and piano, an uplifting chorus, clear English vocals.

[Verse] Morning light comes through the door A little hope, a little more [Chorus] Take this day and make it shine One small step, one dream at a time

AI Studio > Songs and the separate Song Creation page keep their recent generations in separate scopes. Return to the same entry point to find your result.

Synthesize speech

Turn a short piece of text into spoken audio. Choose the spoken language first; available engines and voice controls depend on that choice.

Open this generator
Text-to-speech: language, engine, optional reference voice, narration text and seed.
Text-to-speech: language, engine, optional reference voice, narration text and seed. Open full-size screenshot

Steps

  1. Select Text-to-speech, choose the spoken language and an available TTS engine.
  2. Use the default voice, choose a saved voice, or upload a reference recording when the engine supports it. A reference recording is optional.
  3. Enter up to 200 characters in the text field. Check names, numbers and punctuation; fill in an expression style if the selected engine offers it.
  4. Optionally enter a seed, check the displayed credit cost, then select Generate Speech. Wait for the queue to finish and listen to the audio.

Settings

Reference voice
Upload WAV, MP3 or M4A. Use a clear recording and verify the reference transcript; do not replace it with the new text you want spoken.
Saved voices
To save a new voice, provide its name and confirm the voice authorization shown on the form. Refresh the saved-voice list after saving; voices are filtered by engine.
Seed
Optional whole number from 0 to 2147483647. Leave it blank when you do not need to specify one.
Long scripts
Split scripts into sections of at most 200 characters and generate each section separately.

Example input

Good morning. A fresh cup of coffee and a quiet moment can be the perfect start to your day.

The reference transcript describes the uploaded recording. The text field contains the new narration. These are two different inputs.

Text-to-speech API reference

Preview, reuse & download

The generation queue and recent generations follow the selected tab. Music, songs and speech appear in the queue while processing; images return directly to the image workspace.

Read the task status
Pending means waiting; processing means the task is running. Completed audio can be played in the result panel. A failed task shows an error to check before trying again.
Open & download
Use Open or Download on a result. The file opens in a new tab; save the image or audio through the browser. Copy URL copies the media address.
Reuse an image
Choose a history image, then Reuse & edit or Use as reference. Inspect the new model selection and estimate before generating again.
Find recent work
Recent generations and the queue are stored in this browser. Use the same browser and entry point; clearing browser data or changing devices does not restore this local list.
Clear & delete
Clear in recent generations removes the current tab's local history. Delete in the latest audio-result panel also requests server file deletion. Save anything you need before deleting.

Troubleshooting

Why is image generation disabled?
Enter a prompt, finish any reference uploads and wait for the model catalog and estimate to load. If loading fails, use Retry. Check that an available model supports the attached references.
Why is my song in the wrong language or missing vocals?
Check the vocal language and the Instrumental checkbox. A generated draft can change those settings. Review the lyrics before generating the audio.
Why is speech generation rejected?
Check that the text is nonempty and no longer than 200 characters, the seed is valid and any reference upload has finished. If saving a new voice, complete its authorization checkbox.
Where did my result go?
Select the tab and entry point used to create it, using the same browser. AI Studio and the separate Song Creation page do not share the same recent-results view.
The task failed or is taking a long time.
Check the error shown in the queue, account access and available credits. Avoid repeated submissions while a task is pending. If the problem persists, send support the task ID and error message.
Contact support