AI Studio: images, music, songs & speech
An illustrated guide to generating images, background music, songs and voiceovers, with example prompts, settings and result management.
Open AI StudioChoose a generator
Open Dashboard > AI Studio and select a tab. Prepare a description for images or music, a theme and lyrics for a song, or the text for a voiceover. Check the signed-in account and any displayed credit estimate before submitting.
| Tab | Use it for | Prepare |
|---|---|---|
| Images | Product visuals, illustrations and image edits | A prompt; optional reference images |
| Music | 5–30 seconds of instrumental background music | Style, mood, instruments and duration |
| Songs | 10–240 seconds of vocals or instrumental music | A musical idea, prompt, lyrics and vocal language |
| Text-to-speech | Spoken narration and voiceovers | Up to 200 characters; optional reference voice |
Screenshots show the actual generator forms with example text and demo model data, before submission. They are not generated results. Models, available voices and credit estimates may vary by account and configuration.
Generate images
Describe the image you want, or add reference images to guide an edit. The model menu adapts to the number of references.
Open this generator
Steps
- Select Images and enter the subject, scene, composition and visual style in the prompt box.
- Optionally attach or drag reference images into the composer. Remove any references you no longer need before choosing a model.
- Choose a model, resolution and quality. Wait for the model list and credit estimate to finish loading, then select Generate Image.
- Keep the page open until the image appears. Open the result to inspect it, or use Reuse & edit to bring its prompt and settings back into the composer.
Settings
- References
- The composer displays its attachment limit; each model may accept fewer images. Supported uploads include JPG, PNG, GIF and WebP.
- Resolution & quality
- Choose 1K, 2K or 4K and low, medium or high quality. The credit estimate refreshes when these settings or references change.
- Further edits
- Reuse & edit restores the prompt and settings and adds the image as a reference. Use as reference adds only that image to the current composition.
Example input
A white ceramic coffee cup on a light gray tabletop, soft morning light from the left, a little steam, clean product photography, square composition, no text or logo.
The progress steps describe the waiting process; they are not an exact percentage or a promised completion time.
Generate background music
Use Music for short instrumental background tracks. For lyrics, vocals or longer compositions, use Songs.
Open this generator
Steps
- Select Music and describe the genre, mood, instruments, tempo and intended use.
- Move the duration slider to a value from 5 to 30 seconds in 5-second steps; the initial value is 15 seconds.
- Select Generate Music. Follow the task in the generation queue below the form.
- When ready, listen in the latest result or recent generations and open the audio file to save it.
Settings
- Music prompt
- Up to 2,000 characters. Be specific about instruments and mood, and say instrumental or no vocals when that is your intention.
- Duration
- 5–30 seconds. Match it to the video segment that needs background music.
Example input
Instrumental background music for a coffee product video: light jazz piano, soft brushed drums, relaxed morning mood, steady gentle rhythm, no vocals.
Submitting a task starts generation. Wait for the audio result before evaluating it or submitting another version.
Create songs
Create a song from a musical idea or write the prompt and lyrics yourself. Generating a prompt and lyrics creates a draft; generating the song creates the audio.
Open this generator
Steps
- Select Songs and enter a theme in Song idea. Choose Generate / regenerate song prompt and lyrics if you want a draft; you may also fill in the prompt and lyrics directly.
- Review the song prompt and lyrics. Describe the genre, mood and instruments; organize lyrics with markers such as [Verse] and [Chorus]. You can edit or regenerate the draft.
- Confirm the vocal language, key signature, tempo and duration. Enable Instrumental to disable the lyrics and vocal-language controls.
- Select Generate Song, wait for the queued task to finish, then listen to the resulting audio.
Settings
- Vocal language
- The initial choice follows the song idea, not the website language. Review it manually, especially after generating a new draft.
- Key & tempo
- Choose a key from the menu and a tempo from 30 to 240 BPM.
- Duration
- 10–240 seconds. Draft generation can update the duration, tempo, language and instrumental mode, so review them before submitting.
- Instrumental
- Ignores lyrics and creates music without the vocal-lyrics condition.
Example input
A warm acoustic pop song about starting a new day, gentle guitar and piano, an uplifting chorus, clear English vocals.
[Verse] Morning light comes through the door A little hope, a little more [Chorus] Take this day and make it shine One small step, one dream at a time
AI Studio > Songs and the separate Song Creation page keep their recent generations in separate scopes. Return to the same entry point to find your result.
Synthesize speech
Turn a short piece of text into spoken audio. Choose the spoken language first; available engines and voice controls depend on that choice.
Open this generator
Steps
- Select Text-to-speech, choose the spoken language and an available TTS engine.
- Use the default voice, choose a saved voice, or upload a reference recording when the engine supports it. A reference recording is optional.
- Enter up to 200 characters in the text field. Check names, numbers and punctuation; fill in an expression style if the selected engine offers it.
- Optionally enter a seed, check the displayed credit cost, then select Generate Speech. Wait for the queue to finish and listen to the audio.
Settings
- Reference voice
- Upload WAV, MP3 or M4A. Use a clear recording and verify the reference transcript; do not replace it with the new text you want spoken.
- Saved voices
- To save a new voice, provide its name and confirm the voice authorization shown on the form. Refresh the saved-voice list after saving; voices are filtered by engine.
- Seed
- Optional whole number from 0 to 2147483647. Leave it blank when you do not need to specify one.
- Long scripts
- Split scripts into sections of at most 200 characters and generate each section separately.
Example input
Good morning. A fresh cup of coffee and a quiet moment can be the perfect start to your day.
The reference transcript describes the uploaded recording. The text field contains the new narration. These are two different inputs.
Text-to-speech API referencePreview, reuse & download
The generation queue and recent generations follow the selected tab. Music, songs and speech appear in the queue while processing; images return directly to the image workspace.
- Read the task status
- Pending means waiting; processing means the task is running. Completed audio can be played in the result panel. A failed task shows an error to check before trying again.
- Open & download
- Use Open or Download on a result. The file opens in a new tab; save the image or audio through the browser. Copy URL copies the media address.
- Reuse an image
- Choose a history image, then Reuse & edit or Use as reference. Inspect the new model selection and estimate before generating again.
- Find recent work
- Recent generations and the queue are stored in this browser. Use the same browser and entry point; clearing browser data or changing devices does not restore this local list.
- Clear & delete
- Clear in recent generations removes the current tab's local history. Delete in the latest audio-result panel also requests server file deletion. Save anything you need before deleting.
Troubleshooting
- Why is image generation disabled?
- Enter a prompt, finish any reference uploads and wait for the model catalog and estimate to load. If loading fails, use Retry. Check that an available model supports the attached references.
- Why is my song in the wrong language or missing vocals?
- Check the vocal language and the Instrumental checkbox. A generated draft can change those settings. Review the lyrics before generating the audio.
- Why is speech generation rejected?
- Check that the text is nonempty and no longer than 200 characters, the seed is valid and any reference upload has finished. If saving a new voice, complete its authorization checkbox.
- Where did my result go?
- Select the tab and entry point used to create it, using the same browser. AI Studio and the separate Song Creation page do not share the same recent-results view.
- The task failed or is taking a long time.
- Check the error shown in the queue, account access and available credits. Avoid repeated submissions while a task is pending. If the problem persists, send support the task ID and error message.
