docs.emix.ai
language
language
  • 🇺🇸 English
  • 🇨🇳 Chinese
language
language
  • 🇺🇸 English
  • 🇨🇳 Chinese
Market
File Upload API
Market
File Upload API
  1. Music Generation
  • Getting Started with Emix API (Important)
  • Market
  • Image Models
    • Topaz
      • Topaz - Image Upscale
    • Seedream
      • Seedream3.0 - Text to Image
      • Seedream4.0 - Text to Image
      • Seedream4.0 - Edit
      • Seedream4.5 - Text to Image
      • Seedream4.5 - Edit
      • Seedream5.0 Lite - Text to Image
      • Seedream5.0 Lite - Image to Image
      • Seedream5.0 Pro - Text to Image
      • Seedream5.0 Pro - Image to Image
      • Seedream 5.0 Pro - Layer Decomposition
    • Z-image
      • Z-Image
    • Google
      • Google - imagen4-fast
      • Google - imagen4-ultra
      • Google - imagen4
      • Google - Nano Banana
      • Google - Nano Banana Pro
      • Google - Nano Banana 2
      • Google - Nano Banana 2 Lite
    • Flux-2
      • Flux-2 - Pro Image to Image
      • Flux-2 - Pro Text to Image
      • Flux-2 - Image to Image
      • Flux-2 - Text to Image
    • Grok Imagine
      • Grok Imagine - Text to Image
      • Grok Imagine - image to image
      • Grok Imagine Image 2.0 Text To Image
      • Grok Imagine Image 2.0 Segment Map
      • Grok Imagine Image 2.0 Image Edit
    • GPT Image
      • GPT Image 2.5 Sunburst Text to Image
      • GPT Image 2.5 Sunburst Image to Image
      • GPT Image 2.5 Flare Text to Image
      • GPT Image 2.5 Flare Image to Image
      • GPT Image-1.5 - Text to Image
      • GPT Image-1.5 - Image to Image
      • GPT Image-2 - Text to Image
      • GPT Image 2 - Image To Image
    • Recraft
      • Recraft - Remove Background
      • Recraft - Crisp Upscale
    • Ideogram
      • Ideogram - V3 Reframe
      • Ideogram - Character Edit
      • Ideogram - Character Remix
      • Ideogram - Character
      • Ideogram V3 Text to Image
      • Ideogram V3 Edit
      • Ideogram V3 Remix
    • 4o Image API
      • 4o Image API Quickstart
      • 4o Image Generation Callbacks
      • Generate 4o Image
      • Get 4o Image Details
      • Get Direct Download URL
    • Flux Kontext API
      • Flux Kontext API Quickstart
      • Image Generation or Editing Callbacks
      • Generate or Edit Image
      • Get Image Details
    • Qwen
      • Qwen - Text to Image
      • Qwen - Image to Image
      • Qwen - Image Edit
      • Qwen2 - Image Edit
      • Qwen2 - Text To Image
      • Qwen3 Pro Text to Image
      • Qwen3 Text to Image
      • Qwen3 Pro Image to Image
      • Qwen3 Image to Image
      • Qwen 2.1 - Text to Image
      • Qwen 2.1 - Image to Image
    • Wan
      • Wan 2.7 Image
      • Wan 2.7 Image Pro
  • Video Models
    • OmniHuman
      • Omnihuman 1.5
      • Omnihuman 1.5 Human Identification
      • OmniHuman 1.5 Subject Detection
    • Hailuo
      • Hailuo 2.3 Pro Image to Video
      • Hailuo 2.3 Standard Image to Video
      • Hailuo Pro Text to Video
      • Hailuo Pro Image to Video
      • Hailuo Standard Text to Video
      • Hailuo Standard Image to Video
    • Topaz
      • Topaz - Video Upscale
    • Runway API
      • Runway Image To Video
      • Runway Text To Video
      • Runway Video Extension
    • HappyHorse
      • HappyHorse - text-to-video
      • HappyHorse - image-to-video
      • HappyHorse - reference-to-video
      • HappyHorse - video-edit
      • HappyHorse-1-1 image-to-video
      • HappyHorse-1-1 text-to-video
      • HappyHorse-1-1 reference-to-video
    • Wan
      • Wan - 2.2 A14B Image to Video Turbo
      • Wan - 2.2 A14B Speech to Video Turbo
      • Wan - 2.2 A14B Text to Video Turbo
      • Wan - Animate Move
      • Wan - Animate Replace
      • Wan 2.6 - Image to Video
      • Wan 2.6 - Text to Video
      • Wan 2.6 - Video to Video
      • Wan 2.5 - Image to Video
      • Wan 2.5 - Text to Video
      • Wan 2.7 - Text to Video
      • Wan 2.7 - Image to Video
      • Wan 2.7 - Video Edit
      • Wan 2.7 - Reference to Video
      • Wan 3.0 - Video
      • Wan 3.0 - Video Prime
    • Grok Imagine
      • Grok Imagine Text to Video
      • Grok Imagine Image to Video
      • Grok Imagine - Video Upscale
      • Grok Imagine - Video Extend
      • Grok Imagine Video 1.5 Preview
    • Kling
      • Kling 2.6 Text to Video
      • Kling 2.6 Image to Video
      • Kling - V2.5 Turbo Image to Video Pro
      • Kling - V2.5 Turbo Text to Video Pro
      • Kling AI Avatar Standard
      • Kling AI Avatar Pro
      • Kling V2.1 Master Image to Video
      • Kling V2.1 Master Text to Video
      • Kling V2.1 Pro
      • Kling V2.1 Standard
      • Kling 2.6 motion-control
      • Kling-3.0 motion-control
      • Kling 3.0
      • Kling - V3 Turbo Text to Video
      • Kling - V3 Turbo Image to Video
      • Kling 3.0 Omni Reference To Video
      • Kling 3.0 Omni Transformation
      • Kling 3.0 Omni Image To Video
      • Kling 3.0 Omni Text to Video
    • Bytedance
      • Bytedance Seedance 2.0
      • Bytedance Seedance 2.0 Fast
      • Bytedance Seedance 2.0 Mini
      • Bytedance Seedance 1.5 Pro
      • Bytedance V1 Pro Fast Image to Video
      • Bytedance V1 Pro Image to Video
      • Bytedance - V1 Pro Text to Video
      • Bytedance - V1 Lite Image to Video
      • Bytedance - V1 Lite Text to Video
      • Bytedance Seedance 2.5
    • Veo3.1 API
      • Get 4K Video Callbacks
      • Veo3.1 API Quickstart
      • Veo3.1 Video Generation Callbacks
      • Generate Veo3.1 Video
      • Get Veo3.1 Video Details
      • Get 1080P Video
      • Get 4K Video
      • VEO 3.1 Extend Video
      • VEO 3.1 Text to video
      • VEO 3.1 Image to video
      • VEO 3.1 Reference to vidoe
    • Gemini Omni
      • Gemini Omni 1.1 Flash
      • Gemini Omni Video
      • Gemini Omni Audio
      • Gemini Omni Character
    • Volcengine
      • Volcengine video to video lip sync
    • PixVerse
      • PixVerse V6 Text-to-Video
      • PixVerse V6 Image-to-Video
      • PixVerse V6 First & Last Frame Transition
      • PixVerse V6 Video Extension
      • PixVerse V6 Fusion / Reference-to-Video
    • MiniMax H3
      • MiniMax H3 Text-to-Video
      • MiniMax H3 Image-to-Video
      • MiniMax H3 Reference-to-Video
  • Music Models
    • ElevenLabs
      • elevenlabs/audio-isolation
      • elevenlabs/sound-effect-v2
      • elevenlabs/speech-to-text
      • elevenlabs/text-to-dialogue-v3
      • elevenlabs/text-to-speech-multilingual-v2
      • elevenlabs/text-to-speech-turbo-2-5
    • Suno API
      • Music Generation
        • Music Generation Callbacks
        • Music Extension Callbacks
        • Add Instrumental Callbacks
        • Add Vocals Callbacks
        • Music Cover Generation Callbacks
        • Replace Music Section Callbacks
        • Audio Upload and Extension Callbacks
        • Audio Upload and Cover Callbacks
        • Generate Music
          POST
        • Extend Music
          POST
        • Upload And Cover Audio
          POST
        • Upload And Extend Audio
          POST
        • Add Instrumental to Music
          POST
        • Add Vocals to Music
          POST
        • Get Timestamped Lyrics
          POST
        • Boost Music Style
          POST
        • Generate Music Cover
          POST
        • Replace Music Section
          POST
        • Generate Persona
          POST
        • Generate Mashup Music
          POST
      • Lyrics Generation
        • Lyrics Generation Callbacks
        • Generate Lyrics
      • WAV Conversion
        • Convert to WAV Format
      • Vocal Removal
        • Audio Separation Callbacks
        • Vocal & Instrument Stem Separation
        • Generate MIDI from Audio
      • Music Video Generation
        • Music Video Generation Callbacks
        • Create Music Video
      • Sounds Generation
        • Generate sounds
      • voice
        • Suno Voice Generation Callback
        • Suno Voice Validation Phrase Callback
        • Suno Voice Generate Verification Phrase API
        • Suno Voice Create Custom Voice API
        • Suno Voice Regenerate Verification Phrase
        • Suno Voice Check Availability API
    • Gemini
      • Gemini 3.1 Flash Text to speech
      • Gemini 2.5 Pro Text to Speech
  • Chat Models
    • Grok
      • Grok 4.5
      • Grok 4.3
      • Grok 4.6
    • Codex
      • GPT Codex
    • GPT
      • GPT 5.2
      • GPT 5.4 (response)
      • GPT 5.6 Luna
      • GPT 5.6 Terra
      • GPT 5.6 Sol
      • GPT 5.5 (response)
    • Claude
      • Claude Code + emix.ai Integration Guide
      • Claude Sonnet 5
      • Claude Sonnet 4.5
      • Claude Opus 4.7
      • Claude Opus 4.8
      • Claude Opus 5
      • Claude Fable 5
      • Claude Haiku 4.5
      • Claude Opus 4.5
      • Claude Opus 4.6
      • Claude Sonnet 4.5
      • Claude Sonnet 4.6
    • Gemini
      • Gemini 3.6 Flash
      • Gemini 3.6 Flash (openai)
      • Gemini 3.7 Flash (openai)
      • Gemini 2.5 Pro (openai)
      • Gemini 3 Pro (openai)
      • Gemini 3.1 Pro (openai)
      • Gemini 2.5 Flash (openai)
      • Gemini 3 Flash (openai)
      • Gemini 3.5 Flash
      • Gemini 3.5 Flash (openai)
      • Gemini 3 Flash
      • Gemini 3.7 Flash
      • Gemini 3.8 Flash
      • Gemini 3.8 Flash (openai)
  • Get Task Details
    GET
  1. Music Generation

Generate Music

POST
/api/v1/client/tasks
Generate music with or without lyrics using AI models.

Usage Guide#

This endpoint creates music based on your text prompt
Multiple variations will be generated for each request
You can control detail level with custom mode and instrumental settings

Parameter Details#

Always required: customMode, instrumental, model
In Custom Mode (customMode: true):
title is optional, maximum 80 characters. Only available in this mode
prompt is optional; when provided, it is used strictly as lyrics and sung in the generated track. If lyrics is also provided, lyrics takes priority
At least one of style, lyrics, or negativeTags must be provided; generation is rejected if all are empty
If instrumental: true: generate instrumental music (no vocals)
If instrumental: false: use lyrics as lyrics (fall back to prompt if lyrics is not provided)
Optional: duration, negativeTags, vocalGender, styleWeight, weirdnessConstraint, audioWeight, variety, personaId, personaModel
Character limits by model:
V4: prompt 3000 characters, style 200 characters
V4_5, V4_5PLUS, V4_5ALL, V5, V5_5: prompt 5000 characters, style 1000 characters
V6, V6_MINI, V6_WILD: style 1000 characters; lyrics 5000 characters; negativeTags 1000 characters
title length limit: 80 characters (all models)
In Non-custom Mode (customMode: false):
prompt is optional; it serves as the core idea, and lyrics are generated automatically (not a strict match), maximum 3000 characters
lyrics can be used as a lyrics attachment together with prompt
At least one of imageUrls, videoUrls, audioUrls, style, or lyrics must be provided; generation is rejected if all are empty
imageUrls, videoUrls, and audioUrls are only valid in this mode
Do not pass parameters that are only available in custom mode: title, negativeTags, duration, vocalGender, styleWeight, weirdnessConstraint, audioWeight, variety
Total attachments must not exceed 10: style + lyrics + imageUrls + videoUrls + audioUrls

Optional Parameters#

prompt (string): Description of the desired audio content. Optional. Used as lyrics in custom mode; used as the core idea in non-custom mode (maximum 3000 characters).
lyrics (string): Lyrics content. Optional. V6 maximum 5000 characters. In custom mode, takes priority over prompt as lyrics; in non-custom mode, can be used as a lyrics attachment together with prompt.
imageUrls (array): Image references. Only effective when customMode is false. Up to 5 images, each no more than 10 MB. Supported formats: jpeg, png, webp, bmp.
videoUrls (array): Video references. Only effective when customMode is false. Up to 1 file, each no more than 100 MB, duration no more than 241 seconds. Supported formats: mp4, mov, webm.
audioUrls (array): Audio references. Only effective when customMode is false. Duration must be between 6 seconds and 30 minutes; each file no more than 500 MB.
style (string): Music style specification. See character limits in Parameter Details above.
title (string): Track title. Optional. Only available when customMode is true. Maximum 80 characters. Displayed in player interfaces and filenames.
negativeTags (string): Music styles or traits to exclude from the generated audio. Only available when customMode is true. For V6, V6_MINI, and V6_WILD: maximum 1000 characters.
vocalGender (string): Vocal gender preference. m for male, f for female. Only available when customMode is true. In practice, this only increases probability and cannot guarantee the instruction is followed.
styleWeight (number): Strength of adherence to the specified style. Range 0–1, up to 2 decimal places. Only effective when customMode is true.
weirdnessConstraint (number): Creative/experimental deviation. Range 0–1, up to 2 decimal places. Only effective when customMode is true.
audioWeight (number): Relative weight of audio features. Range 0–1, up to 2 decimal places. Only effective when customMode is true. Not supported when there are no vocals.
variety (number): Diversity of generated results. Integer from 0–4, default 1. Only effective when customMode is true. 0 off (exact style), 1 normal (default, balanced), 2 high (distinct styles), 3 extra (bold exploration), 4 max (maximum variation).
personaId (string): Persona ID or Voice ID. Optional. To generate a Persona ID, see Generate Persona.
personaModel (string): Persona model, style_persona or voice_persona. Only available for V5 (Discontinued), V5.5 (Discontinued), V6, V6_MINI, and V6_WILD.
duration (number): Audio duration in seconds. Range 10–360, default 20. Only available when customMode is true, and only valid when the model is V5_5, V6, V6_MINI, or V6_WILD.

Developer Notes#

Recommendation for new users: Start with customMode: false for simpler usage
Generated files are retained for 14 days
Callback process has three stages: text (text generation), first (first track complete), complete (all tracks complete)

Callbacks

audioGenerated

Request

Authorization
Bearer Token
Provide your bearer token in the
Authorization
header when making requests to protected resources.
Example:
Authorization: Bearer ********************
or
Body Params application/jsonRequired

Examples

Responses

🟢200
application/json
Request successful
Bodyapplication/json

Request Request Example
Shell
JavaScript
Java
Swift
curl --location 'https://api.emix.ai/api/v1/client/tasks' \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '{
    "model": "suno/generate",
    "callBackUrl": "https://api.example.com/callback",
    "input": {
        "prompt": "A calm and relaxing piano track with soft melodies",
        "imageUrls": [
            "https://loremflickr.com/400/400?lock=2768753606921360",
            "https://loremflickr.com/400/400?lock=7539221471970979",
            "https://loremflickr.com/400/400?lock=5550334571658194"
        ],
        "style": "Classical",
        "title": "Peaceful Piano Meditation",
        "customMode": true,
        "instrumental": true,
        "model": "V6",
        "negativeTags": "Heavy Metal, Upbeat Drums",
        "vocalGender": "m",
        "styleWeight": 0.65,
        "weirdnessConstraint": 0.65,
        "audioWeight": 0.65,
        "personaId": "persona_123",
        "personaModel": "voice_persona",
        "duration": 20
    }
}'
Response Response Example
{
    "code": 200,
    "msg": "success",
    "data": {
        "taskId": "5c79****be8e"
    }
}
Previous
Audio Upload and Cover Callbacks
Next
Extend Music
Built with