docs.emix.ai
language
language
  • πŸ‡ΊπŸ‡Έ English
  • πŸ‡¨πŸ‡³ Chinese
language
language
  • πŸ‡ΊπŸ‡Έ English
  • πŸ‡¨πŸ‡³ Chinese
Market
File Upload API
Market
File Upload API
  1. Vocal Removal
  • Getting Started with Emix API (Important)
  • Market
  • Image Models
    • Topaz
      • Topaz - Image Upscale
    • Seedream
      • Seedream3.0 - Text to Image
      • Seedream4.0 - Text to Image
      • Seedream4.0 - Edit
      • Seedream4.5 - Text to Image
      • Seedream4.5 - Edit
      • Seedream5.0 Lite - Text to Image
      • Seedream5.0 Lite - Image to Image
      • Seedream5.0 Pro - Text to Image
      • Seedream5.0 Pro - Image to Image
      • Seedream 5.0 Pro - Layer Decomposition
    • Z-image
      • Z-Image
    • Google
      • Google - imagen4-fast
      • Google - imagen4-ultra
      • Google - imagen4
      • Google - Nano Banana
      • Google - Nano Banana Pro
      • Google - Nano Banana 2
      • Google - Nano Banana 2 Lite
    • Flux-2
      • Flux-2 - Pro Image to Image
      • Flux-2 - Pro Text to Image
      • Flux-2 - Image to Image
      • Flux-2 - Text to Image
    • Grok Imagine
      • Grok Imagine - Text to Image
      • Grok Imagine - image to image
      • Grok Imagine Image 2.0 Text To Image
      • Grok Imagine Image 2.0 Segment Map
      • Grok Imagine Image 2.0 Image Edit
    • GPT Image
      • GPT Image 2.5 Sunburst Text to Image
      • GPT Image 2.5 Sunburst Image to Image
      • GPT Image 2.5 Flare Text to Image
      • GPT Image 2.5 Flare Image to Image
      • GPT Image-1.5 - Text to Image
      • GPT Image-1.5 - Image to Image
      • GPT Image-2 - Text to Image
      • GPT Image 2 - Image To Image
    • Recraft
      • Recraft - Remove Background
      • Recraft - Crisp Upscale
    • Ideogram
      • Ideogram - V3 Reframe
      • Ideogram - Character Edit
      • Ideogram - Character Remix
      • Ideogram - Character
      • Ideogram V3 Text to Image
      • Ideogram V3 Edit
      • Ideogram V3 Remix
    • 4o Image API
      • 4o Image API Quickstart
      • 4o Image Generation Callbacks
      • Generate 4o Image
      • Get 4o Image Details
      • Get Direct Download URL
    • Flux Kontext API
      • Flux Kontext API Quickstart
      • Image Generation or Editing Callbacks
      • Generate or Edit Image
      • Get Image Details
    • Qwen
      • Qwen - Text to Image
      • Qwen - Image to Image
      • Qwen - Image Edit
      • Qwen2 - Image Edit
      • Qwen2 - Text To Image
      • Qwen3 Pro Text to Image
      • Qwen3 Text to Image
      • Qwen3 Pro Image to Image
      • Qwen3 Image to Image
      • Qwen 2.1 - Text to Image
      • Qwen 2.1 - Image to Image
    • Wan
      • Wan 2.7 Image
      • Wan 2.7 Image Pro
  • Video Models
    • OmniHuman
      • Omnihuman 1.5
      • Omnihuman 1.5 Human Identification
      • OmniHuman 1.5 Subject Detection
    • Hailuo
      • Hailuo 2.3 Pro Image to Video
      • Hailuo 2.3 Standard Image to Video
      • Hailuo Pro Text to Video
      • Hailuo Pro Image to Video
      • Hailuo Standard Text to Video
      • Hailuo Standard Image to Video
    • Topaz
      • Topaz - Video Upscale
    • Runway API
      • Runway Image To Video
      • Runway Text To Video
      • Runway Video Extension
    • HappyHorse
      • HappyHorse - text-to-video
      • HappyHorse - image-to-video
      • HappyHorse - reference-to-video
      • HappyHorse - video-edit
      • HappyHorse-1-1 image-to-video
      • HappyHorse-1-1 text-to-video
      • HappyHorse-1-1 reference-to-video
    • Wan
      • Wan - 2.2 A14B Image to Video Turbo
      • Wan - 2.2 A14B Speech to Video Turbo
      • Wan - 2.2 A14B Text to Video Turbo
      • Wan - Animate Move
      • Wan - Animate Replace
      • Wan 2.6 - Image to Video
      • Wan 2.6 - Text to Video
      • Wan 2.6 - Video to Video
      • Wan 2.5 - Image to Video
      • Wan 2.5 - Text to Video
      • Wan 2.7 - Text to Video
      • Wan 2.7 - Image to Video
      • Wan 2.7 - Video Edit
      • Wan 2.7 - Reference to Video
      • Wan 3.0 - Video
      • Wan 3.0 - Video Prime
    • Grok Imagine
      • Grok Imagine Text to Video
      • Grok Imagine Image to Video
      • Grok Imagine - Video Upscale
      • Grok Imagine - Video Extend
      • Grok Imagine Video 1.5 Preview
    • Kling
      • Kling 2.6 Text to Video
      • Kling 2.6 Image to Video
      • Kling - V2.5 Turbo Image to Video Pro
      • Kling - V2.5 Turbo Text to Video Pro
      • Kling AI Avatar Standard
      • Kling AI Avatar Pro
      • Kling V2.1 Master Image to Video
      • Kling V2.1 Master Text to Video
      • Kling V2.1 Pro
      • Kling V2.1 Standard
      • Kling 2.6 motion-control
      • Kling-3.0 motion-control
      • Kling 3.0
      • Kling - V3 Turbo Text to Video
      • Kling - V3 Turbo Image to Video
      • Kling 3.0 Omni Reference To Video
      • Kling 3.0 Omni Transformation
      • Kling 3.0 Omni Image To Video
      • Kling 3.0 Omni Text to Video
    • Bytedance
      • Bytedance Seedance 2.0
      • Bytedance Seedance 2.0 Fast
      • Bytedance Seedance 2.0 Mini
      • Bytedance Seedance 1.5 Pro
      • Bytedance V1 Pro Fast Image to Video
      • Bytedance V1 Pro Image to Video
      • Bytedance - V1 Pro Text to Video
      • Bytedance - V1 Lite Image to Video
      • Bytedance - V1 Lite Text to Video
      • Bytedance Seedance 2.5
    • Veo3.1 API
      • Get 4K Video Callbacks
      • Veo3.1 API Quickstart
      • Veo3.1 Video Generation Callbacks
      • Generate Veo3.1 Video
      • Get Veo3.1 Video Details
      • Get 1080P Video
      • Get 4K Video
      • VEO 3.1 Extend Video
      • VEO 3.1 Text to video
      • VEO 3.1 Image to video
      • VEO 3.1 Reference to vidoe
    • Gemini Omni
      • Gemini Omni 1.1 Flash
      • Gemini Omni Video
      • Gemini Omni Audio
      • Gemini Omni Character
    • Volcengine
      • Volcengine video to video lip sync
    • PixVerse
      • PixVerse V6 Text-to-Video
      • PixVerse V6 Image-to-Video
      • PixVerse V6 First & Last Frame Transition
      • PixVerse V6 Video Extension
      • PixVerse V6 Fusion / Reference-to-Video
    • MiniMax H3
      • MiniMax H3 Text-to-Video
      • MiniMax H3 Image-to-Video
      • MiniMax H3 Reference-to-Video
  • Music Models
    • ElevenLabs
      • elevenlabs/audio-isolation
      • elevenlabs/sound-effect-v2
      • elevenlabs/speech-to-text
      • elevenlabs/text-to-dialogue-v3
      • elevenlabs/text-to-speech-multilingual-v2
      • elevenlabs/text-to-speech-turbo-2-5
    • Suno API
      • Music Generation
        • Music Generation Callbacks
        • Music Extension Callbacks
        • Add Instrumental Callbacks
        • Add Vocals Callbacks
        • Music Cover Generation Callbacks
        • Replace Music Section Callbacks
        • Audio Upload and Extension Callbacks
        • Audio Upload and Cover Callbacks
        • Generate Music
        • Extend Music
        • Upload And Cover Audio
        • Upload And Extend Audio
        • Add Instrumental to Music
        • Add Vocals to Music
        • Get Timestamped Lyrics
        • Boost Music Style
        • Generate Music Cover
        • Replace Music Section
        • Generate Persona
        • Generate Mashup Music
      • Lyrics Generation
        • Lyrics Generation Callbacks
        • Generate Lyrics
      • WAV Conversion
        • Convert to WAV Format
      • Vocal Removal
        • Audio Separation Callbacks
        • Vocalβ€―&β€―Instrument Stem Separation
          POST
        • Generate MIDI from Audio
          POST
      • Music Video Generation
        • Music Video Generation Callbacks
        • Create Music Video
      • Sounds Generation
        • Generate sounds
      • voice
        • Suno Voice Generation Callback
        • Suno Voice Validation Phrase Callback
        • Suno Voice Generate Verification Phrase API
        • Suno Voice Create Custom Voice API
        • Suno Voice Regenerate Verification Phrase
        • Suno Voice Check Availability API
    • Gemini
      • Gemini 3.1 Flash Text to speech
      • Gemini 2.5 Pro Text to Speech
  • Chat Models
    • Grok
      • Grok 4.5
      • Grok 4.3
      • Grok 4.6
    • Codex
      • GPT Codex
    • GPT
      • GPT 5.2
      • GPT 5.4 (response)
      • GPT 5.6 Luna
      • GPT 5.6 Terra
      • GPT 5.6 Sol
      • GPT 5.5 (response)
    • Claude
      • Claude Code + emix.ai Integration Guide
      • Claude Sonnet 5
      • Claude Sonnet 4.5
      • Claude Opus 4.7
      • Claude Opus 4.8
      • Claude Opus 5
      • Claude Fable 5
      • Claude Haiku 4.5
      • Claude Opus 4.5
      • Claude Opus 4.6
      • Claude Sonnet 4.5
      • Claude Sonnet 4.6
    • Gemini
      • Gemini 3.6 Flash
      • Gemini 3.6 Flash (openai)
      • Gemini 3.7 Flash (openai)
      • Gemini 2.5 Pro (openai)
      • Gemini 3 Pro (openai)
      • Gemini 3.1 Pro (openai)
      • Gemini 2.5 Flash (openai)
      • Gemini 3 Flash (openai)
      • Gemini 3.5 Flash
      • Gemini 3.5 Flash (openai)
      • Gemini 3 Flash
      • Gemini 3.7 Flash
      • Gemini 3.8 Flash
      • Gemini 3.8 Flash (openai)
  • Get Task Details
    GET
  1. Vocal Removal

Generate MIDI from Audio

POST
/api/v1/midi/generate
Convert separated audio tracks into MIDI format with detailed note information for each instrument.

Usage Guide#

Convert separated audio tracks into structured MIDI data containing pitch, timing, and velocity information
Requires a completed vocal separation task ID (from the Vocal Removal API)
Generates MIDI note data for multiple detected instruments including drums, bass, guitar, keyboards, and more
Ideal for music transcription, notation, remixing, or educational analysis
Best results on clean, well-separated audio tracks with clear instrument parts

Prerequisites#

Required
You must first use the Vocal & Instrument Stem Separation API to separate your audio before generating MIDI.

Parameter Reference#

NameTypeDescription
taskIdstringRequired. Task ID from a completed vocal separation.
callBackUrlstringRequired. URL to receive MIDI generation completion notifications.
audioIdstringOptional. Specifies which separated audio track to generate MIDI from. This audioId can be obtained from the originData array in the Get Vocal Separation Details endpoint response. Each item in originData contains an id field that can be used here. If not provided, MIDI will be generated from all separated tracks.

Developer Notes#

The callback will contain detailed note data for each detected instrument.
Each note includes: pitch (MIDI note number), start (seconds), end (seconds), velocity (0-1).
Not all instruments may be detected β€” depends on audio content.
Pricing: Check current per-call credit costs at https://emix.ai/pricing.

Request

Authorization
Bearer Token
Provide your bearer token in the
Authorization
header when making requests to protected resources.
Example:
Authorization: Bearer ********************
or
Body Params application/jsonRequired

Examples

Responses

🟒200
application/json
MIDI generation task created successfully
Bodyapplication/json

πŸ”΄500Error
Request Request Example
Shell
JavaScript
Java
Swift
curl --location 'https://api.emix.ai/api/v1/midi/generate' \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '{
  "model": "suno/midi/generate",
  "callBackUrl": "https://example.callback",
  "input": {
    "taskId": "5c79****be8e",
    "audioId": "8ca376e7-******-08aaf2c6dd27"
  }
}'
Response Response Example
200 - ζˆεŠŸη€ΊδΎ‹
{
    "code": 200,
    "msg": "success",
    "data": {
        "taskId": "5c79****be8e"
    }
}
Previous
Vocalβ€―&β€―Instrument Stem Separation
Next
Music Video Generation Callbacks
Built with