Audio Synthesizer Inversion Demo

DDSynth-RL

DDSynth-RL treats synthesizer inversion as conditional generation over discrete synthesizer and MIDI tokens. A masked discrete diffusion model proposes parameter sequences, then GRPO fine-tunes the generator using rewards computed from rendered audio. This demo compares target audio with autoregressive, flow matching, discrete diffusion, and reinforcement-learning fine-tuned variants on both in-domain Dexed and out-of-domain NSynth examples.

Listening Table

In-Domain and OOD Samples