Skip to content
D-CSIL

Research Domain

Creative & Audio Intelligence

Conditional drum generation, 6-stem music source separation on MLX, XTTS voice cloning, and edge TTS — machine learning applied to sound.

OBSERVEDCURRENT EVIDENCE LEVEL

Research Question

How far can locally trained models go in generating and deconstructing music and voice?

The audio lab spans DrumAI-Conditional (conditional drum synthesis), BSRoFormer running 6-stem separation natively on Apple Silicon, StableAudio MLX generation, and Demucs-based DrumSep experiments.

Voice work includes XTTS fine-tuning with three cloned voice profiles built from recorded clips, and the Piper ONNX pipeline for fast TTS on Raspberry Pi-class hardware.