Over a year ago, I posted a video where Gemini and I collaborated on an Ableton Live Session. Music AI is fun! Creating voice-native music software is something I'm really passionate about, so, here is Jamcat.
Jamcat's concept is a voice controlled, generative music performance tool, designed to inspire live, in the moment jam sessions. It uses LLMs, with both text-to-audio and audio-to-audio models to create music in realtime, entirely hands-free.
âĄī¸ Pipecat PhoneLLM Alpha 1 for session control, running on
@modal.
âĄī¸ "StemGen" routing and workload subagent.
âĄī¸ Foundation 1, for melodic stems.
âĄī¸ Various audio models on
@fal (like the excellent
@ElevenLabs Music) for vocals, drums, textures and one-shots.
âĄī¸ ... and of course
@pipecat_ai!
Voice commands can be toggled via push-to-talk (or a foot pedal, if you're wielding an axe), or it can listen the entire time (diarization.)
It makes some mistakes, or sometimes misses a beat, but most of the mistakes in this video were me not paying attention to the bar counts, like a true fake musician.