The next test is a voice, not a face. To build the dataset I have to perform six songs into a microphone, badly, on purpose. Here is the protocol before I run it, including the one mistake that would poison everything.
Every AI pipeline I have built this year came down to the same thing: you are not prompting, you are casting. Ghost Casting proved it with a face, one locked identity held across every angle. The film I cut this spring proved it with a design language, two hundred generations that had to read as one world. So the obvious next question is the one that makes me laugh a little: what happens when the identity is my own voice?
I make music under another name. Which means for once I am not casting a stranger, I am casting me, and the training data has to come from somewhere. Voice models learn from performance, not from reading. So the dataset for this experiment is me doing karaoke alone in a room, on tape, into a good microphone, in registers I have no business attempting.
Six blocks, roughly five minutes each, songs I know cold. Not a script. Scripts make stiff reads and stiff reads make stiff models, the same way a bored actor makes a bored take. The point is spread: a model is worst in the registers it has never heard, so the session deliberately runs from conversational to shouted, laid back to full throated.
Here is the part that surprises people who have only ever recorded podcasts. My interface has beautiful onboard processing, an exciter, a compressor, bass enhancement, presets that make a voice sound expensive instantly. All of it gets turned off.
A model learns whatever is in the signal. If the compressor is on, the compressor becomes part of my voice forever, baked into every future generation whether I want it or not. Flat EQ, no gate, no enhancement, peaks landing between minus twelve and minus six, forty eight kilohertz, twenty four bit, mono, zero plugins on the track. Sound design is a decision you make later. A dataset is a decision you cannot unmake.
If the reference track leaks from headphones into the microphone, the beat becomes part of my voice permanently. Closed backs only, nothing ever from speakers, reference low in the cans, at least two blocks performed a cappella. Record ten seconds of room first and listen back. If you can hear the room, fix the room before you sing a note.
That is the same discipline as a locked design language, or a must-not-change list on a motion prompt. Constraint up front is what makes the volume directable later. Every one of these pipelines has the same shape once you have run a few: decide what cannot vary, then generate freely inside the fence.
I expect the first pass to be uncanny. Close enough to recognize, wrong enough to be unsettling, the audio version of a face that drifts between takes. I expect the sung registers to fail before the spoken ones. I expect at least one block to be unusable because I got excited and clipped it, and I expect to keep that fail in the write up, because field notes that only show wins are advertising.
And the gate at the end is not a vibe. Same as casting: it either sounds like me across registers or it does not, and I would rather kill the model than ship one that almost works. What makes it fun is that this time the person being cast has to sing for it.
Toolchain: a clean multitrack capture into a DAW, ElevenLabs for the voice model, with cleanup and speech tooling around it. Part two will have the audio, the similarity read, and the outtakes.