// ENTRY 004 / 08.14.26 / AI PIPELINE · VOICE

The Voice Lab: Cloning Myself by Singing Karaoke

The next test is a voice, not a face. To build the dataset I have to perform six songs into a microphone, badly, on purpose. Here is the protocol before I run it, including the one mistake that would poison everything.

Every AI pipeline I have built this year came down to the same thing: you are not prompting, you are casting. Ghost Casting proved it with a face, one locked identity held across every angle. The film I cut this spring proved it with a design language, two hundred generations that had to read as one world. So the obvious next question is the one that makes me laugh a little: what happens when the identity is my own voice?

I make music under another name. Which means for once I am not casting a stranger, I am casting me, and the training data has to come from somewhere. Voice models learn from performance, not from reading. So the dataset for this experiment is me doing karaoke alone in a room, on tape, into a good microphone, in registers I have no business attempting.

The dataset is a performance

Six blocks, roughly five minutes each, songs I know cold. Not a script. Scripts make stiff reads and stiff reads make stiff models, the same way a bored actor makes a bored take. The point is spread: a model is worst in the registers it has never heard, so the session deliberately runs from conversational to shouted, laid back to full throated.

// THE SIX BLOCKS · VOICE PLATE v1
01Big range. Conversational up to shouty, wide dynamics, as many registers as one song can hold.
02Melodic, laid back. The sung-rap placement. Relaxed, tonal, in pocket.
03Full throated singing. Power, sustained notes, chest voice. Actual singing, which is where the ego goes to die.
04Soul and softness. Quiet, intimate, close to the mic.
05Rhythmic and percussive. Tight consonants, breath control, the pocket itself.
06A cappella from memory. No reference track at all. The cleanest data in the set.

The rig setting that matters more than the microphone

Here is the part that surprises people who have only ever recorded podcasts. My interface has beautiful onboard processing, an exciter, a compressor, bass enhancement, presets that make a voice sound expensive instantly. All of it gets turned off.

A model learns whatever is in the signal. If the compressor is on, the compressor becomes part of my voice forever, baked into every future generation whether I want it or not. Flat EQ, no gate, no enhancement, peaks landing between minus twelve and minus six, forty eight kilohertz, twenty four bit, mono, zero plugins on the track. Sound design is a decision you make later. A dataset is a decision you cannot unmake.

// THE ONE RULE ABOVE ALL · ZERO BLEED

If the reference track leaks from headphones into the microphone, the beat becomes part of my voice permanently. Closed backs only, nothing ever from speakers, reference low in the cans, at least two blocks performed a cappella. Record ten seconds of room first and listen back. If you can hear the room, fix the room before you sing a note.

That is the same discipline as a locked design language, or a must-not-change list on a motion prompt. Constraint up front is what makes the volume directable later. Every one of these pipelines has the same shape once you have run a few: decide what cannot vary, then generate freely inside the fence.

What I actually expect to happen

I expect the first pass to be uncanny. Close enough to recognize, wrong enough to be unsettling, the audio version of a face that drifts between takes. I expect the sung registers to fail before the spoken ones. I expect at least one block to be unusable because I got excited and clipped it, and I expect to keep that fail in the write up, because field notes that only show wins are advertising.

And the gate at the end is not a vibe. Same as casting: it either sounds like me across registers or it does not, and I would rather kill the model than ship one that almost works. What makes it fun is that this time the person being cast has to sing for it.

Toolchain: a clean multitrack capture into a DAW, ElevenLabs for the voice model, with cleanup and speech tooling around it. Part two will have the audio, the similarity read, and the outtakes.

Entry 004 is the protocol, published before the session runs. Part two follows with results, including the takes that failed. Personal music project, my own voice, no likeness but mine.
← All Field Notes