Three steps, and the second one is where everyone's time goes.
Every homemade karaoke goes through the same three steps, and they are not equally hard:
Forget the old trick of inverting one channel. It only works when the vocal sits dead centre and nothing else does, which is basically never true of a modern mix. The result sounds muffled because along with the vocal you cancelled half of what lives in the centre: kick, snare and bass.
The free standard today is Ultimate Vocal Remover (UVR), which runs on your own machine.
The limit nobody gets around: separation cannot add quality that isn't in the source. Bad input gives you a bad instrumental plus separation artefacts. A live recording, with crowd noise and room reverb, is the worst case there is.
There's a gap between "lyrics on screen" and "timed lyrics". The first is a block of text; the second highlights the right syllable at the right moment, which is what lets somebody sing a song they don't know by heart.
By hand, the classic tool is Karaoke Builder Studio (Windows, buy once). Automatically, a transcription system listens to the isolated vocal and produces the timings. You gain time and lose accuracy: screamed, distorted or heavily processed vocals break any automatic transcription.
If it's just you singing at a computer, a karaoke format is enough. If it has to play on a karaoke machine or a TV, you need CDG or a video with the words burned into the picture.
By hand, budget about an hour per song once you have the hang of it. Automated, the bottleneck becomes processing — separating and transcribing costs machine time, and taking well over the length of the song is normal.
Technically yes, as long as you have the audio file. In practice the result depends on the recording: a clean studio mix comes out well, a live recording with an audience comes out with artefacts, and screamed or heavily distorted vocals break the automatic lyric transcription.
Ultimate Vocal Remover (UVR) is the free standard and runs on your own machine. Avoid Audacity's phase-inversion method: it only works when the vocal is dead centre in the mix and it leaves the instrumental sounding muffled.
By hand, about an hour per song once you know the workflow — timing the lyrics eats nearly all of it. Automated, the time becomes machine processing instead.
StarSinger does those three steps for you. You upload a file you already have, it separates the vocals, transcribes the lyrics with per-syllable timing and gives you a karaoke that plays in the browser, with pitch scoring. Free right now.
The caveats, because they're real: the transcription gets words wrong on screamed or heavily distorted vocals, and processing takes well over the length of the song.
Try it free