A/B listening for the TD-PSOLA pitch engine. Recorded voice, sung baritone and alto, speech. Pitch shifts and formant shifts, side by side against the source.
An opera line, D4 to D5, heavy vibrato. Pitch and formant shifts compared against the source.
Real sung takes through the engine, pitch moved and formants held.
Male and female speech, shifted up and down.
The implementation you just heard, as source. Time-domain PSOLA with the pitch and the vocal tract on separate controls, running inside the audio callback. Plain C++ with the JUCE layer kept separate.