LaVoco: Autoregressive Zero-Shot Voice Conversion

Audio Samples
Anonymous submission to Interspeech 2026
Abstract: We present LaVoco, an autoregressive approach for zero-shot voice conversion built on a pretrained autoregressive TTS backbone. We explore three source content representations: continuous Whisper encoder embeddings projected into the model's embedding space via a learned linear layer, discrete XCodec2 codec tokens and a dual-input variant that combines both. The dual representation proves most robust, achieving the best content preservation and competitive speaker similarity by leveraging complementary semantic and acoustic information while retaining the scalability of next-token prediction. LaVoco is data-efficient, reaching competitive naturalness within a small number of training steps and outperforming diffusion and flow-matching baselines when scaled further. The framework also proves robust to extremely short reference prompts, maintaining high speaker similarity from less than one second of audio.
Setup: We compare LaVoco against Seed-VC, EZ-VC, FreeVC, DDDM-VC, Diff-HierVC, and KNN-VC. Below we present samples from two evaluation tasks: (1) zero-shot voice conversion, where the timbre of a source utterance is converted to match a target speaker, and (2) back-translation, a round-trip reconstruction test measuring cumulative information loss. Samples are selected from our human evaluation (pairwise listening test) where LaVoco achieved the highest combined win rates for speaker similarity and naturalness.

Zero-Shot Voice Conversion

Given a source utterance and a short reference prompt from the target speaker, each model converts the source speech to match the target speaker's timbre while preserving linguistic content.

Sample 1

Source
Reference (target speaker)

LaVoco (Ours)
Seed-VC
EZ-VC
FreeVC
DDDM-VC
Diff-HierVC
KNN-VC

Sample 2

Source
Reference (target speaker)

LaVoco (Ours)
Seed-VC
EZ-VC
FreeVC
DDDM-VC
Diff-HierVC
KNN-VC

Sample 3

Source
Reference (target speaker)

LaVoco (Ours)
Seed-VC
EZ-VC
FreeVC
DDDM-VC
Diff-HierVC
KNN-VC

Sample 4

Source
Reference (target speaker)

LaVoco (Ours)
Seed-VC
EZ-VC
FreeVC
DDDM-VC
Diff-HierVC
KNN-VC

Sample 5

Source
Reference (target speaker)

LaVoco (Ours)
Seed-VC
EZ-VC
FreeVC
DDDM-VC
Diff-HierVC
KNN-VC

Back-Translation Evaluation

A round-trip reconstruction test: we first convert Speaker B's target utterance into Speaker A's voice, then convert the result back into Speaker B's voice using a separate reference clip. The final output is compared against the original to measure cumulative information loss of a double pass through each model.

Pair 1

Ground Truth
LaVoco (Ours)
Seed-VC
EZ-VC

Pair 2

Ground Truth
LaVoco (Ours)
Seed-VC
EZ-VC

Pair 3

Ground Truth
LaVoco (Ours)
Seed-VC
EZ-VC

Pair 4

Ground Truth
LaVoco (Ours)
Seed-VC
EZ-VC