COT-TTS demo gallery

COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning

Recent text-to-speech systems support increasingly controllable and expressive generation, but typically require users to specify how each sentence should be spoken. In realistic dialogue, however, the appropriate speaking manner should be inferred from the preceding context.

Recent text-to-speech (TTS) systems support increasingly controllable and expressive generation, but typically require users to specify how each sentence should be spoken. In realistic dialogue, however, the appropriate speaking manner should be inferred from the preceding context. We introduce COT-TTS, a context-aware reasoning text-to-speech task. Given historical dialogue audio, target text, and reference speech, a system is required to infer the intended speaking manner, generate explicit intermediate reasoning, and synthesize contextually appropriate speech with the specified speaker timbre. Starting from approximately 100K hours of Chinese and English dialogue audio, we construct 9M training samples, including a 1M high-quality subset, together with a source-disjoint benchmark of 800 human-validated samples.

Model

End-to-end autoregressive COT-TTS

We develop 0.6B- and 1.7B-parameter end-to-end autoregressive models that sequentially generate emotion-tagged transcripts, speaking-manner reasoning, and speech tokens. The intermediate reasoning can be inspected and edited before speech synthesis.

End-to-end COT-TTS autoregressive model architecture

Baseline

Task-specific cascaded reference systems

We construct task-specific baselines based on ASR--LLM--TTS, AudioLLM--TTS, and AudioLLM--VC pipelines for side-by-side comparison against the end-to-end models. These cascaded systems highlight how much explicit reasoning and controllable speech generation can be preserved when the task is solved through modular components instead of a unified model.

COT-TTS cascaded baseline architectures

Open-source progress

Current open resources

The public release includes the challenge materials, datasets, baseline resources, model code, training code, and data construction pipeline.

Editable control demos

Controlling duration, rhythm, and expressive intensity

By editing the intermediate CoT text, we can directly steer total duration, speech rhythm, and expressive intensity while keeping the same dialogue context and target text.

Default Control variable Setting 1 Setting 2 Setting 3 Setting 4

Emotion expressiveness

Expressive variation across dialogue contexts

These examples highlight how the system renders different emotions under distinct dialogue situations, with matched English and Chinese cases for each category.

Emotion category Historical audio Target text Output audio

Demo Table

Showing 0 samples

Each row is one COT-TTS case. The model column contains all available systems.

Historical Audio Target Text Model Name