# Confucius4 R2T2

URL: https://interfaze.ai/models/netease-youdaoconfucius4-r2t2

[All models](https://interfaze.ai/models)

Confucius4 R2T2 by netease-youdao, a automatic-speech-recognition model with speech-to-text, multimodal capabilities. Understand and compare speech-to-text, multimodal features, benchmarks, and capabilities.

## Comparison

| Feature | Confucius4 R2T2 | Interfaze |
| --- | --- | --- |
| Input Modalities | audio, text | image, text, audio, video, document |
| Native OCR | No | Yes |
| Long Document Processing | No | Yes |
| Language Support | 11 partial | 162+ |
| Native Speech-to-Text | Yes | Yes |
| Native Object Detection | No | Yes |
| Guardrail Controls | No | Yes |
| Context Input Size | unknown | 1M |
| Tool Calling | No | Tool calling supported + built in browser, code execution and web search |

### Speech-to-Text Capabilities

| Feature | Confucius4 R2T2 | Interfaze |
| --- | --- | --- |
| Time Stamping | No | Yes |
| Speaker Diarization | No | Yes |
| Long Audio Processing | Partial | Yes |
| Audio Processing Speed | 1hr of audio more than 5mins (varies depending on provider and quality output) | 1hr of audio under 30 seconds |
| Intent Recognition | No | Yes |
| Sentiment Analysis | No | Yes |
| Prompting Correction | Yes | Yes |

### Scaling

| Feature | Confucius4 R2T2 | Interfaze |
| --- | --- | --- |
| Scaling | Self-hosted/Provider-hosted with quantization | Unlimited |

[Try Interfaze](https://interfaze.ai/dashboard)[Read the Docs](https://interfaze.ai/docs)

View model card on [Hugging Face](https://huggingface.co/netease-youdao/Confucius4-R2T2)

Confucius4-R2T2 is a low-latency and high-accuracy true streaming Automatic Speech Recognition (ASR) model that features fine-grained and configurable decoding chunks from 80 ms to 2 s. The model operates in append-only output mode: committing transcript text permanently without revising previous words, which is critical for applications where text must be processed or acted upon instantly. This results in a smoother user experience, avoiding disruptive text revisions and visual flickering in real-time applications, such as Real-Time Live Captioning & Subtitling, Downstream NLP Pipelines & LLM Agents, Simultaneous Speech Translation, etc.

R2T2, short for Real Real-Time Transcription, is built upon the Qwen3-ASR model. And it is trained with a unique set of data construction techniques including stable-prefix data, forced time-alignment data, and token-level audio segmentation. Combined with a Longest Stable Prefix (LSP) learning paradigm (tech report will be released soon), R2T2 can dynamically determine when a stable prefix can be safely emitted and when additional audio context is needed. By exposing only stable prefixes, the model provides high-quality context that conditions subsequent predictions while guaranteeing that previously emitted text remains unchanged. Despite its streaming design, R2T2 maintains strong accuracy in offline recognition.

-   **Low-latency and high accuracy streaming recognition** — The model achieves accuracy close to that of offline recognition, with only 200 to 600 milliseconds average latency.
-   **Stable streaming output** — Emitted text is committed as it arrives and remains unchanged.
-   **Configurable low-latency chunking** - Supports decoding chunks from 80 ms to 2 s for different latency/accuracy trade-offs.
-   **No loss in offline accuracy** — Adding streaming support does not degrade offline recognition accuracy.
-   **vLLM backend** — Provides high-throughput inference. A Hugging Face `transformers` backend is also available.
-   **Context and hotword prompts** — Natively supported.
-   **Multilingual support** — Optimized for **Chinese and English**, while also supporting a broad range of additional languages.

Experimental results show that R2T2 achieves state-of-the-art (SOTA) performance in both latency and recognition quality among a range of open-source models, while remaining competitive with leading closed-source systems. The [GitHub repository](https://github.com/netease-youdao/Confucius4-R2T2) provides inference code, a minimal usage example, and a vLLM-based backend supporting both offline and real-time streaming inference.

## Table of Contents

-   [Overview](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#overview)
-   [Demo](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#demo)
    -   [Side-by-side comparison with GPT-Live-Transcribe](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#side-by-side-comparison-with-gpt-live-transcribe)
    -   [Additional resources](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#additional-resources)
-   [Evaluation](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#evaluation)
    -   [Streaming performance](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#streaming-performance)
    -   [Accuracy](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#accuracy)
        -   [English](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#english)
        -   [Chinese](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#chinese)
-   [Installation](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#installation)
    -   [Clone the repository](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#clone-the-repository)
    -   [Option 1: Conda](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#option-1-conda)
    -   [Option 2: uv](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#option-2-uv)
-   [Docker (recommended)](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#docker-recommended)
    -   [1\. Start a container](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#1-start-a-container)
    -   [2\. Run the example inside the container](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#2-run-the-example-inside-the-container)
    -   [3\. Manage the container](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#3-manage-the-container)
-   [Quick Start](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#quick-start)
    -   [Configuration](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#configuration)
-   [Python API](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#python-api)
    -   [Offline transcription (vLLM backend)](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#offline-transcription-vllm-backend)
    -   [Streaming transcription (vLLM backend)](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#streaming-transcription-vllm-backend)
-   [WebSocket Server](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#websocket-server)
    -   [Start and stop the server](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#start-and-stop-the-server)
    -   [WebSocket endpoint](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#websocket-endpoint)
    -   [Message format](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#message-format)
    -   [Example client](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#example-client)
-   [Supported Languages](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#supported-languages)
-   [Community & Contact](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#community--contact)
    -   [WeChat Group](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#wechat-group)
    -   [Discord Server](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#discord-server)
    -   [Business contact](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#business-contact)
    -   [GitHub Issues](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#github-issues)
-   [Acknowledgements](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#acknowledgements)
-   [Citation](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#citation)
-   [License](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#license)

* * *

## Overview

## Demo

### Side-by-side comparison with GPT-Live-Transcribe

### Additional resources

More demonstrations, comparisons, and supporting resources will be added here.

## Evaluation

> If you are an author or maintainer of a model included in these comparisons and have questions or concerns about the results, please feel free to contact us through the [GitHub issue tracker](https://github.com/netease-youdao/Confucius4-R2T2/issues). We are happy to share evaluation details and work with you to verify or correct them.

### Streaming performance

The streaming API supports decoding chunks from 80 ms to 2 s; the figures below show representative WER/latency trade-offs at 160 ms.

### Accuracy

English results use WER (%), and Chinese results use CER (%); lower is better.

※ Pseudo-streaming model: its partial transcript may revise previously emitted text; unmarked models use true streaming, append-only output.

#### English

#### Chinese

## Installation

We recommend using a **fresh, isolated environment**. For local development and source installation, use the **Conda** or **uv** environment below. **Docker** is recommended for quickly running the project with a preconfigured CUDA and runtime environment — see [Docker](https://interfaze.ai/models/netease-youdaoconfucius4-r2t2#docker-recommended).

### Clone the repository

```
git clone https://github.com/netease-youdao/Confucius4-R2T2.git
cd Confucius4-R2T2
```

### Option 1: Conda

```
conda create -n confucius4-r2t2 python=3.12 -y
conda activate confucius4-r2t2


pip install -e .
```

### Option 2: uv

```
uv venv --python 3.12
source .venv/bin/activate


uv pip install -e .
```

Python 3.10+ is supported. Python 3.12 is the version we test against.

vLLM has strict CUDA / PyTorch compatibility requirements. If the install fails to resolve, check the version matrix on the [vLLM website](https://docs.vllm.ai/) and pin a combination that matches your CUDA runtime.

## Docker (recommended)

R2T2 runs out of the box on the official **Qwen3-ASR** Docker image, which already ships every runtime library we need.

Pre-built image: [qwenllm/qwen3-asr](https://hub.docker.com/r/qwenllm/qwen3-asr).

Before you begin, install the [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html) to enable GPU access from Docker. If Docker Hub access is slow or unreliable in your region, you may need to configure a registry mirror.

### 1\. Start a container

```
LOCAL_WORKDIR=/path/to/your/workspace   # host path that will be mounted into the container
HOST_PORT=8000
CONTAINER_PORT=80

docker run --gpus all --name confucius4-r2t2 \
    -v /var/run/docker.sock:/var/run/docker.sock \
    -p $HOST_PORT:$CONTAINER_PORT \
    --mount type=bind,source=$LOCAL_WORKDIR,target=/data/shared/confucius4-r2t2 \
    --shm-size=4gb \
    -it qwenllm/qwen3-asr:latest
```

Your local workspace (`$LOCAL_WORKDIR`) — including a checkout of this repository and the R2T2 checkpoint — will be mounted inside the container at `/data/shared/confucius4-r2t2`. Host port `8000` is mapped to container port `80`; services running inside the container must bind to `0.0.0.0` (not `127.0.0.1`) for port forwarding to work.

### 2\. Run the example inside the container

Once inside the container's shell:

```
cd /data/shared/confucius4-r2t2/Confucius4-R2T2
MODEL_PATH=/data/shared/confucius4-r2t2/Confucius4-R2T2 \
    ./run_example.sh /path/to/audio.wav
```

### 3\. Manage the container

```
docker start confucius4-r2t2
docker exec -it confucius4-r2t2 bash


docker rm -f confucius4-r2t2
```

## Quick Start

Grab any audio file (mono or stereo, any sample rate — it is resampled to 16 kHz internally) and run:

```
./run_example.sh /path/to/audio.wav \
    --model_path /path/to/Confucius4-R2T2 \
    --infer_mode stream_vllm \
    --language Chinese \
    --chunk_size_ms 160
```

Logs are written to `run_example.log` by default. Run `./run_example.sh --help` to see the full flag list.

### Configuration

`run_example.sh` reads the following environment variables (all optional):

| Variable | Default | Description |
| --- | --- | --- |
| MODEL\_PATH | (required) | Path or HF repo id of the R2T2 checkpoint |
| AUDIO | first CLI argument | Path to the input audio file |
| INFER\_MODE | stream\_vllm | stream\_vllm or onetime\_vllm |
| LANGUAGE | Chinese | Language hint (e.g. Chinese, English, …) |
| CHUNK\_SIZE\_MS | 160 | Streaming chunk size (80 ms–2 s supported) |
| UNFIXED\_TOKEN\_NUM | 1 | Number of unfixed trailing tokens (rollback window) |
| CONTEXT | "" | Context / hotword hint prepended to the prompt |
| CUDA\_VISIBLE\_DEVICES | 0 | GPU id(s) to expose |
| LOG\_FILE | run\_example.log | Where to write logs |

You can also call `example.py` directly and pass any of these as flags (`--audio`, `--model_path`, `--infer_mode`, `--language`, `--chunk_size_ms`, `--lookahead_ms`, `--unfixed_token_num`, `--context`).

## Python API

Audio inputs can be passed as a local path, a URL, base64 data, or a `(np.ndarray, sr)` tuple. Batched inference is supported. Remember to wrap vLLM code under `if __name__ == '__main__':` to avoid the `spawn` error described in [vLLM Troubleshooting](https://docs.vllm.ai/en/latest/usage/troubleshooting/#python-multiprocessing).

### Offline transcription (vLLM backend)

```
import librosa
from qwen_asr import Qwen3ASRModel

if __name__ == "__main__":
    asr = Qwen3ASRModel.LLM(
        model="/path/to/Confucius4-R2T2",
        gpu_memory_utilization=0.5,
        max_inference_batch_size=32,
        max_new_tokens=4096,
    )

    wav, sr = librosa.load("path/to/audio.wav", sr=16000, mono=True)

    results = asr.transcribe(
        audio=[(wav, 16000)],
        language=["Chinese"],       # or [None]
        return_time_stamps=False,
    )
    print(results[0].language, results[0].text)
```

### Streaming transcription (vLLM backend)

```
import librosa
from qwen_asr import Qwen3ASRModel

if __name__ == "__main__":
    asr = Qwen3ASRModel.LLM(
        model="/path/to/Confucius4-R2T2",
        gpu_memory_utilization=0.4,
        max_new_tokens=4,           # keep small for low-latency streaming
    )

    wav, sr = librosa.load("path/to/audio.wav", sr=16000, mono=True)

    state = asr.init_streaming_state(
        context="",                 # optional hotword / topic hint
        language="Chinese",         # or None
        unfixed_chunk_num=0,
        unfixed_token_num=1,
        chunk_size_sec=0.16,
    )

    step = int(0.16 * 16000)
    for pos in range(0, len(wav), step):
        seg = wav[pos : pos + step]
        _, text = asr.streaming_transcribe(seg, state, max_new_tokens=2)
        print("text:", text)

    asr.finish_streaming_transcribe(state)
    print("final:", state.text)
```

For a complete streaming example with adaptive `max_new_tokens` and initial-chunk lookahead handling, see [`example.py`](https://github.com/netease-youdao/Confucius4-R2T2/blob/master/example.py).

## WebSocket Server

For real-time, multi-client streaming ASR, the [GitHub repository](https://github.com/netease-youdao/Confucius4-R2T2) ships a ready-to-run WebSocket server (`ws_server.py`), a launcher script (`run_start_server.sh`), and a reference Python client (`ws_client.py`).

### Start and stop the server

```
./run_start_server.sh start \
    --model_path /path/to/Confucius4-R2T2 \
    --vad_model_path /path/to/Stream-VAD \
    --port 8272 \
    --gpu 0


./run_start_server.sh kill


./run_start_server.sh restart \
    --model_path /path/to/Confucius4-R2T2 \
    --vad_model_path /path/to/Stream-VAD \
    --port 8272 \
    --gpu 0
```

| Flag | Env var | Default | Description |
| --- | --- | --- | --- |
| \-m, --model\_path | ASR\_MODEL\_PATH | (required) | Path or HF repo id of the R2T2 checkpoint |
| \-v,--vad\_model\_path | VAD\_MODEL\_PATH | checkpoints/vad/Stream-VAD | Path to the FireRedVAD Stream-VAD model |
| \-p, --port | PORT | 8272 | Port the WebSocket server binds to |
| \-g, --gpu | CUDA\_VISIBLE\_DEVICES | 0 | GPU id(s) exposed to the server process |
| \-h, --host | HOST\_TAG | localhost | Host tag used only in the log file name |

The launcher resolves its own directory, so it can be invoked from anywhere. Logs are written to `nohup_service_ws_<host_tag>_<port>.log` in the current directory. The FireRedVAD model is available from [Hugging Face](https://huggingface.co/FireRedTeam/FireRedVAD/tree/main). We recommend downloading the model files into this repository's `checkpoints` directory:

```
hf download FireRedTeam/FireRedVAD \
    --include "Stream-VAD/*" \
    --local-dir checkpoints/vad


git clone https://huggingface.co/FireRedTeam/FireRedVAD
cp -r FireRedVAD/Stream-VAD checkpoints/vad/
```

Either command leaves the model at `checkpoints/vad/Stream-VAD`, which is exactly what `--vad_model_path` defaults to — so you can drop the flag entirely.

### WebSocket endpoint

| Path | Behavior |
| --- | --- |
| /asr\_stream\_api\_v1 | Streaming ASR. Each message's text is the new (incremental) chunk. |

### Message format

**Client → Server:**

-   Send raw 16 kHz mono PCM as `int16` binary frames (the reference client uses ≈160 ms per frame, i.e. 2560 samples × 2 bytes).
-   Send the string `"YOUDAO_ONETIME_ASR_STREAM_EOS"` to signal end-of-audio; the server will emit any final text and close.

**Server → Client:** JSON messages of the form

```
{
  "status": "success",
  "requestId": "<uuid>",
  "msg": {
    "text": "hello",
    "reset": false,
    "asr_cost_ms": 35.4,
    "total_cost_ms": 42.0
  }
}
```

-   `text` is the newly recognized (incremental) segment since the previous message. Concatenate them client-side to get the full transcript.

### Example client

`ws_client.py` is a minimal example that streams a WAV file to the server and prints the responses.

```
python ws_client.py


python ws_client.py \
    --uri wss://your.host/asr_stream_api_v1 \
    --audio resources/test.wav \
    --save service_ws_test \
    --audio-id test.wav
```

Command-line options:

| Flag | Env var | Default | Description |
| --- | --- | --- | --- |
| \--uri / -u | ASR\_WS\_URI | ws://localhost:8272/asr\_stream\_api\_v1 | WebSocket endpoint to connect to. |
| \--audio / -a | — | built-in sample path | Input audio file (WAV, 16 kHz mono recommended). |
| \--save / -s | — | service\_ws\_test | File to append the final transcript to. |
| \--audio-id | — | basename of --audio | Identifier written next to the result in --save. |

## Supported Languages

R2T2 is optimized for streaming recognition in Chinese and English. Beyond these primary languages, it retains useful cross-lingual streaming capability on languages such as French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Arabic, etc.

## Community & Contact

Join our community to ask questions, share ideas, and connect with other users and developers.

### WeChat Group

Scan the QR code below to join our WeChat group:

### Discord Server

[Join our Discord server](https://discord.gg/GfhaWkCyb)

### Business contact

For high-concurrency, production-grade, domestically deployable, or private deployment solutions, as well as business inquiries and partnership opportunities, please feel free to contact us through the channels below.

-   **Phone:** +86 010-82558901
-   **Email:** [AIcloud\_Business@corp.youdao.com](mailto:AIcloud_Business@corp.youdao.com)

### GitHub Issues

We also welcome discussions in this repository’s [Issues](https://github.com/netease-youdao/Confucius4-R2T2/issues) section. Feel free to ask questions, report bugs, or suggest improvements!

* * *

## Acknowledgements

We sincerely thank the Alibaba Qwen team for open-sourcing the [Qwen3-ASR](https://github.com/QwenLM/Qwen3-ASR) modeling code, which provides the architectural foundation for R2T2.

## Citation

If you use this repository or the R2T2 checkpoint in your research, please cite **Confucius4-R2T2** (this project):

```
@misc{Confucius4-R2T2,
  title        = {Confucius4-R2T2: A Low Latency and High Accuracy Real-Time Speech Recognition Model},
  author       = {NetEase Youdao},
  year         = {2026},
  howpublished = {https://github.com/netease-youdao/Confucius4-R2T2}
}
```

## License

R2T2 uses **dual licensing** to distinguish the source code from the model weights:

-   **Code** in the accompanying GitHub repository is released under the [Apache License 2.0](https://github.com/netease-youdao/Confucius4-R2T2/blob/master/LICENSE) and is free to use, modify, and redistribute (including commercially) under the terms of that license.
-   **Model weights** are released under the [NetEase Model Use License Agreement](https://github.com/netease-youdao/Confucius4-R2T2/blob/master/MODEL_LICENSE).

## Want more deterministic results?

[Try Interfaze](https://interfaze.ai/dashboard)[Read the Docs](https://interfaze.ai/docs)
