# VibeVoice ASR Streaming 7B

URL: https://interfaze.ai/models/microsoftvibevoice-asr-streaming-7b

[All models](https://interfaze.ai/models)

VibeVoice ASR Streaming 7B by microsoft, a automatic-speech-recognition model with speech-to-text, multimodal capabilities. Understand and compare speech-to-text, multimodal features, benchmarks, and capabilities.

## Comparison

| Feature | VibeVoice ASR Streaming 7B | Interfaze |
| --- | --- | --- |
| Input Modalities | text, audio | image, text, audio, video, document |
| Native OCR | No | Yes |
| Long Document Processing | No | Yes |
| Language Support | 10 partial | 162+ |
| Native Speech-to-Text | Yes | Yes |
| Native Object Detection | No | Yes |
| Guardrail Controls | No | Yes |
| Context Input Size | 65.5K | 1M |
| Tool Calling | No | Tool calling supported + built in browser, code execution and web search |

### Speech-to-Text Capabilities

| Feature | VibeVoice ASR Streaming 7B | Interfaze |
| --- | --- | --- |
| Time Stamping | No | Yes |
| Speaker Diarization | Yes | Yes |
| Long Audio Processing | Partial | Yes |
| Audio Processing Speed | 1hr of audio more than 5mins (varies depending on provider and quality output) | 1hr of audio under 30 seconds |
| Intent Recognition | No | Yes |
| Sentiment Analysis | No | Yes |
| Prompting Correction | Yes | Yes |

### Scaling

| Feature | VibeVoice ASR Streaming 7B | Interfaze |
| --- | --- | --- |
| Scaling | Self-hosted/Provider-hosted with quantization | Unlimited |

[Try Interfaze](https://interfaze.ai/dashboard)[Read the Docs](https://interfaze.ai/docs)

View model card on [Hugging Face](https://huggingface.co/microsoft/VibeVoice-ASR-Streaming-7B)

## VibeVoice-ASR-Streaming-7B

[![GitHub](https://img.shields.io/badge/GitHub-Repo-black?logo=github)](https://github.com/microsoft/VibeVoice) [![Live Playground](https://img.shields.io/badge/Live-Playground-green?logo=gradio)](https://aka.ms/vibeasr)

**VibeVoice-ASR-Streaming** is a unified streaming ASR model that transcribes **Who (Speaker)** said **What (Content)**, with support for **Customized Hotwords** and **10 languages**.

➡️ **Code:** [microsoft/VibeVoice](https://github.com/microsoft/VibeVoice) ➡️ **Demo:** [VibeVoice-ASR-Streaming](https://aka.ms/vibeasr)

## 🔥 Key Features

-   **📝 Streaming Speaker-Attributed Transcription**: Continuously transcribes **who** said **what** as speech arrives.
    
-   **👤 Customized Hotwords**: Users can provide customized hotwords, such as names and technical terms, to improve recognition of domain-specific content.
    
-   **🌍 Multilingual Support**: It supports Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish.
    

## Technical Report

📄 [VibeVoice-ASR-Streaming Technical Report](https://arxiv.org/abs/2609.02812)

## Evaluation

## Installation and Usage

Please refer to the [GitHub repository](https://github.com/microsoft/VibeVoice).

## License

This project is licensed under the MIT License.

## Contact

This project was conducted by members of Microsoft Research. We welcome feedback and collaboration from our audience. If you have suggestions, questions, or observe unexpected/offensive behavior in our technology, please contact us at [VibeVoice@microsoft.com](mailto:VibeVoice@microsoft.com). If the team receives reports of undesired behavior or identifies issues independently, we will update this repository with appropriate mitigations.

## Want more deterministic results?

[Try Interfaze](https://interfaze.ai/dashboard)[Read the Docs](https://interfaze.ai/docs)
