-
Notifications
You must be signed in to change notification settings - Fork 2
Expand file tree
/
Copy pathindex.html
More file actions
54 lines (54 loc) · 9.93 KB
/
Copy pathindex.html
File metadata and controls
54 lines (54 loc) · 9.93 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8"><meta name="viewport" content="width=device-width, initial-scale=1">
<title>Qwen Audio | Speech, Voice & Audio Research</title>
<meta name="description" content="Explore Qwen-Audio-3.1-ASR and Qwen-Audio-3.1-Realtime: multilingual speech recognition, audio understanding and natural voice interaction. Discover demos and earlier speech and audio projects.">
<meta name="theme-color" content="#ffffff">
<meta property="og:title" content="Qwen Audio | Speech, Voice & Audio Research">
<meta property="og:description" content="Discover Qwen-Audio-3.1-ASR, Qwen-Audio-3.1-Realtime and the wider speech and audio research family, including the 3.0 project pages.">
<meta property="og:url" content="https://qwenaudio.github.io/">
<meta property="og:image" content="https://qwenaudio.github.io/qwen-audio-3.0-asr/assets/qwen-logo.png">
<link rel="canonical" href="https://qwenaudio.github.io/">
<link rel="icon" href="/qwen-audio-3.0-asr/assets/qwen-logo.png">
<link rel="stylesheet" href="/hub/styles.css"><script src="/hub/main.js" defer></script>
</head>
<body>
<a class="skip" href="#main">Skip to content</a>
<header class="shell masthead">
<a class="brand" href="/" aria-label="Qwen Audio home"><img src="/qwen-audio-3.0-asr/assets/qwen-logo.png" width="32" height="32" alt="">Qwen Audio</a>
<nav aria-label="Main navigation"><a href="#featured" data-zh="Qwen-Audio 系列">Qwen-Audio series</a><a href="#projects" data-zh="所有项目">All projects</a><a href="https://github.com/QwenAudio">GitHub</a></nav>
<button id="language" type="button" aria-label="Switch to Chinese">中文</button>
</header>
<main id="main">
<section class="shell intro"><p class="eyebrow" data-zh="阿里巴巴 · 语音与音频研究">Alibaba · Speech & audio research</p><h1>Qwen Audio<span class="title-dot">.</span></h1>
<p class="lead" data-zh="从聆听、理解到表达,探索 Qwen-Audio 系列与我们的语音音频研究项目。">From listening and understanding to speaking and creating. Explore the Qwen-Audio series and our wider family of speech and audio projects.</p></section>
<section class="shell featured" id="featured" aria-labelledby="featured-title">
<div class="section-heading"><h2 id="featured-title" data-zh="聚焦 Qwen-Audio">Explore Qwen-Audio</h2><span data-zh="模型介绍 · 音频演示 · 技术报告">Models · Audio demos · Technical reports</span></div>
<div class="feature-grid">
<article class="feature asr"><div class="feature-copy"><p class="category" data-zh="语音识别与音频理解">Speech recognition & audio understanding</p><h3><a href="/qwen-audio-3.1-asr/">Qwen-Audio-3.1-ASR</a></h3>
<p data-zh="多语言语音识别,Flash 与 Next 支持非流式说话人分离,Next 进一步支持通用音频理解。">Multilingual speech recognition with non-streaming speaker diarization in Flash and Next, plus general audio understanding in Next.</p>
<div class="facts"><span data-zh="30 种语言">30 languages</span><span data-zh="16 种中文方言">16 Chinese dialects</span></div></div>
<a class="preview" href="/qwen-audio-3.1-asr/" aria-label="Explore Qwen-Audio-3.1-ASR"><img src="/qwen-audio-3.1-asr/assets/launch31/overview.svg" alt="Qwen-Audio-3.1-ASR capability overview"></a>
<div class="feature-footer"><a class="primary" href="/qwen-audio-3.1-asr/#speakers" data-zh="探索 ASR 演示 →">Explore ASR demos →</a><a href="/qwen-audio-3.1-asr/#models" data-zh="模型能力 ↗">Model capabilities ↗</a></div></article>
<article class="feature realtime"><div class="feature-copy"><p class="category" data-zh="实时语音交互">Real-time voice interaction</p><h3><a href="/qwen-audio-3.1-realtime/">Qwen-Audio-3.1-Realtime</a></h3>
<p data-zh="探索语音推理、工具调用与自然交互,以及作为独立运行时扩展的持久语音代理。">Explore spoken reasoning, tool use and natural interaction, alongside a persistent voice agent built as a separate runtime extension.</p>
<div class="facts"><span data-zh="语音推理">Spoken reasoning</span><span data-zh="工具调用">Tool use</span></div></div>
<a class="preview" href="/qwen-audio-3.1-realtime/#agent" aria-label="Explore the persistent voice-agent runtime"><img src="/qwen-audio-3.1-realtime/assets/persistent-voice-harness.png" alt="Persistent voice-agent runtime architecture, a separate extension to Qwen-Audio-3.1-Realtime"></a>
<div class="feature-footer"><a class="primary" href="/qwen-audio-3.1-realtime/#demos" data-zh="聆听交互演示 →">Listen to interaction demos →</a><a href="/qwen-audio-3.1-realtime/#agent" data-zh="语音代理扩展 ↗">Voice-agent extension ↗</a></div></article>
</div></section>
<section class="project-section" id="projects" aria-labelledby="projects-title"><div class="shell"><div class="section-heading"><h2 id="projects-title" data-zh="更多语音与音频项目">More speech & audio projects</h2><span data-zh="理解 · 生成 · 交互">Understand · Generate · Interact</span></div><div class="projects">
<article class="project"><p class="category" data-zh="语音识别 · 3.0">Speech recognition · 3.0</p><h3><a href="/qwen-audio-3.0-asr/">Qwen-Audio-3.0-ASR</a></h3><p data-zh="多语言与方言识别,支持流式转写、上下文识别和热词。">Multilingual and dialect recognition with streaming transcription, contextual recognition and hotwords.</p><div class="project-links"><a href="/qwen-audio-3.0-asr/" data-zh="项目演示 →">Project demos →</a></div></article>
<article class="project"><p class="category" data-zh="语音合成 · 3.0">Speech synthesis · 3.0</p><h3><a href="/qwen-audio-3.0-tts/">Qwen-Audio-3.0-TTS</a></h3><p data-zh="可控语音合成、跨语言音色复刻与长文本生成。">Controllable speech synthesis, cross-lingual voice cloning and long-form generation.</p><div class="project-links"><a href="/qwen-audio-3.0-tts/" data-zh="项目演示 →">Project demos →</a></div></article>
<article class="project"><p class="category" data-zh="语音识别">Speech recognition</p><h3><a href="/funasr/">Fun-ASR</a></h3><p data-zh="面向多语言、方言与专业领域的语音识别。">Speech recognition for multilingual, dialect and domain-specific audio.</p><div class="project-links"><a href="/funasr/" data-zh="项目演示 →">Project demos →</a><a href="https://github.com/QwenAudio/Fun-ASR">GitHub ↗</a></div></article>
<article class="project"><p class="category" data-zh="语音合成">Speech synthesis</p><h3><a href="/cosyvoice3/">CosyVoice 3</a></h3><p data-zh="探索新一代多语言语音生成与音色复刻。">Explore the next generation of multilingual speech generation and voice cloning.</p><div class="project-links"><a href="/cosyvoice3/" data-zh="项目演示 →">Project demos →</a><a href="https://github.com/FunAudioLLM/CosyVoice">GitHub ↗</a></div></article>
<article class="project"><p class="category" data-zh="语音对话">Voice conversation</p><h3><a href="/funaudiochat/">Fun-Audio-Chat</a></h3><p data-zh="连接音频理解与自然语音交互。">Connect audio understanding and natural spoken interaction.</p><div class="project-links"><a href="/funaudiochat/" data-zh="项目演示 →">Project demos →</a><a href="https://github.com/FunAudioLLM/Fun-Audio-Chat">GitHub ↗</a></div></article>
<article class="project"><p class="category" data-zh="语音合成">Speech synthesis</p><h3><a href="/cosyvoice2/">CosyVoice 2</a></h3><p data-zh="流式语音合成、零样本音色复刻与多语言生成。">Streaming synthesis, zero-shot voice cloning and multilingual generation.</p><div class="project-links"><a href="/cosyvoice2/" data-zh="项目演示 →">Project demos →</a></div></article>
<article class="project"><p class="category" data-zh="语音交互">Spoken interaction</p><h3><a href="/minmo/">MinMo</a></h3><p data-zh="以多模态大模型探索实时、自然的语音对话。">Multimodal language models for real-time, natural voice conversation.</p><div class="project-links"><a href="/minmo/" data-zh="项目演示 →">Project demos →</a></div></article>
<article class="project"><p class="category" data-zh="音乐生成">Music generation</p><h3><a href="/inspiremusic/">InspireMusic</a></h3><p data-zh="探索音乐生成、续写与多样化的音乐创作样例。">Explore music generation, continuation and music creation examples.</p><div class="project-links"><a href="/inspiremusic/" data-zh="项目演示 →">Project demos →</a><a href="https://github.com/FunAudioLLM/InspireMusic">GitHub ↗</a></div></article>
<article class="project legacy"><p class="category" data-zh="FunAudioLLM · 理解与生成">FunAudioLLM · Understanding & generation</p><h3><a href="/sensevoice-cosyvoice/">SenseVoice & CosyVoice</a></h3><p data-zh="FunAudioLLM 原始项目:多语言语音理解、情感与事件识别、语音生成及语音交互演示。">The original FunAudioLLM project: multilingual speech understanding, emotion and event recognition, voice generation and voice interaction demos.</p><div class="project-links"><a href="/sensevoice-cosyvoice/" data-zh="项目演示 →">Project demos →</a><a href="https://github.com/QwenAudio/SenseVoice">GitHub ↗</a></div></article>
</div></div></section>
<section class="shell ecosystem"><div><h2 data-zh="把语音研究带入应用">Bring speech research into your application.</h2><p data-zh="通过 FunASR 探索推理、服务部署和应用集成。各模型的可用方式与许可请查看对应项目。">Explore inference, serving and application integration with FunASR. See each project for model availability and licensing.</p></div><a class="primary" href="https://www.funasr.com/" data-zh="FunASR 部署与应用 ↗">FunASR deployment & applications ↗</a></section>
</main>
<footer class="shell"><a class="brand" href="/">Qwen Audio</a><p data-zh="阿里巴巴 · 语音与音频研究">Alibaba · Speech & audio research</p><a href="https://github.com/QwenAudio/QwenAudio.github.io" data-zh="网站源码 ↗">Website source ↗</a></footer>
</body></html>