跳转到主要内容
Qwen-Audio-TTS

非实时语音合成Qwen-Audio-TTS Java SDK参考

本文介绍非实时语音合成Qwen-Audio-TTS的Java SDK调用方法,支持非流式和流式两种调用模式。

前提条件

HttpSpeechSynthesizer 类

包路径:com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesizer 功能:基于HTTP的语音合成,支持非流式和流式两种调用方式。

构造方法

public HttpSpeechSynthesizer()
创建HttpSpeechSynthesizer实例,使用默认配置。SDK会自动从环境变量DASHSCOPE_API_KEY或Constants.apiKey获取API Key。

callAndReturnAudio() - 非流式调用(返回音频数据)

方法签名:
public ByteBuffer callAndReturnAudio(HttpSpeechSynthesisParam param) throws ApiException, NoApiKeyException, InputRequiredException
参数说明:
参数类型说明
paramHttpSpeechSynthesisParam语音合成参数对象,包含模型、文本、音色等配置。
返回值:ByteBuffer,包含完整的音频数据。可通过remaining()获取音频大小(字节)。

call() - 非流式调用(返回音频URL)

方法签名:
public HttpSpeechSynthesisResult call(HttpSpeechSynthesisParam param) throws ApiException, NoApiKeyException, InputRequiredException
参数说明:
参数类型说明
paramHttpSpeechSynthesisParam语音合成参数对象。
返回值:HttpSpeechSynthesisResult对象,通过getAudioInfo().getUrl()获取音频下载URL,URL有效期有限,可通过getAudioInfo().getExpiresAt()获取过期时间。

streamCall() - 流式调用

方法签名:
public void streamCall(HttpSpeechSynthesisParam param, ResultCallback<HttpSpeechSynthesisResult> callback) throws ApiException, NoApiKeyException, InputRequiredException
参数说明:
参数类型说明
paramHttpSpeechSynthesisParam语音合成参数对象。
callbackResultCallback&lt;HttpSpeechSynthesisResult&gt;回调对象,需实现onEvent(接收音频分片)、onComplete(合成完成)、onError(错误处理)三个方法。
该方法为异步调用,音频数据通过回调函数分片返回,适用于对首包延迟有要求的场景。 ResultCallback 回调方法: com.alibaba.dashscope.common.ResultCallback是DashScope SDK提供的通用回调接口,需实现以下三个方法:
方法参数说明
onEventHttpSpeechSynthesisResult result每接收到一个音频分片时触发。通过result.hasAudioData()判断是否包含音频数据,通过result.getAudioDataSize()获取分片大小。
onComplete无语音合成完成时触发,表示所有音频分片已接收完毕。
onErrorException e合成过程中发生错误时触发,可通过e.getMessage()获取错误信息。

HttpSpeechSynthesisParam 类

包路径:com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesisParam 通过Builder模式构建参数对象。 部分参数没有专用的Builder方法,需要通过继承自父类的parameter(String key, Object value)方法或parameters(Map<String, Object>)方法进行设置,详见下表中的说明。
方法类型必填说明
model(String)String是语音合成模型。
text(String)String是待合成文本。 支持 SSML 和 LaTeX 格式输入。将待合成文本替换为对应格式即可。
  • 使用 SSML 时,需同时将 enable&#95;ssml 设置为 true。支持的 SSML 标签及用法,请参见SSML 与 LaTeX。
  • 使用 LaTeX 时,将待合成文本替换为 LaTeX 格式即可,无需额外配置。支持的 LaTeX 语法及用法,请参见LaTeX 公式转语音。
voice(String)String是音色。 取值范围:
format(String)String否音频编码格式。 默认值:mp3。 取值范围:
  • mp3
  • pcm
  • wav
  • opus
sampleRate(int)int否音频采样率(Hz)。 取值范围:8000, 16000, 22050(默认), 24000, 44100, 48000。
volume(int)int否音量。 默认值:50。 取值范围:[0, 100]。
rate(float)float否语速。 默认值:1.0。 取值范围:[0.5, 2.0]。
pitch(float)float否音调。 默认值:1.0。 取值范围:[0.5, 2.0]。
enable_ssmlboolean否是否开启SSML功能。设置为true时,text参数需传入SSML格式文本。支持的SSML标签及用法,请参见SSML 与 LaTeX。SSML 的使用限制(支持的模型、音色和接口),请参见使用限制。 默认值:false。 enable_ssml需要通过HttpSpeechSynthesisParam实例的parameter方法或者parameters方法进行设置: 示例 1 请参见表格下方 示例 2 请参见表格下方
word_timestamp_enabledboolean否是否开启字级别时间戳。 默认值:false。 仅在流式输出模式下可用。支持复刻音色;支持的系统音色请参见Qwen-Audio-TTS音色列表。 word_timestamp_enabled需要通过HttpSpeechSynthesisParam实例的parameter方法或者parameters方法进行设置: 示例 3 请参见表格下方 示例 4 请参见表格下方
seedint否生成时使用的随机数种子,使合成的效果产生变化。在模型版本、文本、音色及其他参数均相同的前提下,使用相同的seed可复现相同的合成结果。 默认值0。 取值范围:[0, 65535]。 seed需要通过HttpSpeechSynthesisParam实例的parameter方法或者parameters方法进行设置: 示例 5 请参见表格下方 示例 6 请参见表格下方
language_hintsList否提示: - 此参数为数组,但当前版本仅处理第一个元素,因此建议只传入一个值。
  • 此参数用于指定语音合成的目标语言,该设置与声音复刻时的样本音频的语种无关。如需设置复刻任务的源语言,请参见声音复刻API参考。
指定语音合成的目标语言,提升合成效果。 当数字、缩写、符号等朗读方式或者小语种合成效果不符合预期时使用,例如:
  • 数字朗读方式不符合预期,“hello, this is 110”读成“hello, this is one one zero”而非“hello, this is 幺幺零”
  • 符号朗读不准确,“@”读成“艾特”而非“at”
  • 小语种合成效果差,合成不自然
  • zh:中文
  • en:英语
  • fr:法语
  • de:德语
  • ja:日语
  • ko:韩语
  • ru:俄语
  • pt:葡萄牙语
  • th:泰语
  • id:印尼语
  • vi:越南语
  • es:西班牙语
  • it:意大利语
  • ms:马来西亚语
  • fil:菲律宾语
  • ar:阿拉伯语
language_hints需要通过HttpSpeechSynthesisParam实例的parameter方法或者parameters方法进行设置: 示例 7 请参见表格下方 示例 8 请参见表格下方
instructionString否设置指令,用于控制方言、情感或角色等合成效果。 具体用法请参见非实时语音合成。 instruction需要通过HttpSpeechSynthesisParam实例的parameter方法或者parameters方法进行设置: 示例 9 请参见表格下方 示例 10 请参见表格下方
bit_rateint否音频码率(单位:kbps)。 默认值:32。 取值范围:[6, 510]。 提示: 仅在format为opus时支持使用该参数。 bit_rate需要通过HttpSpeechSynthesisParam实例的parameter方法或者parameters方法进行设置: 示例 11 请参见表格下方 示例 12 请参见表格下方
enable_aigc_tagboolean否是否在生成的音频中添加AIGC隐性标识。设置为true时,会将隐性标识嵌入到支持格式(wav/mp3/opus)的音频中。 默认值:false。 enable_aigc_tag需要通过HttpSpeechSynthesisParam实例的parameter方法或者parameters方法进行设置: 示例 13 请参见表格下方 示例 14 请参见表格下方
aigc_propagatorString否设置AIGC隐性标识中的 ContentPropagator 字段,用于标识内容的传播者。仅在 enable_aigc_tag 为 true 时生效。 默认值:千问AI平台账号。 aigc_propagator需要通过HttpSpeechSynthesisParam实例的parameter方法或者parameters方法进行设置: 示例 15 请参见表格下方 示例 16 请参见表格下方
aigc_propagate_idString否设置AIGC隐性标识中的 PropagateID 字段,用于唯一标识一次具体的传播行为。仅在 enable_aigc_tag 为 true 时生效。 默认值:本次语音合成请求Request ID。 aigc_propagate_id需要通过HttpSpeechSynthesisParam实例的parameter方法或者parameters方法进行设置: 示例 17 请参见表格下方 示例 18 请参见表格下方
hot_fixMap否文本热修复配置,用于自定义指定词语的发音或对待合成文本进行替换。 参数介绍:
  • pronunciation:自定义发音。指定词语的拼音标注,用于纠正默认发音不准确的情况。
  • replace:文本替换。在语音合成前将指定词语替换为目标文本,替换后的文本将作为实际合成内容。
示例: 示例 19 请参见表格下方 hot_fix需要通过HttpSpeechSynthesisParam实例的parameter方法或者parameters方法进行设置: 示例 20 请参见表格下方 示例 21 请参见表格下方
示例 1(说明):
通过parameter设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("<speak>你好</speak>")
    .voice("longanhuan_v3.6")
    .parameter("enable_ssml", true)
    .build();
示例 2(说明):
通过parameters设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("<speak>你好</speak>")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("enable_ssml", true))
    .build();
示例 3(说明):
通过parameter设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameter("word_timestamp_enabled", true)
    .build();
示例 4(说明):
通过parameters设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("word_timestamp_enabled", true))
    .build();
示例 5(说明):
通过parameter设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameter("seed", 1234)
    .build();
示例 6(说明):
通过parameters设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("seed", 1234))
    .build();
示例 7(说明):
通过parameter设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameter("language_hints", Arrays.asList("zh"))
    .build();
示例 8(说明):
通过parameters设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("language_hints", Arrays.asList("zh")))
    .build();
示例 9(说明):
通过parameter设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameter("instruction", "请用非常开心的语气说话。")
    .build();
示例 10(说明):
通过parameters设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("instruction", "请用非常开心的语气说话。"))
    .build();
示例 11(说明):
通过parameter设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .format("opus")
    .parameter("bit_rate", 32)
    .build();
示例 12(说明):
通过parameters设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .format("opus")
    .parameters(Collections.singletonMap("bit_rate", 32))
    .build();
示例 13(说明):
通过parameter设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameter("enable_aigc_tag", true)
    .build();
示例 14(说明):
通过parameters设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("enable_aigc_tag", true))
    .build();
示例 15(说明):
通过parameter设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameter("enable_aigc_tag", true)
    .parameter("aigc_propagator", "xxxx")
    .build();
示例 16(说明):
通过parameters设置
Map<String, Object> map = new HashMap<>();
map.put("enable_aigc_tag", true);
map.put("aigc_propagator", "xxxx");

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameters(map)
    .build();
示例 17(说明):
通过parameter设置
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameter("enable_aigc_tag", true)
    .parameter("aigc_propagate_id", "xxxx")
    .build();
示例 18(说明):
通过parameters设置
Map<String, Object> map = new HashMap<>();
map.put("enable_aigc_tag", true);
map.put("aigc_propagate_id", "xxxx");

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("我家的后面有一个很大的花园。")
    .voice("longanhuan_v3.6")
    .parameters(map)
    .build();
示例 19(说明):
"hot_fix": {
  "pronunciation": [
    {"天气": "tian1 qi4"}
  ],
  "replace": [
    {"今天": "金天"}
  ]
}
示例 20(说明):
通过parameter设置
Map<String, Object> hotFix = new HashMap<>();

List<Map<String, String>> pronunciation = new ArrayList<>();
Map<String, String> pronItem = new HashMap<>();
pronItem.put("天气", "tian1 qi4");
pronunciation.add(pronItem);
hotFix.put("pronunciation", pronunciation);

List<Map<String, String>> replace = new ArrayList<>();
Map<String, String> replaceItem = new HashMap<>();
replaceItem.put("今天", "金天");
replace.add(replaceItem);
hotFix.put("replace", replace);

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("今天天气真好。")
    .voice("longanhuan_v3.6")
    .parameter("hot_fix", hotFix)
    .build();
示例 21(说明):
通过parameters设置
// 构建hotFix对象同上
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("今天天气真好。")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("hot_fix", hotFix))
    .build();

示例代码

以下示例展示Qwen-Audio-TTS语音合成的非流式和流式调用方式。运行前请确保已设置环境变量DASHSCOPE_API_KEY。
不同模型需使用匹配的音色。更换模型时,请同步更换音色,并确认音色支持目标语言。具体对应关系请参见Qwen-Audio-TTS音色列表。
  • 非流式调用
  • 流式调用
非流式调用会等待服务端合成完成后一次性返回结果。根据返回类型的不同,提供以下两种方式:
  • callAndReturnAudio():返回音频二进制数据(ByteBuffer),适用于直接保存或处理音频的场景。
  • call():返回音频URL,适用于需要通过URL下载音频的场景。
import com.alibaba.dashscope.audio.http_tts.AudioInfo;
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesisParam;
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesisResult;
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesizer;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.InputRequiredException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.utils.Constants;

import java.io.FileOutputStream;
import java.io.IOException;
import java.nio.ByteBuffer;

public class QwenAudioTtsSyncExample {
    static {
        Constants.baseHttpApiUrl = "https://maas.qianwenaiapi.com/api/v1";
    }

    /**
     * 非流式调用示例一:返回音频数据(ByteBuffer)
     * 调用callAndReturnAudio方法,阻塞等待合成完成后返回完整音频数据。
     */
    public static void syncCallReturnAudio() {
        HttpSpeechSynthesizer synthesizer = new HttpSpeechSynthesizer();

        HttpSpeechSynthesisParam param =
            HttpSpeechSynthesisParam.builder()
                .model("qwen-audio-3.0-tts-flash")  // 更换模型时,需同步更换为对应版本的音色
                .text("我家的后面有一个很大的花园。")
                .voice("longanhuan_v3.6")
                .format("wav")
                .sampleRate(24000)
                // 未配置环境变量时,将下行替换为:apiKey("sk-xxx"),即替换为实际的API Key
                .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                // 通过parameter方法设置额外参数
                // .parameter("seed", 1234)
                // .parameter("enable_ssml", true)
                .build();

        try {
            ByteBuffer audioData = synthesizer.callAndReturnAudio(param);
            if (audioData != null && audioData.hasRemaining()) {
                byte[] bytes = new byte[audioData.remaining()];
                audioData.get(bytes);

                try (FileOutputStream fos = new FileOutputStream("sync_output.wav")) {
                    fos.write(bytes);
                    System.out.println("Audio saved to sync_output.wav, size: " + bytes.length + " bytes");
                } catch (IOException e) {
                    System.err.println("Failed to save audio: " + e.getMessage());
                }
            }
        } catch (ApiException | NoApiKeyException | InputRequiredException e) {
            System.err.println("Synthesis failed: " + e.getMessage());
        }
        System.exit(0);
    }

    /**
     * 非流式调用示例二:返回音频URL
     * 调用call方法,返回包含音频URL的结果对象,可通过URL下载音频文件。
     */
    public static void syncCallReturnUrl() {
        HttpSpeechSynthesizer synthesizer = new HttpSpeechSynthesizer();

        HttpSpeechSynthesisParam param =
            HttpSpeechSynthesisParam.builder()
                .model("qwen-audio-3.0-tts-flash")  // 更换模型时,需同步更换为对应版本的音色
                .text("我家的后面有一个很大的花园。")
                .voice("longanhuan_v3.6")
                .format("wav")
                .sampleRate(24000)
                // 未配置环境变量时,将下行替换为:apiKey("sk-xxx"),即替换为实际的API Key
                .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                .build();

        try {
            HttpSpeechSynthesisResult result = synthesizer.call(param);
            System.out.println("Request ID: " + result.getRequestId());

            if (result.hasAudioUrl()) {
                AudioInfo audioInfo = result.getAudioInfo();
                System.out.println("Audio URL: " + audioInfo.getUrl());
                System.out.println("Expires At: " + audioInfo.getExpiresAt());
                System.out.println("Remaining Time: " + audioInfo.getRemainingSeconds() + " seconds");
            }
        } catch (ApiException | NoApiKeyException | InputRequiredException e) {
            System.err.println("Synthesis failed: " + e.getMessage());
        }
    }

    public static void main(String[] args) {
        // 非流式调用示例一:返回音频数据(ByteBuffer)
        // syncCallReturnAudio();
        // 非流式调用示例二:返回音频URL
        syncCallReturnUrl();
    }
}