跳转到主要内容
Qwen-Livetranslate-Flash

音视频翻译-通义千问 API 参考

本文介绍通过 OpenAI 兼容接口调用 qwen3-livetranslate-flash 模型的输入与输出参数。

相关文档:音视频文件翻译-千问
不支持通过 DashScope 接口调用。

OpenAI 兼容

SDK 调用配置的base_url为:https://maas.qianwenaiapi.com/compatible-mode/v1 HTTP 调用配置的endpoint:POST https://maas.qianwenaiapi.com/compatible-mode/v1/chat/completions
您需要已获取与配置 API Key并配置API Key到环境变量。若通过OpenAI SDK进行调用,需要安装SDK。

请求体

import os
from openai import OpenAI

client = OpenAI(
    # 若没有配置环境变量,请用千问AI平台API Key将下行替换为:api_key="sk-xxx",
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    base_url="https://maas.qianwenaiapi.com/compatible-mode/v1",
)

# ----------------音频输入 ----------------
messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "input_audio",
                "input_audio": {
                    "data": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20250211/tixcef/cherry.wav",
                    "format": "wav",
                },
            }
        ],
    }
]

# ----------------视频输入(需取消注释)----------------
# messages = [
#     {
#         "role": "user",
#         "content": [
#             {
#                 "type": "video_url",
#                 "video_url": {
#                     "url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20241115/cqqkru/1.mp4"
#                 },
#             }
#         ],
#     },
# ]

completion = client.chat.completions.create(
    model="qwen3-livetranslate-flash",
    messages=messages,
    modalities=["text", "audio"],
    audio={"voice": "Cherry", "format": "wav"},
    stream=True,
    stream_options={"include_usage": True},
    extra_body={"translation_options": {"source_lang": "zh", "target_lang": "en"}},
)

for chunk in completion:
    print(chunk)
参数类型说明
modelstring(必选) 模型名称。支持的模型:qwen3-livetranslate-flash、qwen3-livetranslate-flash-2025-12-01。
messagesarray(必选) 消息数组,用于向大模型传递上下文。仅支持传入一个 User Message。
User Message(object)
(必选) 用户消息。
User Message
参数类型说明
messages.contentarray(必选) 消息内容。
messages.content.typestring(必选) 可选值:
  • input_audio
输入音频时需设为input_audio。
  • video_url
输入视频文件时需设为video_url。
messages.content.input_audioobject输入的音频信息。当type为input_audio时是必选参数。
messages.content.input_audio.datastring(必选) 音频的 URL 或Base64 Data URL。传入本地文件请参见:输入 Base64 编码的本地文件。
messages.content.input_audio.formatstring(必选) 输入音频的格式,如mp3、wav等。
messages.content.video_urlobject输入的视频文件信息。当type为video_url时是必选参数。
messages.content.video_url.urlstring(必选) 视频文件的公网 URL 或 Base64 Data URL。输入本地视频文件请参见输入 Base64 编码的本地文件。
messages.rolestring(必选) 用户消息的角色,固定为user。
参数类型说明
streamboolean(必选) 默认值为 false 是否以流式方式输出回复。模型仅支持流式输出方式调用,仅可设为true。
stream_optionsobject(可选) 流式输出的配置项,仅在 stream 为 true 时生效。
stream_options.include_usageboolean(可选)默认值为 false 是否在最后一个数据块包含Token消耗信息。 可选值:
  • true:包含;
  • false:不包含。
modalitiesarray(可选)默认值为["text"] 输出数据的模态。可选值:
  • ["text","audio"]:输出文本与音频;
  • ["text"]:仅输出文本。
audioobject(可选) 输出音频的音色与格式。modalities参数需为["text","audio"]。
audio.voicestring(必选) 输出音频的音色。请参见支持的音色。
audio.formatstring(必选) 输出音频的格式,仅支持设定为wav。
max_tokensinteger(可选) 用于限制模型输出的最大 Token 数。若生成内容超过此值,响应将被截断。 默认值与最大值均为模型的最大输出长度,请参见模型选型。
seedinteger(可选) 随机数种子。用于确保在相同输入和参数下生成结果可复现。若调用时传入相同的 seed 且其他参数不变,模型将尽可能返回相同结果。 取值范围:[0,2 31 −1]。
speech_ratefloat(可选)默认值为1.0 控制输出音频的语速。1.0为正常语速,小于1.0为慢速,大于1.0为快速。 取值范围:[0.5, 2.0]。
temperaturefloat(可选)默认值为0.000001 采样温度,控制模型生成内容的多样性。temperature越高,生成的内容更多样,反之更确定。 取值范围: [0, 2) 为了翻译的准确性,不建议修改该值。
top_pfloat(可选)默认值为0.8 核采样的概率阈值,控制模型生成内容的多样性。 top_p越高,生成的内容更多样。反之更确定。 取值范围:(0,1.0] 为了翻译的准确性,不建议修改该值。
presence_penaltyfloat(可选)默认值为0 控制模型生成文本时的内容重复度。 取值范围:[-2.0, 2.0]。正值降低重复度,负值增加重复度。为了翻译的准确性,不建议修改该值。
top_kinteger(可选)默认值为1 生成过程中采样候选集的大小。例如,取值为50时,仅将单次生成中得分最高的50个Token组成随机采样的候选集。取值越大,生成的随机性越高;取值越小,生成的确定性越高。取值为None或当top_k大于100时,表示不启用top_k策略,此时仅有top_p策略生效。 取值需要大于或等于0。为了翻译的准确性,不建议修改该值。 该参数非OpenAI标准参数。通过 Python SDK调用时,请放入 extra_body 对象中,配置方式为:extra_body={"top_k": xxx};通过 Node.js SDK 或 HTTP 方式调用时,请作为顶层参数传递。
repetition_penaltyfloat(可选)默认值为1.05 模型生成时连续序列中的重复度。提高repetition_penalty时可以降低模型生成的重复度,1.0表示不做惩罚。取值大于0即可。为了翻译的准确性,不建议修改该值。 该参数非OpenAI标准参数。通过 Python SDK调用时,请放入 extra_body 对象中,配置方式为:extra_body={"repetition_penalty": xxx};通过 Node.js SDK 或 HTTP 方式调用时,请作为顶层参数传递。
translation_optionsobject(必选) 需配置的翻译参数。
translation_options.source_langstring(可选) 源语言的英文全称,请参见支持的语种。若不设置,模型会自动识别输入的语种。
translation_options.target_langstring(必选) 目标语言的英文全称,请参见支持的语种。 该参数非OpenAI标准参数。通过 Python SDK调用时,请放入 extra_body 对象中,配置方式为:extra_body={"translation_options": xxx};通过 Node.js SDK 或 HTTP 方式调用时,请作为顶层参数传递。

chat响应chunk对象(流式输出)

{
  "id": "chatcmpl-c22a54b8-40cc-4a1d-988b-f84cdf86868f",
  "choices": [
    {
      "delta": {
        "content": " of",
        "function_call": null,
        "refusal": null,
        "role": null,
        "tool_calls": null
      },
      "finish_reason": null,
      "index": 0,
      "logprobs": null
    }
  ],
  "created": 1764755440,
  "model": "qwen3-livetranslate-flash",
  "object": "chat.completion.chunk",
  "service_tier": null,
  "system_fingerprint": null,
  "usage": null
}
参数类型说明
idstring本次调用的唯一标识符。每个chunk对象有相同的 id。
choicesarray模型生成内容的数组。若设置include_usage参数为true,则choices在最后一个chunk中为空数组。
choices.deltaobject请求的增量对象。
choices.delta.contentstring增量消息内容。
choices.delta.reasoning_contentstring该值固定为null。
choices.delta.function_callobject该值固定为null。
choices.delta.audioobject输出的音频信息。
choices.delta.audio.datastring增量的 Base64 音频编码数据。
choices.delta.audio.expires_atinteger创建请求时的时间戳。
choices.delta.audio.idstring输出音频的唯一标识符。
choices.delta.refusalobject该参数当前固定为null。
choices.delta.rolestring增量消息对象的角色,只在第一个chunk中有值。
choices.delta.tool_callsarray该值固定为null。
choices.finish_reasonstring模型停止生成的原因。有以下情况:
  • 自然停止输出时为stop;
  • 生成未结束时为null;
  • 生成长度过长而结束为length。
choices.indexinteger当前响应在choices数组中的索引,固定为0。
choices.logprobsobject该值固定为null。
createdinteger本次请求被创建时的时间戳。每个chunk有相同的时间戳。
modelstring本次请求使用的模型。
objectstring始终为chat.completion.chunk。
service_tierstring该值固定为null。
system_fingerprintstring该值固定为null。
usageobject本次请求消耗的Token。只在include_usage为true时,在最后一个chunk显示。
usage.completion_tokensinteger模型输出的 Token 数。
usage.prompt_tokensinteger输入 Token 数。
usage.total_tokensinteger总 Token 数,为prompt_tokens与completion_tokens的总和。
usage.completion_tokens_detailsobject输出 Token 的详细信息。
usage.completion_tokens_details.audio_tokensinteger输出的音频 Token 数。
usage.completion_tokens_details.reasoning_tokensinteger该值固定为null。
usage.completion_tokens_details.text_tokensinteger输出文本 Token 数。
usage.prompt_tokens_detailsobject输入 Token的细粒度分类。
usage.prompt_tokens_details.audio_tokensinteger输入音频的 Token 数。 > 视频文件中的音频 Token 数通过本参数返回。
usage.prompt_tokens_details.text_tokensinteger输入文本的 Token 数。该值固定为0。
usage.prompt_tokens_details.video_tokensinteger输入视频的 Token 数。