跳转到主要内容
实时多模态

服务端事件

本文介绍 Qwen-Omni-Realtime API 的服务端事件,包括工具调用(Function Calling)相关事件。

error

服务端返回的错误信息。
{
  "event_id": "event_RoUu4T8yExPMI37GKwaOC",
  "type": "error",
  "error": {
    "type": "invalid_request_error",
    "code": "invalid_value",
    "message": "Invalid modalities: ['audio']. Supported combinations are: ['text'] and ['audio', 'text'].",
    "param": "session.modalities"
  }
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为error。
errorobject错误的详细信息。
error.typestring错误类型。
error.codestring错误码。
error.messagestring错误信息。
error.paramstring与错误相关的参数,如session.modalities。

session.created

客户端连接后,服务端返回的第一个事件,包含本次连接的默认配置信息。
{
    "event_id": "event_RdvlSpbBb2ssyBjYrDHjt",
    "type": "session.created",
    "session": {
        "object": "realtime.session",
        "model": "qwen3-omni-flash-realtime",
        "modalities": [
            "text",
            "audio"
        ],
        "voice": "Cherry",
        "input_audio_format": "pcm",
        "output_audio_format": "pcm",
        "input_audio_transcription": {
            "model": "qwen3-asr-flash-realtime"
        },
        "turn_detection": {
            // 取值为server_vad或semantic_vad(qwen3.8-omni-flash-realtime和qwen3.5-omni-realtime系列支持)
            "type": "server_vad",
            "threshold": 0.5,
            "prefix_padding_ms": 300,
            "silence_duration_ms": 800,
            "create_response": true,
            "interrupt_response": true
        },
        "enable_search": false,
        "search_options": {},
        "tools": [],
        "temperature": 0.8,
        "id": "sess_Ov7GOXoNXhNjlxXtOGKQS"
    }
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为session.created。
sessionobject会话的配置信息。
session.objectstring固定为realtime.session。
session.modelstring使用的模型。
session.modalitiesarray模型输出模态设置。
session.voicestring模型生成音频的音色。
session.input_audio_formatstring用户输入音频的格式,当前仅支持设为pcm。输入音频要求为16 kHz采样率的PCM音频流。
session.output_audio_formatstring模型输出音频的格式,当前仅支持设为pcm。输出音频为24 kHz采样率的PCM音频流。当前不支持自定义输出采样率。
session.input_audio_transcriptionobject语音转录的配置。
session.input_audio_transcription.modelstring语音转录模型,固定为qwen3-asr-flash-realtime,不支持修改。
session.turn_detectionobject语音活动检测(VAD)的配置。
session.turn_detection.typestringVAD类型。取值:server_vad(默认值)或 semantic_vad。详情请参见客户端事件。
session.turn_detection.thresholdfloatVAD检测阈值。
session.turn_detection.silence_duration_msinteger检测语音停止的静音持续时间。
session.turn_detection.idle_timeout_msinteger静默超时时间(毫秒)。仅在 server_vad 模式下,使用 qwen3.5-omni-plus-realtime 或 qwen3.5-omni-flash-realtime 模型时返回。
session.enable_searchboolean是否启用联网搜索功能。Qwen3.8-Omni-Flash-Realtime 和 Qwen3.5-Omni-Realtime 系列模型支持。
session.search_optionsobject联网搜索选项配置。
session.temperaturefloat模型的温度参数。

session.updated

收到用户的 session.update 请求后,若处理成功,则返回此事件;若出错,则返回 error 事件。
{
    "event_id": "event_X1HsXS4b4uptp6yo1LgKd",
    "type": "session.updated",
    "session": {
        "id": "sess_Aih6vAcY5Ddt6jwFx1tCa",
        "object": "realtime.session",
        "model": "qwen3.5-omni-flash-realtime",
        "modalities": [
            "text",
            "audio"
        ],
        "audio": {
            "input": {
                "format": {
                    "type": "pcm",
                    "sample_rate": 16000
                }
            },
            "output": {
                "format": {
                    "type": "wav",
                    "sample_rate": 24000
                }
            }
        },
        "instructions": "你是个人助理小云,请你准确且友好地解答用户的问题,始终以乐于助人的态度回应。",
        "voice": "Tina",
        "input_audio_format": "pcm",
        "output_audio_format": "pcm",
        "input_audio_transcription": {
            "model": "qwen3-asr-flash-realtime"
        },
        "turn_detection": {
            // 取值为server_vad或semantic_vad(qwen3.8-omni-flash-realtime和qwen3.5-omni-realtime系列支持)
            "type": "server_vad",
            "threshold": 0.1,
            "prefix_padding_ms": 500,
            "silence_duration_ms": 900,
            "create_response": true,
            "interrupt_response": true
        },
        "enable_search": false,
        "search_options": {},
        "tools": [
            {
                "type": "function",
                "function": {
                    "name": "get_current_weather",
                    "description": "当你想查询指定城市的天气时非常有用。",
                    "parameters": {
                        "type": "object",
                        "properties": {
                            "location": {"type": "string", "description": "城市名称"}
                        },
                        "required": ["location"]
                    }
                }
            }
        ],
        "temperature": 0.8,
        "max_response_output_token": "inf",
        "max_tokens": 16384,
        "repetition_penalty": 1.05,
        "presence_penalty": 0.0,
        "top_k": 50,
        "top_p": 1.0,
        "seed":-1
    }
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为session.updated。
sessionobject会话的配置信息。
session.temperaturefloat模型的温度参数。
session.modalitiesarray模型输出模态设置。
session.voicestring模型生成音频的音色。
session.instructionsstring模型的目标与角色。
session.audioobject回显的音频格式配置。若客户端传入了 session.audio.input.format / session.audio.output.format,服务端将在 session.updated 中按相同嵌套结构回显。未使用嵌套字段的客户端,服务端事件结构保持原有行为。
session.audio.input.format.typestring用户输入音频格式。可选值:pcm(默认值)、wav。
session.audio.input.format.sample_rateinteger用户输入音频采样率,单位为 Hz。
session.audio.output.format.typestring模型输出音频格式。可选值:pcm(默认值)、wav。
session.audio.output.format.sample_rateinteger模型输出音频采样率,单位为 Hz。
session.input_audio_formatstring历史兼容字段,回显客户端配置的输入音频格式。
session.output_audio_formatstring历史兼容字段,回显客户端配置的输出音频格式。
session.input_audio_transcriptionobject语音转录的配置。
session.input_audio_transcription.modelstring语音转录模型,固定为qwen3-asr-flash-realtime,不支持修改。
session.turn_detectionobject语音活动检测(VAD)的配置。
session.turn_detection.typestringVAD类型。取值:server_vad(默认值)或 semantic_vad。详情请参见客户端事件。
session.turn_detection.thresholdfloatVAD检测阈值。
session.turn_detection.silence_duration_msinteger检测语音停止的静音持续时间。
session.turn_detection.idle_timeout_msinteger静默超时时间(毫秒)。仅在 server_vad 模式下,使用 qwen3.5-omni-plus-realtime 或 qwen3.5-omni-flash-realtime 模型时返回。
session.enable_searchboolean(可选) 是否启用联网搜索功能。Qwen3.8-Omni-Flash-Realtime 和 Qwen3.5-Omni-Realtime 系列模型支持。
session.search_optionsobject(可选) 联网搜索选项配置。
session.toolsarray(可选) 工具定义列表。以下属性描述 type="function" 的 Function Calling 工具。Qwen3.8-Omni-Flash-Realtime 的 MCP 配置字段见客户端工具配置;MCP 连接地址和凭证不会在 session.updated 中回显。
session.tools.typestring(必选) 固定为 function。
session.tools.function.namestring(必选) 自定义的工具函数名称,建议使用与函数相同的名称,如get_current_weather或get_current_time。
session.tools.function.descriptionstring(可选) 对工具函数功能的描述,大模型会参考该字段来选择是否使用该工具函数。
session.tools.function.parametersobject(可选) 对工具函数入参的描述,大模型会参考该字段来进行入参的提取。如果工具函数不需要输入参数,则无需指定。
session.tools.function.parameters.typestring(必选) 固定为 object。
session.tools.function.parameters.propertiesobject(可选) 描述各入参的名称、数据类型与描述。Key 值为入参的名称,Value 值为包含数据类型(type)与描述(description)的对象。
session.tools.function.parameters.requiredarray(可选) 指定哪些入参为必填项。
session.top_pfloat核采样的概率阈值。
session.top_kinteger模型生成过程中,采样候选集的大小。
session.max_tokensinteger模型在本次请求返回的最大 Token 数。
session.repetition_penaltyfloat控制模型生成时,连续序列中的重复度*。*
session.presence_penaltyfloat控制模型在生成内容时的重复度。
session.seedinteger模型在每次请求时,运行结果一致性程度。

input_audio_buffer.speech_started

在 VAD 模式下,当服务端在音频缓冲区中检测到语音开始时,会返回此事件。
若服务端尚未检测到语音,则每次向缓冲区添加音频时都可能触发此事件。
{
    "event_id": "event_Pvp8nEhsQuGCQbFJ9x58n",
    "type": "input_audio_buffer.speech_started",
    "audio_start_ms": 3647,
    "item_id": "item_YbAiGvK2H7YaS34o4R6Ba"
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为input_audio_buffer.speech_started。
audio_start_msinteger从音频开始写入缓冲区到首次检测到语音所经过的毫秒数。
item_idstring语音停止时将创建的用户消息项的 ID。 > 用户消息项用于将用户输入追加到对话历史,供模型后续推理与生成使用。

input_audio_buffer.speech_stopped

在 VAD 模式下,当音频缓冲区中检测到语音结束时,服务端会返回此事件。 同时,服务端还会返回一个 conversation.item.created 事件,以创建对应的用户消息项。
{
    "event_id": "event_UhQiqNVRsgUiq4KUS5Xb5",
    "type": "input_audio_buffer.speech_stopped",
    "audio_end_ms": 4453,
    "item_id": "item_YbAiGvK2H7YaS34o4R6Ba"
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为input_audio_buffer.speech_stopped。
audio_end_msinteger语音停止时刻距会话开始经过的毫秒数。
item_idstring将创建的用户消息项的 ID。

input_audio_buffer.committed

当输入音频缓冲区被提交时返回此事件。
  • 在VAD模式下,当检测到用户说话结束时,服务端会自动提交音频缓冲区并返回此事件。
  • 在 Manual 模式下,当客户端发送input_audio_buffer.commit事件后,服务端返回此事件。
{
    "event_id": "event_Iy6sUzL1nmdFgshFYxJEz",
    "type": "input_audio_buffer.committed",
    "item_id": "item_YbAiGvK2H7YaS34o4R6Ba"
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为input_audio_buffer.committed。
item_idstring将创建的用户消息项的 ID。

input_audio_buffer.cleared

客户端发送input_audio_buffer.clear事件后,服务端将返回此事件。
{
  "event_id": "event_RoUu4T8yExPMI37GKwaOC",
  "type": "input_audio_buffer.cleared"
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为input_audio_buffer.cleared。

conversation.item.created

当对话项创建时返回此事件。 示例:
{
    "event_id": "event_JEfkrr9gO3Ny7Xcv9bGVd",
    "type": "conversation.item.created",
    "item": {
        "id": "item_YbAiGvK2H7YaS34o4R6Ba",
        "object": "realtime.item",
        "type": "message",
        "status": "in_progress",
        "role": "assistant",
        "content": [
            {
                "type": "input_audio"
            }
        ]
    }
}
// 工具调用场景
{
    "event_id": "event_S1hkaIQgcuQD8OEdOpGHQ",
    "type": "conversation.item.created",
    "item": {
        "id": "item_FEG9qJGNkPcdf4et3p7BV",
        "object": "realtime.item",
        "type": "function_call",
        "status": "in_progress",
        "call_id": "call_bc0a7fb7235840f69ecfe4",
        "name": "get_current_weather",
        "arguments": ""
    }
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为conversation.item.created。
response_idstring(条件必填) 仅当 item.type 为 mcp_approval_request 时返回;生成本次 MCP 调用的父 Response ID。
item_idstring(条件必填) 仅当 item.type 为 mcp_approval_request 时返回;审批请求 item ID,与 item.id 相同。
previous_item_idstring(条件必填) 仅当 item.type 为 mcp_approval_request 时返回;被审批的 mcp_call item ID。
itemobject要添加到对话中的项。
item.idstring对话项的唯一ID。
item.objectstringmessage 和 function_call 项固定为 realtime.item;MCP 项若返回该字段,值也为 realtime.item。
item.statusstring若返回该字段,表示对话项的状态。mcp_call 初始项必填,值为 in_progress。
item.rolestring仅消息项包含,表示消息的角色。
item.contentarray[object]消息的内容。当 type 为 message 时存在。下列字段对应本页示例中的数组元素。
item.content.typestring内容片段的类型;本页示例为 input_audio。
item.typestring对话项的类型包括 message(常规消息)和 function_call(工具调用)。Qwen3.8-Omni-Flash-Realtime 还可返回 mcp_list_tools、mcp_call、mcp_approval_request,事件示例见MCP 对话项。
item.namestring当 type 为 function_call 时,被调用的函数名称。
item.call_idstring当 type 为 function_call 时,本次函数调用的唯一 ID。
item.argumentsstring当 type 为 function_call 时,函数调用的参数(JSON 字符串)。
type=mcp_list_tools 该类型的 item.id 与工具发现状态事件的 item_id 相同。
参数类型说明
item.server_labelstring(必填) MCP Server 标识。
item.toolsarray[object](必填) 经校验和 allowed_tools 过滤的工具定义,发现失败时为空数组。
item.tools.namestring(必填) 工具原始名称,1~64 位,仅允许字母、数字、下划线、点和连字符。
item.tools.descriptionstring(可选) MCP Server 提供的工具说明。
item.tools.input_schemaobject(必填) MCP Server 提供的 JSON Schema。服务端不执行完整的 JSON Schema 语义校验,最终由 MCP Server 校验工具参数。
item.tools.input_schema.typestring(必填) 根节点固定为 object。
item.tools.input_schema.propertiesobject(可选) 按自定义名称列出各参数。例如下方 mcp_list_tools 示例 中的 city 是参数名,其 type 和 description 分别说明数据类型和用途;各参数的 JSON Schema 由 MCP Server 定义。
item.tools.input_schema.requiredarray[string](可选) 必填参数名列表。 additionalProperties、$defs、oneOf、anyOf、allOf 等其他标准 JSON Schema 字段按原始 schema 保留并传递。
item.tools.annotationsobject(可选) MCP Server 提供的 ToolAnnotations。以下是 MCP 标准中的可选提示字段,实际返回内容由 MCP Server 决定。
item.tools.annotations.titlestring(可选) 工具的显示名称。
item.tools.annotations.readOnlyHintboolean(可选) 提示工具是否只读。
item.tools.annotations.destructiveHintboolean(可选) 提示工具是否可能执行破坏性修改。
item.tools.annotations.idempotentHintboolean(可选) 提示重复调用是否不会产生额外效果。
item.tools.annotations.openWorldHintboolean(可选) 提示工具是否可能与外部系统交互。这些字段均为提示,不应作为安全判断依据。
item.errorobject(可选) 工具发现失败时出现。
item.error.typestring(必填) 固定为 tool_execution_error。
item.error.messagestring(必填) 面向客户端的安全错误描述,不包含上游敏感响应体。
type=mcp_call(初始项)
参数类型说明
item.objectstring(必填) 固定为 realtime.item。
item.statusstring(必填) 初始为 in_progress。
item.server_labelstring(必填) 实际执行工具的 MCP Server 标识。
item.namestring(必填) 工具原始名称。
item.call_idstring(必填) 本次调用的唯一 ID。
item.argumentsstring(必填) 初始通常为空字符串;完整参数以后续事件及最终项为准。
type=mcp_approval_request 审批请求的 item.id 应在审批回复中原样填入 approval_request_id。
参数类型说明
item.server_labelstring(必填) 待执行工具所属的 MCP Server 标识。
item.namestring(必填) 待执行工具的原始名称。
item.argumentsstring(必填) 待审批调用的完整参数 JSON 字符串。
item.call_idstring(必填) 被审批的工具调用 ID。

conversation.item.input_audio_transcription.delta

开启输入音频转录后,此事件会在用户说话过程中高频发送,用于展示实时识别的中间结果。您可以通过拼接 text + stash 获取当前最完整的句子预览。
event_idstring本次事件唯一标识符。typestring事件类型,固定为conversation.item.input_audio_transcription.delta。item_idstring关联的对话项 ID。content_indexinteger包含音频的内容部分的索引。textstring已确认的文本前缀。这是当前句子中,模型已确认不会再变更的部分。stashstring预识别的文本后缀。这是紧跟在已确认部分之后,模型仍在处理、可能会被修正的临时草稿。languagestring被识别音频的语种。emotionstring被识别音频的情感。可选值:neutral(平静)、happy(愉快)、sad(悲伤)、angry(愤怒)、surprised(惊讶)、disgusted(厌恶)、fearful(恐惧)。
{
    "event_id": "event_C7jzoeSFuiwOZS6tR14yx",
    "type": "conversation.item.input_audio_transcription.delta",
    "item_id": "item_ThVYhLHOdeXb4bBSvzSFF",
    "content_index": 0,
    "text": "",
    "stash": "今天天气怎么样?",
    "language": "zh",
    "emotion": "neutral",
    "obfuscation": "ABEXGYmxdmc97u"
}
在任何时刻,要获取当前最完整的句子预览,都需要将这两个字段拼接起来:实时预览句子 = text + stash。
假设用户正在说:"今天天气不错,阳光明媚。"以下是您可能会收到的事件流以及如何解读它们:

时间点

用户说话进度

API 响应 (text 和 stash)

客户端 UI 应显示 (text + stash)

T1

"今天……"

text: ""

stash: "今天"

今天

T2

"……天气……"

text: ""

stash: "今天天气"

今天天气

T3

"……不错"

text: "今天"

stash: "天气不错"

今天天气不错

(注意,"今天"已被确认并移入text)

T4

(短暂停顿)

text: "今天天气不错,"

stash: ""

今天天气不错,

(前半句完全确认)

T5

"……阳光……"

text: "今天天气不错,"

stash: "阳光"

今天天气不错,阳光

T6

"……明媚。"

text: "今天天气不错,"

stash: "阳光明媚。"

今天天气不错,阳光明媚。

T7

(结束说话)

-

使用 conversation.item.input_audio_transcription.completed 的 transcript 内容作为最终结果。

conversation.item.input_audio_transcription.completed

此事件表示用户音频写入缓冲区后生成的转录结果。其转录由内置的语音识别模型(固定为 qwen3-asr-flash-realtime)处理,不支持修改。
语音识别模型生成的转录文本可能与 Qwen-Omni-Realtime 模型的理解存在差异,仅供参考。
{
    "event_id": "event_FrrZcxiDfTB9LD9p4pVng",
    "type": "conversation.item.input_audio_transcription.completed",
    "item_id": "item_YbAiGvK2H7YaS34o4R6Ba",
    "content_index": 0,
    "transcript": "喂,你好。"
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为conversation.item.input_audio_transcription.completed。
item_idstring用户消息项的 ID。
content_indexinteger当前固定为0。
transcriptstring转录的文本内容。

conversation.item.input_audio_transcription.failed

启用输入音频转录后,若用户音频转录失败,服务端会返回此事件。此事件独立于 error 事件,便于客户端识别。
{
  "type": "conversation.item.input_audio_transcription.failed",
  "item_id": "<item_id>",
  "content_index": 0,
  "error": {
    "code": "<code>",
    "message": "<message>",
    "param": "<param>"
  }
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为conversation.item.input_audio_transcription.failed。
item_idstring用户消息项的 ID。
content_indexinteger当前固定为0。
errorobject错误信息。
error.codestring错误码。
error.messagestring错误消息。
error.paramstring错误相关的参数。

response.created

当服务端生成新的模型响应时,会返回此事件。
{
    "event_id": "event_XuDavMzQN3KKepqGu3KRh",
    "type": "response.created",
    "response": {
        "id": "resp_HaVOPdbmX6vifiV5pAfJY",
        "object": "realtime.response",
        "conversation_id": "conv_FjJaccpnvwHNo9cPVuzGc",
        "status": "in_progress",
        "modalities": [
            "text",
            "audio"
        ],
        "voice": "Cherry",
        "output_audio_format": "pcm",
        "output": []
    }
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.created。
responseobject响应对象。
response.idstring响应的唯一 ID。
response.conversation_idstring当前会话的唯一ID。
response.objectstring对象类型,此事件下固定为realtime.response。
response.statusstring响应的状态。在[completed, failed, in_progress, or incomplete]范围内。
response.modalitiesarray响应的模态。
response.voicestring模型生成音频的音色。
response.outputarray此事件下目前为空。

response.done

响应生成完成后,服务端会返回此事件。事件中的 response 对象包含除原始音频数据外的全部输出项。 示例:
{
    "event_id": "event_CSaxRRYLvbrfexDXAEuDG",
    "type": "response.done",
    "response": {
        "id": "resp_HaVOPdbmX6vifiV5pAfJY",
        "object": "realtime.response",
        "conversation_id": "conv_FjJaccpnvwHNo9cPVuzGc",
        "status": "completed",
        "modalities": [
            "text",
            "audio"
        ],
        "voice": "Cherry",
        "output_audio_format": "pcm",
        "output": [
            {
                "id": "item_Ls6MtCUWO7LM4E59QziNv",
                "object": "realtime.item",
                "type": "message",
                "status": "completed",
                "role": "assistant",
                "content": [
                    {
                        "type": "audio",
                        "transcript": "你好呀!有什么我可以帮你的吗?"
                    }
                ]
            }
        ],
        "usage": {
            "total_tokens": 377,
            "input_tokens": 336,
            "output_tokens": 41,
            "input_tokens_details": {
                "text_tokens": 228,
                "audio_tokens": 108
            },
            "output_tokens_details": {
                "text_tokens": 9,
                "audio_tokens": 32
            },
            "plugins": {
                "search": {
                    "count": 1,
                    "strategy": "agent"
                }
            }
        }
    }
}
// 工具调用场景
{
    "event_id": "event_T1EFAJp43X2DWtDRmxTtx",
    "type": "response.done",
    "response": {
        "id": "resp_TucN5QgymL5MA8vkJvFlS",
        "object": "realtime.response",
        "conversation_id": "conv_SEDZESRlefT8WvLSmEn6E",
        "status": "completed",
        "modalities": ["text", "audio"],
        "voice": "Ethan",
        "output_audio_format": "pcm",
        "output": [
            {
                "id": "item_FEG9qJGNkPcdf4et3p7BV",
                "object": "realtime.item",
                "type": "function_call",
                "status": "completed",
                "call_id": "call_bc0a7fb7235840f69ecfe4",
                "name": "get_current_weather",
                "arguments": " {\"location\": \"杭州\"}"
            }
        ],
        "usage": {
            "total_tokens": 567,
            "input_tokens": 524,
            "output_tokens": 43,
            "input_tokens_details": {
                "text_tokens": 487,
                "audio_tokens": 37
            },
            "output_tokens_details": {
                "text_tokens": 43
            }
        }
    }
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.done。
responseobject响应对象。
response.idstring响应的唯一 ID。
response.conversation_idstring当前会话的唯一ID。
response.objectstring对象类型,此事件下固定为realtime.response。
response.statusstring响应的状态。
response.modalitiesarray响应的模态。
response.voicestring模型生成音频的音色。
response.outputarray响应的输出项数组;每个元素是对话项对象。
response.output.idstring响应输出对应的ID。
response.output.typestring输出项的类型,可选值包括 message(常规消息)、function_call(工具调用),Qwen3.8-Omni-Flash-Realtime 还支持 mcp_call,事件示例见MCP 对话项。
response.output.objectstringmessage 和 function_call 项固定为 realtime.item;MCP 项若返回该字段,值也为 realtime.item。
response.output.statusstring输出项的状态。
response.output.rolestring仅消息项包含,表示消息的角色。
response.output.contentarray输出项的内容。当 type 为 message 时存在;每个元素是输出内容对象。
response.output.content.typestring输出内容的类型。输出为纯文本时,为text;输出包含音频时,为audio。
response.output.content.textstring输出的文本内容。
response.output.content.transcriptstring音频转录为文字后的内容。
response.output.namestring当 type 为 function_call 时,被调用的函数名称。
response.output.call_idstring当 type 为 function_call 时,函数调用的唯一 ID。
response.output.argumentsstring当 type 为 function_call 时,函数调用的完整参数(JSON 字符串)。
type=mcp_call(最终项)
参数类型说明
response.output.statusstring(必填) completed 或 failed。
response.output.server_labelstring(必填) 实际执行工具的 MCP Server 标识。
response.output.namestring(必填) 工具原始名称。
response.output.call_idstring(必填) 本次调用的唯一 ID。
response.output.argumentsstring(必填) 完整参数 JSON 字符串。
response.outputstring(可选) MCP tools/call.result 对象序列化后的 JSON 字符串;部分失败场景也可能出现。
response.output.errorobject(可选) 调用失败时出现。
response.output.error.typestring(必填) 固定为 tool_execution_error。
response.output.error.messagestring(必填) 面向客户端的安全错误描述,不包含上游敏感响应体。 MCP Server 返回 isError=true 时最终状态为 failed,也可能保留原始 output。
参数类型说明
response.usageobject本次响应的 Token 消耗信息。
response.usage.total_tokensinteger本次响应消耗的总 Token 数。
response.usage.input_tokensinteger输入 Token 数。
response.usage.output_tokensinteger输出 Token 数。
response.usage.input_tokens_detailsobject输入 Token 的分项详情,包含 text_tokens(文本 Token 数)和 audio_tokens(音频 Token 数)。
response.usage.output_tokens_detailsobject输出 Token 的分项详情,包含 text_tokens(文本 Token 数)和 audio_tokens(音频 Token 数)。
response.usage.pluginsobject(可选) 插件使用计量信息。启用联网搜索(enable_search)时返回。
response.usage.plugins.searchobject联网搜索计量信息。
response.usage.plugins.search.countinteger搜索次数。
response.usage.plugins.search.strategystring搜索策略。

response.text.delta

当输出模态仅包含文本,且模型增量生成新的文本时,服务端将返回此事件。
{
    "delta": "喂",
    "event_id": "event_TH49MauuPmRo1RGaMSlP7",
    "type": "response.text.delta",
    "response_id": "resp_PrRSvPVpnCExdUOGHHLuP",
    "item_id": "item_L8IRm9kRXFpxoOjDqDC96",
    "output_index": 0,
    "content_index": 0
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.text.delta。
deltastring返回的增量文本。
response_idstring回复的ID。
item_idstring消息项ID,可以关联同一个消息项。
output_indexinteger响应中输出项的索引, 目前固定为 0。
content_indexinteger响应中输出项中内部部分的索引, 目前固定为 0。

response.text.done

当输出模态仅包含文本,且模型生成的文本结束时,服务端将返回此事件。
当响应中断、不完整或取消时,也会返回此事件。
{
  "event_id": "event_B1lIeE2Nac33zn5V7h2mm",
  "type": "response.text.done",
  "response_id": "resp_B1lIdtjF4Noqpn5NOjznj",
  "item_id": "item_B1lIdJsAJlJiFs8ztWpJt",
  "output_index": 0,
  "content_index": 0,
  "text": "How can I assist you today?"
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.text.done。
response_idstring响应的ID。
item_idstring消息项ID。
output_indexinteger响应输出项的索引。
content_indexinteger响应输出项的索引。
textstring模型输出的完整文本。

response.audio.delta

当输出模态包含音频,且模型增量生成新的音频数据时,服务端将返回此事件。
{
  "event_id": "event_B1osWMZBtrEQbiIwW0qHQ",
  "type": "response.audio.delta",
  "response_id": "resp_P79OOMs8LnrXVpiIHUCKR",
  "item_id": "item_OFaPGtzfWCPyGzxnuEX9i",
  "output_index": 0,
  "content_index": 0,
  "delta": "{base64 audio}"
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.audio.delta。
response_idstring响应的ID。
item_idstring消息项ID。
output_indexinteger响应输出项的索引。
content_indexinteger响应输出项的索引。
deltastring模型增量输出的音频数据,使用Base64编码。

response.audio.done

当输出模态包含音频,且模型完成生成音频数据时,服务端将返回此事件。
当响应中断、不完整或取消时,也会返回此事件。
{
    "event_id": "event_Le1TDl7VfyHQxl47DtGxI",
    "type": "response.audio.done",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.audio.done。
response_idstring响应的ID。
item_idstring消息项ID。
output_indexinteger响应输出项的索引。
content_indexinteger响应输出项的索引。

response.audio_transcript.delta

当输出模态包含音频,且模型增量生成新的音频对应的文本时,服务端将返回 response.audio_transcript.delta 事件。
{
    "event_id": "event_BksW7fOwnyavZdDxIzZYM",
    "type": "response.audio_transcript.delta",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0,
    "delta": "有什么"
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.audio_transcript.delta。
response_idstring响应的ID。
item_idstring消息项ID。
output_indexinteger响应输出项的索引。
content_indexinteger响应输出项的索引。
deltastring增量文本。

response.audio_transcript.done

当输出模态包含音频,且模型完成音频转录后,服务端将返回 response.audio_transcript.done 事件。
{
    "event_id": "event_X49tL2WerT4WjxcmH16lS",
    "type": "response.audio_transcript.done",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0,
    "transcript": "你好呀!有什么我可以帮你的吗?"
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.audio_transcript.done。
response_idstring响应的ID。
item_idstring消息项ID。
output_indexinteger响应输出项的索引。
content_indexinteger响应输出项的索引。
transcriptstring完整文本。

response.function_call_arguments.delta

当模型以流式方式生成函数调用的参数字符串时,每产生一段新内容,服务端推送一次本事件。客户端应按接收顺序将各事件中的 delta 字段拼接,得到与当前进度一致的参数文本;完整内容以随后的 response.function_call_arguments.done 为准。
{
    "event_id": "event_SlKoJyEbPEqLq14DSM1u5",
    "type": "response.function_call_arguments.delta",
    "response_id": "resp_JnTOsWXlFhKcFohZbtfz6",
    "item_id": "item_Rhcms7CauTNsQprV5S4Hr",
    "output_index": 0,
    "call_id": "call_2be200f4cafe419b9530dd",
    "delta": " {\"location\": \"北京\"}"
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.function_call_arguments.delta。
response_idstring响应的ID。
item_idstring消息项ID。
output_indexinteger该响应中输出项的索引。
call_idstring本次函数调用的唯一 ID,与同一轮中的 done 事件保持一致。
deltastring本段新增的参数字符串片段(增量)。需按顺序拼接。

response.function_call_arguments.done

函数调用参数已全部生成完毕。本事件中的 arguments 为完整的参数字符串。客户端可在收到本事件后解析参数并调用本地工具函数;应以本事件中的完整 arguments 为准,而非 delta 拼接结果。
{
    "event_id": "event_X6suLyuL5agdH7r6koesM",
    "type": "response.function_call_arguments.done",
    "response_id": "resp_JnTOsWXlFhKcFohZbtfz6",
    "item_id": "item_Rhcms7CauTNsQprV5S4Hr",
    "output_index": 0,
    "name": "get_current_weather",
    "call_id": "call_2be200f4cafe419b9530dd",
    "arguments": " {\"location\": \"北京\"}"
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.function_call_arguments.done。
response_idstring响应的ID。
item_idstring消息项ID。
output_indexinteger该响应中输出项的索引。
call_idstring本次函数调用的唯一 ID。
namestring被调用的函数名称。
argumentsstring函数调用的完整参数,一般以 JSON 字符串形式表示。

response.output_item.added

在响应生成过程中创建新项目时,服务端返回此事件。项目类型可以是 message(常规消息)、function_call(工具调用),Qwen3.8-Omni-Flash-Realtime 还支持 mcp_call,事件示例见MCP 对话项。
{
    "event_id": "event_DsCO341DEVtiATtCB6BUY",
    "type": "response.output_item.added",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "output_index": 0,
    "item": {
        "id": "item_Ls6MtCUWO7LM4E59QziNv",
        "object": "realtime.item",
        "type": "message",
        "status": "in_progress",
        "role": "assistant",
        "content": []
    }
}
// 工具调用场景
{
    "event_id": "event_HXmKt5pGoiRtXx7Hq7zpN",
    "type": "response.output_item.added",
    "response_id": "resp_TucN5QgymL5MA8vkJvFlS",
    "output_index": 0,
    "item": {
        "id": "item_FEG9qJGNkPcdf4et3p7BV",
        "object": "realtime.item",
        "type": "function_call",
        "status": "in_progress",
        "call_id": "call_bc0a7fb7235840f69ecfe4",
        "name": "get_current_weather",
        "arguments": ""
    }
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.output_item.added。
response_idstring响应的ID。
output_indexinteger响应输出项的索引。
itemobject输出项信息。
item.idstring输出项的唯一ID。
item.objectstringmessage 和 function_call 项固定为 realtime.item;MCP 项若返回该字段,值也为 realtime.item。
item.statusstring输出项的状态。
item.rolestring仅消息项包含,表示消息的角色。
item.contentarray消息的内容。当 type 为 message 时存在。本页示例中初始为空数组;完成项的内容见 response.output_item.done 示例。
item.typestring输出项的类型。可选值包括 message(常规消息)、function_call(工具调用),Qwen3.8-Omni-Flash-Realtime 还支持 mcp_call,事件示例见MCP 对话项。
item.namestring当 type 为 function_call 时,被调用的函数名称。
item.call_idstring当 type 为 function_call 时,本次函数调用的唯一 ID。
item.argumentsstring当 type 为 function_call 时,函数调用的参数(JSON 字符串)。在 added 事件中初始为空字符串。
type=mcp_call(初始项)
参数类型说明
item.objectstring(必填) 固定为 realtime.item。
item.statusstring(必填) 初始为 in_progress。
item.server_labelstring(必填) 实际执行工具的 MCP Server 标识。
item.namestring(必填) 工具原始名称。
item.call_idstring(必填) 本次调用的唯一 ID。
item.argumentsstring(必填) 初始通常为空字符串;完整参数以 response.mcp_call_arguments.done 和最终项为准。

response.output_item.done

当新的项目输出完成时,服务端返回此事件。
{
    "event_id": "event_MEu5nlLw1LsOguHiehIP8",
    "type": "response.output_item.done",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "output_index": 0,
    "item": {
        "id": "item_Ls6MtCUWO7LM4E59QziNv",
        "object": "realtime.item",
        "type": "message",
        "status": "completed",
        "role": "assistant",
        "content": [
            {
                "type": "audio",
                "text": "你好呀!有什么我可以帮你的吗?"
            }
        ]
    }
}
// 工具调用场景
{
    "event_id": "event_FHspdfAnCyjuME3mmAwSY",
    "type": "response.output_item.done",
    "response_id": "resp_TucN5QgymL5MA8vkJvFlS",
    "output_index": 0,
    "item": {
        "id": "item_FEG9qJGNkPcdf4et3p7BV",
        "object": "realtime.item",
        "type": "function_call",
        "status": "completed",
        "call_id": "call_bc0a7fb7235840f69ecfe4",
        "name": "get_current_weather",
        "arguments": " {\"location\": \"杭州\"}"
    }
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.output_item.done。
response_idstring响应的ID。
output_indexinteger响应输出项的索引。
itemobject输出项信息。
item.idstring输出项的唯一ID。
item.objectstringmessage 和 function_call 项固定为 realtime.item;MCP 项若返回该字段,值也为 realtime.item。
item.statusstring输出项的状态。
item.rolestring仅消息项包含,表示消息的角色。
item.contentarray[object]消息的内容。当 type 为 message 时存在。下列字段对应本页示例中的数组元素。
item.content.typestring内容片段的类型;本页示例为 audio。
item.content.textstring本页示例中的文本内容。
item.typestring输出项的类型。可选值包括 message(常规消息)、function_call(工具调用),Qwen3.8-Omni-Flash-Realtime 还支持 mcp_call,事件示例见MCP 对话项。
item.namestring当 type 为 function_call 时,被调用的函数名称。
item.call_idstring当 type 为 function_call 时,本次函数调用的唯一 ID。
item.argumentsstring当 type 为 function_call 时,函数调用的完整参数(JSON 字符串)。
type=mcp_call(最终项)
参数类型说明
item.statusstring(必填) completed 或 failed。
item.server_labelstring(必填) 实际执行工具的 MCP Server 标识。
item.namestring(必填) 工具原始名称。
item.call_idstring(必填) 本次调用的唯一 ID。
item.argumentsstring(必填) 完整参数 JSON 字符串。
item.outputstring(可选) MCP tools/call.result 对象序列化后的 JSON 字符串;部分失败场景也可能出现。
item.errorobject(可选) 调用失败时出现。
item.error.typestring(必填) 固定为 tool_execution_error。
item.error.messagestring(必填) 面向客户端的安全错误描述,不包含上游敏感响应体。 MCP Server 返回 isError=true 时最终状态为 failed,也可能保留原始 output。

response.content_part.added

在响应生成过程中,向助手消息项中添加新内容部分时,服务端返回此事件。
{
    "event_id": "event_AVBOmrgY3C8bjlRajfSUT",
    "type": "response.content_part.added",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0,
    "part": {
        "type": "audio",
        "text": ""
    }
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.content_part.added。
response_idstring响应的ID。
item_idstring消息项ID。
output_indexinteger响应输出项的索引,目前固定为 0。
content_indexinteger响应输出项中内部部分的索引, 目前固定为 0。
partobject输出项信息。
part.typestring内容部分的类型。
part.textstring内容部分的文本。

response.content_part.done

在助手消息项中的内容部分完成流式传输时,服务端返回此事件。
{
    "event_id": "event_Il8HD19v58Qr5IBkw7LtN",
    "type": "response.content_part.done",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0,
    "part": {
        "type": "audio",
        "text": "你好呀!有什么我可以帮你的吗?"
    }
}
参数类型说明
event_idstring本次事件唯一标识符。
typestring事件类型,固定为response.content_part.done。
response_idstring响应的ID。
item_idstring消息项ID。
output_indexinteger响应输出项的索引,目前固定为 0。
content_indexinteger该项内容数组中内容部分的索引,目前固定为 0。
partobject输出项信息。
part.typestring内容部分的类型。
part.textstring内容部分的文本。

MCP 事件

本节介绍 qwen3.8-omni-flash-realtime 的 WebSocket 事件与字段。基础事件与字段见前述事件说明。表中“必填”表示所属对象出现时必须提供。ID 均为不透明字符串,不应依赖其长度、前缀或生成规则。arguments、output 的外层类型为 string,读取内容时需要再次解析 JSON。MCP 连接、工具发现和调用受服务配额及超时限制。

mcp_list_tools.*

以下三个事件具有相同的字段结构:
事件 type触发时机
mcp_list_tools.in_progress开始发现某个 MCP Server 的工具
mcp_list_tools.completed工具发现成功并完成 allowed_tools 过滤
mcp_list_tools.failed工具发现失败、超时或结果超过服务限制
字段路径类型必填说明
event_idstring是本条服务端事件的唯一 ID
typestring是上表中的固定事件类型
item_idstring是本次工具发现对应的 mcp_list_tools item ID
mcp_list_tools.in_progress 可能早于 session.updated 到达。对于 completed 和 failed,服务端会先发送包含最终工具列表或错误的 conversation.item.created,再发送相同 item_id 的状态事件。 示例:
{
  "event_id": "opaque_event_id",
  "type": "mcp_list_tools.completed",
  "item_id": "opaque_item_id"
}

response.mcp_call_arguments.delta

字段路径类型必填说明
event_idstring是本条服务端事件的唯一 ID
typestring是固定为 response.mcp_call_arguments.delta
response_idstring是父 Response ID
item_idstring是当前 mcp_call item ID
output_indexinteger是当前 item 在父 Response output 数组中的零基索引
deltastring是工具参数 JSON 字符串的本次增量片段,按事件顺序拼接
obfuscationstring否可选混淆字符串;客户端可以忽略,不影响 delta 拼接和参数解析
示例:
{
  "event_id": "opaque_event_id",
  "type": "response.mcp_call_arguments.delta",
  "response_id": "opaque_response_id",
  "item_id": "opaque_item_id",
  "output_index": 0,
  "delta": "{\"city\":\"杭",
  "obfuscation": "opaque-value"
}

response.mcp_call_arguments.done

字段路径类型必填说明
event_idstring是本条服务端事件的唯一 ID
typestring是固定为 response.mcp_call_arguments.done
response_idstring是父 Response ID
item_idstring是当前 mcp_call item ID
output_indexinteger是当前 item 在父 Response output 数组中的零基索引
argumentsstring是完整工具参数,内容为 JSON 字符串;客户端应以本字段为准
示例:
{
  "event_id": "opaque_event_id",
  "type": "response.mcp_call_arguments.done",
  "response_id": "opaque_response_id",
  "item_id": "opaque_item_id",
  "output_index": 0,
  "arguments": "{\"city\":\"杭州\"}"
}

response.mcp_call.*

事件 type触发时机
response.mcp_call.in_progress参数生成完成,且审批通过或无需审批,开始调用 MCP Server
response.mcp_call.completedMCP 工具调用成功
response.mcp_call.failed连接、协议、工具业务错误、审批拒绝、审批超时、调用超时或取消
三个事件均使用以下字段:
字段路径类型必填说明
event_idstring是本条服务端事件的唯一 ID
typestring是上表中的固定事件类型
item_idstring是当前 mcp_call item ID
output_indexinteger是当前 item 在父 Response output 数组中的零基索引
示例:
{
  "event_id": "opaque_event_id",
  "type": "response.mcp_call.in_progress",
  "item_id": "opaque_item_id",
  "output_index": 0
}

MCP 对话项

以下对象是事件中 item 的取值类型,适用于 Qwen3.8-Omni-Flash-Realtime。

mcp_list_tools

该对象通过 conversation.item.created.item 返回;字段见 conversation.item.created 的 type=mcp_list_tools 分支。 示例:
{
  "event_id": "opaque_event_id",
  "type": "conversation.item.created",
  "item": {
    "id": "opaque_item_id",
    "type": "mcp_list_tools",
    "server_label": "amap",
    "tools": [
      {
        "name": "maps_weather",
        "description": "查询天气",
        "input_schema": {
          "type": "object",
          "properties": {
            "city": {
              "type": "string",
              "description": "城市名"
            }
          },
          "required": ["city"]
        }
      }
    ]
  }
}

初始 mcp_call

模型选择 MCP 工具时,该对象出现在 response.output_item.added.item 中,也可能通过 conversation.item.created.item 返回;字段见 response.output_item.added 和 conversation.item.created 的对应 item 分支。 response.output_item.added 示例:
{
  "event_id": "opaque_event_id",
  "type": "response.output_item.added",
  "response_id": "opaque_response_id",
  "output_index": 0,
  "item": {
    "id": "opaque_item_id",
    "object": "realtime.item",
    "type": "mcp_call",
    "status": "in_progress",
    "call_id": "opaque_call_id",
    "server_label": "amap",
    "name": "maps_weather",
    "arguments": ""
  }
}

mcp_approval_request

该对象通过 conversation.item.created 返回;事件顶层字段和审批请求 item 字段见 conversation.item.created。 示例:
{
  "event_id": "opaque_event_id",
  "type": "conversation.item.created",
  "response_id": "opaque_response_id",
  "item_id": "opaque_approval_id",
  "previous_item_id": "opaque_item_id",
  "item": {
    "id": "opaque_approval_id",
    "type": "mcp_approval_request",
    "server_label": "amap",
    "name": "maps_weather",
    "arguments": "{\"city\":\"杭州\"}",
    "call_id": "opaque_call_id"
  }
}

最终 mcp_call

MCP 调用进入终态后,该对象出现在 response.output_item.done.item 中,并进入父 response.done.response.output 的最终快照;字段见 response.output_item.done 和 response.done 的对应分支。 成功示例:
{
  "event_id": "opaque_event_id",
  "type": "response.output_item.done",
  "response_id": "opaque_response_id",
  "output_index": 0,
  "item": {
    "id": "opaque_item_id",
    "type": "mcp_call",
    "status": "completed",
    "call_id": "opaque_call_id",
    "server_label": "amap",
    "name": "maps_weather",
    "arguments": "{\"city\":\"杭州\"}",
    "output": "{\"content\":[{\"type\":\"text\",\"text\":\"...\"}],\"isError\":false}"
  }
}
失败示例:
{
  "event_id": "opaque_event_id",
  "type": "response.output_item.done",
  "response_id": "opaque_response_id",
  "output_index": 0,
  "item": {
    "id": "opaque_item_id",
    "type": "mcp_call",
    "status": "failed",
    "call_id": "opaque_call_id",
    "server_label": "amap",
    "name": "maps_weather",
    "arguments": "{\"city\":\"杭州\"}",
    "error": {
      "type": "tool_execution_error",
      "message": "MCP tool call failed (call_timeout)."
    }
  }
}

MCP 错误

MCP 错误可能发生在工具发现或工具调用阶段。工具发现失败时的 error 字段见 conversation.item.created 的 mcp_list_tools 分支;工具调用失败时见 response.output_item.done 的 mcp_call 分支。
相关文档:实时(Qwen-Omni-Realtime)。