Skip to main content
POST
CometAPI 通过单一的 OpenAI 兼容接口,将聊天补全路由到多个提供商,包括 OpenAI、Claude 和 Gemini。通过更改 model 参数即可切换模型;只需将 base_url 设置为 https://api.cometapi.com/v1,大多数 OpenAI 兼容 SDK 都可使用。
不同模型提供商之间的请求参数和响应字段可能存在显著差异。当需要完整的参数列表或提供商特定行为时,请查阅所用模型背后提供商的官方文档。例如,reasoning_effort 仅适用于推理模型(o-series、GPT-5.1+),且部分模型不支持 logprobsn > 1。
对于 OpenAI Pro 模型、o-series 推理模型和 Codex 模型,请改用 响应 端点。这些模型系列在响应 API 上的支持更完整。

消息角色

对于较新模型(GPT-4.1、GPT-5 系列和 o-series),指令消息应优先使用 developer 而非 system。两者均可使用,但 developer 提供更强的指令遵循能力。

发送多模态输入

许多模型除文本外还支持图像和音频。要发送多模态消息,请对 content 使用数组格式:
detail 参数控制图像分析深度:
  • low — 更快,使用更少的 Token(固定成本)
  • high — 详细分析,消耗更多 Token
  • auto — 由模型决定(默认)

流式传输响应

要接收增量输出,请将 stream 设为 true。响应将以 服务器发送事件(SSE),其中每个事件包含一个 chat.completion.chunk 对象:
要在流式响应中包含 Token 用量统计,请将 stream_options.include_usage 设为 true。用量数据会出现在 [DONE] 之前的最后一个块中。

请求结构化输出

要强制模型返回与特定 schema 匹配的有效 JSON,请使用 response_format
JSON Schema 模式(json_schema)保证输出与您的 schema 完全匹配。JSON Object 模式(json_object)仅保证返回有效 JSON,不强制其结构。

调用工具和函数

要使模型能够调用外部函数,请提供工具定义:
当模型决定调用工具时,响应将包含 finish_reason: "tool_calls",且 message.tool_calls 数组将包含函数名称和参数。随后执行该函数,并将结果作为带有匹配 tool_call_idtool 消息发回。

跨提供商说明

  • max_tokens — 旧版参数。适用于大多数模型,但已不推荐用于较新的 OpenAI 模型。
  • max_completion_tokens — 推荐用于 GPT-4.1、GPT-5 系列和 o-series 模型的参数。推理模型必须使用该参数,因为它同时包含输出 Token 和推理 Token。
CometAPI 会在路由到不同提供商时自动处理映射。
  • system — 传统的指令角色。适用于所有模型。
  • developer — 随 o1 模型引入。为较新模型提供更强的指令遵循能力。在较旧模型上会回退为 system 行为。
对于面向 GPT-4.1+ 或 o-series 模型的新项目,请使用 developer

常见问题

如何处理速率限制?

遇到 429 Too Many Requests 时,请实施指数退避:

如何维护对话上下文?

messages 数组中包含完整的对话历史:

finish_reason 的含义是什么?

如何控制成本?

  1. 使用 max_completion_tokens 限制输出长度。
  2. 使用 gpt-5.6-terra 平衡智能性与成本,或使用 gpt-5.6-luna 处理高效的大规模工作负载。
  3. 保持 Prompt 简洁,避免冗余上下文。
  4. usage 响应字段中监控 Token 用量。

授权

Authorization
string
header
必填

Bearer token authentication. Use your CometAPI key.

请求体

application/json
model
string
默认值:gpt-5.6-sol
必填

Model ID to use for this request. See the Models page for current options.

示例:

"gpt-4.1"

messages
object[]
必填

A list of messages forming the conversation. Each message has a role (system, user, assistant, or developer) and content (text string or multimodal content array).

stream
boolean

If true, partial response tokens are delivered incrementally via server-sent events (SSE). The stream ends with a data: [DONE] message.

temperature
number
默认值:1

Sampling temperature between 0 and 2. Higher values (e.g., 0.8) produce more random output; lower values (e.g., 0.2) make output more focused and deterministic. Recommended to adjust this or top_p, but not both.

必填范围: 0 <= x <= 2
top_p
number
默认值:1

Nucleus sampling parameter. The model considers only the tokens whose cumulative probability reaches top_p. For example, 0.1 means only the top 10% probability tokens are considered. Recommended to adjust this or temperature, but not both.

必填范围: 0 <= x <= 1
n
integer
默认值:1

Number of completion choices to generate for each input message. Defaults to 1.

stop
string

Up to 4 sequences where the API will stop generating further tokens. Can be a string or an array of strings.

max_tokens
integer

Maximum number of tokens to generate in the completion. The total of input + output tokens is capped by the model's context length.

presence_penalty
number
默认值:0

Number between -2.0 and 2.0. Positive values penalize tokens based on whether they have already appeared, encouraging the model to explore new topics.

必填范围: -2 <= x <= 2
frequency_penalty
number
默认值:0

Number between -2.0 and 2.0. Positive values penalize tokens proportionally to how often they have appeared, reducing verbatim repetition.

必填范围: -2 <= x <= 2
logit_bias
object

A JSON object mapping token IDs to bias values from -100 to 100. The bias is added to the model's logits before sampling. Values between -1 and 1 subtly adjust likelihood; -100 or 100 effectively ban or force selection of a token.

user
string

A unique identifier for your end-user. Helps with abuse detection and monitoring.

max_completion_tokens
integer

An upper bound for the number of tokens to generate, including visible output tokens and reasoning tokens. Use this instead of max_tokens for GPT-4.1+, GPT-5 series, and o-series models.

response_format
object

Specifies the output format. Use {"type": "json_object"} for JSON mode, or {"type": "json_schema", "json_schema": {...}} for strict structured output.

tools
object[]

A list of tools the model may call. Currently supports function type tools.

tool_choice
默认值:auto

Controls how the model selects tools. auto (default): model decides. none: no tools. required: must call a tool.

logprobs
boolean
默认值:false

Whether to return log probabilities of the output tokens.

top_logprobs
integer

Number of most likely tokens to return at each position (0-20). Requires logprobs to be true.

必填范围: 0 <= x <= 20
reasoning_effort
enum<string>

Controls the reasoning effort for o-series and GPT-5.1+ models.

可用选项:
low,
medium,
high
stream_options
object

Options for streaming. Only valid when stream is true.

service_tier
enum<string>

Specifies the processing tier.

可用选项:
auto,
default,
flex,
priority

响应

Successful chat completion response.

id
string

Unique completion identifier.

示例:

"chatcmpl-abc123"

object
enum<string>

Object type. Non-streaming responses use chat.completion.

可用选项:
chat.completion
示例:

"chat.completion"

created
integer

Unix timestamp of creation.

示例:

1774412483

model
string

The model used (may include version suffix).

示例:

"gpt-5.4-2026-03-05"

choices
object[]

Array of completion choices.

usage
object

Token accounting for this request. Billing uses these counts.

service_tier
string

Service tier that processed the request, when the provider reports one.

示例:

"default"

system_fingerprint
string | null

Provider backend configuration fingerprint, when the provider reports one.

示例:

"fp_490a4ad033"