チャット補完を作成
CometAPI の POST /v1/chat/completions を使用して、ストリーミング、temperature、max_tokens を制御しながら、複数メッセージの会話をチャットモデルに送信します。
model パラメータを変更してモデルを切り替えます。ほとんどの OpenAI 互換 SDK は、base_url を https://api.cometapi.com/v1 に設定することで使用できます。
メッセージのロール
マルチモーダル入力を送信する
多くのモデルは、テキストに加えて画像と音声をサポートしています。マルチモーダルメッセージを送信するには、content に配列形式を使用します。
detail パラメータは画像分析の詳細度を制御します。
low— より高速で、使用トークン数が少ない(固定コスト)high— 詳細な分析で、より多くのトークンを消費auto— モデルが決定(デフォルト)
レスポンスをストリーミングする
段階的な出力を受け取るには、stream を true に設定します。レスポンスは Server-Sent Events(SSE))として配信され、各イベントには chat.completion.chunk オブジェクトが含まれます。
構造化出力をリクエストする
特定のスキーマに一致する有効な JSON をモデルに返させるには、response_format を使用します。
json_schema)では、出力がスキーマに完全に一致することが保証されます。JSON Object モード(json_object)では有効な JSON のみが保証され、構造は強制されません。ツールと関数を呼び出す
モデルが外部関数を呼び出せるようにするには、ツール定義を指定します。finish_reason: "tool_calls" が含まれ、message.tool_calls 配列には関数名と引数が格納されます。次に関数を実行し、一致する tool_call_id を含む tool メッセージとして結果を送り返します。
プロバイダー間の注意事項
プロバイダー間のパラメータ対応状況
プロバイダー間のパラメータ対応状況
max_tokens と max_completion_tokens
max_tokens と max_completion_tokens
max_tokens— 従来のパラメータです。ほとんどのモデルで動作しますが、新しい OpenAI モデルでは非推奨です。max_completion_tokens— GPT-4.1、GPT-5 シリーズ、および o-series モデルに推奨されるパラメータです。出力トークンと推論トークンの両方を含むため、推論モデルでは必須です。
system と developer のロール
system と developer のロール
system— 従来の指示用ロールです。すべてのモデルで動作します。developer— o1 モデルで導入されました。新しいモデルに対してより強力な指示追従を提供します。古いモデルではsystemの動作にフォールバックします。
developer を使用してください。よくある質問
レート制限を処理するには?
429 Too Many Requests が発生した場合は、指数バックオフを実装してください。
会話コンテキストを維持するには?
会話履歴全体をmessages 配列に含めます。
finish_reason とは何を意味しますか?
コストを管理するには?
- 出力長を制限するには
max_completion_tokensを使用します。 - 知能とコストのバランスには
gpt-5.6-terraを使用し、効率的な大量ワークロードにはgpt-5.6-lunaを使用します。 - プロンプトは簡潔にし、重複したコンテキストを避けてください。
usageレスポンスフィールドでトークン使用量を監視します。
承認
Bearer token authentication. Use your CometAPI key.
ボディ
Model ID to use for this request. See the Models page for current options.
"gpt-4.1"
A list of messages forming the conversation. Each message has a role (system, user, assistant, or developer) and content (text string or multimodal content array).
If true, partial response tokens are delivered incrementally via server-sent events (SSE). The stream ends with a data: [DONE] message.
Sampling temperature between 0 and 2. Higher values (e.g., 0.8) produce more random output; lower values (e.g., 0.2) make output more focused and deterministic. Recommended to adjust this or top_p, but not both.
0 <= x <= 2Nucleus sampling parameter. The model considers only the tokens whose cumulative probability reaches top_p. For example, 0.1 means only the top 10% probability tokens are considered. Recommended to adjust this or temperature, but not both.
0 <= x <= 1Number of completion choices to generate for each input message. Defaults to 1.
Up to 4 sequences where the API will stop generating further tokens. Can be a string or an array of strings.
Maximum number of tokens to generate in the completion. The total of input + output tokens is capped by the model's context length.
Number between -2.0 and 2.0. Positive values penalize tokens based on whether they have already appeared, encouraging the model to explore new topics.
-2 <= x <= 2Number between -2.0 and 2.0. Positive values penalize tokens proportionally to how often they have appeared, reducing verbatim repetition.
-2 <= x <= 2A JSON object mapping token IDs to bias values from -100 to 100. The bias is added to the model's logits before sampling. Values between -1 and 1 subtly adjust likelihood; -100 or 100 effectively ban or force selection of a token.
A unique identifier for your end-user. Helps with abuse detection and monitoring.
An upper bound for the number of tokens to generate, including visible output tokens and reasoning tokens. Use this instead of max_tokens for GPT-4.1+, GPT-5 series, and o-series models.
Specifies the output format. Use {"type": "json_object"} for JSON mode, or {"type": "json_schema", "json_schema": {...}} for strict structured output.
A list of tools the model may call. Currently supports function type tools.
Controls how the model selects tools. auto (default): model decides. none: no tools. required: must call a tool.
Whether to return log probabilities of the output tokens.
Number of most likely tokens to return at each position (0-20). Requires logprobs to be true.
0 <= x <= 20Controls the reasoning effort for o-series and GPT-5.1+ models.
low, medium, high Options for streaming. Only valid when stream is true.
Specifies the processing tier.
auto, default, flex, priority レスポンス
Successful chat completion response.
Unique completion identifier.
"chatcmpl-abc123"
Object type. Non-streaming responses use chat.completion.
chat.completion "chat.completion"
Unix timestamp of creation.
1774412483
The model used (may include version suffix).
"gpt-5.4-2026-03-05"
Array of completion choices.
Token accounting for this request. Billing uses these counts.
Service tier that processed the request, when the provider reports one.
"default"
Provider backend configuration fingerprint, when the provider reports one.
"fp_490a4ad033"