채팅 완성 생성
CometAPI POST /v1/chat/completions를 사용하여 스트리밍, temperature 및 max_tokens 제어 기능으로 여러 메시지로 구성된 대화를 채팅 모델에 전송합니다.
model 매개변수를 변경하여 모델을 전환할 수 있으며, 대부분의 OpenAI 호환 SDK는 base_url을 https://api.cometapi.com/v1으로 설정하면 작동합니다.
메시지 역할
멀티모달 입력 전송
많은 모델이 텍스트와 함께 이미지 및 오디오를 지원합니다. 멀티모달 메시지를 전송하려면content에 배열 형식을 사용하세요:
detail 매개변수는 이미지 분석 깊이를 제어합니다:
low— 더 빠르고 더 적은 토큰을 사용합니다(고정 비용)high— 세부 분석, 더 많은 토큰 소비auto— 모델이 결정합니다(기본값)
응답 스트리밍
증분 출력을 수신하려면stream을 true으로 설정하세요. 응답은 다음 형식으로 전달됩니다 서버 전송 이벤트(SSE), 각 이벤트에는 chat.completion.chunk 객체가 포함됩니다:
구조화된 출력 요청
모델이 특정 스키마와 일치하는 유효한 JSON을 반환하도록 강제하려면response_format을 사용하세요:
json_schema)는 출력이 스키마와 정확히 일치하도록 보장합니다. JSON Object 모드(json_object)는 유효한 JSON만 보장하며 구조는 강제하지 않습니다.도구 및 함수 호출
모델이 외부 함수를 호출하도록 하려면 도구 정의를 제공하세요:finish_reason: "tool_calls"이 포함되고 message.tool_calls 배열에는 함수 이름과 인수가 포함됩니다. 그런 다음 함수를 실행하고 일치하는 tool_call_id을 포함한 tool 메시지로 결과를 다시 전송합니다.
제공업체 간 참고 사항
제공업체별 매개변수 지원
제공업체별 매개변수 지원
max_tokens와 max_completion_tokens 비교
max_tokens와 max_completion_tokens 비교
max_tokens— 레거시 매개변수입니다. 대부분의 모델에서 작동하지만 최신 OpenAI 모델에서는 더 이상 사용되지 않습니다.max_completion_tokens— GPT-4.1, GPT-5 시리즈 및 o-series 모델에 권장되는 매개변수입니다. 출력 토큰과 추론 토큰을 모두 포함하므로 추론 모델에는 필수입니다.
system 역할과 developer 역할 비교
system 역할과 developer 역할 비교
system— 기존 지침 역할입니다. 모든 모델에서 작동합니다.developer— o1 모델과 함께 도입되었습니다. 최신 모델에서 더 강력한 지침 준수를 제공합니다. 이전 모델에서는system동작으로 대체됩니다.
developer을 사용하세요.FAQ
속도 제한은 어떻게 처리하나요?
429 Too Many Requests이 발생하면 지수 백오프를 구현하세요:
대화 컨텍스트는 어떻게 유지하나요?
전체 대화 기록을messages 배열에 포함하세요:
finish_reason은 무엇을의미하나요?
비용은 어떻게 제어하나요?
- 출력 길이를 제한하려면
max_completion_tokens을 사용하세요. - 지능과 비용의 균형을 위해
gpt-5.6-terra을 사용하거나, 효율적인 대량 워크로드에는gpt-5.6-luna을 사용하세요. - 프롬프트는 간결하게 유지하고 중복된 컨텍스트는 피하세요.
usage응답 필드에서 토큰 사용량을 모니터링하세요.
인증
Bearer token authentication. Use your CometAPI key.
본문
Model ID to use for this request. See the Models page for current options.
"gpt-4.1"
A list of messages forming the conversation. Each message has a role (system, user, assistant, or developer) and content (text string or multimodal content array).
If true, partial response tokens are delivered incrementally via server-sent events (SSE). The stream ends with a data: [DONE] message.
Sampling temperature between 0 and 2. Higher values (e.g., 0.8) produce more random output; lower values (e.g., 0.2) make output more focused and deterministic. Recommended to adjust this or top_p, but not both.
0 <= x <= 2Nucleus sampling parameter. The model considers only the tokens whose cumulative probability reaches top_p. For example, 0.1 means only the top 10% probability tokens are considered. Recommended to adjust this or temperature, but not both.
0 <= x <= 1Number of completion choices to generate for each input message. Defaults to 1.
Up to 4 sequences where the API will stop generating further tokens. Can be a string or an array of strings.
Maximum number of tokens to generate in the completion. The total of input + output tokens is capped by the model's context length.
Number between -2.0 and 2.0. Positive values penalize tokens based on whether they have already appeared, encouraging the model to explore new topics.
-2 <= x <= 2Number between -2.0 and 2.0. Positive values penalize tokens proportionally to how often they have appeared, reducing verbatim repetition.
-2 <= x <= 2A JSON object mapping token IDs to bias values from -100 to 100. The bias is added to the model's logits before sampling. Values between -1 and 1 subtly adjust likelihood; -100 or 100 effectively ban or force selection of a token.
A unique identifier for your end-user. Helps with abuse detection and monitoring.
An upper bound for the number of tokens to generate, including visible output tokens and reasoning tokens. Use this instead of max_tokens for GPT-4.1+, GPT-5 series, and o-series models.
Specifies the output format. Use {"type": "json_object"} for JSON mode, or {"type": "json_schema", "json_schema": {...}} for strict structured output.
A list of tools the model may call. Currently supports function type tools.
Controls how the model selects tools. auto (default): model decides. none: no tools. required: must call a tool.
Whether to return log probabilities of the output tokens.
Number of most likely tokens to return at each position (0-20). Requires logprobs to be true.
0 <= x <= 20Controls the reasoning effort for o-series and GPT-5.1+ models.
low, medium, high Options for streaming. Only valid when stream is true.
Specifies the processing tier.
auto, default, flex, priority 응답
Successful chat completion response.
Unique completion identifier.
"chatcmpl-abc123"
Object type. Non-streaming responses use chat.completion.
chat.completion "chat.completion"
Unix timestamp of creation.
1774412483
The model used (may include version suffix).
"gpt-5.4-2026-03-05"
Array of completion choices.
Token accounting for this request. Billing uses these counts.
Service tier that processed the request, when the provider reports one.
"default"
Provider backend configuration fingerprint, when the provider reports one.
"fp_490a4ad033"