Skip to main content
POST
Python
Sử dụng endpoint này để chuyển âm thanh thành văn bản bằng ngôn ngữ gốc. Endpoint này phù hợp cho ghi chú cuộc họp, tin nhắn thoại, lập chỉ mục media, phụ đề và các quy trình hỗ trợ cần văn bản có thể tìm kiếm.

Yêu cầu đầu tiên

Gửi một tệp âm thanh được hỗ trợ cùng với modelfile. Hãy giữ tệp đầu tiên ngắn trong khi bạn xác thực việc xử lý tải lên, xác thực và phân tích phản hồi.

Đọc phản hồi

Phản hồi mặc định bao gồm text đã được phiên âm. Nếu bạn yêu cầu định dạng phản hồi khác, hãy đảm bảo client của bạn phân tích đúng định dạng đó thay vì giả định cấu trúc JSON mặc định.

Các bước tiếp theo

  • Sử dụng Create Speech khi bạn cần đầu ra chuyển văn bản thành giọng nói.
  • Sử dụng Create Translation khi đầu ra đích cần là tiếng Anh.

Ủy quyền

Authorization
string
header
bắt buộc

Bearer token authentication. Use your CometAPI key.

Nội dung

multipart/form-data
file
file
bắt buộc

The audio file to transcribe. Supported formats: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm.

model
string
mặc định:whisper-1
bắt buộc

The speech-to-text model to use. Choose a current speech model from the Models page.

language
string

The language of the input audio in ISO-639-1 format (e.g., en, zh, ja). Supplying the language improves accuracy and latency.

prompt
string

Optional text to guide the model's style or continue a previous audio segment. The prompt should match the audio language.

response_format
enum<string>
mặc định:json

The output format for the transcription.

Tùy chọn có sẵn:
json,
text,
srt,
verbose_json,
vtt
temperature
number
mặc định:0

Sampling temperature between 0 and 1. Higher values produce more random output; lower values are more focused. When set to 0, the model auto-adjusts temperature using log probability.

Phạm vi bắt buộc: 0 <= x <= 1

Phản hồi

200 - application/json

The transcription result.

text
string
bắt buộc

The transcribed text.