Create Transcription
Sử dụng CometAPI POST /v1/audio/transcriptions để chuyển âm thanh thành văn bản bằng model phiên âm và định dạng phản hồi đã chọn.
Yêu cầu đầu tiên
Gửi một tệp âm thanh được hỗ trợ cùng vớimodel và file. Hãy giữ tệp đầu tiên ngắn trong khi bạn xác thực việc xử lý tải lên, xác thực và phân tích phản hồi.
Đọc phản hồi
Phản hồi mặc định bao gồmtext đã được phiên âm. Nếu bạn yêu cầu định dạng phản hồi khác, hãy đảm bảo client của bạn phân tích đúng định dạng đó thay vì giả định cấu trúc JSON mặc định.
Các bước tiếp theo
- Sử dụng Create Speech khi bạn cần đầu ra chuyển văn bản thành giọng nói.
- Sử dụng Create Translation khi đầu ra đích cần là tiếng Anh.
Ủy quyền
Bearer token authentication. Use your CometAPI key.
Nội dung
The audio file to transcribe. Supported formats: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm.
The speech-to-text model to use. Choose a current speech model from the Models page.
The language of the input audio in ISO-639-1 format (e.g., en, zh, ja). Supplying the language improves accuracy and latency.
Optional text to guide the model's style or continue a previous audio segment. The prompt should match the audio language.
The output format for the transcription.
json, text, srt, verbose_json, vtt Sampling temperature between 0 and 1. Higher values produce more random output; lower values are more focused. When set to 0, the model auto-adjusts temperature using log probability.
0 <= x <= 1Phản hồi
The transcription result.
The transcribed text.