> ## Documentation Index
> Fetch the complete documentation index at: https://apidoc.cometapi.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Create Transcription

> Sử dụng CometAPI POST /v1/audio/transcriptions để chuyển âm thanh thành văn bản bằng model phiên âm và định dạng phản hồi đã chọn.

Sử dụng endpoint này để chuyển âm thanh thành văn bản bằng ngôn ngữ gốc. Endpoint này phù hợp cho ghi chú cuộc họp, tin nhắn thoại, lập chỉ mục media, phụ đề và các quy trình hỗ trợ cần văn bản có thể tìm kiếm.

## Yêu cầu đầu tiên

Gửi một tệp âm thanh được hỗ trợ cùng với `model` và `file`. Hãy giữ tệp đầu tiên ngắn trong khi bạn xác thực việc xử lý tải lên, xác thực và phân tích phản hồi.

## Đọc phản hồi

Phản hồi mặc định bao gồm `text` đã được phiên âm. Nếu bạn yêu cầu định dạng phản hồi khác, hãy đảm bảo client của bạn phân tích đúng định dạng đó thay vì giả định cấu trúc JSON mặc định.

## Các bước tiếp theo

* Sử dụng [Create Speech](/api/audio/create-speech) khi bạn cần đầu ra chuyển văn bản thành giọng nói.
* Sử dụng [Create Translation](/api/audio/create-translation) khi đầu ra đích cần là tiếng Anh.


## OpenAPI

````yaml api/openapi/audio/post-create-transcription.openapi.json POST /v1/audio/transcriptions
openapi: 3.1.0
info:
  title: Create transcription API
  version: 1.0.0
servers:
  - url: https://api.cometapi.com
security:
  - bearerAuth: []
paths:
  /v1/audio/transcriptions:
    post:
      summary: Create transcription
      operationId: create_transcription
      requestBody:
        required: true
        content:
          multipart/form-data:
            schema:
              type: object
              properties:
                file:
                  format: binary
                  type: string
                  description: >-
                    The audio file to transcribe. Supported formats: flac, mp3,
                    mp4, mpeg, mpga, m4a, ogg, wav, webm.
                model:
                  type: string
                  description: >-
                    The speech-to-text model to use. Choose a current speech
                    model from the [Models page](/overview/models).
                  default: whisper-1
                language:
                  type: string
                  description: >-
                    The language of the input audio in ISO-639-1 format (e.g.,
                    `en`, `zh`, `ja`). Supplying the language improves accuracy
                    and latency.
                prompt:
                  type: string
                  description: >-
                    Optional text to guide the model's style or continue a
                    previous audio segment. The prompt should match the audio
                    language.
                response_format:
                  type: string
                  description: The output format for the transcription.
                  enum:
                    - json
                    - text
                    - srt
                    - verbose_json
                    - vtt
                  default: json
                temperature:
                  type: number
                  description: >-
                    Sampling temperature between 0 and 1. Higher values produce
                    more random output; lower values are more focused. When set
                    to 0, the model auto-adjusts temperature using log
                    probability.
                  minimum: 0
                  maximum: 1
                  default: 0
              required:
                - file
                - model
      responses:
        '200':
          description: The transcription result.
          content:
            application/json:
              schema:
                type: object
                required:
                  - text
                properties:
                  text:
                    type: string
                    description: The transcribed text.
              examples:
                Default:
                  summary: Transcription result
                  value:
                    text: Hello, welcome to CometAPI.
      x-codeSamples:
        - lang: python
          label: Create transcription
          source: |-
            import os
            from openai import OpenAI

            client = OpenAI(
                api_key=os.environ["COMETAPI_KEY"],
                base_url="https://api.cometapi.com/v1"
            )

            audio_file = open("audio.mp3", "rb")
            transcription = client.audio.transcriptions.create(
                model="whisper-1",
                file=audio_file
            )
            print(transcription.text)
        - lang: javascript
          label: Create transcription
          source: |-
            import OpenAI from "openai";
            import fs from "fs";

            const client = new OpenAI({
              apiKey: process.env.COMETAPI_KEY,
              baseURL: "https://api.cometapi.com/v1"
            });

            const transcription = await client.audio.transcriptions.create({
              model: "whisper-1",
              file: fs.createReadStream("audio.mp3")
            });
            console.log(transcription.text);
        - lang: shell
          label: Create transcription
          source: |-
            curl -X POST https://api.cometapi.com/v1/audio/transcriptions \
              -H "Authorization: Bearer $COMETAPI_KEY" \
              -F model="whisper-1" \
              -F file="@audio.mp3"
components:
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: Bearer token authentication. Use your CometAPI key.

````