> ## Documentation Index
> Fetch the complete documentation index at: https://apidoc.cometapi.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Create Speech

> Gunakan CometAPI POST /v1/audio/speech untuk mengonversi teks menjadi audio dengan model text-to-speech dan format output yang dipilih.

Gunakan endpoint ini untuk mengubah teks menjadi file audio melalui audio API yang kompatibel dengan OpenAI. Endpoint ini cocok untuk narasi, prompt suara singkat, fitur pembacaan nyaring, dan alur kerja lain ketika aplikasi Anda sudah memiliki teks dan memerlukan output suara.

## Permintaan pertama

Mulailah dengan tiga field: `model`, `input`, dan `voice`. Buat permintaan pertama tetap singkat agar Anda dapat memverifikasi autentikasi, format audio, dan penanganan file sebelum menyesuaikan kecepatan atau format output.

## Membaca respons

Respons berupa audio biner, bukan JSON. Dalam contoh SDK, tulis respons ke file seperti `output.mp3`. Dalam klien HTTP langsung, simpan body respons dan atur ekstensi file agar sesuai dengan `response_format` yang diminta.

## Langkah berikutnya

* Gunakan [Create Transcription](/api/audio/create-transcription) ketika Anda perlu mengubah suara kembali menjadi teks.
* Gunakan [Create Translation](/api/audio/create-translation) ketika Anda memerlukan teks bahasa Inggris dari audio non-Inggris.


## OpenAPI

````yaml api/openapi/audio/post-create-speech.openapi.json POST /v1/audio/speech
openapi: 3.1.0
info:
  title: Create speech API
  version: 1.0.0
servers:
  - url: https://api.cometapi.com
security:
  - bearerAuth: []
paths:
  /v1/audio/speech:
    post:
      summary: Create speech
      operationId: create_speech
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              required:
                - model
                - input
                - voice
              properties:
                model:
                  type: string
                  description: >-
                    The TTS model to use. Choose a current speech model from the
                    [Models page](/overview/models).
                  default: tts-1
                input:
                  type: string
                  description: >-
                    The text to generate audio for. Maximum length is 4096
                    characters.
                  maxLength: 4096
                voice:
                  type: string
                  description: The voice to use for speech synthesis.
                  enum:
                    - alloy
                    - ash
                    - ballad
                    - coral
                    - echo
                    - fable
                    - onyx
                    - nova
                    - sage
                    - shimmer
                  default: alloy
                response_format:
                  type: string
                  description: The audio output format.
                  enum:
                    - mp3
                    - opus
                    - aac
                    - flac
                    - wav
                    - pcm
                  default: mp3
                speed:
                  type: number
                  description: >-
                    The speed of the generated audio. Select a value between
                    0.25 and 4.0.
                  minimum: 0.25
                  maximum: 4
                  default: 1
            examples:
              Default:
                summary: Standard TTS (tts-1)
                value:
                  model: tts-1
                  input: The quick brown fox jumped over the lazy dog.
                  voice: alloy
              gpt_4o_mini_tts:
                summary: GPT-4o mini TTS (gpt-4o-mini-tts)
                value:
                  model: gpt-4o-mini-tts
                  input: The quick brown fox jumped over the lazy dog.
                  voice: alloy
      responses:
        '200':
          description: The audio file content.
          content:
            audio/mpeg:
              schema:
                type: string
                format: binary
      x-codeSamples:
        - lang: python
          label: Create speech
          source: |-
            import os
            from openai import OpenAI

            client = OpenAI(
                api_key=os.environ["COMETAPI_KEY"],
                base_url="https://api.cometapi.com/v1"
            )

            response = client.audio.speech.create(
                model="tts-1",
                voice="alloy",
                input="The quick brown fox jumped over the lazy dog."
            )

            response.stream_to_file("output.mp3")
        - lang: javascript
          label: Create speech
          source: |-
            import OpenAI from "openai";
            import fs from "fs";

            const client = new OpenAI({
              apiKey: process.env.COMETAPI_KEY,
              baseURL: "https://api.cometapi.com/v1"
            });

            const response = await client.audio.speech.create({
              model: "tts-1",
              voice: "alloy",
              input: "The quick brown fox jumped over the lazy dog."
            });

            const buffer = Buffer.from(await response.arrayBuffer());
            fs.writeFileSync("output.mp3", buffer);
        - lang: shell
          label: Create speech
          source: |-
            curl -X POST https://api.cometapi.com/v1/audio/speech \
              -H "Authorization: Bearer $COMETAPI_KEY" \
              -H "Content-Type: application/json" \
              -d '{
                "model": "tts-1",
                "input": "The quick brown fox jumped over the lazy dog.",
                "voice": "alloy"
              }' \
              --output output.mp3
components:
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: Bearer token authentication. Use your CometAPI key.

````