> ## Documentation Index
> Fetch the complete documentation index at: https://apidoc.cometapi.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Create Speech

> Usa CometAPI POST /v1/audio/speech para convertir texto en audio con un modelo de texto a voz y un formato de salida seleccionados.

Usa este endpoint para convertir texto en un archivo de audio mediante la API de audio compatible con OpenAI. Es adecuado para narración, indicaciones de voz cortas, funciones de lectura en voz alta y otros flujos de trabajo en los que tu aplicación ya tiene texto y necesita salida de voz.

## Primera solicitud

Comienza con tres campos: `model`, `input` y `voice`. Mantén la primera solicitud breve para que puedas verificar la autenticación, el formato de audio y el manejo de archivos antes de ajustar la velocidad o el formato de salida.

## Lee la respuesta

La respuesta es audio binario, no JSON. En los ejemplos de SDK, escribe la respuesta en un archivo como `output.mp3`. En clientes HTTP directos, guarda el cuerpo de la respuesta y establece la extensión del archivo para que coincida con el `response_format` solicitado.

## Siguientes pasos

* Usa [Create Transcription](/api/audio/create-transcription) cuando necesites convertir voz de nuevo en texto.
* Usa [Create Translation](/api/audio/create-translation) cuando necesites texto en inglés a partir de audio que no esté en inglés.


## OpenAPI

````yaml api/openapi/audio/post-create-speech.openapi.json POST /v1/audio/speech
openapi: 3.1.0
info:
  title: Create speech API
  version: 1.0.0
servers:
  - url: https://api.cometapi.com
security:
  - bearerAuth: []
paths:
  /v1/audio/speech:
    post:
      summary: Create speech
      operationId: create_speech
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              required:
                - model
                - input
                - voice
              properties:
                model:
                  type: string
                  description: >-
                    The TTS model to use. Choose a current speech model from the
                    [Models page](/overview/models).
                  default: tts-1
                input:
                  type: string
                  description: >-
                    The text to generate audio for. Maximum length is 4096
                    characters.
                  maxLength: 4096
                voice:
                  type: string
                  description: The voice to use for speech synthesis.
                  enum:
                    - alloy
                    - ash
                    - ballad
                    - coral
                    - echo
                    - fable
                    - onyx
                    - nova
                    - sage
                    - shimmer
                  default: alloy
                response_format:
                  type: string
                  description: The audio output format.
                  enum:
                    - mp3
                    - opus
                    - aac
                    - flac
                    - wav
                    - pcm
                  default: mp3
                speed:
                  type: number
                  description: >-
                    The speed of the generated audio. Select a value between
                    0.25 and 4.0.
                  minimum: 0.25
                  maximum: 4
                  default: 1
            examples:
              Default:
                summary: Standard TTS (tts-1)
                value:
                  model: tts-1
                  input: The quick brown fox jumped over the lazy dog.
                  voice: alloy
              gpt_4o_mini_tts:
                summary: GPT-4o mini TTS (gpt-4o-mini-tts)
                value:
                  model: gpt-4o-mini-tts
                  input: The quick brown fox jumped over the lazy dog.
                  voice: alloy
      responses:
        '200':
          description: The audio file content.
          content:
            audio/mpeg:
              schema:
                type: string
                format: binary
      x-codeSamples:
        - lang: python
          label: Create speech
          source: |-
            import os
            from openai import OpenAI

            client = OpenAI(
                api_key=os.environ["COMETAPI_KEY"],
                base_url="https://api.cometapi.com/v1"
            )

            response = client.audio.speech.create(
                model="tts-1",
                voice="alloy",
                input="The quick brown fox jumped over the lazy dog."
            )

            response.stream_to_file("output.mp3")
        - lang: javascript
          label: Create speech
          source: |-
            import OpenAI from "openai";
            import fs from "fs";

            const client = new OpenAI({
              apiKey: process.env.COMETAPI_KEY,
              baseURL: "https://api.cometapi.com/v1"
            });

            const response = await client.audio.speech.create({
              model: "tts-1",
              voice: "alloy",
              input: "The quick brown fox jumped over the lazy dog."
            });

            const buffer = Buffer.from(await response.arrayBuffer());
            fs.writeFileSync("output.mp3", buffer);
        - lang: shell
          label: Create speech
          source: |-
            curl -X POST https://api.cometapi.com/v1/audio/speech \
              -H "Authorization: Bearer $COMETAPI_KEY" \
              -H "Content-Type: application/json" \
              -d '{
                "model": "tts-1",
                "input": "The quick brown fox jumped over the lazy dog.",
                "voice": "alloy"
              }' \
              --output output.mp3
components:
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: Bearer token authentication. Use your CometAPI key.

````