Skip to main content
POST
cURL
This endpoint does not return an Operation. It responds as soon as the voice is queued, with the voice resource itself (state: "PVC_VOICE_STATE_QUEUED"). Poll Get a PVC voice to track progress — there’s no separate operation ID to look up.
Requires at least 600 seconds (10 minutes) of cumulative sample audio, measured after any trims are applied. Calling Train again while the voice is already PVC_VOICE_STATE_QUEUED or PVC_VOICE_STATE_TRAINING is a no-op — it returns the current voice unchanged rather than starting a second run. Training starts are rate-limited per plan; exceeding your plan’s rate returns 429.

How long training takes

Wall time scales with how much audio you uploaded, not how many samples it’s split across. As a rough guide, training an hour of audio takes on the order of 10-30 minutes; expect it to take longer under heavy platform demand.
Underlying model training is capped at roughly 2800 seconds of audio and 1000 MB of source data per run — audio beyond that is silently truncated rather than rejected. Keep total sample audio comfortably under this if you’re uploading long-form recordings.

Authorizations

Authorization
string
header
required

Your API key. Read permissions are required for GET endpoints. Write permissions are required for POST, PATCH, and DELETE endpoints.

For Basic authentication, please populate Basic $INWORLD_API_KEY. You can create a key in one command with the Inworld CLI: inworld workspace add-key.

Path Parameters

voiceId
string
required

Voice ID of the draft PVC voice to train.

Body

application/json

The body is of type object.

Response

A successful response.

A Professional Voice Clone resource.

name
string
read-only

Resource name. Format: workspaces/{workspace}/pvcVoices/{voice}.

voiceId
string
read-only

Voice ID, derived from displayName at creation time. Use this value as {voiceId} on every other PVC endpoint, and as the voiceId in TTS synthesis requests once the voice is PVC_VOICE_STATE_READY.

displayName
string

The human-readable name shown anywhere the voice is listed or selected.

languageCode
string

The voice's language as a BCP-47-shaped locale string, e.g. en-US. Immutable after creation.

state
enum<string>

Lifecycle state of a PVC voice.

  • PVC_VOICE_STATE_DRAFT: Editable. Samples can be added, trimmed, or removed, and metadata can be updated.
  • PVC_VOICE_STATE_QUEUED: Training requested; waiting for a training slot.
  • PVC_VOICE_STATE_TRAINING: Actively training.
  • PVC_VOICE_STATE_READY: Training succeeded. Usable for TTS synthesis; permanent — cannot be deleted through this API.
  • PVC_VOICE_STATE_FAILED: Training failed. Editing the voice (e.g. renaming it, or adding/removing a sample) returns it to PVC_VOICE_STATE_DRAFT with its remaining samples intact.
Available options:
PVC_VOICE_STATE_UNSPECIFIED,
PVC_VOICE_STATE_DRAFT,
PVC_VOICE_STATE_QUEUED,
PVC_VOICE_STATE_TRAINING,
PVC_VOICE_STATE_READY,
PVC_VOICE_STATE_FAILED
failure
object

Populated on a PVC voice when its state is PVC_VOICE_STATE_FAILED.

incarnationId
string
read-only

Identifier that stays stable across edits to the same voice, and changes each time it is retrained. Use it to tell two reads of the same voiceId apart across a retrain.

samples
object[]
read-only

Audio samples currently attached to the voice.

createTime
string<date-time>
read-only
updateTime
string<date-time>
read-only