Fields
| Field | Type | Required | Description | Example |
|---|---|---|---|---|
Channel | *int64 | :heavy_minus_sign: | Zero-based audio channel index for the word, present when the provider transcribes channels separately | 0 |
Confidence | *float64 | :heavy_minus_sign: | Provider confidence for the word from 0 to 1, present when the provider returns per-word confidence | 0.98 |
End | float64 | :heavy_check_mark: | Word end time in seconds | 0.4 |
Speaker | *int64 | :heavy_minus_sign: | Speaker index for the word, present when the provider returns diarization data | 0 |
SpeakerLabel | *string | :heavy_minus_sign: | Provider speaker label for the word, present when the provider labels speakers with a string | speaker_0 |
Start | float64 | :heavy_check_mark: | Word start time in seconds | 0 |
Type | *components.STTWordType | :heavy_minus_sign: | Kind of entry; omitted or “word” for spoken words, “audio_event” for non-speech sounds the provider tags with timestamps | word |
Word | string | :heavy_check_mark: | The transcribed word, or the event tag such as “(laughter)” when type is audio_event | Hello |