Speech-to-text for one immutable audio or video artifact, returning normalized text and timed segments for audio-to-text or video transcription workflows.