Files
dsh-skills/skills/lab-speech-to-text/SKILL.md

2.4 KiB

name, description
name description
lab-speech-to-text Transcribe speech from audio files (wav, mp3, ogg, m4a) using the user's home-lab Whishper speech-to-text server (GPU-backed). Use when the user wants to convert audio/voice to text, transcribe a recording, or subtitle.

Lab Speech-to-Text (Whishper)

The user's home lab runs a Whishper speech-to-text server at http://192.168.31.159:8082 (server2, CUDA GPU, small model). Use it to transcribe audio into text. No authentication is required — the API is open on the LAN.

Workflow

Use run_code to upload the audio, poll the transcription, then return the recognized text.

1. Upload the audio

POST /api/transcriptions with multipart form-data. Field files = the audio blob (mime: audio/wav, audio/mpeg, audio/ogg, or audio/mp4). Optional fields: language (e.g. ru, en), modelSize.

Response:

{ "id": "<uuid>", "status": 0, "task": "transcribe", "device": "cuda", "result": { "text": "" } }

2. Poll for the result

GET {base}/api/transcriptions/{id}. A status of -1 means the job is still queued or running. When status >= 0 and result.text is non-empty, transcription is done.

3. Return the text

The transcription text is in result.text. Return it (trimmed). Also note result.language and result.duration if useful.

Policy

  • Timeout: poll no longer than ~120s for typical clips. Longer audio may exceed that; report a timeout rather than looping forever.
  • File source: read the audio via ctx.fs.readBytes(target, signal, maxBytes) (cap ~20MB). Prefer a local path to the file.
  • If the user wants a language hint, pass it as language (e.g. ru).

Example (run_code)

const base = 'http://192.168.31.159:8082'
// read the audio file bytes first
const fd = new FormData()
fd.append('files', new Blob([audioBytes], { type: 'audio/wav' }), 'note.wav')
fd.append('language', 'ru') // optional
const up = await (await fetch(base + '/api/transcriptions', { method: 'POST', body: fd })).json()
if (!up.id) throw new Error('no id: ' + JSON.stringify(up))
let text = null
for (let i = 0; i < 80 && !text; i++) {
  const s = await (await fetch(base + '/api/transcriptions/' + up.id)).json()
  const t = s?.result?.text
  if (typeof t === 'string' && t.length > 0 && s.status !== -1) text = t
  else await new Promise(sl => setTimeout(sl, 1500))
}
if (!text) throw new Error('Transcription timed out')
return text.trim()