2.4 KiB
name, description
| name | description |
|---|---|
| lab-speech-to-text | Transcribe speech from audio files (wav, mp3, ogg, m4a) using the user's home-lab Whishper speech-to-text server (GPU-backed). Use when the user wants to convert audio/voice to text, transcribe a recording, or subtitle. |
Lab Speech-to-Text (Whishper)
The user's home lab runs a Whishper speech-to-text server at http://192.168.31.159:8082 (server2, CUDA GPU, small model). Use it to transcribe audio into text. No authentication is required — the API is open on the LAN.
Workflow
Use run_code to upload the audio, poll the transcription, then return the recognized text.
1. Upload the audio
POST /api/transcriptions with multipart form-data. Field files = the audio blob (mime: audio/wav, audio/mpeg, audio/ogg, or audio/mp4). Optional fields: language (e.g. ru, en), modelSize.
Response:
{ "id": "<uuid>", "status": 0, "task": "transcribe", "device": "cuda", "result": { "text": "" } }
2. Poll for the result
GET {base}/api/transcriptions/{id}. A status of -1 means the job is still queued or running. When status >= 0 and result.text is non-empty, transcription is done.
3. Return the text
The transcription text is in result.text. Return it (trimmed). Also note result.language and result.duration if useful.
Policy
- Timeout: poll no longer than ~120s for typical clips. Longer audio may exceed that; report a timeout rather than looping forever.
- File source: read the audio via
ctx.fs.readBytes(target, signal, maxBytes)(cap ~20MB). Prefer a local path to the file. - If the user wants a language hint, pass it as
language(e.g.ru).
Example (run_code)
const base = 'http://192.168.31.159:8082'
// read the audio file bytes first
const fd = new FormData()
fd.append('files', new Blob([audioBytes], { type: 'audio/wav' }), 'note.wav')
fd.append('language', 'ru') // optional
const up = await (await fetch(base + '/api/transcriptions', { method: 'POST', body: fd })).json()
if (!up.id) throw new Error('no id: ' + JSON.stringify(up))
let text = null
for (let i = 0; i < 80 && !text; i++) {
const s = await (await fetch(base + '/api/transcriptions/' + up.id)).json()
const t = s?.result?.text
if (typeof t === 'string' && t.length > 0 && s.status !== -1) text = t
else await new Promise(sl => setTimeout(sl, 1500))
}
if (!text) throw new Error('Transcription timed out')
return text.trim()