Files
dsh-skills/skills/lab-pdf-ocr/SKILL.md

2.7 KiB

name, description
name description
lab-pdf-ocr Extract text from PDF documents using the user's home-lab Docling OCR server (GPU-backed). Use when the user wants to recognize text, OCR a PDF, or convert a scanned or structured PDF to markdown text.

Lab PDF OCR (Docling)

The user's home lab runs a Docling OCR server at http://192.168.31.159:5001 (server2, CUDA GPU). Use it to extract text from PDFs. No authentication is required — the API is open on the LAN.

Workflow

Use run_code to upload the PDF, poll the task, then return the extracted markdown text.

1. Upload the PDF

POST /v1/convert/file/async with multipart form-data. Field files = the PDF blob (Content-Type application/pdf); options = {} (JSON).

Response:

{ "task_id": "<uuid>", "task_type": "convert", "task_status": "pending" }

2. Poll for completion

Poll GET {base}/v1/status/poll/{task_id}?wait=2 every ~1.2s until task_status is success. If it is failed, fetch the result to get the error.

3. Fetch the result

GET {base}/v1/result/{task_id} returns:

{
  "document": { "filename": "scan.pdf", "md_content": "<recognized markdown text>" },
  "status": "success",
  "errors": []
}

Return document.md_content — the recognized markdown text.

Policy

  • Timeout: poll no longer than ~120s for typical PDFs. Large or heavily scanned PDFs may take longer; if the budget is exceeded, report a timeout.
  • File source: read the PDF via ctx.fs.readBytes(target, signal, maxBytes) (cap ~20MB).
  • Trim leading/trailing whitespace when returning the text. The md_content is the authoritative result — ignore html_content/json_content unless the user specifically wants structure.

Example (run_code)

const base = 'http://192.168.31.159:5001'
const bytes = await tools.read({ file_path: '/path/to/file.pdf' }) // read via fs tool
// build multipart FormData with the pdf blob
const fd = new FormData()
fd.append('files', new Blob([bytes], { type: 'application/pdf' }), 'file.pdf')
fd.append('options', JSON.stringify({}))
const r = await (await fetch(base + '/v1/convert/file/async', { method: 'POST', body: fd })).json()
if (!r.task_id) throw new Error('no task_id: ' + JSON.stringify(r))
for (let i = 0; i < 120; i++) {
  const s = await (await fetch(base + '/v1/status/poll/' + r.task_id + '?wait=2')).json()
  if (s.task_status === 'success') {
    const res = await (await fetch(base + '/v1/result/' + r.task_id)).json()
    return (res.document?.md_content ?? '').trim()
  }
  if (s.task_status === 'failed' || s.task_status === 'error') throw new Error('Docling failed: ' + s.task_status)
  await new Promise(sl => setTimeout(sl, 1200))
}
throw new Error('Docling OCR timed out')