2.7 KiB
2.7 KiB
name, description
| name | description |
|---|---|
| lab-pdf-ocr | Extract text from PDF documents using the user's home-lab Docling OCR server (GPU-backed). Use when the user wants to recognize text, OCR a PDF, or convert a scanned or structured PDF to markdown text. |
Lab PDF OCR (Docling)
The user's home lab runs a Docling OCR server at http://192.168.31.159:5001 (server2, CUDA GPU). Use it to extract text from PDFs. No authentication is required — the API is open on the LAN.
Workflow
Use run_code to upload the PDF, poll the task, then return the extracted markdown text.
1. Upload the PDF
POST /v1/convert/file/async with multipart form-data. Field files = the PDF blob (Content-Type application/pdf); options = {} (JSON).
Response:
{ "task_id": "<uuid>", "task_type": "convert", "task_status": "pending" }
2. Poll for completion
Poll GET {base}/v1/status/poll/{task_id}?wait=2 every ~1.2s until task_status is success. If it is failed, fetch the result to get the error.
3. Fetch the result
GET {base}/v1/result/{task_id} returns:
{
"document": { "filename": "scan.pdf", "md_content": "<recognized markdown text>" },
"status": "success",
"errors": []
}
Return document.md_content — the recognized markdown text.
Policy
- Timeout: poll no longer than ~120s for typical PDFs. Large or heavily scanned PDFs may take longer; if the budget is exceeded, report a timeout.
- File source: read the PDF via
ctx.fs.readBytes(target, signal, maxBytes)(cap ~20MB). - Trim leading/trailing whitespace when returning the text. The
md_contentis the authoritative result — ignorehtml_content/json_contentunless the user specifically wants structure.
Example (run_code)
const base = 'http://192.168.31.159:5001'
const bytes = await tools.read({ file_path: '/path/to/file.pdf' }) // read via fs tool
// build multipart FormData with the pdf blob
const fd = new FormData()
fd.append('files', new Blob([bytes], { type: 'application/pdf' }), 'file.pdf')
fd.append('options', JSON.stringify({}))
const r = await (await fetch(base + '/v1/convert/file/async', { method: 'POST', body: fd })).json()
if (!r.task_id) throw new Error('no task_id: ' + JSON.stringify(r))
for (let i = 0; i < 120; i++) {
const s = await (await fetch(base + '/v1/status/poll/' + r.task_id + '?wait=2')).json()
if (s.task_status === 'success') {
const res = await (await fetch(base + '/v1/result/' + r.task_id)).json()
return (res.document?.md_content ?? '').trim()
}
if (s.task_status === 'failed' || s.task_status === 'error') throw new Error('Docling failed: ' + s.task_status)
await new Promise(sl => setTimeout(sl, 1200))
}
throw new Error('Docling OCR timed out')