68 lines
2.7 KiB
Markdown
68 lines
2.7 KiB
Markdown
---
|
|
name: lab-pdf-ocr
|
|
description: Extract text from PDF documents using the user's home-lab Docling OCR server (GPU-backed). Use when the user wants to recognize text, OCR a PDF, or convert a scanned or structured PDF to markdown text.
|
|
---
|
|
|
|
# Lab PDF OCR (Docling)
|
|
|
|
The user's home lab runs a Docling OCR server at `http://192.168.31.159:5001` (server2, CUDA GPU). Use it to extract text from PDFs. No authentication is required — the API is open on the LAN.
|
|
|
|
## Workflow
|
|
|
|
Use `run_code` to upload the PDF, poll the task, then return the extracted markdown text.
|
|
|
|
### 1. Upload the PDF
|
|
|
|
POST `/v1/convert/file/async` with multipart form-data. Field `files` = the PDF blob (Content-Type `application/pdf`); `options` = `{}` (JSON).
|
|
|
|
Response:
|
|
|
|
```json
|
|
{ "task_id": "<uuid>", "task_type": "convert", "task_status": "pending" }
|
|
```
|
|
|
|
### 2. Poll for completion
|
|
|
|
Poll `GET {base}/v1/status/poll/{task_id}?wait=2` every ~1.2s until `task_status` is `success`. If it is `failed`, fetch the result to get the error.
|
|
|
|
### 3. Fetch the result
|
|
|
|
`GET {base}/v1/result/{task_id}` returns:
|
|
|
|
```json
|
|
{
|
|
"document": { "filename": "scan.pdf", "md_content": "<recognized markdown text>" },
|
|
"status": "success",
|
|
"errors": []
|
|
}
|
|
```
|
|
|
|
Return `document.md_content` — the recognized markdown text.
|
|
|
|
## Policy
|
|
- **Timeout:** poll no longer than ~120s for typical PDFs. Large or heavily scanned PDFs may take longer; if the budget is exceeded, report a timeout.
|
|
- **File source:** read the PDF via `ctx.fs.readBytes(target, signal, maxBytes)` (cap ~20MB).
|
|
- Trim leading/trailing whitespace when returning the text. The `md_content` is the authoritative result — ignore `html_content`/`json_content` unless the user specifically wants structure.
|
|
|
|
### Example (run_code)
|
|
|
|
```js
|
|
const base = 'http://192.168.31.159:5001'
|
|
const bytes = await tools.read({ file_path: '/path/to/file.pdf' }) // read via fs tool
|
|
// build multipart FormData with the pdf blob
|
|
const fd = new FormData()
|
|
fd.append('files', new Blob([bytes], { type: 'application/pdf' }), 'file.pdf')
|
|
fd.append('options', JSON.stringify({}))
|
|
const r = await (await fetch(base + '/v1/convert/file/async', { method: 'POST', body: fd })).json()
|
|
if (!r.task_id) throw new Error('no task_id: ' + JSON.stringify(r))
|
|
for (let i = 0; i < 120; i++) {
|
|
const s = await (await fetch(base + '/v1/status/poll/' + r.task_id + '?wait=2')).json()
|
|
if (s.task_status === 'success') {
|
|
const res = await (await fetch(base + '/v1/result/' + r.task_id)).json()
|
|
return (res.document?.md_content ?? '').trim()
|
|
}
|
|
if (s.task_status === 'failed' || s.task_status === 'error') throw new Error('Docling failed: ' + s.task_status)
|
|
await new Promise(sl => setTimeout(sl, 1200))
|
|
}
|
|
throw new Error('Docling OCR timed out')
|
|
``` |