chore: add shared skills catalog (19 skills), installers, manifest, validator
This commit is contained in:
68
skills/lab-pdf-ocr/SKILL.md
Normal file
68
skills/lab-pdf-ocr/SKILL.md
Normal file
@@ -0,0 +1,68 @@
|
||||
---
|
||||
name: lab-pdf-ocr
|
||||
description: Extract text from PDF documents using the user's home-lab Docling OCR server (GPU-backed). Use when the user wants to recognize text, OCR a PDF, or convert a scanned or structured PDF to markdown text.
|
||||
---
|
||||
|
||||
# Lab PDF OCR (Docling)
|
||||
|
||||
The user's home lab runs a Docling OCR server at `http://192.168.31.159:5001` (server2, CUDA GPU). Use it to extract text from PDFs. No authentication is required — the API is open on the LAN.
|
||||
|
||||
## Workflow
|
||||
|
||||
Use `run_code` to upload the PDF, poll the task, then return the extracted markdown text.
|
||||
|
||||
### 1. Upload the PDF
|
||||
|
||||
POST `/v1/convert/file/async` with multipart form-data. Field `files` = the PDF blob (Content-Type `application/pdf`); `options` = `{}` (JSON).
|
||||
|
||||
Response:
|
||||
|
||||
```json
|
||||
{ "task_id": "<uuid>", "task_type": "convert", "task_status": "pending" }
|
||||
```
|
||||
|
||||
### 2. Poll for completion
|
||||
|
||||
Poll `GET {base}/v1/status/poll/{task_id}?wait=2` every ~1.2s until `task_status` is `success`. If it is `failed`, fetch the result to get the error.
|
||||
|
||||
### 3. Fetch the result
|
||||
|
||||
`GET {base}/v1/result/{task_id}` returns:
|
||||
|
||||
```json
|
||||
{
|
||||
"document": { "filename": "scan.pdf", "md_content": "<recognized markdown text>" },
|
||||
"status": "success",
|
||||
"errors": []
|
||||
}
|
||||
```
|
||||
|
||||
Return `document.md_content` — the recognized markdown text.
|
||||
|
||||
## Policy
|
||||
- **Timeout:** poll no longer than ~120s for typical PDFs. Large or heavily scanned PDFs may take longer; if the budget is exceeded, report a timeout.
|
||||
- **File source:** read the PDF via `ctx.fs.readBytes(target, signal, maxBytes)` (cap ~20MB).
|
||||
- Trim leading/trailing whitespace when returning the text. The `md_content` is the authoritative result — ignore `html_content`/`json_content` unless the user specifically wants structure.
|
||||
|
||||
### Example (run_code)
|
||||
|
||||
```js
|
||||
const base = 'http://192.168.31.159:5001'
|
||||
const bytes = await tools.read({ file_path: '/path/to/file.pdf' }) // read via fs tool
|
||||
// build multipart FormData with the pdf blob
|
||||
const fd = new FormData()
|
||||
fd.append('files', new Blob([bytes], { type: 'application/pdf' }), 'file.pdf')
|
||||
fd.append('options', JSON.stringify({}))
|
||||
const r = await (await fetch(base + '/v1/convert/file/async', { method: 'POST', body: fd })).json()
|
||||
if (!r.task_id) throw new Error('no task_id: ' + JSON.stringify(r))
|
||||
for (let i = 0; i < 120; i++) {
|
||||
const s = await (await fetch(base + '/v1/status/poll/' + r.task_id + '?wait=2')).json()
|
||||
if (s.task_status === 'success') {
|
||||
const res = await (await fetch(base + '/v1/result/' + r.task_id)).json()
|
||||
return (res.document?.md_content ?? '').trim()
|
||||
}
|
||||
if (s.task_status === 'failed' || s.task_status === 'error') throw new Error('Docling failed: ' + s.task_status)
|
||||
await new Promise(sl => setTimeout(sl, 1200))
|
||||
}
|
||||
throw new Error('Docling OCR timed out')
|
||||
```
|
||||
Reference in New Issue
Block a user