pdf-vision
Gemini vision-powered PDF to markdown converter. Handles scanned docs, multi-column layouts, tables, footnotes, flowcharts, and degraded documents that text-based extraction destroys. Uses two-tier model routing (cheap model for clean digital pages, capable model for everything else) with per-chunk confidence scoring, anti-hallucination detection, and continual learning from user corrections. Use when a user needs to convert, extract, analyze, or process any PDF document — especially scanned documents, government forms, legal contracts, academic papers, or anything where pypdf/pdfplumber returns garbage or nothing.
What this skill does
# pdf-vision
Vision-powered PDF processing that sees documents the way humans do — not as coordinates and font metadata, but as structured content with meaning.
## When to Use
- Converting PDFs to clean markdown (especially scanned, multi-column, or complex layouts)
- Processing documents that pypdf/pdfplumber/Acrobat garble (tables, flowcharts, footnotes)
- Batch processing document archives
- Extracting structured data from government forms, legal contracts, academic papers
- Any PDF task where text-based extraction fails or returns nothing
## Quick Start
```bash
# Install dependencies
pip install pymupdf google-genai
# Set API key
export GEMINI_API_KEY=your-key
# Analyze a PDF (preflight — no OCR, just document analysis)
python scripts/preflight.py document.pdf
# Convert PDF to markdown
python scripts/ocr_pipeline.py document.pdf
# Convert with custom output path
python scripts/ocr_pipeline.py document.pdf -o output.md
# Convert with specific chunk size
python scripts/ocr_pipeline.py document.pdf -c 8
# Learn from a correction
python scripts/learn.py original.md corrected.md
```
## How It Works
### 1. Preflight Analysis (~$0.005)
Samples 8 pages from beginning, middle, and end of the document. Sends to Gemini flash-lite to detect: document type, language, column layout, footnotes, tables, scan quality, running headers/footers, text density. Configures the entire pipeline automatically.
**Why scattered sampling:** A 123-page Latin manuscript with an English preface fools a first-5-pages sample. Sampling beginning + middle + end correctly detects the real document characteristics.
### 2. Two-Tier Model Routing
Routes documents to the right model based on difficulty:
| Difficulty | Model | Cost/1M tokens (in/out) |
|-----------|-------|------------------------|
| Clean digital | ~~gemini-2.5-flash-lite~~ | $0.10 / $0.40 |
| Everything else | ~~gemini-3.1-flash-lite-preview~~ | $0.25 / $1.50 |
**Why two tiers:** We tested three models on the same 10 Latin manuscript pages. Gemini 3.1 Flash Lite extracted 4x more content (214K vs 50K chars) than 3.0 Flash Preview, while costing half as much. The cheap model handles clean digital docs fine. Everything else goes to 3.1.
### 3. Adaptive Chunking
Chunk size based on document density, not fixed. Dense scholarly text: 6-10 pages. Standard docs: 12-15. Each chunk includes continuation context ("Pages 13-24 of 126. Continue from previous.") for cross-page coherence.
### 4. Per-Chunk Confidence Scoring
Every chunk scored 0.0-1.0 based on: output length vs expected (from preflight word density), truncation detection, garbled text runs, unbalanced markdown, and hallucination detection. Low-confidence chunks flagged in YAML frontmatter for human review.
### 5. Anti-Hallucination Guard
Image-heavy pages (maps, charts, photos) can cause models to fabricate plausible text. Prompt instructs the model to output only `*(Map/image omitted)*` for image pages. Filler-phrase detector flags generic boilerplate in short chunks (e.g., "is well-positioned to support").
### 6. Boundary Smoothing
AI-powered pass that detects and fixes broken sentences, duplicate headers, and artifacts at chunk boundaries.
### 7. Continual Learning
Correct a mistake, run `python scripts/learn.py original.md corrected.md`. Extracts patterns via Gemini, saves to `~/.pdf-vision/corrections.json`. Next run: matching corrections injected as "Known Issues" in the prompt. Works for structural patterns; correctly ignores corrections that contradict what the model sees on the page.
## Output Format
```markdown
---
source_file: document.pdf
pages: 126
processing_date: 2026-03-07T14:30:00Z
models_used:
gemini-2.5-flash-lite: 80 pages
gemini-3.1-flash-lite-preview: 46 pages
total_cost: $0.12
avg_confidence: 0.91
low_confidence_pages: [72, 103]
document_type: government-form
language: english
---
# Document Title
[Clean markdown content with preserved structure...]
```
## Customization Points
```
OCR Model (clean pages): ~~gemini-2.5-flash-lite~~
OCR Model (all other pages): ~~gemini-3.1-flash-lite-preview~~
Default language: ~~auto-detect~~
Max chunk size: ~~15 pages~~
Min confidence threshold: ~~0.7~~
Flag for human review below: ~~0.6 confidence~~
Output format: ~~markdown~~
Boundary smoothing: ~~enabled~~
```
## Domain Presets
Configure pdf-vision for your industry:
- Legal: ~~disabled~~ — Preserve clause numbering (1.1, 1.1.1), extract defined terms, detect signature blocks, never summarize clauses
- Medical: ~~disabled~~ — Normalize drug names, format ICD/CPT codes, flag HIPAA content
- Government: ~~disabled~~ — Preserve form field blanks (____), extract checkbox states ([X]/[ ]), keep cross-references (ITB-clause 9.2), preserve tender numbering
- Academic: ~~disabled~~ — Preserve LaTeX equations, extract bibliography, link footnotes bidirectionally
- Financial: ~~disabled~~ — Extract financial tables, detect GAAP/IFRS terminology, preserve audit structure
- Latin Manuscript: ~~disabled~~ — Preserve column markers (col. 1125), keep editorial apparatus ([add. ed.], [om. ed.]), preserve ALL-CAPS chapter headings (CAP. XL), italicize vernacular glosses
## Dependencies
```
pymupdf>=1.24.0
google-genai>=1.0.0
```
Related in AI Agents
skill-development
IncludedComprehensive meta-skill for creating, managing, validating, auditing, and distributing Claude Code skills and slash commands (unified in v2.1.3+). Provides skill templates, creation workflows, validation patterns, audit checklists, naming conventions, YAML frontmatter guidance, progressive disclosure examples, and best practices lookup. Use when creating new skills, validating existing skills, auditing skill quality, understanding skill architecture, needing skill templates, learning about YAML frontmatter requirements, progressive disclosure patterns, tool restrictions (allowed-tools), skill composition, skill naming conventions, troubleshooting skill activation issues, creating custom slash commands, configuring command frontmatter, using command arguments ($ARGUMENTS, $1, $2), bash execution in commands, file references in commands, command namespacing, plugin commands, MCP slash commands, Skill tool configuration, or deciding between skills vs slash commands. Delegates to docs-management skill for official documentation.
reprompter
IncludedTransform messy prompts into well-structured, effective prompts — single or multi-agent. Use when: "reprompt", "reprompt this", "clean up this prompt", "structure my prompt", rough text needing XML tags and best practices, "reprompter teams", "repromptception", "run with quality", "smart run", "smart agents", multi-agent tasks, audits, parallel work, anything going to agent teams. Don't use when: simple Q&A, pure chat, immediate execution-only tasks. See "Don't Use When" section for details. Outputs: Structured XML/Markdown prompt, quality score (before/after), optional team brief + per-agent sub-prompts, agent team output files. Success criteria: Single mode quality score ≥ 7/10; Repromptception per-agent prompt quality score 8+/10; all required sections present, actionable and specific.
adaptive-compaction
IncludedAdaptive add-on policy and recovery layer that decides WHEN to compact, prune, snapshot, or fork -- replacing fixed-percent auto-compaction across Claude Code, Codex, and MCP-capable hosts. Trigger on auto-compact timing or damage: "when should I compact", "is it safe to compact now or start a fresh session", "auto-compact fires too early/mid-task", "switching to an unrelated task but the window still has space", "context rot", "answers get worse the longer the session runs", "the agent forgot the plan or my decisions after it summarized", "add a layer on top that manages context without changing the agent", raising autoCompactWindow to give the policy room, or installing/tuning a cross-tool compaction policy or PreCompact hook -- even when "compaction" is never said but the problem is context-window pressure or post-summarization memory loss. Do NOT use to summarize a conversation, build RAG, write a summarization prompt (decides WHEN not HOW), or answer max-context-length trivia.
agent-skill-creator
IncludedCreate cross-platform agent skills from workflow descriptions. Activates when users ask to create an agent, automate a repetitive workflow, create a custom skill, or need advanced agent creation. Triggers on phrases like create agent for, automate workflow, create skill for, every day I have to, daily I need to, turn process into agent, need to automate, create a cross-platform skill, validate this skill, export this skill, migrate this skill. Supports single skills, multi-agent suites, transcript processing, template-based creation, interactive configuration, cross-platform export, and spec validation.
llm-wiki
IncludedUse when building or maintaining a persistent personal knowledge base (second brain) in Obsidian where an LLM incrementally ingests sources, updates entity/concept pages, maintains cross-references, and keeps a synthesis current. Triggers include "second brain", "Obsidian wiki", "personal knowledge management", "ingest this paper/article/book", "build a research wiki", "compound knowledge", "Memex", or whenever the user wants knowledge to accumulate across sessions instead of being re-derived by RAG on every query.
skill-master
IncludedAgent Skills authoring, evaluation, and optimization. Create, edit, validate, benchmark, and improve skills following the agentskills.io specification. Use when designing SKILL.md files, structuring skill folders (references, scripts, assets), ingesting external documentation into skills, running trigger evals, benchmarking skill quality, optimizing descriptions, or performing blind A/B comparisons. Keywords: agentskills.io, SKILL.md, skill authoring, eval, benchmark, trigger optimization.