Create, merge, split, extract text/tables from PDFs, fill forms, add watermarks, and convert to/from other formats. USE WHEN pdf, PDF file, merge PDF, extract tables, fill form, split PDF.
What this skill does
# PDF Processing Guide
## ๐ฏ Load Full PAI Context
**Before starting any task with this skill, load complete PAI context:**
`read ~/.claude/PAI/SKILL.md`
This provides access to:
- Complete contact list (Angela, Bunny, Saลกa, Greg, team members)
- Stack preferences (TypeScript>Python, bun>npm, uv>pip)
- Security rules and repository safety protocols
- Response format requirements (structured emoji format)
- Voice IDs for agent routing (ElevenLabs)
- Personal preferences and operating instructions
## When to Activate This Skill
### Direct PDF Task Triggers
- User wants to **create** a new PDF document
- User wants to **merge**, **combine**, or **concatenate** multiple PDFs
- User wants to **split** or **separate** a PDF into individual pages/sections
- User mentions "**extract text from PDF**", "**PDF text extraction**"
- User mentions "**extract tables from PDF**", "**PDF tables**"
- User wants to "**fill PDF form**", "**PDF form filling**"
- User mentions "**OCR**", "**scanned PDF**", or "**scan to text**"
- User wants to add **watermarks**, **password protection**, or **encryption**
- User wants to **extract images** from a PDF
- User wants to **rotate pages** or manipulate PDF structure
### Contextual Triggers
- User provides a **.pdf file path** for processing
- User mentions form filling automation or batch PDF processing
- User needs to process PDFs programmatically at scale
## ๐ PDF Workflow Routing
This skill supports multiple PDF processing workflows:
### Creation Workflow
**Trigger:** "create PDF", "generate PDF", "make PDF", "PDF from data"
**Tools:** reportlab (Python)
**Documentation:** Lines 136-181 (SKILL.md)
**Use Cases:**
- Creating new PDFs from scratch
- Generating reports programmatically
- Multi-page documents with text and graphics
- PDF generation from templates or data
### Merge/Split Workflow
**Trigger:** "merge PDFs", "combine PDFs", "split PDF", "separate pages"
**Tools:** pypdf (Python), qpdf (CLI)
**Documentation:** Lines 46-68 (SKILL.md), Lines 199-211 (qpdf)
**Use Cases:**
- Combining multiple PDFs into one document
- Splitting PDFs into individual pages or ranges
- Reorganizing PDF page order
- Extracting specific page ranges
### Text Extraction Workflow
**Trigger:** "extract text", "PDF to text", "read PDF content"
**Tools:** pdfplumber (Python), pdftotext (CLI)
**Documentation:** Lines 95-103 (pdfplumber), Lines 186-196 (pdftotext)
**Use Cases:**
- Extracting text while preserving layout
- Converting PDFs to plain text
- Batch text extraction from multiple PDFs
- Metadata extraction
### Table Extraction Workflow
**Trigger:** "extract tables", "PDF tables", "table data from PDF"
**Tools:** pdfplumber + pandas (Python)
**Documentation:** Lines 106-133 (SKILL.md)
**Use Cases:**
- Extracting structured table data to Excel/CSV
- Financial data extraction from PDF reports
- Converting PDF tables to dataframes
- Multi-table extraction and combination
### Form Filling Workflow
**Trigger:** "fill PDF form", "PDF form filling", "complete PDF form"
**Tools:** pdf-lib (JavaScript) or pypdf (Python)
**Documentation:** forms.md (complete guide)
**Use Cases:**
- Programmatic form completion
- Batch form processing
- Template-based PDF generation
- Form field population from data sources
### OCR Workflow
**Trigger:** "OCR", "scanned PDF", "extract text from scan", "image to text"
**Tools:** pytesseract + pdf2image (Python)
**Documentation:** Lines 227-244 (SKILL.md)
**Use Cases:**
- Extracting text from scanned documents
- Processing image-based PDFs
- Converting scanned forms to editable text
- Legacy document digitization
### Manipulation Workflow
**Trigger:** "watermark", "password protect", "encrypt PDF", "rotate pages", "extract images"
**Tools:** pypdf (Python), pdfimages (CLI)
**Documentation:** Lines 246-288 (SKILL.md)
**Use Cases:**
- Adding watermarks to PDFs
- Password protection and encryption
- Page rotation and transformation
- Image extraction from PDFs
## Overview
This guide covers essential PDF processing operations using Python libraries and command-line tools. For advanced features, JavaScript libraries, and detailed examples, see reference.md. If you need to fill out a PDF form, read forms.md and follow its instructions.
## Quick Start
```python
from pypdf import PdfReader, PdfWriter
# Read a PDF
reader = PdfReader("document.pdf")
print(f"Pages: {len(reader.pages)}")
# Extract text
text = ""
for page in reader.pages:
text += page.extract_text()
```
## Python Libraries
### pypdf - Basic Operations
#### Merge PDFs
```python
from pypdf import PdfWriter, PdfReader
writer = PdfWriter()
for pdf_file in ["doc1.pdf", "doc2.pdf", "doc3.pdf"]:
reader = PdfReader(pdf_file)
for page in reader.pages:
writer.add_page(page)
with open("merged.pdf", "wb") as output:
writer.write(output)
```
#### Split PDF
```python
reader = PdfReader("input.pdf")
for i, page in enumerate(reader.pages):
writer = PdfWriter()
writer.add_page(page)
with open(f"page_{i+1}.pdf", "wb") as output:
writer.write(output)
```
#### Extract Metadata
```python
reader = PdfReader("document.pdf")
meta = reader.metadata
print(f"Title: {meta.title}")
print(f"Author: {meta.author}")
print(f"Subject: {meta.subject}")
print(f"Creator: {meta.creator}")
```
#### Rotate Pages
```python
reader = PdfReader("input.pdf")
writer = PdfWriter()
page = reader.pages[0]
page.rotate(90) # Rotate 90 degrees clockwise
writer.add_page(page)
with open("rotated.pdf", "wb") as output:
writer.write(output)
```
### pdfplumber - Text and Table Extraction
#### Extract Text with Layout
```python
import pdfplumber
with pdfplumber.open("document.pdf") as pdf:
for page in pdf.pages:
text = page.extract_text()
print(text)
```
#### Extract Tables
```python
with pdfplumber.open("document.pdf") as pdf:
for i, page in enumerate(pdf.pages):
tables = page.extract_tables()
for j, table in enumerate(tables):
print(f"Table {j+1} on page {i+1}:")
for row in table:
print(row)
```
#### Advanced Table Extraction
```python
import pandas as pd
with pdfplumber.open("document.pdf") as pdf:
all_tables = []
for page in pdf.pages:
tables = page.extract_tables()
for table in tables:
if table: # Check if table is not empty
df = pd.DataFrame(table[1:], columns=table[0])
all_tables.append(df)
# Combine all tables
if all_tables:
combined_df = pd.concat(all_tables, ignore_index=True)
combined_df.to_excel("extracted_tables.xlsx", index=False)
```
### reportlab - Create PDFs
#### Basic PDF Creation
```python
from reportlab.lib.pagesizes import letter
from reportlab.pdfgen import canvas
c = canvas.Canvas("hello.pdf", pagesize=letter)
width, height = letter
# Add text
c.drawString(100, height - 100, "Hello World!")
c.drawString(100, height - 120, "This is a PDF created with reportlab")
# Add a line
c.line(100, height - 140, 400, height - 140)
# Save
c.save()
```
#### Create PDF with Multiple Pages
```python
from reportlab.lib.pagesizes import letter
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer, PageBreak
from reportlab.lib.styles import getSampleStyleSheet
doc = SimpleDocTemplate("report.pdf", pagesize=letter)
styles = getSampleStyleSheet()
story = []
# Add content
title = Paragraph("Report Title", styles['Title'])
story.append(title)
story.append(Spacer(1, 12))
body = Paragraph("This is the body of the report. " * 20, styles['Normal'])
story.append(body)
story.append(PageBreak())
# Page 2
story.append(Paragraph("Page 2", styles['Heading1']))
story.append(Paragraph("Content for page 2", styles['Normal']))
# Build PDF
doc.build(story)
```
## Command-Line Tools
### pdftotext (poppler-utils)
```bash
# Extract text
pdftotext input.pdf output.txt
# Extract text preserving layout
pdftotext -layout input.pdf output.txt
# Extract specificRelated in Writing & Docs
jax-development
IncludedUse this skill when the user is writing, debugging, profiling, refactoring, reviewing, benchmarking, parallelising, exporting, or explaining JAX code, or when they mention JAX, jax.numpy, jit, grad, value_and_grad, vmap, scan, lax, random keys, pytrees, jax.Array, sharding, Mesh, PartitionSpec, NamedSharding, pmap, shard_map, Pallas, XLA, StableHLO, checkify, profiler, or the JAX repo. It helps turn NumPy or PyTorch-style code into pure functional JAX, fix tracer/control-flow/shape/PRNG bugs, remove recompiles and host-device syncs, choose transforms and sharding strategies, inspect jaxpr/lowering/IR, and benchmark compiled code correctly.
nature-article-writer
IncludedDrafts, rewrites, diagnostically critiques, and style-calibrates primary research manuscripts for Nature and Nature Portfolio journals. Use when the user wants a Nature-style title, summary paragraph or abstract, introduction, results, discussion, methods, figure legends, presubmission enquiry, cover letter, reviewer response, or when a scientific draft sounds generic, jargon-heavy, structurally weak, or AI-ish and needs precise, broad-reader-friendly prose without inventing data, analyses, or references. Best for primary research articles and letters rather than reviews or press releases unless explicitly adapting one.
deckrd
IncludedDocument-driven framework that derives requirements, specifications, implementation plans, and executable tasks from goals through structured AI dialogue. Use when user says "write requirements", "create spec", "plan implementation", "derive tasks", "structure this feature", "break down into tasks", or "document this module". Also use for reverse engineering existing code into docs (/deckrd rev). Do NOT use for direct code writing โ use /deckrd-coder after tasks are generated. Do NOT use when the user only wants to run or fix existing code without planning.
clinical-decision-support
IncludedGenerate professional clinical decision support (CDS) documents for pharmaceutical and clinical research settings, including patient cohort analyses (biomarker-stratified with outcomes) and treatment recommendation reports (evidence-based guidelines with decision algorithms). Supports GRADE evidence grading, statistical analysis (hazard ratios, survival curves, waterfall plots), biomarker integration, and regulatory compliance. Outputs publication-ready LaTeX/PDF format optimized for drug development, clinical research, and evidence synthesis.
handling-sf-data
IncludedSalesforce data operations with 130-point scoring. Use this skill to create, update, delete, bulk import/export, generate test data, and clean up org records using sf CLI and anonymous Apex. TRIGGER when: user creates test data, performs bulk import/export, uses sf data CLI commands, needs data factory patterns for Apex tests, or needs to seed/clean records in a Salesforce org. DO NOT TRIGGER when: SOQL query writing only (use querying-soql), Apex test execution (use running-apex-tests), or metadata deployment (use deploying-metadata).
accelint-ac-to-playwright
IncludedConvert and validate acceptance criteria for Playwright test automation. Use when user asks to (1) review/evaluate/check if AC are ready for automation, (2) assess if AC can be converted as-is, (3) validate AC quality for Playwright, (4) turn AC into tests, (5) generate tests from acceptance criteria, (6) convert .md bullets or .feature Gherkin files to Playwright specs, (7) create test automation from requirements. Handles both bullet-style markdown and Gherkin syntax with JSON test plan generation and validation.