playwright-scraper
Playwright web scraping: dynamic content, auth flows, pagination, data extraction, screenshots
What this skill does
## playwright-scraper
### Purpose
This skill enables web scraping using Playwright, a Node.js library for browser automation. It focuses on handling dynamic content, authentication flows, pagination, data extraction, and screenshots to reliably scrape modern websites.
### When to Use
Use this skill for scraping sites with JavaScript-rendered content (e.g., React or Angular apps), sites requiring login (e.g., dashboards), handling multi-page results (e.g., search results), or capturing visual data (e.g., screenshots for verification). Avoid for static HTML sites where simpler tools like requests suffice.
### Key Capabilities
- Dynamically load and interact with content using Playwright's browser control.
- Manage authentication flows, such as logging in via forms or API tokens.
- Handle pagination by navigating pages, clicking "next" buttons, or parsing URLs.
- Extract data using selectors, with options for JSON output or file saves.
- Capture screenshots or full-page PDFs for debugging or reporting.
- Supports headless or visible browser modes for flexibility.
### Usage Patterns
Always initialize a browser context first, then create pages for navigation. Use async patterns for reliability. For authenticated scraping, handle cookies or sessions per context. Structure scripts to loop through pages for pagination and use try-catch for flaky elements. Pass configurations via JSON files or environment variables for reusability.
### Common Commands/API
Use Playwright's Node.js API. Install via `npm install playwright`. Key methods include:
- Launch browser: `const browser = await playwright.chromium.launch({ headless: true });`
- Navigate page: `const page = await browser.newPage(); await page.goto('https://example.com');`
- Handle auth: `await page.fill('#username', process.env.USERNAME); await page.fill('#password', process.env.PASSWORD); await page.click('#login');`
- Extract data: `const data = await page.evaluate(() => document.querySelector('#target').innerText); console.log(data);`
- Pagination: `while (await page.$('#next-button')) { await page.click('#next-button'); await page.waitForSelector('.item'); }`
- Take screenshot: `await page.screenshot({ path: 'screenshot.png' });`
CLI flags for running scripts: Use `npx playwright test` with flags like `--headed` for visible mode or `--timeout 30000` for extended waits.
### Integration Notes
Integrate by importing Playwright in Node.js projects. For auth, use environment variables like `$PLAYWRIGHT_USERNAME` and `$PLAYWRIGHT_PASSWORD` to avoid hardcoding. Configuration format: Use a JSON file for settings, e.g., `{ "url": "https://target.com", "selector": "#data-element" }`. Pass it via script args: `node scraper.js --config config.json`. For larger systems, chain with tools like Puppeteer (if migrating) or export data to databases via `page.evaluate` results. Ensure compatibility with Node.js 14+ and handle proxy settings with `browser.launch({ proxy: { server: 'http://myproxy.com:8080' } })`.
### Error Handling
Anticipate common errors like timeout on dynamic loads or selector failures. Use `page.waitForSelector` with timeouts: `await page.waitForSelector('#element', { timeout: 10000 }).catch(err => console.error('Element not found:', err));`. For network issues, wrap `page.goto` in try-catch: `try { await page.goto(url, { waitUntil: 'networkidle' }); } catch (e) { console.error('Navigation failed:', e.message); await browser.close(); }`. Handle authentication failures by checking for error elements: `if (await page.$('#error-message')) { throw new Error('Login failed'); }`. Log errors with details and retry up to 3 times using a loop.
### Concrete Usage Examples
1. **Scraping a logged-in dashboard:** First, set env vars: `export PLAYWRIGHT_USERNAME='[email protected]'` and `export PLAYWRIGHT_PASSWORD='securepass'`. Then, run: `const browser = await playwright.chromium.launch(); const page = await browser.newPage(); await page.goto('https://dashboard.com/login'); await page.fill('#username', process.env.PLAYWRIGHT_USERNAME); await page.fill('#password', process.env.PLAYWRIGHT_PASSWORD); await page.click('#submit'); const data = await page.evaluate(() => document.querySelector('#dashboard-data').innerText); console.log(data); await browser.close();` This extracts data from a protected page.
2. **Handling pagination on a search site:** Script: `const browser = await playwright.chromium.launch(); const page = await browser.newPage(); await page.goto('https://search.com?q=query'); let items = []; while (true) { items.push(...await page.$$eval('.result-item', elements => elements.map(el => el.innerText))); const nextButton = await page.$('#next-page'); if (!nextButton) break; await nextButton.click(); await page.waitForTimeout(2000); } console.log(items); await browser.close();` This collects results across multiple pages.
### Graph Relationships
- Related to: "selenium-automation" (alternative browser automation tool)
- Depends on: "node-runtime" (for Playwright execution)
- Complements: "data-extraction" (for post-processing scraped data)
- In cluster: "community" (shared with other open-source tools)
Related in Writing & Docs
jax-development
IncludedUse this skill when the user is writing, debugging, profiling, refactoring, reviewing, benchmarking, parallelising, exporting, or explaining JAX code, or when they mention JAX, jax.numpy, jit, grad, value_and_grad, vmap, scan, lax, random keys, pytrees, jax.Array, sharding, Mesh, PartitionSpec, NamedSharding, pmap, shard_map, Pallas, XLA, StableHLO, checkify, profiler, or the JAX repo. It helps turn NumPy or PyTorch-style code into pure functional JAX, fix tracer/control-flow/shape/PRNG bugs, remove recompiles and host-device syncs, choose transforms and sharding strategies, inspect jaxpr/lowering/IR, and benchmark compiled code correctly.
nature-article-writer
IncludedDrafts, rewrites, diagnostically critiques, and style-calibrates primary research manuscripts for Nature and Nature Portfolio journals. Use when the user wants a Nature-style title, summary paragraph or abstract, introduction, results, discussion, methods, figure legends, presubmission enquiry, cover letter, reviewer response, or when a scientific draft sounds generic, jargon-heavy, structurally weak, or AI-ish and needs precise, broad-reader-friendly prose without inventing data, analyses, or references. Best for primary research articles and letters rather than reviews or press releases unless explicitly adapting one.
deckrd
IncludedDocument-driven framework that derives requirements, specifications, implementation plans, and executable tasks from goals through structured AI dialogue. Use when user says "write requirements", "create spec", "plan implementation", "derive tasks", "structure this feature", "break down into tasks", or "document this module". Also use for reverse engineering existing code into docs (/deckrd rev). Do NOT use for direct code writing — use /deckrd-coder after tasks are generated. Do NOT use when the user only wants to run or fix existing code without planning.
clinical-decision-support
IncludedGenerate professional clinical decision support (CDS) documents for pharmaceutical and clinical research settings, including patient cohort analyses (biomarker-stratified with outcomes) and treatment recommendation reports (evidence-based guidelines with decision algorithms). Supports GRADE evidence grading, statistical analysis (hazard ratios, survival curves, waterfall plots), biomarker integration, and regulatory compliance. Outputs publication-ready LaTeX/PDF format optimized for drug development, clinical research, and evidence synthesis.
handling-sf-data
IncludedSalesforce data operations with 130-point scoring. Use this skill to create, update, delete, bulk import/export, generate test data, and clean up org records using sf CLI and anonymous Apex. TRIGGER when: user creates test data, performs bulk import/export, uses sf data CLI commands, needs data factory patterns for Apex tests, or needs to seed/clean records in a Salesforce org. DO NOT TRIGGER when: SOQL query writing only (use querying-soql), Apex test execution (use running-apex-tests), or metadata deployment (use deploying-metadata).
accelint-ac-to-playwright
IncludedConvert and validate acceptance criteria for Playwright test automation. Use when user asks to (1) review/evaluate/check if AC are ready for automation, (2) assess if AC can be converted as-is, (3) validate AC quality for Playwright, (4) turn AC into tests, (5) generate tests from acceptance criteria, (6) convert .md bullets or .feature Gherkin files to Playwright specs, (7) create test automation from requirements. Handles both bullet-style markdown and Gherkin syntax with JSON test plan generation and validation.