scrape-posts
Scrape new articles from Milan Jovanovic's blog (November 2025+). Optimized - pre-filters from listing page, only scrapes new articles.
What this skill does
# Scrape Milan Jovanovic Blog Posts
Scrape new articles from Milan Jovanovic's .NET blog with **optimized pre-filtering**. Parses dates from listing page to avoid unnecessary per-article scraping.
## Arguments
- `--force`: Re-scrape all articles (compare content hash to skip unchanged)
- `--since YYYY-MM-DD`: Custom date filter (default: 2025-11-01)
- `--limit N`: Limit number of articles (for testing)
- `--dry-run`: Preview what would be scraped without saving
## Optimized Workflow
### Step 1: Invoke Skill
Invoke the `milan-jovanovic:milan-jovanovic-blog` skill to load context and access scripts.
### Step 2: Pre-Filter from Listing Page (OPTIMIZATION)
**Key efficiency optimization:** Parse dates from listing page BEFORE scraping individual articles.
1. Scrape the blog listing page using `firecrawl_scrape`:
```text
URL: https://www.milanjovanovic.tech/blog
Format: markdown
```
2. Save listing content to temp file (e.g., `.claude/temp/milan-listing.md`)
3. Run pre-filter script to identify articles needing scraping:
```bash
# Normal mode - only new articles
python scripts/core/check_new_articles.py .claude/temp/milan-listing.md --json --since 2025-11-01
# Force mode - include existing for re-check
python scripts/core/check_new_articles.py .claude/temp/milan-listing.md --json --force --since 2025-11-01
```
4. Parse JSON output to get `to_scrape` list. If empty, skip to Step 5 (no scraping needed).
### Step 3: Scrape Only Needed Articles
For each article in `to_scrape`:
1. **For articles with `in_index: false`** (new):
- Scrape full article with `firecrawl_scrape`
- Extract publication date from metadata
- Clean promotional content
- Save to `canonical/milanjovanovic-tech/blog/{slug}.md`
2. **For articles with `in_index: true`** (force mode re-check):
- Scrape full article with `firecrawl_scrape`
- Clean promotional content
- Generate content hash
- Compare to `content_hash` from pre-filter output
- If unchanged, skip writing (log as "skipped - unchanged")
- If changed, save updated content
### Step 4: Update Index
After scraping completes:
```bash
python scripts/management/refresh_index.py
```
### Step 5: Report Statistics
Report:
- Articles found on listing page
- Articles needing scraping (new + force re-check)
- Articles skipped (already indexed, not in force mode)
- Articles skipped (unchanged content hash, force mode)
- Articles filtered (before cutoff date)
- Any errors
## Content Cleanup Patterns
The scraper removes these promotional patterns:
**Footer patterns (stop processing):**
- "Whenever you're ready, there are X ways I can help you"
- "Become a Better .NET Software Engineer"
- "Hi, I'm Milan"
**Sponsor patterns (remove section):**
- AuthKit/WorkOS mentions
- "Sponsor this newsletter" links
- Incident response sponsor content
**Inline patterns (remove):**
- Reading time ("5 min read")
- "Manage read history" links
- Empty image placeholders
## Efficiency Gains
| Scenario | Without Optimization | With Optimization |
|----------|----------------------|-------------------|
| No new articles | 10+ firecrawl requests | 1-2 requests |
| 1 new article | 10+ firecrawl requests | 2-3 requests |
| Force (unchanged) | 10+ requests | 10+ requests but skips writes |
**Why this matters:** Firecrawl has API costs and rate limits. Pre-filtering saves 80-90% of requests when articles haven't changed.
## Example Usage
```text
/milan-jovanovic:scrape-posts
/milan-jovanovic:scrape-posts --limit 3 --dry-run
/milan-jovanovic:scrape-posts --force
/milan-jovanovic:scrape-posts --since 2025-12-01
```
## Troubleshooting
### Firecrawl Not Available
If firecrawl MCP is not connected, the command will fail. Ensure the firecrawl MCP server is configured and running.
### Date Parsing Issues
If listing page dates can't be parsed, the script logs them in `no_date` category. These articles are skipped unless you provide a specific URL.
### Pre-Filter Shows 0 Articles
If `check_new_articles.py` shows 0 articles to scrape:
- All articles are already indexed (use `--force` to re-check)
- All articles are before the cutoff date (adjust `--since`)
- Listing page format changed (check regex patterns in script)
Related in Writing & Docs
jax-development
IncludedUse this skill when the user is writing, debugging, profiling, refactoring, reviewing, benchmarking, parallelising, exporting, or explaining JAX code, or when they mention JAX, jax.numpy, jit, grad, value_and_grad, vmap, scan, lax, random keys, pytrees, jax.Array, sharding, Mesh, PartitionSpec, NamedSharding, pmap, shard_map, Pallas, XLA, StableHLO, checkify, profiler, or the JAX repo. It helps turn NumPy or PyTorch-style code into pure functional JAX, fix tracer/control-flow/shape/PRNG bugs, remove recompiles and host-device syncs, choose transforms and sharding strategies, inspect jaxpr/lowering/IR, and benchmark compiled code correctly.
nature-article-writer
IncludedDrafts, rewrites, diagnostically critiques, and style-calibrates primary research manuscripts for Nature and Nature Portfolio journals. Use when the user wants a Nature-style title, summary paragraph or abstract, introduction, results, discussion, methods, figure legends, presubmission enquiry, cover letter, reviewer response, or when a scientific draft sounds generic, jargon-heavy, structurally weak, or AI-ish and needs precise, broad-reader-friendly prose without inventing data, analyses, or references. Best for primary research articles and letters rather than reviews or press releases unless explicitly adapting one.
deckrd
IncludedDocument-driven framework that derives requirements, specifications, implementation plans, and executable tasks from goals through structured AI dialogue. Use when user says "write requirements", "create spec", "plan implementation", "derive tasks", "structure this feature", "break down into tasks", or "document this module". Also use for reverse engineering existing code into docs (/deckrd rev). Do NOT use for direct code writing — use /deckrd-coder after tasks are generated. Do NOT use when the user only wants to run or fix existing code without planning.
clinical-decision-support
IncludedGenerate professional clinical decision support (CDS) documents for pharmaceutical and clinical research settings, including patient cohort analyses (biomarker-stratified with outcomes) and treatment recommendation reports (evidence-based guidelines with decision algorithms). Supports GRADE evidence grading, statistical analysis (hazard ratios, survival curves, waterfall plots), biomarker integration, and regulatory compliance. Outputs publication-ready LaTeX/PDF format optimized for drug development, clinical research, and evidence synthesis.
handling-sf-data
IncludedSalesforce data operations with 130-point scoring. Use this skill to create, update, delete, bulk import/export, generate test data, and clean up org records using sf CLI and anonymous Apex. TRIGGER when: user creates test data, performs bulk import/export, uses sf data CLI commands, needs data factory patterns for Apex tests, or needs to seed/clean records in a Salesforce org. DO NOT TRIGGER when: SOQL query writing only (use querying-soql), Apex test execution (use running-apex-tests), or metadata deployment (use deploying-metadata).
accelint-ac-to-playwright
IncludedConvert and validate acceptance criteria for Playwright test automation. Use when user asks to (1) review/evaluate/check if AC are ready for automation, (2) assess if AC can be converted as-is, (3) validate AC quality for Playwright, (4) turn AC into tests, (5) generate tests from acceptance criteria, (6) convert .md bullets or .feature Gherkin files to Playwright specs, (7) create test automation from requirements. Handles both bullet-style markdown and Gherkin syntax with JSON test plan generation and validation.