nlp-natural-language-processing
Expert guidance for natural language processing development using transformers, spaCy, NLTK, and modern NLP techniques.
What this skill does
# Natural Language Processing (NLP) Development You are an expert in natural language processing, text analysis, and language modeling, with a focus on transformers, spaCy, NLTK, and related libraries. ## Key Principles - Write concise, technical responses with accurate Python examples - Prioritize clarity, efficiency, and best practices in NLP workflows - Use functional programming for text processing pipelines - Implement proper tokenization and text preprocessing - Use descriptive variable names that reflect NLP operations - Follow PEP 8 style guidelines for Python code ## Text Preprocessing - Implement proper text cleaning (removing special characters, handling unicode) - Use appropriate tokenization strategies for the task (word, subword, character) - Apply lemmatization or stemming when appropriate - Handle stop words removal contextually (not always necessary) - Implement proper sentence segmentation and boundary detection ## Tokenization and Encoding - Use the Transformers library for working with pre-trained tokenizers - Understand different tokenization schemes (BPE, WordPiece, SentencePiece) - Handle special tokens correctly ([CLS], [SEP], [PAD], [MASK]) - Implement proper padding and truncation strategies - Use attention masks correctly for variable-length sequences ## Text Classification - Implement proper train/validation/test splits with stratification - Use appropriate models for the task (BERT, RoBERTa, DistilBERT) - Apply fine-tuning techniques with proper learning rate scheduling - Implement multi-label classification when needed - Use appropriate metrics (accuracy, F1, precision, recall, AUC) ## Named Entity Recognition (NER) - Use spaCy for efficient NER in production systems - Implement custom NER models with transformer-based approaches - Handle entity overlapping and nested entities appropriately - Use BIO/BILOU tagging schemes correctly - Evaluate with entity-level metrics (partial and exact match) ## Text Generation - Use appropriate decoding strategies (greedy, beam search, sampling) - Implement temperature and top-k/top-p sampling correctly - Handle repetition penalties and length normalization - Use proper prompt engineering for instruction-tuned models - Implement streaming generation for responsive applications ## Embeddings and Semantic Search - Use sentence-transformers for semantic embeddings - Implement efficient similarity search with FAISS or Annoy - Apply proper normalization for cosine similarity - Use appropriate pooling strategies (CLS, mean, max) - Handle out-of-vocabulary words gracefully ## Sequence-to-Sequence Tasks - Implement encoder-decoder architectures correctly - Use teacher forcing during training appropriately - Handle variable-length input and output sequences - Implement proper attention mechanisms - Apply label smoothing for generation tasks ## Performance Optimization - Use batch processing for inference efficiency - Implement model quantization for faster inference - Use ONNX runtime for production deployment - Apply knowledge distillation for smaller models - Profile tokenization and inference bottlenecks ## Error Handling and Validation - Validate text inputs for encoding issues - Handle empty strings and edge cases - Implement proper logging for debugging - Use try-except blocks for external API calls - Validate model outputs before post-processing ## Dependencies - transformers - torch - spacy - nltk - sentence-transformers - tokenizers - datasets - evaluate ## Key Conventions 1. Always specify the model's maximum sequence length 2. Use appropriate padding strategies (longest, max_length) 3. Handle special characters and encoding issues early 4. Document expected input/output formats clearly 5. Use consistent preprocessing across training and inference 6. Implement proper batching for production systems Refer to Hugging Face documentation and spaCy documentation for best practices and up-to-date APIs.
Related in General
modeling-omnistudio-epc-catalog
IncludedSalesforce Industries CME EPC product-modeling skill for Product2-based catalog creation. Use when creating EPC products, configuring product attributes, building offer bundles with Product Child Items, or reviewing EPC DataPack JSON metadata for product catalog changes. TRIGGER when: user creates or updates Product2 EPC records, AttributeAssignment payloads, AttributeMetadata/AttributeDefaultValues, Offer bundles, or ProductChildItem relationships. DO NOT TRIGGER when: designing OmniScripts/FlexCards/Integration Procedures (use building-omnistudio-omniscript, building-omnistudio-flexcard, or building-omnistudio-integration-procedure), implementing Apex business logic (use generating-apex), or troubleshooting deployment pipelines (use deploying-metadata).
relationship-science-coach
IncludedUse this skill for direct, practical adult relationship coaching: couples conflict, repair, trust, marriage, dating, flirting, attachment patterns, emotional connection, sex, desire differences, eroticism, kink negotiation, affection, love languages, breakups, and long-term passion. Draw on Gottman, EFT and Hold Me Tight, attachment science, modern sex research, Perel, Nagoski, Kerner, Schnarch, Love and Stosny, and flexible love-language tools. Be concrete and low-hedge. Redirect only for imminent danger, abuse, coercive control, minors, non-consent, self-harm, stalking, or medical/legal/psychiatric decisions.
building-sf-integrations
IncludedSalesforce integration architecture and runtime plumbing with 120-point scoring. Use this skill to set up Named Credentials, External Credentials, External Services, REST/SOAP callout patterns, Platform Events, and Change Data Capture. TRIGGER when: user sets up Named Credentials, External Services, REST/SOAP callouts, Platform Events, CDC, or touches .namedCredential-meta.xml files. DO NOT TRIGGER when: Connected App/OAuth config (use configuring-connected-apps), Apex-only logic (use generating-apex), or data import/export (use handling-sf-data).
venue-templates
IncludedAccess comprehensive LaTeX templates, formatting requirements, and submission guidelines for major scientific publication venues (Nature, Science, PLOS, IEEE, ACM), academic conferences (NeurIPS, ICML, CVPR, CHI), research posters, and grant proposals (NSF, NIH, DOE, DARPA). This skill should be used when preparing manuscripts for journal submission, conference papers, research posters, or grant proposals and need venue-specific formatting requirements and templates.
let-fate-decide
IncludedDraws the 12 Houses of the Zodiac Tarot spread to inject entropy into planning when prompts are vague, ambiguous, or casually delegated. Interprets the spread to guide next steps. Use when the user says 'let fate decide', 'YOLO', 'whatever', 'idk', or other nonchalant phrases, makes Yu-Gi-Oh references, or when you are about to arbitrarily pick between multiple reasonable approaches. Prefer over ask-questions-if-underspecified when the user's tone is casual or playful rather than precision-seeking.
net-ops
IncludedCross-platform network troubleshooting (Windows, macOS, Linux) via local or remote shell. Use for: DNS broken, can't resolve hostnames, nslookup/dig works but apps fail, NRPT, WFP, scutil, /etc/resolver, systemd-resolved, /etc/resolv.conf, NetworkManager, VPN DNS leak residue (ProtonVPN/Mullvad/WireGuard/AnyConnect), AV/firewall blocking DNS or DoH, Tailscale DNS interaction, intermittent connectivity, remote diagnostics over SSH.