pandas-best-practices
Best practices for Pandas data manipulation, analysis, and DataFrame operations in Python
What this skill does
# Pandas Best Practices Expert guidelines for Pandas development, focusing on data manipulation, analysis, and efficient DataFrame operations. ## Code Style and Structure - Write concise, technical responses with accurate Python examples - Prioritize reproducibility in data analysis workflows - Use functional programming; avoid unnecessary classes - Prefer vectorized operations over explicit loops - Use descriptive variable names reflecting data content - Follow PEP 8 style guidelines ## DataFrame Creation and I/O - Use `pd.read_csv()`, `pd.read_excel()`, `pd.read_json()` with appropriate parameters - Specify `dtype` parameter to ensure correct data types on load - Use `parse_dates` for automatic datetime parsing - Set `index_col` when the data has a natural index column - Use `chunksize` for reading large files incrementally ## Data Selection - Use `.loc[]` for label-based indexing - Use `.iloc[]` for integer position-based indexing - Avoid chained indexing (e.g., `df['col'][0]`) - use `.loc` or `.iloc` instead - Use boolean indexing for conditional selection: `df[df['col'] > value]` - Use `.query()` method for complex filtering conditions ## Method Chaining - Prefer method chaining for data transformations when possible - Use `.pipe()` for applying custom functions in a chain - Chain operations like `.assign()`, `.query()`, `.groupby()`, `.agg()` - Keep chains readable by breaking across multiple lines ## Data Cleaning and Validation ### Missing Data - Check for missing data with `.isna()` and `.info()` - Handle missing data appropriately: `.fillna()`, `.dropna()`, or imputation - Use `pd.NA` for nullable integer and boolean types - Document decisions about missing data handling ### Data Quality Checks - Implement data quality checks at the beginning of analysis - Validate data types with `.dtypes` and convert as needed - Check for duplicates with `.duplicated()` and handle appropriately - Use `.describe()` for quick statistical overview ### Type Conversion - Use `.astype()` for explicit type conversion - Use `pd.to_datetime()` for date parsing - Use `pd.to_numeric()` with `errors='coerce'` for safe numeric conversion - Utilize categorical data types for low-cardinality string columns ## Grouping and Aggregation ### GroupBy Operations - Use `.groupby()` for efficient aggregation operations - Specify aggregation functions with `.agg()` for multiple operations - Use named aggregation for clearer output column names - Consider `.transform()` for broadcasting results back to original shape ### Pivot Tables and Reshaping - Use `.pivot_table()` for multi-dimensional aggregation - Use `.melt()` to convert wide to long format - Use `.pivot()` to convert long to wide format - Use `.stack()` and `.unstack()` for hierarchical index manipulation ## Performance Optimization ### Memory Efficiency - Use categorical data types for low-cardinality strings - Downcast numeric types when appropriate - Use `pd.eval()` and `.eval()` for large expression evaluation ### Computation Speed - Use vectorized operations instead of `.apply()` with row-wise functions - Prefer built-in aggregation functions over custom ones - Use `.values` or `.to_numpy()` for NumPy operations when faster ### Avoiding Common Pitfalls - Avoid iterating with `.iterrows()` - use vectorized operations - Don't modify DataFrames while iterating - Be aware of SettingWithCopyWarning - use `.copy()` when needed - Avoid growing DataFrames row by row - collect in list and create once ## Time Series Operations - Use `DatetimeIndex` for time series data - Leverage `.resample()` for time-based aggregation - Use `.shift()` and `.diff()` for lag operations - Use `.rolling()` and `.expanding()` for window calculations ## Merging and Joining - Use `.merge()` for SQL-style joins - Specify `how` parameter: 'inner', 'outer', 'left', 'right' - Use `validate` parameter to check join cardinality - Use `.concat()` for stacking DataFrames ## Key Conventions - Import as `import pandas as pd` - Use `snake_case` for column names when possible - Document data sources and transformations - Keep notebooks reproducible with clear cell execution order
Related in General
modeling-omnistudio-epc-catalog
IncludedSalesforce Industries CME EPC product-modeling skill for Product2-based catalog creation. Use when creating EPC products, configuring product attributes, building offer bundles with Product Child Items, or reviewing EPC DataPack JSON metadata for product catalog changes. TRIGGER when: user creates or updates Product2 EPC records, AttributeAssignment payloads, AttributeMetadata/AttributeDefaultValues, Offer bundles, or ProductChildItem relationships. DO NOT TRIGGER when: designing OmniScripts/FlexCards/Integration Procedures (use building-omnistudio-omniscript, building-omnistudio-flexcard, or building-omnistudio-integration-procedure), implementing Apex business logic (use generating-apex), or troubleshooting deployment pipelines (use deploying-metadata).
relationship-science-coach
IncludedUse this skill for direct, practical adult relationship coaching: couples conflict, repair, trust, marriage, dating, flirting, attachment patterns, emotional connection, sex, desire differences, eroticism, kink negotiation, affection, love languages, breakups, and long-term passion. Draw on Gottman, EFT and Hold Me Tight, attachment science, modern sex research, Perel, Nagoski, Kerner, Schnarch, Love and Stosny, and flexible love-language tools. Be concrete and low-hedge. Redirect only for imminent danger, abuse, coercive control, minors, non-consent, self-harm, stalking, or medical/legal/psychiatric decisions.
building-sf-integrations
IncludedSalesforce integration architecture and runtime plumbing with 120-point scoring. Use this skill to set up Named Credentials, External Credentials, External Services, REST/SOAP callout patterns, Platform Events, and Change Data Capture. TRIGGER when: user sets up Named Credentials, External Services, REST/SOAP callouts, Platform Events, CDC, or touches .namedCredential-meta.xml files. DO NOT TRIGGER when: Connected App/OAuth config (use configuring-connected-apps), Apex-only logic (use generating-apex), or data import/export (use handling-sf-data).
venue-templates
IncludedAccess comprehensive LaTeX templates, formatting requirements, and submission guidelines for major scientific publication venues (Nature, Science, PLOS, IEEE, ACM), academic conferences (NeurIPS, ICML, CVPR, CHI), research posters, and grant proposals (NSF, NIH, DOE, DARPA). This skill should be used when preparing manuscripts for journal submission, conference papers, research posters, or grant proposals and need venue-specific formatting requirements and templates.
let-fate-decide
IncludedDraws the 12 Houses of the Zodiac Tarot spread to inject entropy into planning when prompts are vague, ambiguous, or casually delegated. Interprets the spread to guide next steps. Use when the user says 'let fate decide', 'YOLO', 'whatever', 'idk', or other nonchalant phrases, makes Yu-Gi-Oh references, or when you are about to arbitrarily pick between multiple reasonable approaches. Prefer over ask-questions-if-underspecified when the user's tone is casual or playful rather than precision-seeking.
net-ops
IncludedCross-platform network troubleshooting (Windows, macOS, Linux) via local or remote shell. Use for: DNS broken, can't resolve hostnames, nslookup/dig works but apps fail, NRPT, WFP, scutil, /etc/resolver, systemd-resolved, /etc/resolv.conf, NetworkManager, VPN DNS leak residue (ProtonVPN/Mullvad/WireGuard/AnyConnect), AV/firewall blocking DNS or DoH, Tailscale DNS interaction, intermittent connectivity, remote diagnostics over SSH.