databricks-2025
ADF + Databricks 2025 integration patterns. PROACTIVELY activate for: (1) Databricks Job activity in ADF, (2) DatabricksJob (preview) vs DatabricksNotebook activity, (3) ServiceNow V2 connector, (4) ADF managed identity authentication for Databricks, (5) Databricks serverless linked services, (6) Snowflake V2 connector, (7) Databricks job parameters and outputs, (8) MFA enforcement and authentication updates, (9) Unity Catalog integration, (10) Delta Live Tables orchestration from ADF. Provides: Databricks linked service templates (PAT, MSI, serverless), DatabricksJob activity examples, parameter passing recipes, and authentication migration guidance.
What this skill does
# Azure Data Factory Databricks Integration 2025 ## Databricks Job Activity (Recommended 2025) **CRITICAL UPDATE (2025):** The Databricks Job activity is now the **ONLY recommended method** for orchestrating Databricks in ADF. Microsoft strongly recommends migrating from legacy Notebook, Python, and JAR activities. ### Quick Reference - **Activity type:** `DatabricksJob` (NOT `DatabricksSparkJob` or `DatabricksNotebook`) - **Parameter property:** `jobParameters` (NOT `parameters`) - **Linked service auth:** Managed Identity (`"authentication": "MSI"`) recommended - **Cluster config:** Do NOT specify cluster properties in linked service; the Databricks Job controls compute ### Why Databricks Job Activity? | Feature | Notebook Activity (Legacy) | Job Activity (2025) | |---------|---------------------------|---------------------| | Compute | Must configure cluster in linked service | Serverless by default | | Workflow tasks | Single notebook | Multi-task DAGs (notebook, Python, SQL, DLT) | | Retry | ADF-level only | Job-level + task-level | | Repair runs | Not supported | Rerun failed tasks only | | Git integration | Limited | Full Databricks Git support + DABs | | Lineage | None | Built-in data lineage | | If/Else logic | Must use ADF control flow | Native If/Else task types | ### Benefits Summary 1. **Serverless Execution** -- No cluster specification needed; automatic serverless compute with faster startup and lower costs 2. **Advanced Workflow Features** -- Run As, Task Values, Conditional Execution, AI/BI Tasks, Repair Runs, Notifications, Queuing 3. **Centralized Job Management** -- Jobs defined once in Databricks workspace; single source of truth with Git-backed versioning 4. **Cost Optimization** -- Serverless compute (pay only for execution), job clusters (auto-terminating), spot instance support For complete JSON examples of Job activity, linked service, and pipeline configurations, see `references/databricks-job-examples.md`. ## Connectors and Enhancements (2025+) ### ServiceNow V2 Connector (RECOMMENDED - V1 End of Support) **ServiceNow V1 connector is at End of Support. Migrate to V2 immediately.** | Feature | V1 | V2 | |---------|----|----| | Linked service type | `ServiceNow` | `ServiceNowV2` | | Source type | `ServiceNowSource` | `ServiceNowV2Source` | | Query builder | Custom | Aligns with ServiceNow condition builder | | Performance | Standard | Enhanced extraction | | OData support | No | Yes | **Migration steps:** Update linked service type to `ServiceNowV2`, update source type to `ServiceNowV2Source`, test queries in ServiceNow UI condition builder, adjust timeouts. ### Enhanced PostgreSQL Connector Improved performance with 2025 SSL enhancements: `enableSsl: true`, `sslMode: "Require"`. ### Enhanced Snowflake Connector Improved performance with KeyPair authentication support and Key Vault secret integration. ### Managed Identity for Azure Storage New managed identity support for Azure Table Storage and Azure Files connectors (system-assigned and user-assigned). ### Mapping Data Flows - Spark 3.3 Spark 3.3 now powers Mapping Data Flows with 30% faster processing, Adaptive Query Execution (AQE), dynamic partition pruning, improved caching, and better column statistics. ### Azure DevOps Server 2022 Support Git integration now supports on-premises Azure DevOps Server 2022 via the `hostName` property. For complete JSON examples of all connectors, see `references/connector-examples.md`. ## Managed Identity 2025 Best Practices ### User-Assigned vs System-Assigned | Scenario | Recommendation | |----------|---------------| | Single ADF, simple setup | System-assigned | | Multiple data factories | User-assigned (shared identity) | | Complex multi-environment | User-assigned | | Granular permission control | User-assigned | | Identity lifecycle independence | User-assigned | Use ADF's centralized **Credentials** feature to consolidate Microsoft Entra ID-based credentials across multiple linked services. ### MFA Enforcement (Enforced Since October 2025) Azure MFA is mandatory for all interactive user logins. Impact on ADF: - Managed identities are **UNAFFECTED** -- no MFA required for service accounts - Service principals with certificate auth are the recommended alternative to secrets - All interactive user logins require MFA ### Principle of Least Privilege | Resource | Source Role | Sink Role | |----------|-----------|-----------| | Storage Blob | `Storage Blob Data Reader` | `Storage Blob Data Contributor` | | SQL Database | `db_datareader` | `db_datareader` + `db_datawriter` | | Key Vault | `Get` secrets only | `Get` secrets only | For complete managed identity JSON examples, see `references/connector-examples.md`. ## Best Practices (2025) 1. **Use Databricks Job Activity (MANDATORY)** -- Stop using Notebook, Python, JAR activities. Define workflows in Databricks workspace with serverless compute. 2. **Managed Identity Authentication (MANDATORY)** -- Use managed identities for ALL Azure resources. Leverage Credentials feature for consolidation. MFA-compliant since October 2025. 3. **Monitor Job Execution** -- Track Databricks Job run IDs from ADF output, log parameters for auditability, set up alerts for failures, leverage built-in lineage. 4. **Optimize Spark 3.3 (Data Flows)** -- Enable AQE, use 4-8 partitions per core, broadcast joins for small dimensions, dynamic partition pruning. ## Resources - [Databricks Job Activity](https://learn.microsoft.com/azure/data-factory/transform-data-using-databricks-spark-job) - [ADF Connectors](https://learn.microsoft.com/azure/data-factory/connector-overview) - [Managed Identity Authentication](https://learn.microsoft.com/azure/data-factory/data-factory-service-identity) - [Mapping Data Flows](https://learn.microsoft.com/azure/data-factory/concepts-data-flow-overview) ## Progressive Disclosure References - **Databricks Job Examples**: `references/databricks-job-examples.md` - Complete JSON for Job activity, linked services, pipeline, and Databricks workspace job definition - **Connector Examples**: `references/connector-examples.md` - Complete JSON for ServiceNow V2, PostgreSQL, Snowflake, Azure Storage MI, Mapping Data Flows, and Azure DevOps Server
Related in General
modeling-omnistudio-epc-catalog
IncludedSalesforce Industries CME EPC product-modeling skill for Product2-based catalog creation. Use when creating EPC products, configuring product attributes, building offer bundles with Product Child Items, or reviewing EPC DataPack JSON metadata for product catalog changes. TRIGGER when: user creates or updates Product2 EPC records, AttributeAssignment payloads, AttributeMetadata/AttributeDefaultValues, Offer bundles, or ProductChildItem relationships. DO NOT TRIGGER when: designing OmniScripts/FlexCards/Integration Procedures (use building-omnistudio-omniscript, building-omnistudio-flexcard, or building-omnistudio-integration-procedure), implementing Apex business logic (use generating-apex), or troubleshooting deployment pipelines (use deploying-metadata).
relationship-science-coach
IncludedUse this skill for direct, practical adult relationship coaching: couples conflict, repair, trust, marriage, dating, flirting, attachment patterns, emotional connection, sex, desire differences, eroticism, kink negotiation, affection, love languages, breakups, and long-term passion. Draw on Gottman, EFT and Hold Me Tight, attachment science, modern sex research, Perel, Nagoski, Kerner, Schnarch, Love and Stosny, and flexible love-language tools. Be concrete and low-hedge. Redirect only for imminent danger, abuse, coercive control, minors, non-consent, self-harm, stalking, or medical/legal/psychiatric decisions.
building-sf-integrations
IncludedSalesforce integration architecture and runtime plumbing with 120-point scoring. Use this skill to set up Named Credentials, External Credentials, External Services, REST/SOAP callout patterns, Platform Events, and Change Data Capture. TRIGGER when: user sets up Named Credentials, External Services, REST/SOAP callouts, Platform Events, CDC, or touches .namedCredential-meta.xml files. DO NOT TRIGGER when: Connected App/OAuth config (use configuring-connected-apps), Apex-only logic (use generating-apex), or data import/export (use handling-sf-data).
venue-templates
IncludedAccess comprehensive LaTeX templates, formatting requirements, and submission guidelines for major scientific publication venues (Nature, Science, PLOS, IEEE, ACM), academic conferences (NeurIPS, ICML, CVPR, CHI), research posters, and grant proposals (NSF, NIH, DOE, DARPA). This skill should be used when preparing manuscripts for journal submission, conference papers, research posters, or grant proposals and need venue-specific formatting requirements and templates.
let-fate-decide
IncludedDraws the 12 Houses of the Zodiac Tarot spread to inject entropy into planning when prompts are vague, ambiguous, or casually delegated. Interprets the spread to guide next steps. Use when the user says 'let fate decide', 'YOLO', 'whatever', 'idk', or other nonchalant phrases, makes Yu-Gi-Oh references, or when you are about to arbitrarily pick between multiple reasonable approaches. Prefer over ask-questions-if-underspecified when the user's tone is casual or playful rather than precision-seeking.
net-ops
IncludedCross-platform network troubleshooting (Windows, macOS, Linux) via local or remote shell. Use for: DNS broken, can't resolve hostnames, nslookup/dig works but apps fail, NRPT, WFP, scutil, /etc/resolver, systemd-resolved, /etc/resolv.conf, NetworkManager, VPN DNS leak residue (ProtonVPN/Mullvad/WireGuard/AnyConnect), AV/firewall blocking DNS or DoH, Tailscale DNS interaction, intermittent connectivity, remote diagnostics over SSH.