Skip to content

Latest commit

 

History

21 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Databricks Asset Bundle CI/CD Databricks Asset Bundle CI/CD Databricks Asset Bundle CI/CD

Databricks Asset Bundle (DAB) - Customer Data Pipeline

Databricks PySpark GitHub Actions License: MIT An enterprise-ready Databricks Asset Bundle (DAB) template demonstrating production data pipeline orchestration using Delta Live Tables (DLT), PySpark Pipelines, Unity Catalog, and automated CI/CD with GitHub Actions.

📋 Table of Contents


🌟 Overview

This repository provides a modular, declarative structure for managing data engineering workloads on Databricks. Utilizing Databricks Asset Bundles (DABs), infrastructure and code are treated as unified software artifacts—enabling seamless deployment across Development, QA, and Production environments.

🏗️ Architecture

The pipeline implements a streaming Medallion Architecture using PySpark Delta Live Tables:

flowchart LR
    subgraph Source["Data Source"]
        TPCH["tpch.customer (Sample Data)"]
    end
    subgraph Medallion["Medallion Pipeline"]
        Bronze["Bronze Layer\n(raw_customer)\n+ Ingestion Timestamp"]
        Silver["Silver Layer\n(customer_silver)\n+ Regex Sanitization\n+ Schema Formatting\n+ Processing Timestamp"]
    end
    subgraph Governance["Unity Catalog"]
        UC["Catalog: ${var.catalog_name}\nSchema: ${var.schema_name}"]
    end
    TPCH -->|Stream Read| Bronze
    Bronze -->|Stream Read & Transform| Silver
    Medallion --- Governance
Loading

🚀 Getting Started

1. Clone the Repository

git clone https://github.com/hridoy1335/databricks_asset_bundle.git
cd databricks_asset_bundle

2. Authenticate Databricks CLI

Authenticate with your Databricks target workspace:

databricks configure --host https://<your-databricks-workspace-instance>

3. Validate the Bundle Configuration

Verify the DAB configuration files for your target environment (dev or qa):

# Validate development environment
databricks bundle validate -t dev
# Validate QA environment
databricks bundle validate -t qa

4. Deploy to Databricks Workspace

Deploy assets to your workspace root directory:

# Deploy to Development
databricks bundle deploy -t dev

5. Run the Pipeline Job

Trigger the deployed job pipeline:

databricks bundle run dab_job -t dev

🔄 CI/CD Pipeline

Automated continuous integration and deployment are configured using GitHub Actions (.github/workflows/main.yml).

Workflow Steps:

  1. Validate: Automatically runs databricks bundle validate -t qa on pull requests or commits to main.
  2. Deploy QA: Executes databricks bundle deploy -t qa upon validation success.

Required GitHub Secrets:

To enable CI/CD deployment, configure the following repository secrets under Settings ➔ Secrets and variables ➔ Actions:

  • DATABRICKS_HOST_QA: Databricks instance URL for QA workspace.
  • DATABRICKS_TOKEN_QA: Personal Access Token (PAT) or Service Principal Token for QA workspace.
  • (Optional) DATABRICKS_HOST_PROD & DATABRICKS_TOKEN_PROD for production deployments.

📜 License

This project is licensed under the MIT License.

About

databricks_asset_bundle

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages