Skip to content

Repository files navigation

Testbricks

image

workflow

workflow

Databricks notebooks are awsome way for interactive development. The inbuilt IDE is fantastic, stable, fully featured and has the right AI assistance levels. So what's the problem ?

Testing is the main problem, more specifically Unit testing.

Testbricks is my effort to decouple a typical Databricks Stack -

  • Delta lake and Unity Catalog etc for storing data
  • Notebooks with Pyspark code to transform said data
  • Databricks jobs to orchestrate a bunch of notebooks

Quickstart

pip install testbricks

PySpark needs a JDK on PATH (Java 8+). Create a SparkProxy, import dbutils, and run a Databricks workflow JSON with LocalWorkflowRunner:

from testbricks import SparkProxy, LocalWorkflowRunner
from testbricks.dbutils import dbutils

# CSV tables live under {base_path}/{schema}/{table}.csv
# e.g. ./data/bronze/customers.csv  →  spark.read.table("bronze.customers")
spark = SparkProxy("./data")

# Optional: use dbutils in the driver process (notebooks get it automatically)
dbutils.widgets.text("filter_country", "USA")
print(dbutils.widgets.get("filter_country"))

runner = LocalWorkflowRunner(
    source_dir="./notebooks",          # local .py files named after the notebook
    workflow_json_path="./workflow.json",
    base_path="./data",                # same catalog root as SparkProxy
)
runner.run_workflow(extra_globals={"spark": spark})

workflow.json is a Databricks job export with a tasks list. Notebook paths are resolved to {source_dir}/{last_path_segment}.py:

{
  "tasks": [
    {
      "task_key": "enrich_customers",
      "notebook_task": {
        "notebook_path": "/Workspace/jobs/enrich_customers",
        "base_parameters": { "filter_country": "USA" }
      }
    },
    {
      "task_key": "build_summary",
      "depends_on": [{ "task_key": "enrich_customers" }],
      "notebook_task": {
        "notebook_path": "/Workspace/jobs/build_summary"
      }
    }
  ]
}

A notebook such as notebooks/enrich_customers.py can use the usual Databricks names — spark is injected via extra_globals, and dbutils is injected by the runner:

dbutils.widgets.text("filter_country", "ALL")
country = dbutils.widgets.get("filter_country")

df = spark.read.option("header", "true").option("inferSchema", "true").table("bronze.customers")
if country != "ALL":
    df = df.filter(df.country == country)

df.write.mode("overwrite").saveAsTable("silver.customers_enriched")

Spark write modes and schema options

Table writes (saveAsTable / insertInto) stay CSV-backed. File writes (csv / parquet / json / save) use native Spark under base_path. format("delta").save(path) is stored as parquet (no Delta log).

Write Missing table Existing table
default / overwrite create replace rows
append create append rows (exact column set, unless mergeSchema)
error / errorIfExists create raise AnalysisException
ignore create no-op (file and temp view unchanged)
insertInto raise AnalysisException append, or replace when overwrite=True / mode("overwrite")

Schema flags:

Option Effect
overwriteSchema=true + overwrite replace the CSV even when columns change
overwriteSchema=false (default) + overwrite raise SchemaMismatchError on incompatible schema change
mergeSchema=true + append union missing columns with nulls
mergeSchema=false (default) + append raise SchemaMismatchError if columns differ
same column, incompatible types always raise SchemaMismatchError (merge only adds columns)

partitionBy columns must exist on the DataFrame. option("replaceWhere", "<predicate>") with mode("overwrite") deletes matching stored rows then appends the new frame. bucketBy / sortBy are accepted no-ops (bucketing is not simulated). df.writeTo(table).using(...).create() / .replace() / .append() maps onto the same table writer; createOrReplace and overwritePartitions raise NotImplementedError with a saveAsTable hint.

Repair-and-rerun

LocalWorkflowRunner.run_workflow can re-run a subgraph without executing the full DAG:

runner.run_workflow(extra_globals={"spark": spark}, only=["build_summary"])
runner.run_workflow(extra_globals={"spark": spark}, from_task="enrich_customers")
  • only=["task_a", ...] runs those task keys in topological order.
  • from_task="task_a" runs that task and every downstream dependent.
  • Tasks not selected are treated as already SUCCESS so run_if / depends_on still resolve.

Overwrite table writes (mode("overwrite") / default) are naturally idempotent on re-run. append writes are not: a repair will insert another copy of the rows. For append tables, use a repair policy in the notebook (skip if the target already has the batch, dedupe on a key, or pass a flag that switches the write to overwrite for that run).

Lakeflow JSON extras the runner understands:

  • dbutils.jobs.taskValues.set / .get across tasks (shared immediately; isolated notebook.run commits on return)
  • run_if and depends_on[].outcome (ALL_SUCCESS, ALL_FAILED, AT_LEAST_ONE_SUCCESS, ALL_DONE, NONE_FAILED, AT_LEAST_ONE_FAILED); ineligible tasks are SKIPPED
  • condition_task (EQUAL / EQUAL_TO, NOT_EQUAL, GREATER_THAN, …) with {{tasks.<key>.values.<name>}} operands
  • for_each_task sequential expansion of inputs with {{input}} parameters
  • max_retries / min_retry_interval_millis (retry notebook and for_each tasks); timeout_seconds is accepted and logged, not enforced

Key Modules

  1. SparkProxy - A Spark proxy that manipulates incoming Delta table reads and writes and redirects them to interactions with CSV files stored locally
  2. LocalWorkflowRunner - A notebook orchestrator that takes the notebook .py files as defined in a Databricks Workflow JSON file and executes them as per the DAG definition. Databricks comment magics %run, %sh (# %sh / # MAGIC %sh, including %sh -e), and %fs (# %fs / # MAGIC %fs) work in those notebooks.
  3. dbutils - A drop in replacement for Databricks dbutils. Supports fs (ls/put/cp/mv/rm/mkdirs), widgets (including combobox/multiselect/getAll), secrets (get/list/listScopes), notebook (exit / %run), library.restartPython, and data.summarize.

About

Run your PySpark Databricks workflow on a local developer environment seamlessly. Built using Cursor Start

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages