Skip to main content

Setting up Tuva on Databricks

This walkthrough runs Tuva Core 1.0 on the small synthetic dataset using a Databricks SQL warehouse. For your own data, follow Getting Started to create a connector project that supplies the Input Layer, then use the Databricks connection and storage setup below.

Prerequisites​

  • Git, Python 3.12, and a Bash-compatible shell for the commands below.
  • A Databricks SQL warehouse you can use, plus its server hostname and HTTP path from Connection details.
  • A development catalog and schema where your dbt identity can create and modify tables and views. With Unity Catalog, it needs USE CATALOG, USE SCHEMA, and the relevant object privileges; it also needs CREATE SCHEMA if dbt will create Tuva's output schemas.
  • Databricks authentication configured for your identity. The example uses a personal access token supplied through an environment variable. OAuth and service-principal options are described in the Databricks connection documentation.
  • Access from Databricks compute to Tuva's public S3 asset paths under s3://tuva-public-resources/. Tuva's Databricks loader uses COPY INTO from S3, including when the Databricks workspace is hosted on another cloud.

For Unity Catalog, ask your administrator to configure the applicable external location and grant READ FILES, along with access to the target catalog and schemas. Public object availability does not by itself grant access through your workspace's storage controls. See Databricks COPY INTO permissions.

1. Get the tagged integration project​

git clone --branch v1.0.0 --single-branch https://github.com/tuva-health/tuva-core.git
cd tuva-core

The integration_tests project maps the package's synthetic data to the Input Layer. It installs this local Core checkout and all eight standalone marts at their v1.0.0 tags. The build selection below runs Core and the synthetic Input Layer first.

2. Install dbt and the adapter​

These versions match the Databricks configuration in Tuva's 1.0 release CI:

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install "dbt-core==1.11.14" "dbt-databricks==1.12.4" "sqlparse==0.6.0"

This is a tested configuration, not the minimum supported version. Core's declared dbt range is >=1.10.5,<3.0.0. This walkthrough uses dbt Core; see runtime coverage for the separate Core 2 and Fusion validation scope. The Databricks adapter uses Python APIs and does not require an ODBC driver.

3. Configure a local profile​

Create a separate profile directory to preserve any existing dbt profiles:

mkdir -p .dbt-profile
export DBT_PROFILES_DIR="$PWD/.dbt-profile"

Save this as .dbt-profile/profiles.yml:

default:
target: dev
outputs:
dev:
type: databricks
host: "{{ env_var('DATABRICKS_HOST') }}"
http_path: "{{ env_var('DATABRICKS_HTTP_PATH') }}"
catalog: "{{ env_var('DATABRICKS_CATALOG') }}"
schema: "{{ env_var('DATABRICKS_SCHEMA') }}"
token: "{{ env_var('DBT_ENV_SECRET_DATABRICKS_TOKEN') }}"
threads: 4

Set the referenced environment variables in your shell or secrets manager. The hostname must omit https://; use the HTTP path from the SQL warehouse's connection details. Choose a development catalog and schema that your identity can write to. Keep the token out of the project and version control. The profile name default matches the integration project's dbt_project.yml.

If you use OAuth instead, replace the token configuration with the documented OAuth fields for your adapter. Authentication to Databricks and permission to read the S3 assets are separate prerequisites.

4. Install packages and check the connection​

Run these commands from the tuva-core directory. The helper selects the integration_tests project and your local profile:

./scripts/dbt-local deps
./scripts/dbt-local debug
./scripts/dbt-local parse --no-partial-parse

dbt deps downloads package code. It does not load Tuva's data assets. dbt debug should report a successful Databricks connection. For connection errors, check the hostname, HTTP path, authentication, compute access, and catalog/schema privileges.

5. Build Core with synthetic data​

./scripts/dbt-local build --full-refresh \
--select package:integration_tests package:the_tuva_project

The first build loads the selected package seed assets, runs unit tests, materializes models, and runs data tests in dependency order. The integration defaults enable claims, clinical data, and provider attribution, with synthetic_data_size: small. The resulting relations use the configured catalog and Tuva's output schemas. The build uses Databricks compute and writes tables, so run it in your development environment.

To build all eight installed marts as well, use the complete project selection:

./scripts/dbt-local build --full-refresh

For a connector project, install Core and only the standalone packages you need as described in Getting Started.

Optional Data Quality​

Data Quality is disabled by default. To include Input Data Quality and the optional logical failure-key relation in the synthetic Core build:

./scripts/dbt-local build --full-refresh \
--select package:integration_tests package:the_tuva_project \
--vars '{data_quality_enabled: true, enable_data_quality_failure_keys: true}'

Use native YAML booleans as shown. dbt test runs unit and data tests; it does not materialize the Data Quality models. Use dbt build and the documented Structural Data Quality and Logical Data Quality selections when working with those relations.