Setting up Tuva on Databricks
This walkthrough runs Tuva Core 1.0 on the small synthetic dataset using a Databricks SQL warehouse. For your own data, follow Getting Started to create a connector project that supplies the Input Layer, then use the Databricks connection and storage setup below.
Prerequisites
- Git, Python 3.12, and a Bash-compatible shell for the commands below.
- A Databricks SQL warehouse you can use, plus its server hostname and HTTP path from Connection details.
- A development catalog and schema where your dbt identity can create and
modify tables and views. With Unity Catalog, it needs
USE CATALOG,USE SCHEMA, and the relevant object privileges; it also needsCREATE SCHEMAif dbt will create Tuva's output schemas. - Databricks authentication configured for your identity. The example uses a personal access token supplied through an environment variable. OAuth and service-principal options are described in the Databricks connection documentation.
- Access from Databricks compute to Tuva's public S3 asset paths under
s3://tuva-public-resources/. Tuva's Databricks loader usesCOPY INTOfrom S3, including when the Databricks workspace is hosted on another cloud.
For Unity Catalog, ask your administrator to configure the applicable external
location and grant READ FILES, along with access to the target catalog and
schemas. Public object availability does not by itself grant access through
your workspace's storage controls. See Databricks COPY INTO permissions.
1. Get the tagged integration project
git clone --branch v1.0.0 --single-branch https://github.com/tuva-health/tuva-core.git
cd tuva-core
The integration_tests project maps the package's synthetic data to the Input
Layer. It installs this local Core checkout and all eight standalone marts at
their v1.0.0 tags. The build selection below runs Core and the synthetic
Input Layer first.
2. Install dbt and the adapter
These versions match the Databricks configuration in Tuva's 1.0 release CI:
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install "dbt-core==1.11.14" "dbt-databricks==1.12.4" "sqlparse==0.6.0"
This is a tested configuration, not the minimum supported version. Core's
declared dbt range is >=1.10.5,<3.0.0. This walkthrough uses dbt Core; see
runtime coverage
for the separate Core 2 and Fusion validation scope. The
Databricks adapter
uses Python APIs and does not require an ODBC driver.
3. Configure a local profile
Create a separate profile directory to preserve any existing dbt profiles:
mkdir -p .dbt-profile
export DBT_PROFILES_DIR="$PWD/.dbt-profile"
Save this as .dbt-profile/profiles.yml:
default:
target: dev
outputs:
dev:
type: databricks
host: "{{ env_var('DATABRICKS_HOST') }}"
http_path: "{{ env_var('DATABRICKS_HTTP_PATH') }}"
catalog: "{{ env_var('DATABRICKS_CATALOG') }}"
schema: "{{ env_var('DATABRICKS_SCHEMA') }}"
token: "{{ env_var('DBT_ENV_SECRET_DATABRICKS_TOKEN') }}"
threads: 4
Set the referenced environment variables in your shell or secrets manager.
The hostname must omit https://; use the HTTP path from the SQL warehouse's
connection details. Choose a development catalog and schema that your identity
can write to. Keep the token out of the project and version control. The
profile name default matches the integration project's dbt_project.yml.
If you use OAuth instead, replace the token configuration with the documented OAuth fields for your adapter. Authentication to Databricks and permission to read the S3 assets are separate prerequisites.
4. Install packages and check the connection
Run these commands from the tuva-core directory. The helper selects the
integration_tests project and your local profile:
./scripts/dbt-local deps
./scripts/dbt-local debug
./scripts/dbt-local parse --no-partial-parse
dbt deps downloads package code. It does not load Tuva's data assets.
dbt debug should report a successful Databricks connection. For connection
errors, check the hostname, HTTP path, authentication, compute access, and
catalog/schema privileges.
5. Build Core with synthetic data
./scripts/dbt-local build --full-refresh \
--select package:integration_tests package:the_tuva_project
The first build loads the selected package seed assets, runs unit tests,
materializes models, and runs data tests in dependency order. The integration
defaults enable claims, clinical data, and provider attribution, with
synthetic_data_size: small. The resulting relations use the configured
catalog and Tuva's output schemas. The build uses Databricks compute and writes
tables, so run it in your development environment.
To build all eight installed marts as well, use the complete project selection:
./scripts/dbt-local build --full-refresh
For a connector project, install Core and only the standalone packages you need as described in Getting Started.
Optional Data Quality
Data Quality is disabled by default. To include Input Data Quality and the optional logical failure-key relation in the synthetic Core build:
./scripts/dbt-local build --full-refresh \
--select package:integration_tests package:the_tuva_project \
--vars '{data_quality_enabled: true, enable_data_quality_failure_keys: true}'
Use native YAML booleans as shown. dbt test runs unit and data tests; it does
not materialize the Data Quality models. Use dbt build and the documented
Structural Data Quality and
Logical Data Quality selections when
working with those relations.