Skip to main content

🏁 Getting Started

There are two main ways to get started with Tuva Core:

  1. Run Tuva with Synthetic Data: run the local integration_tests project against versioned Tuva synthetic data from the Tuva Core package.
  2. Run Tuva with Your Data: build or use a dbt connector that maps your warehouse source data to the Tuva Input Layer and import Tuva Core into that connector.

Run Tuva with Synthetic Data​

The synthetic workflow runs from the integration_tests dbt project in the tuva-core repo.

The integration_tests project does two things:

  • Maps synthetic data that ships with the Tuva Core package into the Tuva Input Layer
  • Installs the local Tuva Core package and all eight standalone data marts at their immutable release tags. The build selection determines which packages run.

Step 1: Setup​

To run Tuva you need to have dbt installed and you need a data warehouse. If you have neither, follow the instructions below to set up dbt and DuckDB locally. Alternatively skip this step if you will configure your own data warehouse.

Clone the repo and move into it:

git clone --branch v1.0.0 https://github.com/tuva-health/tuva-core.git
cd tuva-core

Create a Python environment and install dbt with DuckDB:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install dbt-duckdb

Create a separate local profile directory so an existing profile is preserved. DuckDB should use one thread when loading seeds.

mkdir -p .dbt-profile
export DBT_PROFILES_DIR="$PWD/.dbt-profile"
cat > "$DBT_PROFILES_DIR/profiles.yml" <<'EOF'
default:
target: dev
outputs:
dev:
type: duckdb
path: /tmp/tuva_core.duckdb
threads: 1
EOF

Step 2: Run Tuva​

The helper script runs dbt from integration_tests and uses your local profile.

./scripts/dbt-local deps
./scripts/dbt-local build --full-refresh --select package:integration_tests package:the_tuva_project

This builds the synthetic Input Layer plus Tuva Core. It does not run the optional data mart packages.

The integration project's packages.yml installs local Core plus all eight standalone data marts at v1.0.0. The commands above select only Core and the synthetic Input Layer. To build the complete installed package graph, run ./scripts/dbt-local build --full-refresh. A connector project can install only the optional packages it needs, as shown in Optionally Run Data Marts.

The default synthetic payload is small. To run the larger synthetic dataset, pass synthetic_data_size: large:

./scripts/dbt-local build --full-refresh \
--select package:integration_tests package:the_tuva_project \
--vars '{synthetic_data_size: large}'

The current synthetic workflow enables claims, clinical, and provider attribution by default in integration_tests/dbt_project.yml. Data quality is opt-in:

vars:
claims_enabled: true
clinical_enabled: true
provider_attribution_enabled: true
data_quality_enabled: false
synthetic_data_size: small

integration_tests/dbt_project.yml is the best place to see the full commented list of supported Tuva Core vars.

Run Tuva with Your Data​

Use this path when you want to run Tuva on your own claims, EHR, ADT, HIE, lab, or other healthcare data. Tuva Core does not read raw source systems directly. Your dbt project must first map those sources into the Tuva Input Layer.

Step 1: Create A Connector Project​

A connector is a dbt project that transforms raw source tables into the Tuva Input Layer. Start from the Connector Template when you are building a new connector.

Use the connector project to:

  • define raw source tables in models/_sources.yml;
  • stage and clean source fields where needed;
  • expose final Input Layer models with the exact Tuva table and column names;
  • set the Tuva vars that match the domains you have mapped.

For a detailed walkthrough, see Building a Connector.

Step 2: Import Tuva Core​

The latest published stable release of Tuva is 1.0.0.

For Tuva 1.0, add the canonical repository and immutable release tag to your connector project's packages.yml:

packages:
- git: "https://github.com/tuva-health/tuva-core.git"
revision: "v1.0.0"

This is dbt's supported Git package syntax and does not require dbt Hub indexing. The tag must exist in the repository; check its Releases page if dbt cannot resolve it. Use a published release tag for a production install, rather than the mutable main branch.

The dbt package identity remains the_tuva_project. Use one installation method per package and avoid duplicate Git and dbt Hub entries.

Run dbt deps after changing dependencies and keep package-lock.yml with your connector project.

Step 3: Map The Input Layer​

Create a root dbt model with the exact name and contract for every Input Layer table in each enabled domain. Domain flags enable a group of tables, not individual tables.

  • Claims: eligibility, medical_claim, and pharmacy_claim.
  • Clinical: appointment, condition, encounter, immunization, lab_result, location, medication, observation, patient, practitioner, and procedure.
  • Provider attribution: provider_attribution, when provider_attribution_enabled: true.

If a source has no records for a required model, supply an empty relation with the declared columns and types so dbt can resolve it. An empty model satisfies the shape of the contract; it does not pass Structural Data Quality's population checks. Review those findings before using the outputs for analytics.

Provider attribution can calculate assignments from claims and incorporate payer-supplied assignments. Enabling it requires claims and the provider_attribution Input Layer model. See Provider Attribution for its setup and population requirements.

Step 4: Set Vars​

Set the ref behavior flag and enable only the domains you have mapped. Use native YAML booleans, as shown below; quoted values such as "true" and "false" are rejected.

flags:
require_ref_searches_node_package_before_root: true

vars:
claims_enabled: true
clinical_enabled: false
provider_attribution_enabled: false

# Optional: build queryable data_quality result tables.
data_quality_enabled: true

Use these rules:

  • set claims_enabled: true when payer claims Input Layer tables are mapped;
  • set clinical_enabled: true when provider clinical Input Layer tables are mapped;
  • set provider_attribution_enabled: true only when claims are enabled and provider attribution is mapped;
  • set data_quality_enabled: true when you want Tuva to build Structural and Logical Data Quality result tables.

For the full var reference, see dbt Variables.

Step 5: Run Tuva Core​

After mapping the Input Layer, follow the Data Quality tutorial to load assets, build the connector inputs and Core wrappers, and inspect Structural and Logical findings. An intentionally empty input remains a population readiness finding; revisit source availability, enabled domains, and affected analytics before treating the implementation as ready.

Then build Core and run its native tests:

dbt build --select +package:the_tuva_project

This builds Tuva Core on top of your mapped source data and loads the required Tuva Data Assets.

Step 6: Optionally Run Data Marts​

Data marts are separate dbt packages that run on top of Tuva Core. Add only the packages you need to your connector project's packages.yml.

For example:

packages:
- git: "https://github.com/tuva-health/tuva-core.git"
revision: "v1.0.0"
- git: "https://github.com/tuva-health/cms_chronic_conditions.git"
revision: "v1.0.0"
- git: "https://github.com/tuva-health/cms_hcc.git"
revision: "v1.0.0"
- git: "https://github.com/tuva-health/quality_measures.git"
revision: "v1.0.0"

The root connector owns Core and each optional dependency. The complete 1.0 installation example lists all eight packages, including Semantic Layer's required dependencies. Then run Tuva Core plus the selected packages and their connector ancestors:

dbt deps
dbt build --select +package:the_tuva_project package:cms_chronic_conditions package:cms_hcc package:quality_measures

Explore the Tuva Data Model​

After the build completes, inspect the generated schemas and tables with your preferred SQL client:

select schema_name
from information_schema.schemata
order by 1;

select count(*) from core.patient;
select count(*) from core.medical_claim;
select count(*) from core.condition;

For warehouse-specific setup, see Supported Data Warehouses.