> ## Documentation Index
> Fetch the complete documentation index at: https://docs.variable.global/llms.txt
> Use this file to discover all available pages before exploring further.

# Self-hosted deployment

> How to deploy Variable in a self-hosted environment

<Note>
  **Enterprise feature:** [contact us](mailto:hello@variable.co) if you're interested in self-hosting Variable
</Note>

# Overview

The Variable application is comprised of 4 main components: a UI, a stateless API server,
a Neo4j database, and file storage for uploaded files. The UI and API are each delivered as a Docker container,
while the database can be run as a Docker container or as a managed service. The file storage
can be any object storage service, such as AWS S3, Google Cloud Storage, or a self-hosted solution.

# Components

1. **UI**: A small nginx container that serves the static files for the Variable UI.
2. **Server**: A stateless API server. It can be scaled horizontally by running multiple instances behind a load balancer.
3. **Database**: A Neo4j database that stores all the data for Variable. It can be run as a standalone service or as part of a managed database solution.
4. **File Storage**: For any files uploaded to Variable, such as images or documents. This can be any object storage service, such as AWS S3, Google Cloud Storage, or a self-hosted solution.
5. **Malware Scanning**: An optional component that scans uploaded files for malware before they are stored. This can be integrated with services like [ClamAV](https://www.clamav.net) or other malware scanning solutions.

# Authentication

User authentication is typically done via SSO. Currently supported providers are:

* Microsoft / Azure AD
* Google Workspace

Password authentication can also be enabled via the `ENABLE_PASSWORD_AUTH` environment variable.

# Deployment

The UI, Server and Database can be deployed using any Docker environment, either locally or in the cloud (e.g. Google Compute Engine/Cloud Run).
File storage can be configured to use any object storage service, or a self-hosted solution.

# Configuration

Each component will require configuration to connect to the other components.

## API Configuration

These are the environment variables that need to be set:

* **VARIABLE\_APP\_URL**: The base URL of the Variable application, which is used to generate links and handle redirects.
* **SYSTEM\_ADMIN\_EMAILS**: Semicolon-separated list of system admin email addresses. The first email is used as the initial admin user for database seeding.
* **DB\_NEO4J\_URL**: The URI for the Neo4j database (e.g. bolt://localhost:7687).
* **NEO4J\_AUTH**: The username/password for the Neo4j database.
* **JWT\_SECRET**: A secret used to sign JSON Web Tokens.
* **STORAGE\_PROVIDER**: The storage provider to use (e.g. "gcs" for Google Cloud Storage, "azure" for Azure Blob Storage).
* **VARIABLE\_PUBLIC\_BUCKET\_UPLOAD**: The bucket name for public file uploads.
* **VARIABLE\_PUBLIC\_BUCKET\_READ**: The bucket name for public file reads.
* **VARIABLE\_PRIVATE\_BUCKET\_UPLOAD**: The bucket name for private file uploads.
* **VARIABLE\_PRIVATE\_BUCKET\_READ**: The bucket name for private file reads.
* **LOG\_LEVEL**: Controls server log verbosity. Valid values: `error`, `warn`, `info`, `debug`. Defaults to `info` (or `debug` in dev/local environments unless explicitly set).

## Auth Configuration

Set these environment variables to configure authentication:

* **ENABLE\_PASSWORD\_AUTH**: Set to "false" to disable password authentication.
* **OAUTH\_ONLY**: Set to "true" to only allow OAuth authentication.
* **SESSION\_SECRET**: A secret used to sign the session cookie that carries OAuth sign-in state. Use a long random value, and the same value on every API instance.
* **GOOGLE\_CLIENT\_ID**: The client ID for Google OAuth.
* **GOOGLE\_CLIENT\_SECRET**: The client secret for Google OAuth.
* **MICROSOFT\_CLIENT\_ID**: The client ID for Microsoft/Azure AD OAuth.
* **MICROSOFT\_CLIENT\_SECRET**: The client secret for Microsoft/Azure AD OAuth.
* **MICROSOFT\_TENANT\_ID**: The tenant for Microsoft/Azure AD OAuth. Accepts a tenant GUID, a verified domain (e.g. `contoso.com` or `contoso.onmicrosoft.com`), or one of the multi-tenant aliases `common`, `organizations`, or `consumers`.
* **MICROSOFT\_CALLBACK\_URL**: The callback URL for Microsoft OAuth (e.g. `https://your-domain.com/api/auth/microsoft/callback`).

OAuth sign-in keeps its state in that session cookie between the redirect to the provider and the callback, so the cookie has to make the round trip:

* **Browsers must reach Variable over HTTPS.** The cookie is `Secure`, so a browser on plain HTTP doesn't keep it.
* **Sign-in must start on the callback's host.** Callbacks go to the host of `VARIABLE_APP_URL`, or of `GOOGLE_CALLBACK_URL` / `MICROSOFT_CALLBACK_URL` when set. A cookie set on any other hostname isn't sent with the callback.
* **Every API instance needs the same `SESSION_SECRET`.** A cookie signed by one instance fails the signature check on another.

If the cookie doesn't make the round trip, sign-in returns to the sign-in page with an error, and the API logs `Unable to verify authorization request state.` with a `hint` naming the likely cause.

If TLS ends at a load balancer or reverse proxy, have the proxy that forwards `/api` to the server send `X-Forwarded-Proto: https`. Sign-in works without it, but the API otherwise treats every request as plain HTTP.

## Email Configuration

Variable sends transactional email for things like user invitations, password resets, and data requests. Self-hosted deployments deliver email via SMTP. If SMTP is not configured, email sending is disabled and email-dependent flows will not work: password reset, and setting a password after an invitation, which needs the link in the invitation email. Invited people can still sign in with Google, Microsoft or SSO where those are set up.

Configure these to route email through your own SMTP server (e.g. Office 365, Google Workspace SMTP relay, Amazon SES SMTP, Postfix):

* **EMAIL\_ENABLED**: Set to `true` to enable outbound email. Defaults to `false`. When `false`, the app runs normally but no email is sent.
* **EMAIL\_DEFAULT\_SENDER**: The `From` address used on outbound mail (e.g. `hello@your-domain.com`). Required for email to be considered configured.
* **EMAIL\_ALLOW\_DOMAINS** (optional): Semicolon-separated allowlist of recipient domains. When set to anything other than `*`, mail to recipients outside the listed domains is silently dropped - useful for staging environments or for locking outbound mail to internal domains during a rollout. Set to `*` (or leave unset) to allow all domains.
* **SMTP\_HOST**: SMTP server hostname (e.g. `smtp.office365.com`, `smtp-relay.gmail.com`).
* **SMTP\_PORT**: SMTP server port. Defaults to `587`.
* **SMTP\_SECURE**: Set to `true` to use implicit TLS (typically port `465`). Defaults to `false`, which uses STARTTLS on port `587`.
* **SMTP\_USER**: SMTP username.
* **SMTP\_PASSWORD**: SMTP password or app-specific token.

SMTP is considered configured when `SMTP_HOST`, `SMTP_USER`, and `SMTP_PASSWORD` are all set.

```bash .env SMTP example theme={"system"}
EMAIL_ENABLED=true
EMAIL_DEFAULT_SENDER="hello@your-domain.com"
SMTP_HOST=smtp.your-provider.com
SMTP_PORT=587
SMTP_SECURE=false
SMTP_USER=your-smtp-user
SMTP_PASSWORD=your-smtp-password
```

## Database Configuration

Variable uses **Neo4j 5.26** (LTS) with [APOC Core](https://neo4j.com/docs/apoc/current/). APOC Core ships inside the official Neo4j 5.x Docker image and is activated via the `NEO4J_PLUGINS` environment variable - no separate download is required.

### Docker volumes

If hosting the database as a Docker container, you will need to have the following volumes mounted:

* **/data**: This is where the Neo4j database files will be stored.
* **/logs**: This is where the Neo4j logs will be stored.
* **/import**: This is where you can place any initial data files to be imported into the database.

### Required Neo4j settings

These environment variables must be set on the Neo4j container for Variable to function correctly:

```yaml theme={"system"}
environment:
  # APOC Core activation (required - ships inside the Neo4j 5.x image)
  - NEO4J_PLUGINS=["apoc"]
  - NEO4J_dbms_security_procedures_allowlist=apoc.*
  - NEO4J_dbms_security_procedures_unrestricted=apoc.*

  # APOC features used by Variable
  - NEO4J_apoc_import_file_enabled=true
  - NEO4J_apoc_trigger_enabled=true
  - NEO4J_apoc_import_file_use__neo4j__config=true

  # Transaction timeout (default 120s recommended)
  - NEO4J_db_transaction_timeout=120s

  # Lenient relationship creation (required)
  - NEO4J_dbms_cypher_lenient__create__relationship=true
```

### Recommended Neo4j settings

These settings are not strictly required but are recommended for production deployments:

```yaml theme={"system"}
environment:
  # Memory - Neo4j 5 enforces transaction memory limits by default; disable for now
  - NEO4J_db_memory_transaction_total_max=0
  - NEO4J_dbms_memory_transaction_total_max=0

  # Query logging - helps diagnose slow queries
  - NEO4J_db_logs_query_threshold=100ms
```

<Warning>
  When disabling transaction memory limits (`=0`), ensure your Neo4j container has an explicit memory constraint (Docker `mem_limit`, Kubernetes resource limits, or VM-level cap) so that a single expensive query cannot consume all host memory and trigger an OOM kill.
</Warning>

### Example Docker Compose service

```yaml docker-compose.yaml theme={"system"}
services:
  db:
    image: neo4j:5.26.31
    cpu_count: 4
    ports:
      - "7474:7474"
      - "7687:7687"
    volumes:
      - ./neo4j/import:/import
      - ./neo4j/data:/data
      - ./neo4j/logs:/logs
    environment:
      - NEO4J_AUTH=neo4j/<your-password>
      - NEO4J_PLUGINS=["apoc"]
      - NEO4J_db_transaction_timeout=120s
      - NEO4J_dbms_security_procedures_allowlist=apoc.*
      - NEO4J_dbms_security_procedures_unrestricted=apoc.*
      - NEO4J_db_logs_query_threshold=100ms
      - NEO4J_apoc_import_file_enabled=true
      - NEO4J_apoc_export_file_enabled=true
      - NEO4J_apoc_trigger_enabled=true
      - NEO4J_apoc_import_file_use__neo4j__config=true
      - NEO4J_dbms_cypher_lenient__create__relationship=true
      - NEO4J_db_memory_transaction_total_max=0
      - NEO4J_dbms_memory_transaction_total_max=0
```

<Note>
  The `NEO4J_AUTH` value must match what you configure in the API server's `NEO4J_AUTH` environment variable. The format is `username/password`.
</Note>

## URL Mapping

The load balancer or reverse proxy should be configured to map the following URLs:

The UI should be served at `<APP_URL>` and should point to the Variable UI container.

The API should be served at `<APP_URL>/api`

Public files should be served at `<APP_URL>/uploads`

## File Storage Configuration

The application is set up to upload files to the an "upload" bucket where the files are scanned for malware
before being moved to a "read" bucket.

## Architecture Diagram

```mermaid theme={"system"}
flowchart TD
    LB(Load Balancer)

    subgraph Serverless
    UI(Variable UI)
    API(API Server)
    MSS(Malware Scanner Optional)
    end

    subgraph Compute
    DB[(Neo4j Server)]
    DBL[Neo4j Logs]
    DBD[Neo4j Data]
    DBI[Neo4j Import]
    end

    subgraph File Storage
    PBU@{ shape: docs, label: "PUBLIC_BUCKET_UPLOAD" }
    PRU@{ shape: docs, label: "PRIVATE_BUCKET_UPLOAD" }
    PUR@{ shape: docs, label: "PUBLIC_BUCKET_READ" }
    PRR@{ shape: docs, label: "PRIVATE_BUCKET_READ" }
    end

    LB -->|URL Mapping:/*| UI
    LB -->|URL Mapping:/api| API
    LB -->|URL Mapping:/upload| PUR
    API <-->|Database Queries| DB
    DB <--> DBD
    DB --> DBL
    DBI --> DB
    API -->|Upload| PBU
    API -->|Upload| PRU
    PRR -->|Read| API
    PBU -->|Malware Scan| MSS
    PRU -->|Malware Scan| MSS
    MSS -->|Scan Result OK| PUR
    MSS -->|Scan Result OK| PRR
```

## Terraform examples

This is not a complete Terraform configuration, but it provides a starting point to see how the architecture pieces above can be set up in Terraform.

### Security Headers (locals.tf)

Define your security headers in a `locals.tf` file so they can be reused across multiple backend services and buckets without duplication.

```terraform locals.tf lines expandable theme={"system"}
locals {
  csp_directives = join("", [
    "default-src 'self';",
    "script-src 'self';",
    "style-src 'self' 'unsafe-inline';",
    "img-src 'self' data: https://storage.googleapis.com;",
    "font-src 'self';",
    "connect-src 'self';",
    "frame-src 'self' blob: https://storage.googleapis.com;",
    "frame-ancestors 'self';",
    "object-src 'none';",
    "base-uri 'self';",
    "form-action 'self' https://accounts.google.com https://login.microsoftonline.com;",
    "upgrade-insecure-requests",
  ])

  security_response_headers = [
    "Strict-Transport-Security: max-age=31536000; includeSubDomains",
    "Content-Security-Policy: ${local.csp_directives}",
    "X-Content-Type-Options: nosniff",
    "Referrer-Policy: strict-origin-when-cross-origin",
    "Permissions-Policy: geolocation=(self), microphone=(), camera=()",
  ]
}
```

Replace `https://storage.googleapis.com` with your storage domain. Add additional origins as needed (see [CSP section](#content-security-policy-csp)).

### Load Balancer and URL Mapping

```terraform load-balancer.tf lines expandable theme={"system"}
# Backend Service for the API
resource "google_compute_backend_service" "variable_api" {
  name = "backend-service-variable-api"

  log_config {
    enable = true
  }

  backend {
    group = google_compute_region_network_endpoint_group.variable_api.id
  }
}

# Backend Service for the UI (nginx container, e.g. on Cloud Run)
resource "google_compute_backend_service" "variable_ui" {
  name = "backend-service-variable-ui"

  custom_response_headers = local.security_response_headers

  log_config {
    enable = true
  }

  enable_cdn = true
  cdn_policy {
    # USE_ORIGIN_HEADERS only caches responses the origin explicitly marks
    # cacheable via Cache-Control, so dynamic HTML is not cached while
    # the content-hashed static assets served by nginx are.
    cache_mode = "USE_ORIGIN_HEADERS"

    # Negative caching off: the UI uses try_files $uri /index.html so 404
    # is effectively impossible, and we do not want transient 5xx from
    # cold starts or deploys cached at the edge.
    negative_caching = false

    # Content-hashed asset filenames already make the cache key unique
    # per asset version, so query strings would only fragment the cache.
    cache_key_policy {
      include_host         = true
      include_protocol     = true
      include_query_string = false
    }
  }

  backend {
    group = google_compute_region_network_endpoint_group.variable_ui.id
  }
}

# Backend Bucket for uploads
resource "google_compute_backend_bucket" "variable_uploads" {
  name        = "backend-service-variable-uploads"
  bucket_name = google_storage_bucket.variable_uploads.name
  enable_cdn  = true

  cdn_policy {
    cache_mode = "CACHE_ALL_STATIC"
  }
}

# URL Map for the Variable application
resource "google_compute_url_map" "variable" {
  name        = "url-map-variable"
  description = "Maps the incoming request to either the UI container or the API on /api routes."

  default_service = google_compute_backend_service.variable_ui.id

  host_rule {
    hosts        = ["*"]
    path_matcher = "all"
  }

  path_matcher {
    name            = "all"
    default_service = google_compute_backend_service.variable_ui.id

    path_rule {
      paths   = ["/api", "/api/*"]
      service = google_compute_backend_service.variable_api.id
    }

    path_rule {
      paths   = ["/uploads", "/uploads/*"]
      service = google_compute_backend_bucket.variable_uploads.id
    }
  }
}

resource "google_compute_target_https_proxy" "variable" {
  name    = "ssl-lb-variable"
  url_map = google_compute_url_map.variable.id
  certificate_map = "..."
}
```

### UI Server

The UI is served by a small nginx container (published as `variable-ui`) that hosts the built static assets. Deploy it the same way as the API - any Docker environment works (Cloud Run, GKE, GCE, etc.). The example below uses Cloud Run.

```terraform ui.tf lines expandable theme={"system"}
resource "google_cloud_run_service" "variable_ui" {
  name     = "variable-ui"
  location = var.region

  autogenerate_revision_name = true

  template {
    spec {
      containers {
        resources {
          limits = {
            cpu    = "1000m"
            memory = "512Mi"
          }
        }
        image = "gcr.io/.../variable-ui:latest" # TODO: Update with the actual image URL
      }
    }
  }
}

resource "google_cloud_run_service_iam_member" "variable_ui_public" {
  location = google_cloud_run_service.variable_ui.location
  service  = google_cloud_run_service.variable_ui.name
  role     = "roles/run.invoker"
  member   = "allUsers"
}

resource "google_compute_region_network_endpoint_group" "variable_ui" {
  name                  = "neg-variable-ui"
  region                = var.region
  network_endpoint_type = "SERVERLESS"

  cloud_run {
    service = google_cloud_run_service.variable_ui.name
  }
}
```

### Storage

#### Uploads

```terraform uploads.tf lines expandable theme={"system"}
resource "google_storage_bucket" "public_uploads_read" {
  name     = "public-uploads-read"
  location = var.europe

  versioning { enabled = true }

  website {
    main_page_suffix = "index.html"
    not_found_page   = "index.html"
  }
}

resource "google_storage_bucket" "public_uploads_write" {
  name     = "public-uploads-write"
  location = var.region

  public_access_prevention = "enforced"

  versioning { enabled = true }
}

resource "google_storage_bucket" "private_uploads_read" {
  name     = "private-uploads-read"
  location = var.europe
  project  = var.project_id

  public_access_prevention = "enforced"

  versioning { enabled = true }
}

resource "google_storage_bucket" "private_uploads_write" {
  name     = "private-uploads-write"
  location = var.region

  public_access_prevention = "enforced"

  versioning { enabled = true }
}
```

### API Server

```terraform api.tf lines expandable theme={"system"}
resource "google_cloud_run_service" "variable_api" {
  name     = "variable-api"
  location = var.region

  autogenerate_revision_name = true

  template {
    spec {
      containers {
        resources {
          limits = {
            cpu    = "4000m"
            memory = "4Gi"
          }
        }
        image = "gcr.io/.../variable-api:latest" # TODO: Update with the actual image URL
      }
    }
  }
}
```

### Google Compute for Neo4j

```terraform database.tf lines expandable theme={"system"}
module "gce-container" {
  source  = "terraform-google-modules/container-vm/google"
  version = "~> 3.0"

  container = {
    image = "docker.io/neo4j:5.26.31"

    env = [
      {
        name  = "NEO4J_db_transaction_timeout"
        value = "120s"
      },
      {
        name  = "NEO4J_dbms_security_procedures_allowlist"
        value = "apoc.*"
      },
      {
        name  = "NEO4J_dbms_security_procedures_unrestricted"
        value = "apoc.*"
      },
      {
        name  = "NEO4J_server_default__listen__address"
        value = "0.0.0.0"
      },
      {
        name  = "NEO4J_server_bolt_listen__address"
        value = "0.0.0.0:7687"
      },
      {
        name  = "NEO4J_server_http_listen__address"
        value = "0.0.0.0:7474"
      },
      {
        name  = "NEO4J_apoc_export_file_enabled"
        value = "true"
      },
      {
        name  = "NEO4J_apoc_import_file_enabled"
        value = "true"
      },
      {
        name  = "NEO4J_apoc_import_file_use__neo4j__config"
        value = "true"
      },
      {
        name  = "NEO4J_apoc_trigger_enabled"
        value = "true"
      },
      {
        name  = "NEO4J_dbms_cypher_lenient__create__relationship"
        value = "true"
      },
      {
        name  = "NEO4J_PLUGINS"
        value = "[\"apoc\"]"
      },
      {
        name  = "NEO4J_AUTH"
        value = "neo4j/<your-password>"
      },
    ]

    # APOC Core ships inside the Neo4j 5.x image - no separate plugin disk needed.
    volumeMounts = [
      {
        mountPath = "/cache"
        name      = "tempfs-0"
        readOnly  = false
      },
      {
        mountPath = "/data"
        name      = "data-disk-0"
        readOnly  = false
      },
    ]
  }

  volumes = [
    {
      name = "tempfs-0"

      emptyDir = {
        medium = "Memory"
      }
    },
    {
      name = "data-disk-0"

      gcePersistentDisk = {
        pdName = "data-disk-0"
        fsType = "ext4"
      }
    },
  ]

  restart_policy = "Always"
}

resource "google_compute_disk" "pd_neo4j" {
  name     = "neo4j-data-disk"
  zone     = var.zone
  type     = "pd-ssd" # SSD persistent disk
  size     = 25
}

resource "google_compute_instance" "neo4j_vm" {
  name = "neo4j"
  machine_type = "n2-standard-8"
  zone         = var.zone

  shielded_instance_config {
    enable_secure_boot = true
  }
  allow_stopping_for_update = false

  scheduling {
    provisioning_model = "STANDARD"
    preemptible        = false
    automatic_restart  = true

    instance_termination_action = null
  }

  boot_disk {
    initialize_params {
      size  = 10
      image = module.gce-container.source_image
    }
  }

  attached_disk {
    source      = google_compute_disk.pd_neo4j.self_link
    device_name = "data-disk-0"
    mode        = "READ_WRITE"
  }

  network_interface {
    subnetwork = module.vpc-module.subnets["${var.region}/serverless-subnet"].name
    network_ip = google_compute_address.neo4j.address
  }

  metadata = {
    gce-container-declaration = module.gce-container.metadata_value
    google-logging-enabled    = true
    block-project-ssh-keys    = true
  }

  labels = {
    container-vm = module.gce-container.vm_container_label
  }

  tags = ["neo4j-instance"]

  service_account {
    email  = google_service_account.neo4j.email
    scopes = ["cloud-platform"]
  }
}

resource "google_compute_address" "neo4j" {
  name         = "int-ip-neo4j"
  description  = "The IP address used for the Neo4j VM in the network."
  address_type = "INTERNAL"
  address      = "..."
}
```

## Security Headers

Your load balancer, reverse proxy, or the UI's nginx container should set the following security headers on responses served to the browser. These are typically configured on the outermost layer that serves the UI - either the cloud load balancer / CDN, or the nginx container itself if it is the outermost hop.

### Content Security Policy (CSP)

The CSP header restricts which origins the browser is allowed to load resources from. A recommended baseline for Variable:

```
Content-Security-Policy: default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline'; img-src 'self' data: https://<your-storage-domain>; font-src 'self'; connect-src 'self'; frame-src 'self' blob: https://<your-storage-domain>; frame-ancestors 'self'; object-src 'none'; base-uri 'self'; form-action 'self' https://accounts.google.com https://login.microsoftonline.com; upgrade-insecure-requests
```

Replace `<your-storage-domain>` with the domain of your file storage service (e.g. `https://storage.googleapis.com` for GCS, or `https://<account>.blob.core.windows.net` for Azure Blob Storage).

Depending on which third-party integrations you enable, you may need to add additional origins:

| Integration | Directives to update |
| - | - |
| **SSO provider** (e.g. Auth0) | `frame-src`, `connect-src`, `form-action` (add your OAuth provider's authorize URL origin) |
| **Analytics** (e.g. Segment) | `script-src` (`cdn.segment.com`), `connect-src` (`api.segment.io`, `cdn.segment.com`) |
| **Payments** (e.g. Stripe) | `script-src` (`js.stripe.com`), `connect-src` (`api.stripe.com`, `m.stripe.com`), `frame-src` (`js.stripe.com`) |
| **Google Maps** | `script-src`, `connect-src`, `img-src` (`maps.googleapis.com`) |
| **Status page widget** | `script-src`, `connect-src`, `frame-src` (your statuspage domain) |
| **Google Fonts** | `style-src` (`fonts.googleapis.com`), `font-src` (`fonts.gstatic.com`) |

### Other Recommended Headers

| Header | Recommended Value | Purpose |
| - | - | - |
| `Strict-Transport-Security` | `max-age=31536000; includeSubDomains` | Enforce HTTPS for one year |
| `X-Content-Type-Options` | `nosniff` | Prevent MIME-type sniffing |
| `Referrer-Policy` | `strict-origin-when-cross-origin` | Limit referrer information |
| `Permissions-Policy` | `geolocation=(self), microphone=(), camera=()` | Restrict browser feature access |

### Example: nginx

<Warning>
  If you are deploying behind a cloud load balancer (e.g. GCP Cloud Load Balancing, Azure Application Gateway, AWS ALB/CloudFront), security headers are typically set at the load balancer or CDN level - **not** in nginx. Setting them in both places can cause duplicate headers. Only use the nginx configuration below if nginx is your outermost reverse proxy.
</Warning>

<Note>
  Nginx's `add_header` directive is not inherited into `location` blocks that define their own `add_header`. Define your security headers in a shared snippet file and `include` it in each location block.
</Note>

```nginx /etc/nginx/snippets/security-headers.conf theme={"system"}
add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;
add_header Content-Security-Policy "default-src 'self';script-src 'self';style-src 'self' 'unsafe-inline';img-src 'self' data:;font-src 'self';connect-src 'self';frame-src 'self' blob:;frame-ancestors 'self';object-src 'none';base-uri 'self';form-action 'self' https://accounts.google.com https://login.microsoftonline.com;upgrade-insecure-requests" always;
add_header X-Content-Type-Options "nosniff" always;
add_header Referrer-Policy "strict-origin-when-cross-origin" always;
add_header Permissions-Policy "geolocation=(self), microphone=(), camera=()" always;
```

```nginx theme={"system"}
location / {
    try_files $uri /index.html;
    add_header Cache-Control "no-cache, no-store, must-revalidate";
    include snippets/security-headers.conf;
}

location = /env-config.js {
    add_header Cache-Control "no-cache, no-store, must-revalidate";
    include snippets/security-headers.conf;
}

location ~* \.(?:css|js)$ {
    expires 1y;
    add_header Cache-Control "public, immutable";
    include snippets/security-headers.conf;
    access_log off;
}
```

If your storage domain is separate from your app domain, add it to `img-src` and `frame-src` (e.g. `img-src 'self' data: https://<your-storage-domain>`).

# Onboarding

After deploying the containers and configuring environment variables, follow these steps to initialize your Variable instance.

## Step 1: Start the Containers

Start all services using Docker Compose (or your container orchestration tool). The API server will automatically run any pending database migrations on startup.

```bash theme={"system"}
docker compose up -d
```

## Step 2: Populate Base Data

Navigate to your Variable instance URL in a browser. When the UI detects an empty database, it will guide you through the setup process automatically. This creates:

* Database constraints and indexes
* GHG Protocol scopes (Scope 1, 2, 3)
* Emission factor data sources (DEFRA, Ember, ecoinvent, etc.)
* LCA stage definitions (A1-A3, B1-B7, C1-C4, D)
* Product taxonomy hierarchy
* Geographic locations and electricity grid data
* Units of measurement
* An admin user (using the first email from `SYSTEM_ADMIN_EMAILS`) with full access
* Your company with an Enterprise subscription

## Step 3: Load Emission Factors

Load your emission factor data using either a [data package](#data-packages-recommended) (recommended) or [URL-based import](#url-based-import). See the [Loading Emission Factors](#loading-emission-factors) section below for details.

## Step 4: Sign In

Navigate to your Variable instance URL and sign in. If you configured OAuth (Google or Azure AD), the SSO flow will be used. If you enabled password authentication, use the admin email to create your account.

The admin user (the first email in `SYSTEM_ADMIN_EMAILS`) is automatically granted global admin privileges and assigned as the owner of the company created in Step 2.

# Loading Emission Factors

There are two ways to load emission factors into your Variable instance: **data packages** (recommended) and **URL-based import**.

## Data packages (recommended)

A data package is a directory of CSV files that is mounted directly into the API server container. This is the simplest approach for self-hosted deployments because the data is loaded from the local filesystem with no external network access required.

### 1. Pull your data package

Variable delivers data packages via [Google Artifact Registry](https://cloud.google.com/artifact-registry). You will receive a service account key file (`key.json`) and your registry URL from Variable.

**Install ORAS**

[ORAS](https://oras.land) is the tool used to pull data packages from the registry.

```bash theme={"system"}
# macOS
brew install oras

# Linux
ORAS_VERSION="1.2.0"
curl -LO "https://github.com/oras-project/oras/releases/download/v${ORAS_VERSION}/oras_${ORAS_VERSION}_linux_amd64.tar.gz"
tar -xzf "oras_${ORAS_VERSION}_linux_amd64.tar.gz" -C /usr/local/bin/ oras
```

**Authenticate**

```bash theme={"system"}
cat key.json | oras login <region>-docker.pkg.dev -u _json_key --password-stdin
```

**Pull and extract**

```bash theme={"system"}
# Create the data package directory
mkdir -p data-packages

# Pull the latest package
oras pull <region>-docker.pkg.dev/variable-global/<your-company>/variable-data:latest \
  --output data-packages/

# Extract
tar -xzf data-packages/variable-data-<your-company>-<version>.tar.gz -C data-packages/
```

To pin to a specific version, replace `latest` with the version tag (e.g. `2026.1`).

After extraction, your `data-packages/` directory will contain the `emissionFactors/` and `impacts/` subdirectories ready to be mounted.

### 3. Mount the volume

In your `docker-compose.yaml`, the data package directory is mounted into the API server container and configured via the `DATA_PACKAGE_DIR` environment variable:

```yaml theme={"system"}
server:
  volumes:
    - ./data-packages:/data-packages
  environment:
    - DATA_PACKAGE_DIR=/data-packages
```

### 4. Install the data package

1. Start your containers and sign in to Variable.
2. Navigate to **Admin > Data sources**.
3. Click the **Import** dropdown and select **Install emission factors**.
4. The installation will begin. The Data Sources table refreshes automatically during installation, so you can watch the emission factor counts increase as data is loaded.

Depending on the size of the dataset, installation may take several minutes. Both emission factors and impact data (if present) are installed in a single operation.

<Note>If the **Install emission factors** option does not appear, verify that CSV files are present in the `emissionFactors/` directory inside your data package volume and that `DATA_PACKAGE_DIR` is set correctly. Check the server logs for the `Data package dir:` message on startup.</Note>

## URL-based import

For cases where you prefer to host emission factor files on an external server or cloud storage bucket, you can import them by providing a URL.

1. Sign in to your Variable instance and navigate to **Admin > Data sources**.
2. Click the **Import** dropdown and select **Import emission factors**.
3. Provide an **authenticated URL** to the CSV file. This URL must be accessible from the server (e.g., a pre-signed S3 URL, a GCS signed URL, or a URL with an access token).
4. Preview the data and click **Import** to load the emission factors into the database.

To import impact data separately, use **Import impacts** from the same dropdown.

## Supported data sources

The following data sources are supported:

* DEFRA 2022
* Ember 2022, 2024, 2025
* ecoinvent 3.8, 3.11
* Idemat 2023
* IPCC AR6
* AIB Residual Grid Mix
* US EPA eGRID 2022
* WIOD

<Note>Variable will provide the emission factor CSV files as part of your license.</Note>

## Recommended Infrastructure

For a self-hosted deployment, we recommend the following infrastructure:

* **API Server**: 2 vCPUs, 4 GB RAM
* **Database**: 4 vCPUs, 16 GB RAM (or use a managed database service)
* **File Storage**: Use a scalable object storage service like AWS S3, Google Cloud, etc.
* **Malware Scanning**: Use a service like ClamAV or integrate with a third-party malware scanning service.

# Health Checks

The API server exposes an unauthenticated health check endpoint that load balancers, container orchestrators, and uptime monitors can use to verify the process is running.

## Endpoint

```http theme={"system"}
GET /api/health
```

No authentication is required. A healthy response is a `200 OK` with a JSON body:

```json theme={"system"}
{
  "message": "OK",
  "apiVersion": "1.2.3",
  "uiVersion": "1.2.3",
  "uptime": 1234.56,
  "supportEnabled": true,
  "timestamp": "2026-04-22T12:00:00.000Z"
}
```

| Field | Description |
| - | - |
| `message` | Always `"OK"` when the API is reachable. |
| `apiVersion` | Version of the API server package. |
| `uiVersion` | Version of the UI, sourced from `VARIABLE_APP_VERSION` (falls back to the root package version). |
| `uptime` | Seconds since the API process started. |
| `supportEnabled` | `true` when the support integration is configured. |
| `timestamp` | Current server time in ISO 8601. |

<Note>
  This is a basic API reachability check: it verifies the API process can respond to HTTP requests, but it does not query the Neo4j database or file storage. `/api/health` is exempt from the [maintenance mode](#maintenance-mode) gate, so it keeps returning `200 OK` even while maintenance is active - safe to use as a Kubernetes `livenessProbe` or load-balancer healthcheck without special handling. If you need readiness validation for downstream dependencies, add a separate probe - for example, a script that calls `/api/health` and then runs a trivial Cypher query against Neo4j.
</Note>

## Response codes

| Status | Meaning |
| - | - |
| `200 OK` | API is reachable. |
| Connection refused / timeout | API process is down or unreachable. |

## Example configurations

<Note>
  The examples below use port `8181` (the default). Replace it with whatever value you configured for `API_PORT`.
</Note>

**curl**

```bash theme={"system"}
curl -fsS https://<your-domain>/api/health
```

**Docker Compose**

The official Variable API image is based on `node:slim` and does not ship with `curl` or `wget`, so the example below uses Node to make the HTTP request. If you build your own image, use whichever HTTP client is available in it.

```yaml theme={"system"}
services:
  server:
    healthcheck:
      test:
        - "CMD"
        - "node"
        - "-e"
        - "fetch('http://localhost:8181/api/health').then(r => process.exit(r.ok ? 0 : 1)).catch(() => process.exit(1))"
      interval: 30s
      timeout: 5s
      retries: 3
      start_period: 30s
```

**Kubernetes**

```yaml theme={"system"}
readinessProbe:
  httpGet:
    path: /api/health
    port: 8181
  initialDelaySeconds: 30
  periodSeconds: 30
  timeoutSeconds: 5
  failureThreshold: 3
```

**GCP Load Balancer (Terraform)**

```terraform theme={"system"}
resource "google_compute_health_check" "variable_api" {
  name = "variable-api-health"

  http_health_check {
    request_path = "/api/health"
    port         = 8181
  }

  check_interval_sec  = 30
  timeout_sec         = 5
  healthy_threshold   = 1
  unhealthy_threshold = 3
}
```

# Maintenance Mode

When performing upgrades, database migrations, or other operations that require the application to be temporarily unavailable, you can put Variable into **maintenance mode**. While maintenance mode is active, the API returns `503 Service Unavailable` with a user-friendly message for every non-exempt request, and the UI renders a dedicated maintenance screen instead of the application.

Routine windows are configured from the admin UI and stored in the database, so toggling maintenance on or off no longer requires a redeploy. A single environment variable remains for genuine emergencies.

## Activation paths

Three independent paths activate the gate; if any is active, the API blocks. Subsystem detail lives in [`packages/api/src/maintenance/README.md`](https://github.com/variable-co/cascade/blob/main/packages/api/src/maintenance/README.md).

| Path | Source | When to use |
| - | - | - |
| **Scheduled window** | Admin → Maintenance windows | Recurring planned downtime (daily, weekly, or monthly window in an IANA timezone). |
| **Manual mode** | Admin → Maintenance windows ("Start now") | One-off "start now / end now" maintenance without scheduling. |
| **Panic switch** | `VARIABLE_MAINTENANCE_MODE=true` env var | The database is down or unreachable. Short-circuits Neo4j entirely, so the gate keeps responding even with no working backing store. Requires a redeploy or restart to clear. |

Scheduled and manual mode propagate across API replicas within \~30s of a save (cache poll interval). The panic switch is per-replica and propagates only when each replica restarts with the new env var.

## Configuring scheduled windows

Sign in as a system admin (an email listed in `SYSTEM_ADMIN_EMAILS`) and open **Admin → Maintenance windows**.

* **Recurrence**: `Daily`, `Weekly`, or `Monthly`. Selecting `No schedule` clears the window.
* **Weekday** (weekly only) or **Day of month** (monthly only, 1–28 so every month fires).
* **Window start**: `HH:mm` 24-hour wall-clock time in the chosen timezone.
* **Duration**: minutes the window stays active (1 minute to 7 days).
* **Timezone**: any IANA zone (e.g. `Europe/London`, `America/New_York`). The host's clock has no effect on the result.

The window runs as `[start, start + duration)` — a request exactly at the end-time is **not** blocked. On spring-forward DST transitions, if the start lands in the skipped hour the day's window is dropped; on fall-back transitions, a window that crosses the duplicated hour can run \~1h longer than its configured duration.

The schedule fields are validated server-side on save and surfaced as inline errors in the form.

## Starting maintenance immediately

Use **Start now** on the same page to enable manual mode without scheduling. The button requires typing `START` to confirm — there is no undo, and active sessions begin seeing 503s within \~30s. Manual mode has no server-side auto-expiry; it runs until a system admin ends it from the maintenance screen.

## Ending a window early

Once maintenance is active, the admin form is no longer reachable — every non-auth route, including `/admin/maintenance`, is replaced by the maintenance screen. System admins see an **End maintenance** button on that screen; clicking it ends the current window immediately (whether it was started manually or by the schedule) and leaves future scheduled occurrences untouched.

Active browser sessions that loaded during maintenance reload automatically once the next request succeeds, so users recover without action.

## Panic switch (`VARIABLE_MAINTENANCE_MODE`)

The env var is a fallback for the case where the admin UI itself is unreachable — most commonly because Neo4j is down. Setting `VARIABLE_MAINTENANCE_MODE=true` blocks every non-exempt request without touching Neo4j and without consulting the scheduled config.

```bash .env theme={"system"}
VARIABLE_MAINTENANCE_MODE=true
```

Apply it the same way as any other env-var change:

```bash theme={"system"}
# After updating VARIABLE_MAINTENANCE_MODE in .env:
docker compose up -d --no-deps server
```

On platforms like Cloud Run, update the service's environment variables and deploy a new revision. The previous `VARIABLE_MAINTENANCE_END_DATE` env var has been removed — scheduled windows now carry their own end time.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.