Skip to main content
Enterprise feature: contact us if you’re interested in self-hosting Variable

Overview

The Variable application is comprised of 4 main components: a UI, a stateless API server, a Neo4j database, and file storage for uploaded files. The UI and API are each delivered as a Docker container, while the database can be run as a Docker container or as a managed service. The file storage can be any object storage service, such as AWS S3, Google Cloud Storage, or a self-hosted solution.

Components

  1. UI: A small nginx container that serves the static files for the Variable UI.
  2. Server: A stateless API server. It can be scaled horizontally by running multiple instances behind a load balancer.
  3. Database: A Neo4j database that stores all the data for Variable. It can be run as a standalone service or as part of a managed database solution.
  4. File Storage: For any files uploaded to Variable, such as images or documents. This can be any object storage service, such as AWS S3, Google Cloud Storage, or a self-hosted solution.
  5. Malware Scanning: An optional component that scans uploaded files for malware before they are stored. This can be integrated with services like ClamAV or other malware scanning solutions.

Authentication

User authentication is typically done via SSO. Currently supported providers are:
  • Microsoft / Azure AD
  • Google Workspace
Password authentication can also be enabled via the ENABLE_PASSWORD_AUTH environment variable.

Deployment

The UI, Server and Database can be deployed using any Docker environment, either locally or in the cloud (e.g. Google Compute Engine/Cloud Run). File storage can be configured to use any object storage service, or a self-hosted solution.

Configuration

Each component will require configuration to connect to the other components.

API Configuration

These are the environment variables that need to be set:
  • VARIABLE_APP_URL: The base URL of the Variable application, which is used to generate links and handle redirects.
  • SYSTEM_ADMIN_EMAILS: Semicolon-separated list of system admin email addresses. The first email is used as the initial admin user for database seeding.
  • DB_NEO4J_URL: The URI for the Neo4j database (e.g. bolt://localhost:7687).
  • NEO4J_AUTH: The username/password for the Neo4j database.
  • JWT_SECRET: A secret used to sign JSON Web Tokens.
  • STORAGE_PROVIDER: The storage provider to use (e.g. “gcs” for Google Cloud Storage, “azure” for Azure Blob Storage).
  • VARIABLE_PUBLIC_BUCKET_UPLOAD: The bucket name for public file uploads.
  • VARIABLE_PUBLIC_BUCKET_READ: The bucket name for public file reads.
  • VARIABLE_PRIVATE_BUCKET_UPLOAD: The bucket name for private file uploads.
  • VARIABLE_PRIVATE_BUCKET_READ: The bucket name for private file reads.
  • LOG_LEVEL: Controls server log verbosity. Valid values: error, warn, info, debug. Defaults to info (or debug in dev/local environments unless explicitly set).

Auth Configuration

Set these environment variables to configure authentication:
  • ENABLE_PASSWORD_AUTH: Set to “false” to disable password authentication.
  • OAUTH_ONLY: Set to “true” to only allow OAuth authentication.
  • SESSION_SECRET: A secret used to sign the session cookie that carries OAuth sign-in state. Use a long random value, and the same value on every API instance.
  • GOOGLE_CLIENT_ID: The client ID for Google OAuth.
  • GOOGLE_CLIENT_SECRET: The client secret for Google OAuth.
  • MICROSOFT_CLIENT_ID: The client ID for Microsoft/Azure AD OAuth.
  • MICROSOFT_CLIENT_SECRET: The client secret for Microsoft/Azure AD OAuth.
  • MICROSOFT_TENANT_ID: The tenant for Microsoft/Azure AD OAuth. Accepts a tenant GUID, a verified domain (e.g. contoso.com or contoso.onmicrosoft.com), or one of the multi-tenant aliases common, organizations, or consumers.
  • MICROSOFT_CALLBACK_URL: The callback URL for Microsoft OAuth (e.g. https://your-domain.com/api/auth/microsoft/callback).
OAuth sign-in keeps its state in that session cookie between the redirect to the provider and the callback, so the cookie has to make the round trip:
  • Browsers must reach Variable over HTTPS. The cookie is Secure, so a browser on plain HTTP doesn’t keep it.
  • Sign-in must start on the callback’s host. Callbacks go to the host of VARIABLE_APP_URL, or of GOOGLE_CALLBACK_URL / MICROSOFT_CALLBACK_URL when set. A cookie set on any other hostname isn’t sent with the callback.
  • Every API instance needs the same SESSION_SECRET. A cookie signed by one instance fails the signature check on another.
If the cookie doesn’t make the round trip, sign-in returns to the sign-in page with an error, and the API logs Unable to verify authorization request state. with a hint naming the likely cause. If TLS ends at a load balancer or reverse proxy, have the proxy that forwards /api to the server send X-Forwarded-Proto: https. Sign-in works without it, but the API otherwise treats every request as plain HTTP.

Email Configuration

Variable sends transactional email for things like user invitations, password resets, and data requests. Self-hosted deployments deliver email via SMTP. If SMTP is not configured, email sending is disabled and email-dependent flows will not work: password reset, and setting a password after an invitation, which needs the link in the invitation email. Invited people can still sign in with Google, Microsoft or SSO where those are set up. Configure these to route email through your own SMTP server (e.g. Office 365, Google Workspace SMTP relay, Amazon SES SMTP, Postfix):
  • EMAIL_ENABLED: Set to true to enable outbound email. Defaults to false. When false, the app runs normally but no email is sent.
  • EMAIL_DEFAULT_SENDER: The From address used on outbound mail (e.g. hello@your-domain.com). Required for email to be considered configured.
  • EMAIL_ALLOW_DOMAINS (optional): Semicolon-separated allowlist of recipient domains. When set to anything other than *, mail to recipients outside the listed domains is silently dropped - useful for staging environments or for locking outbound mail to internal domains during a rollout. Set to * (or leave unset) to allow all domains.
  • SMTP_HOST: SMTP server hostname (e.g. smtp.office365.com, smtp-relay.gmail.com).
  • SMTP_PORT: SMTP server port. Defaults to 587.
  • SMTP_SECURE: Set to true to use implicit TLS (typically port 465). Defaults to false, which uses STARTTLS on port 587.
  • SMTP_USER: SMTP username.
  • SMTP_PASSWORD: SMTP password or app-specific token.
SMTP is considered configured when SMTP_HOST, SMTP_USER, and SMTP_PASSWORD are all set.
.env SMTP example

Database Configuration

Variable uses Neo4j 5.26 (LTS) with APOC Core. APOC Core ships inside the official Neo4j 5.x Docker image and is activated via the NEO4J_PLUGINS environment variable - no separate download is required.

Docker volumes

If hosting the database as a Docker container, you will need to have the following volumes mounted:
  • /data: This is where the Neo4j database files will be stored.
  • /logs: This is where the Neo4j logs will be stored.
  • /import: This is where you can place any initial data files to be imported into the database.

Required Neo4j settings

These environment variables must be set on the Neo4j container for Variable to function correctly:
These settings are not strictly required but are recommended for production deployments:
When disabling transaction memory limits (=0), ensure your Neo4j container has an explicit memory constraint (Docker mem_limit, Kubernetes resource limits, or VM-level cap) so that a single expensive query cannot consume all host memory and trigger an OOM kill.

Example Docker Compose service

docker-compose.yaml
The NEO4J_AUTH value must match what you configure in the API server’s NEO4J_AUTH environment variable. The format is username/password.

URL Mapping

The load balancer or reverse proxy should be configured to map the following URLs: The UI should be served at <APP_URL> and should point to the Variable UI container. The API should be served at <APP_URL>/api Public files should be served at <APP_URL>/uploads

File Storage Configuration

The application is set up to upload files to the an “upload” bucket where the files are scanned for malware before being moved to a “read” bucket.

Architecture Diagram

Terraform examples

This is not a complete Terraform configuration, but it provides a starting point to see how the architecture pieces above can be set up in Terraform.

Security Headers (locals.tf)

Define your security headers in a locals.tf file so they can be reused across multiple backend services and buckets without duplication.
locals.tf
Replace https://storage.googleapis.com with your storage domain. Add additional origins as needed (see CSP section).

Load Balancer and URL Mapping

load-balancer.tf

UI Server

The UI is served by a small nginx container (published as variable-ui) that hosts the built static assets. Deploy it the same way as the API - any Docker environment works (Cloud Run, GKE, GCE, etc.). The example below uses Cloud Run.
ui.tf

Storage

Uploads

uploads.tf

API Server

api.tf

Google Compute for Neo4j

database.tf

Security Headers

Your load balancer, reverse proxy, or the UI’s nginx container should set the following security headers on responses served to the browser. These are typically configured on the outermost layer that serves the UI - either the cloud load balancer / CDN, or the nginx container itself if it is the outermost hop.

Content Security Policy (CSP)

The CSP header restricts which origins the browser is allowed to load resources from. A recommended baseline for Variable:
Replace <your-storage-domain> with the domain of your file storage service (e.g. https://storage.googleapis.com for GCS, or https://<account>.blob.core.windows.net for Azure Blob Storage). Depending on which third-party integrations you enable, you may need to add additional origins:

Example: nginx

If you are deploying behind a cloud load balancer (e.g. GCP Cloud Load Balancing, Azure Application Gateway, AWS ALB/CloudFront), security headers are typically set at the load balancer or CDN level - not in nginx. Setting them in both places can cause duplicate headers. Only use the nginx configuration below if nginx is your outermost reverse proxy.
Nginx’s add_header directive is not inherited into location blocks that define their own add_header. Define your security headers in a shared snippet file and include it in each location block.
/etc/nginx/snippets/security-headers.conf
If your storage domain is separate from your app domain, add it to img-src and frame-src (e.g. img-src 'self' data: https://<your-storage-domain>).

Onboarding

After deploying the containers and configuring environment variables, follow these steps to initialize your Variable instance.

Step 1: Start the Containers

Start all services using Docker Compose (or your container orchestration tool). The API server will automatically run any pending database migrations on startup.

Step 2: Populate Base Data

Navigate to your Variable instance URL in a browser. When the UI detects an empty database, it will guide you through the setup process automatically. This creates:
  • Database constraints and indexes
  • GHG Protocol scopes (Scope 1, 2, 3)
  • Emission factor data sources (DEFRA, Ember, ecoinvent, etc.)
  • LCA stage definitions (A1-A3, B1-B7, C1-C4, D)
  • Product taxonomy hierarchy
  • Geographic locations and electricity grid data
  • Units of measurement
  • An admin user (using the first email from SYSTEM_ADMIN_EMAILS) with full access
  • Your company with an Enterprise subscription

Step 3: Load Emission Factors

Load your emission factor data using either a data package (recommended) or URL-based import. See the Loading Emission Factors section below for details.

Step 4: Sign In

Navigate to your Variable instance URL and sign in. If you configured OAuth (Google or Azure AD), the SSO flow will be used. If you enabled password authentication, use the admin email to create your account. The admin user (the first email in SYSTEM_ADMIN_EMAILS) is automatically granted global admin privileges and assigned as the owner of the company created in Step 2.

Loading Emission Factors

There are two ways to load emission factors into your Variable instance: data packages (recommended) and URL-based import. A data package is a directory of CSV files that is mounted directly into the API server container. This is the simplest approach for self-hosted deployments because the data is loaded from the local filesystem with no external network access required.

1. Pull your data package

Variable delivers data packages via Google Artifact Registry. You will receive a service account key file (key.json) and your registry URL from Variable. Install ORAS ORAS is the tool used to pull data packages from the registry.
Authenticate
Pull and extract
To pin to a specific version, replace latest with the version tag (e.g. 2026.1). After extraction, your data-packages/ directory will contain the emissionFactors/ and impacts/ subdirectories ready to be mounted.

3. Mount the volume

In your docker-compose.yaml, the data package directory is mounted into the API server container and configured via the DATA_PACKAGE_DIR environment variable:

4. Install the data package

  1. Start your containers and sign in to Variable.
  2. Navigate to Admin > Data sources.
  3. Click the Import dropdown and select Install emission factors.
  4. The installation will begin. The Data Sources table refreshes automatically during installation, so you can watch the emission factor counts increase as data is loaded.
Depending on the size of the dataset, installation may take several minutes. Both emission factors and impact data (if present) are installed in a single operation.
If the Install emission factors option does not appear, verify that CSV files are present in the emissionFactors/ directory inside your data package volume and that DATA_PACKAGE_DIR is set correctly. Check the server logs for the Data package dir: message on startup.

URL-based import

For cases where you prefer to host emission factor files on an external server or cloud storage bucket, you can import them by providing a URL.
  1. Sign in to your Variable instance and navigate to Admin > Data sources.
  2. Click the Import dropdown and select Import emission factors.
  3. Provide an authenticated URL to the CSV file. This URL must be accessible from the server (e.g., a pre-signed S3 URL, a GCS signed URL, or a URL with an access token).
  4. Preview the data and click Import to load the emission factors into the database.
To import impact data separately, use Import impacts from the same dropdown.

Supported data sources

The following data sources are supported:
  • DEFRA 2022
  • Ember 2022, 2024, 2025
  • ecoinvent 3.8, 3.11
  • Idemat 2023
  • IPCC AR6
  • AIB Residual Grid Mix
  • US EPA eGRID 2022
  • WIOD
Variable will provide the emission factor CSV files as part of your license.
For a self-hosted deployment, we recommend the following infrastructure:
  • API Server: 2 vCPUs, 4 GB RAM
  • Database: 4 vCPUs, 16 GB RAM (or use a managed database service)
  • File Storage: Use a scalable object storage service like AWS S3, Google Cloud, etc.
  • Malware Scanning: Use a service like ClamAV or integrate with a third-party malware scanning service.

Health Checks

The API server exposes an unauthenticated health check endpoint that load balancers, container orchestrators, and uptime monitors can use to verify the process is running.

Endpoint

No authentication is required. A healthy response is a 200 OK with a JSON body:
This is a basic API reachability check: it verifies the API process can respond to HTTP requests, but it does not query the Neo4j database or file storage. /api/health is exempt from the maintenance mode gate, so it keeps returning 200 OK even while maintenance is active - safe to use as a Kubernetes livenessProbe or load-balancer healthcheck without special handling. If you need readiness validation for downstream dependencies, add a separate probe - for example, a script that calls /api/health and then runs a trivial Cypher query against Neo4j.

Response codes

Example configurations

The examples below use port 8181 (the default). Replace it with whatever value you configured for API_PORT.
curl
Docker Compose The official Variable API image is based on node:slim and does not ship with curl or wget, so the example below uses Node to make the HTTP request. If you build your own image, use whichever HTTP client is available in it.
Kubernetes
GCP Load Balancer (Terraform)

Maintenance Mode

When performing upgrades, database migrations, or other operations that require the application to be temporarily unavailable, you can put Variable into maintenance mode. While maintenance mode is active, the API returns 503 Service Unavailable with a user-friendly message for every non-exempt request, and the UI renders a dedicated maintenance screen instead of the application. Routine windows are configured from the admin UI and stored in the database, so toggling maintenance on or off no longer requires a redeploy. A single environment variable remains for genuine emergencies.

Activation paths

Three independent paths activate the gate; if any is active, the API blocks. Subsystem detail lives in packages/api/src/maintenance/README.md. Scheduled and manual mode propagate across API replicas within ~30s of a save (cache poll interval). The panic switch is per-replica and propagates only when each replica restarts with the new env var.

Configuring scheduled windows

Sign in as a system admin (an email listed in SYSTEM_ADMIN_EMAILS) and open Admin → Maintenance windows.
  • Recurrence: Daily, Weekly, or Monthly. Selecting No schedule clears the window.
  • Weekday (weekly only) or Day of month (monthly only, 1–28 so every month fires).
  • Window start: HH:mm 24-hour wall-clock time in the chosen timezone.
  • Duration: minutes the window stays active (1 minute to 7 days).
  • Timezone: any IANA zone (e.g. Europe/London, America/New_York). The host’s clock has no effect on the result.
The window runs as [start, start + duration) — a request exactly at the end-time is not blocked. On spring-forward DST transitions, if the start lands in the skipped hour the day’s window is dropped; on fall-back transitions, a window that crosses the duplicated hour can run ~1h longer than its configured duration. The schedule fields are validated server-side on save and surfaced as inline errors in the form.

Starting maintenance immediately

Use Start now on the same page to enable manual mode without scheduling. The button requires typing START to confirm — there is no undo, and active sessions begin seeing 503s within ~30s. Manual mode has no server-side auto-expiry; it runs until a system admin ends it from the maintenance screen.

Ending a window early

Once maintenance is active, the admin form is no longer reachable — every non-auth route, including /admin/maintenance, is replaced by the maintenance screen. System admins see an End maintenance button on that screen; clicking it ends the current window immediately (whether it was started manually or by the schedule) and leaves future scheduled occurrences untouched. Active browser sessions that loaded during maintenance reload automatically once the next request succeeds, so users recover without action.

Panic switch (VARIABLE_MAINTENANCE_MODE)

The env var is a fallback for the case where the admin UI itself is unreachable — most commonly because Neo4j is down. Setting VARIABLE_MAINTENANCE_MODE=true blocks every non-exempt request without touching Neo4j and without consulting the scheduled config.
.env
Apply it the same way as any other env-var change:
On platforms like Cloud Run, update the service’s environment variables and deploy a new revision. The previous VARIABLE_MAINTENANCE_END_DATE env var has been removed — scheduled windows now carry their own end time.