Enterprise feature: contact us if you’re interested in self-hosting Variable
Overview
The Variable application is comprised of 4 main components: a UI, a stateless API server, a Neo4j database, and file storage for uploaded files. The UI and API are each delivered as a Docker container, while the database can be run as a Docker container or as a managed service. The file storage can be any object storage service, such as AWS S3, Google Cloud Storage, or a self-hosted solution.Components
- UI: A small nginx container that serves the static files for the Variable UI.
- Server: A stateless API server. It can be scaled horizontally by running multiple instances behind a load balancer.
- Database: A Neo4j database that stores all the data for Variable. It can be run as a standalone service or as part of a managed database solution.
- File Storage: For any files uploaded to Variable, such as images or documents. This can be any object storage service, such as AWS S3, Google Cloud Storage, or a self-hosted solution.
- Malware Scanning: An optional component that scans uploaded files for malware before they are stored. This can be integrated with services like ClamAV or other malware scanning solutions.
Authentication
User authentication is typically done via SSO. Currently supported providers are:- Microsoft / Azure AD
- Google Workspace
ENABLE_PASSWORD_AUTH environment variable.
Deployment
The UI, Server and Database can be deployed using any Docker environment, either locally or in the cloud (e.g. Google Compute Engine/Cloud Run). File storage can be configured to use any object storage service, or a self-hosted solution.Configuration
Each component will require configuration to connect to the other components.API Configuration
These are the environment variables that need to be set:- VARIABLE_APP_URL: The base URL of the Variable application, which is used to generate links and handle redirects.
- SYSTEM_ADMIN_EMAILS: Semicolon-separated list of system admin email addresses. The first email is used as the initial admin user for database seeding.
- DB_NEO4J_URL: The URI for the Neo4j database (e.g. bolt://localhost:7687).
- NEO4J_AUTH: The username/password for the Neo4j database.
- JWT_SECRET: A secret used to sign JSON Web Tokens.
- STORAGE_PROVIDER: The storage provider to use (e.g. “gcs” for Google Cloud Storage, “azure” for Azure Blob Storage).
- VARIABLE_PUBLIC_BUCKET_UPLOAD: The bucket name for public file uploads.
- VARIABLE_PUBLIC_BUCKET_READ: The bucket name for public file reads.
- VARIABLE_PRIVATE_BUCKET_UPLOAD: The bucket name for private file uploads.
- VARIABLE_PRIVATE_BUCKET_READ: The bucket name for private file reads.
- LOG_LEVEL: Controls server log verbosity. Valid values:
error,warn,info,debug. Defaults toinfo(ordebugin dev/local environments unless explicitly set).
Auth Configuration
Set these environment variables to configure authentication:- ENABLE_PASSWORD_AUTH: Set to “false” to disable password authentication.
- OAUTH_ONLY: Set to “true” to only allow OAuth authentication.
- SESSION_SECRET: A secret used to sign the session cookie that carries OAuth sign-in state. Use a long random value, and the same value on every API instance.
- GOOGLE_CLIENT_ID: The client ID for Google OAuth.
- GOOGLE_CLIENT_SECRET: The client secret for Google OAuth.
- MICROSOFT_CLIENT_ID: The client ID for Microsoft/Azure AD OAuth.
- MICROSOFT_CLIENT_SECRET: The client secret for Microsoft/Azure AD OAuth.
- MICROSOFT_TENANT_ID: The tenant for Microsoft/Azure AD OAuth. Accepts a tenant GUID, a verified domain (e.g.
contoso.comorcontoso.onmicrosoft.com), or one of the multi-tenant aliasescommon,organizations, orconsumers. - MICROSOFT_CALLBACK_URL: The callback URL for Microsoft OAuth (e.g.
https://your-domain.com/api/auth/microsoft/callback).
- Browsers must reach Variable over HTTPS. The cookie is
Secure, so a browser on plain HTTP doesn’t keep it. - Sign-in must start on the callback’s host. Callbacks go to the host of
VARIABLE_APP_URL, or ofGOOGLE_CALLBACK_URL/MICROSOFT_CALLBACK_URLwhen set. A cookie set on any other hostname isn’t sent with the callback. - Every API instance needs the same
SESSION_SECRET. A cookie signed by one instance fails the signature check on another.
Unable to verify authorization request state. with a hint naming the likely cause.
If TLS ends at a load balancer or reverse proxy, have the proxy that forwards /api to the server send X-Forwarded-Proto: https. Sign-in works without it, but the API otherwise treats every request as plain HTTP.
Email Configuration
Variable sends transactional email for things like user invitations, password resets, and data requests. Self-hosted deployments deliver email via SMTP. If SMTP is not configured, email sending is disabled and email-dependent flows will not work: password reset, and setting a password after an invitation, which needs the link in the invitation email. Invited people can still sign in with Google, Microsoft or SSO where those are set up. Configure these to route email through your own SMTP server (e.g. Office 365, Google Workspace SMTP relay, Amazon SES SMTP, Postfix):- EMAIL_ENABLED: Set to
trueto enable outbound email. Defaults tofalse. Whenfalse, the app runs normally but no email is sent. - EMAIL_DEFAULT_SENDER: The
Fromaddress used on outbound mail (e.g.hello@your-domain.com). Required for email to be considered configured. - EMAIL_ALLOW_DOMAINS (optional): Semicolon-separated allowlist of recipient domains. When set to anything other than
*, mail to recipients outside the listed domains is silently dropped - useful for staging environments or for locking outbound mail to internal domains during a rollout. Set to*(or leave unset) to allow all domains. - SMTP_HOST: SMTP server hostname (e.g.
smtp.office365.com,smtp-relay.gmail.com). - SMTP_PORT: SMTP server port. Defaults to
587. - SMTP_SECURE: Set to
trueto use implicit TLS (typically port465). Defaults tofalse, which uses STARTTLS on port587. - SMTP_USER: SMTP username.
- SMTP_PASSWORD: SMTP password or app-specific token.
SMTP_HOST, SMTP_USER, and SMTP_PASSWORD are all set.
.env SMTP example
Database Configuration
Variable uses Neo4j 5.26 (LTS) with APOC Core. APOC Core ships inside the official Neo4j 5.x Docker image and is activated via theNEO4J_PLUGINS environment variable - no separate download is required.
Docker volumes
If hosting the database as a Docker container, you will need to have the following volumes mounted:- /data: This is where the Neo4j database files will be stored.
- /logs: This is where the Neo4j logs will be stored.
- /import: This is where you can place any initial data files to be imported into the database.
Required Neo4j settings
These environment variables must be set on the Neo4j container for Variable to function correctly:Recommended Neo4j settings
These settings are not strictly required but are recommended for production deployments:Example Docker Compose service
docker-compose.yaml
The
NEO4J_AUTH value must match what you configure in the API server’s NEO4J_AUTH environment variable. The format is username/password.URL Mapping
The load balancer or reverse proxy should be configured to map the following URLs: The UI should be served at<APP_URL> and should point to the Variable UI container.
The API should be served at <APP_URL>/api
Public files should be served at <APP_URL>/uploads
File Storage Configuration
The application is set up to upload files to the an “upload” bucket where the files are scanned for malware before being moved to a “read” bucket.Architecture Diagram
Terraform examples
This is not a complete Terraform configuration, but it provides a starting point to see how the architecture pieces above can be set up in Terraform.Security Headers (locals.tf)
Define your security headers in alocals.tf file so they can be reused across multiple backend services and buckets without duplication.
locals.tf
https://storage.googleapis.com with your storage domain. Add additional origins as needed (see CSP section).
Load Balancer and URL Mapping
load-balancer.tf
UI Server
The UI is served by a small nginx container (published asvariable-ui) that hosts the built static assets. Deploy it the same way as the API - any Docker environment works (Cloud Run, GKE, GCE, etc.). The example below uses Cloud Run.
ui.tf
Storage
Uploads
uploads.tf
API Server
api.tf
Google Compute for Neo4j
database.tf
Security Headers
Your load balancer, reverse proxy, or the UI’s nginx container should set the following security headers on responses served to the browser. These are typically configured on the outermost layer that serves the UI - either the cloud load balancer / CDN, or the nginx container itself if it is the outermost hop.Content Security Policy (CSP)
The CSP header restricts which origins the browser is allowed to load resources from. A recommended baseline for Variable:<your-storage-domain> with the domain of your file storage service (e.g. https://storage.googleapis.com for GCS, or https://<account>.blob.core.windows.net for Azure Blob Storage).
Depending on which third-party integrations you enable, you may need to add additional origins:
Other Recommended Headers
Example: nginx
Nginx’s
add_header directive is not inherited into location blocks that define their own add_header. Define your security headers in a shared snippet file and include it in each location block./etc/nginx/snippets/security-headers.conf
img-src and frame-src (e.g. img-src 'self' data: https://<your-storage-domain>).
Onboarding
After deploying the containers and configuring environment variables, follow these steps to initialize your Variable instance.Step 1: Start the Containers
Start all services using Docker Compose (or your container orchestration tool). The API server will automatically run any pending database migrations on startup.Step 2: Populate Base Data
Navigate to your Variable instance URL in a browser. When the UI detects an empty database, it will guide you through the setup process automatically. This creates:- Database constraints and indexes
- GHG Protocol scopes (Scope 1, 2, 3)
- Emission factor data sources (DEFRA, Ember, ecoinvent, etc.)
- LCA stage definitions (A1-A3, B1-B7, C1-C4, D)
- Product taxonomy hierarchy
- Geographic locations and electricity grid data
- Units of measurement
- An admin user (using the first email from
SYSTEM_ADMIN_EMAILS) with full access - Your company with an Enterprise subscription
Step 3: Load Emission Factors
Load your emission factor data using either a data package (recommended) or URL-based import. See the Loading Emission Factors section below for details.Step 4: Sign In
Navigate to your Variable instance URL and sign in. If you configured OAuth (Google or Azure AD), the SSO flow will be used. If you enabled password authentication, use the admin email to create your account. The admin user (the first email inSYSTEM_ADMIN_EMAILS) is automatically granted global admin privileges and assigned as the owner of the company created in Step 2.
Loading Emission Factors
There are two ways to load emission factors into your Variable instance: data packages (recommended) and URL-based import.Data packages (recommended)
A data package is a directory of CSV files that is mounted directly into the API server container. This is the simplest approach for self-hosted deployments because the data is loaded from the local filesystem with no external network access required.1. Pull your data package
Variable delivers data packages via Google Artifact Registry. You will receive a service account key file (key.json) and your registry URL from Variable.
Install ORAS
ORAS is the tool used to pull data packages from the registry.
latest with the version tag (e.g. 2026.1).
After extraction, your data-packages/ directory will contain the emissionFactors/ and impacts/ subdirectories ready to be mounted.
3. Mount the volume
In yourdocker-compose.yaml, the data package directory is mounted into the API server container and configured via the DATA_PACKAGE_DIR environment variable:
4. Install the data package
- Start your containers and sign in to Variable.
- Navigate to Admin > Data sources.
- Click the Import dropdown and select Install emission factors.
- The installation will begin. The Data Sources table refreshes automatically during installation, so you can watch the emission factor counts increase as data is loaded.
If the Install emission factors option does not appear, verify that CSV files are present in the
emissionFactors/ directory inside your data package volume and that DATA_PACKAGE_DIR is set correctly. Check the server logs for the Data package dir: message on startup.URL-based import
For cases where you prefer to host emission factor files on an external server or cloud storage bucket, you can import them by providing a URL.- Sign in to your Variable instance and navigate to Admin > Data sources.
- Click the Import dropdown and select Import emission factors.
- Provide an authenticated URL to the CSV file. This URL must be accessible from the server (e.g., a pre-signed S3 URL, a GCS signed URL, or a URL with an access token).
- Preview the data and click Import to load the emission factors into the database.
Supported data sources
The following data sources are supported:- DEFRA 2022
- Ember 2022, 2024, 2025
- ecoinvent 3.8, 3.11
- Idemat 2023
- IPCC AR6
- AIB Residual Grid Mix
- US EPA eGRID 2022
- WIOD
Variable will provide the emission factor CSV files as part of your license.
Recommended Infrastructure
For a self-hosted deployment, we recommend the following infrastructure:- API Server: 2 vCPUs, 4 GB RAM
- Database: 4 vCPUs, 16 GB RAM (or use a managed database service)
- File Storage: Use a scalable object storage service like AWS S3, Google Cloud, etc.
- Malware Scanning: Use a service like ClamAV or integrate with a third-party malware scanning service.
Health Checks
The API server exposes an unauthenticated health check endpoint that load balancers, container orchestrators, and uptime monitors can use to verify the process is running.Endpoint
200 OK with a JSON body:
This is a basic API reachability check: it verifies the API process can respond to HTTP requests, but it does not query the Neo4j database or file storage.
/api/health is exempt from the maintenance mode gate, so it keeps returning 200 OK even while maintenance is active - safe to use as a Kubernetes livenessProbe or load-balancer healthcheck without special handling. If you need readiness validation for downstream dependencies, add a separate probe - for example, a script that calls /api/health and then runs a trivial Cypher query against Neo4j.Response codes
Example configurations
The examples below use port
8181 (the default). Replace it with whatever value you configured for API_PORT.node:slim and does not ship with curl or wget, so the example below uses Node to make the HTTP request. If you build your own image, use whichever HTTP client is available in it.
Maintenance Mode
When performing upgrades, database migrations, or other operations that require the application to be temporarily unavailable, you can put Variable into maintenance mode. While maintenance mode is active, the API returns503 Service Unavailable with a user-friendly message for every non-exempt request, and the UI renders a dedicated maintenance screen instead of the application.
Routine windows are configured from the admin UI and stored in the database, so toggling maintenance on or off no longer requires a redeploy. A single environment variable remains for genuine emergencies.
Activation paths
Three independent paths activate the gate; if any is active, the API blocks. Subsystem detail lives inpackages/api/src/maintenance/README.md.
Scheduled and manual mode propagate across API replicas within ~30s of a save (cache poll interval). The panic switch is per-replica and propagates only when each replica restarts with the new env var.
Configuring scheduled windows
Sign in as a system admin (an email listed inSYSTEM_ADMIN_EMAILS) and open Admin → Maintenance windows.
- Recurrence:
Daily,Weekly, orMonthly. SelectingNo scheduleclears the window. - Weekday (weekly only) or Day of month (monthly only, 1–28 so every month fires).
- Window start:
HH:mm24-hour wall-clock time in the chosen timezone. - Duration: minutes the window stays active (1 minute to 7 days).
- Timezone: any IANA zone (e.g.
Europe/London,America/New_York). The host’s clock has no effect on the result.
[start, start + duration) — a request exactly at the end-time is not blocked. On spring-forward DST transitions, if the start lands in the skipped hour the day’s window is dropped; on fall-back transitions, a window that crosses the duplicated hour can run ~1h longer than its configured duration.
The schedule fields are validated server-side on save and surfaced as inline errors in the form.
Starting maintenance immediately
Use Start now on the same page to enable manual mode without scheduling. The button requires typingSTART to confirm — there is no undo, and active sessions begin seeing 503s within ~30s. Manual mode has no server-side auto-expiry; it runs until a system admin ends it from the maintenance screen.
Ending a window early
Once maintenance is active, the admin form is no longer reachable — every non-auth route, including/admin/maintenance, is replaced by the maintenance screen. System admins see an End maintenance button on that screen; clicking it ends the current window immediately (whether it was started manually or by the schedule) and leaves future scheduled occurrences untouched.
Active browser sessions that loaded during maintenance reload automatically once the next request succeeds, so users recover without action.
Panic switch (VARIABLE_MAINTENANCE_MODE)
The env var is a fallback for the case where the admin UI itself is unreachable — most commonly because Neo4j is down. Setting VARIABLE_MAINTENANCE_MODE=true blocks every non-exempt request without touching Neo4j and without consulting the scheduled config.
.env
VARIABLE_MAINTENANCE_END_DATE env var has been removed — scheduled windows now carry their own end time.