docverse Helm values reference#

Helm values reference table for the docverse application.

Key

Type

Default

Description

affinity

object

{}

Affinity rules for the docverse deployment pod

cloudsql.enabled

bool

false

Enable the Cloud SQL Auth Proxy sidecar, used with Cloud SQL databases on Google Cloud

cloudsql.image.pullPolicy

string

"IfNotPresent"

Pull policy for Cloud SQL Auth Proxy images

cloudsql.image.repository

string

"gcr.io/cloudsql-docker/gce-proxy"

Cloud SQL Auth Proxy image to use

cloudsql.image.tag

string

"1.38.3"

Cloud SQL Auth Proxy tag to use

cloudsql.instanceConnectionName

string

""

Instance connection name for a Cloud SQL PostgreSQL instance

cloudsql.resources

object

See values.yaml

Resource requests and limits for Cloud SQL Auth Proxy

cloudsql.serviceAccount

string

""

The Google service account that has an IAM binding to the docverse Kubernetes service accounts and has the cloudsql.client role

config.arqRedisUrl

string

Points to embedded Redis

URL for Redis arq queue database

config.cdnPurgeEnabled

bool

false

Whether a long-profile edition publish is followed by a purge of the project’s hostname from the Cloudflare edge cache. Leave this off until the Docverse Worker edge-caches edition responses: today a purge invalidates nothing while still spending calls against Cloudflare’s per-account purge rate limit (5 per minute on the Free plan), which a keeper-sync backfill across many hostnames exceeds within seconds. Re-enabling is tracked in lsst-sqre/docverse#683.

config.credentialKeyRotation

bool

false

Set true during a credential-encryption (Fernet) key rotation to deliver the retired key (DOCVERSE_CREDENTIAL_ENCRYPTION_KEY_RETIRED) to all pods so existing credentials can still be decrypted. Set back to false and remove the Vault key once all credentials have been re-encrypted.

config.databaseUrl

string

""

Database URL for PostgreSQL

config.githubAppId

string

nil

GitHub App ID for Docverse to use when accessing GitHub repositories. If not set, Docverse will operate in a limited mode without GitHub integration.

config.keeperSync.copyRetryDelaySeconds

float

30

Seconds the keeper-sync worker waits before re-running a build copy that failed on a transport error at either end (an R2 outage that outlasted the per-object budget, or an LTD S3 timeout during a download). The copy is re-run exactly once; any other failure fails the edition as before.

config.keeperSync.enabled

bool

false

Enable the Keeper-sync worker that consumes the docverse:sync-queue arq queue. Requires the docverse image to provide docverse.worker.main.KeeperSyncWorkerSettings.

config.keeperSync.jobTimeoutSeconds

int

3600

Per-job timeout, in seconds, for keeper-sync arq jobs. The server derives its two neighbours from this value: the slice budget (sliceBudgetSeconds) to this less 600 s and the stuck-run reaper threshold (config.reaperThresholds.keeperSyncSeconds) to this plus 1800 s, so the three read as one ladder, budget < timeout < reaper. Raising it is also the escape hatch for an edition too large to copy inside one slice, since the budget follows it up.

config.keeperSync.sliceBudgetSeconds

string

Derived by the server (jobTimeoutSeconds - 600)

How long, in seconds, one keeper_sync_project job keeps starting editions before it completes its row and hands the rest of the project to a continuation job, so a project of any size converges across a chain of jobs that each finish inside jobTimeoutSeconds. Checked before each edition, never mid-copy. Leave unset to let the server derive it as jobTimeoutSeconds less a 600 s margin (3000 s at the default timeout); an explicit value must be greater than 0 and less than jobTimeoutSeconds or the server refuses to start. Set it low (for example 300) on a development environment to exercise the continuation chain on a mid-size project.

config.keeperSync.uploadConcurrency

int

32

Process-wide cap on concurrent presigned uploads across every running keeper-sync job in the sync worker. Each concurrent upload holds one outbound connection, and the copy client keeps that many (plus 10 of headroom) connections alive, so this bounds the NAT source ports one sync-worker pod consumes. 32 fits the default 64 static Cloud NAT ports per GKE node; raise it toward DOCVERSE_KEEPER_SYNC_MAX_JOBS x DOCVERSE_KEEPER_SYNC_COPY_CONCURRENCY (80 at the defaults) once the Cloud NAT allocation is raised and the NAT drop logs stay at zero through a full backfill.

config.keeperSync.uploadMaxAttempts

int

6

Attempts, including the first, that the keeper-sync worker spends on one presigned upload of a copied object before failing the build copy. Transport failures and retryable statuses (429, 5xx) share the budget; backoff starts at 0.5 s and doubles. At the default of 6 an object rides out an R2 connect outage of roughly 65 s.

config.keeperSync.uploadMaxBackoffSeconds

float

30

Ceiling, in seconds, on any single wait between attempts of a keeper-sync presigned upload, including a server-requested Retry-After. At the default attempt count the exponential backoff peaks at 8 s, so this only bites for a longer Retry-After.

config.logLevel

string

"INFO"

Logging level

config.logProfile

string

"production"

Logging profile (production for JSON, development for human-friendly)

config.maintenance.editionReconcileEnabled

bool

true

Enable the edition reconcile loop that re-drives editions whose recorded publish state has drifted from what the CDN serves. The cron stays registered either way; disabling only makes each tick a no-op.

config.maintenance.editionReconcileMaxActionsPerJob

int

100

Maximum number of republish plus unpublish actions one per-org edition_reconcile job applies per tick; the remainder is reported as capped and picked up on the next tick.

config.maintenance.enabled

bool

false

Enable the maintenance worker that consumes the docverse:maintenance-queue arq queue. Requires the docverse image to provide docverse.worker.main.MaintenanceWorkerSettings.

config.maintenance.gitRefAuditEnabled

bool

false

Whether to enable auditing the git ref lifecycle rule. Enabling this will cause docverse to make GitHub API calls to determine if the git ref associated with an edition still exists.

config.maintenance.jobTimeoutSeconds

int

3600

Per-job timeout, in seconds, for maintenance-pool jobs (lifecycle evaluation, git ref audits, and purgatory cleanup).

config.maintenance.purgatoryCleanupCronHour

int

3

UTC hour at which the daily purgatory_cleanup dispatcher cron runs.

config.maintenance.purgatoryCleanupCronMinute

int

23

UTC minute of purgatoryCleanupCronHour at which the daily purgatory_cleanup dispatcher cron runs.

config.maintenance.purgatoryCleanupEnabled

bool

true

Whether to run the daily purgatory_cleanup sweep that permanently deletes the object-store content (unpacked tree and staging tarball) of soft-deleted builds once the organization’s purgatory retention has elapsed and stamps date_purged on the row. The cron is registered either way, so flipping this does not require a worker restart.

config.maintenance.purgatoryCleanupMaxBuildsPerJob

int

500

Cap on the builds a single per-org purgatory_cleanup job reclaims per daily tick, oldest deletion first; anything past the cap is picked up by the next tick.

config.mallocArenaMax

string

"2"

Value of MALLOC_ARENA_MAX for every Docverse process. Capping glibc malloc arenas at 2 halved the resident memory the Keeper-sync worker keeps after a burst of build copies (lsst-sqre/docverse#760, dev run 4 on lsst-sqre/docverse#751) at no visible cost. Set to the empty string to leave glibc’s default.

config.memoryDiagnostics.enabled

bool

false

Enable the in-process memory sampler in every Docverse process (API, worker, Keeper-sync worker, maintenance worker). When on, each process logs one Memory sample line per interval carrying its resident size, resident high-water mark, and garbage-collector counts, so a pod’s memory growth can be read from its logs without attaching a profiler (the pods run on a read-only root filesystem with every capability dropped, so nothing can be attached).

config.memoryDiagnostics.intervalSeconds

int

60

Seconds between memory samples.

config.memoryDiagnostics.topN

int

10

Number of allocation sites, ranked by growth since the previous sample, included in each sample when tracemallocEnabled is on.

config.memoryDiagnostics.tracemallocEnabled

bool

false

Also trace Python allocations with tracemalloc and add the traced heap size, its peak, the live object count, and the top allocation sites by growth since the previous sample to each Memory sample line. Roughly doubles the heap a process needs, so leave it off in production and enable it on a development environment while chasing a leak. Has no effect unless enabled is also set.

config.memoryDiagnostics.tracemallocFrames

int

5

Stack depth tracemalloc records per allocation. Deeper traces attribute growth to the calling code rather than to the library that allocated, at more overhead per allocation.

config.metrics.application

string

"docverse"

Name under which to log metrics. Generally there is no reason to change this.

config.metrics.enabled

bool

false

Whether to enable sending application metrics events to Sasquatch over Kafka. When disabled, Docverse uses a no-op metrics manager.

config.metrics.events.topicPrefix

string

"lsst.square.metrics.events"

Topic prefix for events. It may sometimes be useful to change this in development environments.

config.metrics.schemaManager.registryUrl

string

Sasquatch in the local cluster

URL of the Confluent-compatible schema registry server

config.pathPrefix

string

"/docverse/api"

URL path prefix

config.reaperThresholds.buildProcessingSeconds

int

28800

Stuck-run reaper threshold, in seconds, for build_processing jobs.

config.reaperThresholds.dashboardBuildSeconds

int

1800

Stuck-run reaper threshold, in seconds, for dashboard_build jobs.

config.reaperThresholds.dashboardSyncSeconds

int

21600

Stuck-run reaper threshold, in seconds, for dashboard_sync jobs.

config.reaperThresholds.editionReconcileSeconds

string

Derived by the server (jobTimeoutSeconds + 1800)

Stuck-run reaper threshold, in seconds, for edition_reconcile jobs. Leave unset to let the server derive it as config.maintenance.jobTimeoutSeconds plus a 1800 s margin; an explicit value must be strictly greater than that timeout or the server refuses to start.

config.reaperThresholds.keeperSyncSeconds

string

Derived by the server (config.keeperSync.jobTimeoutSeconds + 1800)

Stuck-run reaper threshold, in seconds, for keeper-sync jobs. Leave unset to let the server derive it as config.keeperSync.jobTimeoutSeconds plus a 1800 s margin (5400 s at the default timeout). The keeper-sync functions run with a single attempt, so arq has already cancelled any job that reached its timeout and a row still in progress past that point is dead; the old flat 21600 s pin only kept the project parked behind the active-job index for hours. An explicit value only lowers the threshold: anything above timeout + 1800 is capped there and the sync worker logs a warning at startup naming both values.

config.reaperThresholds.lifecycleSeconds

int

21600

Stuck-run reaper threshold, in seconds, for lifecycle_eval and git_ref_audit jobs (maintenance pool).

config.reaperThresholds.publishEditionSeconds

int

14400

Stuck-run reaper threshold, in seconds, for publish_edition jobs.

config.reaperThresholds.purgatoryCleanupSeconds

int

21600

Stuck-run reaper threshold, in seconds, for purgatory_cleanup jobs (maintenance pool).

config.sentry.enabled

bool

false

Whether to send error reports and tracing data to Sentry. Requires the sentry-dsn secret to be set in Vault.

config.sentry.tracesSampleRate

float

0

The percentage of requests that should be traced. This should be a float between 0 and 1.

config.slackAlerts

bool

false

Whether to send Slack alerts for unexpected failures

config.superadminUsers

list

["jonathansick"]

Usernames that have super admin (de facto admin in all organizations)

config.updateSchema

bool

false

Whether to run Alembic schema migrations on install/upgrade

global.host

string

Set by Argo CD

Host name for ingress

global.repertoireUrl

string

Set by Argo CD

Base URL for Repertoire discovery API

global.vaultSecretsPath

string

Set by Argo CD

Base path for Vault secrets

image.pullPolicy

string

"IfNotPresent"

Pull policy for the docverse image

image.pythonModule

string

"docverse_server"

Top-level Python module of the server in this image. Images built after the DM-55658 packaging restructure use “docverse_server”; earlier images use “docverse”. Must match the deployed image tag.

image.repository

string

"ghcr.io/lsst-sqre/docverse"

Image to use in the docverse deployment

image.tag

string

The appVersion of the chart

Tag of image to use

ingress.annotations

object

{}

Additional annotations for the ingress rule

maintenanceWorker.affinity

object

{}

Affinity rules for the maintenance worker pod

maintenanceWorker.nodeSelector

object

{}

Node selection rules for the maintenance worker pod

maintenanceWorker.podAnnotations

object

{}

Annotations for the maintenance worker pod

maintenanceWorker.replicaCount

int

1

Number of maintenance worker pods to start

maintenanceWorker.resources

object

See values.yaml

Resource limits and requests for the maintenance worker pod

maintenanceWorker.tolerations

list

[]

Tolerations for the maintenance worker pod

nodeSelector

object

{}

Node selection rules for the docverse deployment pod

podAnnotations

object

{}

Annotations for the docverse deployment pod

redis.affinity

object

{}

Affinity rules for the Redis pod

redis.nodeSelector

object

{}

Node selection rules for the Redis pod

redis.persistence.enabled

bool

true

Whether to persist Redis storage. Setting this to false will use emptyDir and lose data on every restart.

redis.persistence.size

string

"1Gi"

Amount of persistent storage to request

redis.persistence.storageClass

string

""

Class of storage to request

redis.persistence.volumeClaimName

string

""

Use an existing PVC, not dynamic provisioning. If this is set, the size, storageClass, and accessMode settings are ignored.

redis.podAnnotations

object

{}

Pod annotations for the Redis pod

redis.resources

object

See values.yaml

Resource limits and requests for the Redis pod

redis.resources.requests.cpu

string

"50m"

GKE Autopilot requires a minimum CPU request of 50m

redis.tolerations

list

[]

Tolerations for the Redis pod

replicaCount.api

int

1

Number of API deployment pods to start

replicaCount.worker

int

1

Number of worker deployment pods to start

resources

object

See values.yaml

Resource limits and requests for the docverse deployment pod

resources.requests.cpu

string

"50m"

GKE Autopilot requires a minimum CPU request of 50m

syncWorker.affinity

object

{}

Affinity rules for the Keeper-sync worker pod

syncWorker.extraEnv

list

[]

Additional environment variables for the Keeper-sync worker container, as a list of name/value pairs. Meant for runtime experiments that need no chart change (lsst-sqre/docverse#753).

syncWorker.nodeSelector

object

{}

Node selection rules for the Keeper-sync worker pod

syncWorker.podAnnotations

object

{}

Annotations for the Keeper-sync worker pod

syncWorker.replicaCount

int

1

Number of Keeper-sync worker pods to start

syncWorker.resources

object

See values.yaml

Resource limits and requests for the Keeper-sync worker pod

syncWorker.resources.limits.memory

string

"2Gi"

Higher than the other workers because keeper-sync buffers whole LTD objects in memory while copying builds, across several concurrent jobs. Measured idle RSS after a 477-project backfill is about 840Mi (roundtable-prod, 2026-09-23), so the previous 1Gi limit left almost no room for the next wave.

syncWorker.tolerations

list

[]

Tolerations for the Keeper-sync worker pod

tolerations

list

[]

Tolerations for the docverse deployment pod

workerResources

object

See values.yaml

Resource limits and requests for the docverse worker pod

workerResources.limits.memory

string

"1Gi"

The worker idles at roughly 480Mi after a keeper-sync backfill and was OOM-killed at 512Mi while running ten concurrent publish_edition and dashboard_build jobs (roundtable-prod, 2026-09-23). A killed worker strands in-flight publishes until the 4 h reaper runs, so leave headroom for a full burst.

workerResources.requests.cpu

string

"50m"

GKE Autopilot requires a minimum CPU request of 50m

This page was last modified on .