Blixt Documentation v2.9

Four layers

BlixtFS builds its configuration from four layers. Each overrides the one before it:

  1. Built-in defaults.
  2. A YAML file, passed with --config /data/config.yaml.
  3. Environment variables, such as DATABASE_PASSWORD or AUTH_TOKEN.
  4. Command-line flags, such as --default_cloud aws.

Put the bulk of your configuration in the file, and use environment variables for secrets so they stay out of it. A file only needs the values you want to change. Anything you leave out keeps its default.

The configuration file

The file is organised by section. A typical single-node file needs only cloud:

cloud:
  gcp:
    project: my-project
    credentials: /credentials
    buckets:
      - bucket: gs://training-data
        read_only: true
        lazy: true
      - bucket: gs://results

The sections you are most likely to change:

Section Controls
general Default cloud, worker threads
cloud Providers, credentials, buckets, change notifications, event log
disk Where data, cache and logs live (default /data)
cache Read cache configuration
indexing Indexing workers and consistency-check policy
filesystem Default owner and group for files written by other tools
protocol.nfs NFS protocol gateway
protocol.smb SMB protocol gateway
logging What gets logged, and in which format

Configuration settings lists the important settings in each section.

Printing the effective configuration

Every server logs its effective configuration at startup, after all four layers have been applied. Check the log first when a setting doesn’t seem to take effect.

Configuration in scale-out deployments

In a scale-out deployment you configure the config server. It distributes the configuration to every other server, which fetches it at startup and receives updates while running. Settings that describe one particular process, such as its IP addresses, go in the local section, which is never distributed.

Cloud credentials are the exception: the config server never distributes them. Every server that talks to a bucket needs its own credentials, through workload identity, a mounted file or environment variables.

After changing the configuration in Kubernetes, restart the config server:

kubectl -n bfs rollout restart deploy/config