Configuration
How BlixtFS reads its settings, and how configuration reaches every server in a scale-out deployment.
Four layers
BlixtFS builds its configuration from four layers. Each overrides the one before it:
- Built-in defaults.
- A YAML file, passed with
--config /data/config.yaml. - Environment variables, such as
DATABASE_PASSWORDorAUTH_TOKEN. - Command-line flags, such as
--default_cloud aws.
Put the bulk of your configuration in the file, and use environment variables for secrets so they stay out of it. A file only needs the values you want to change. Anything you leave out keeps its default.
The configuration file
The file is organised by section. A typical single-node file needs only cloud:
cloud:
gcp:
project: my-project
credentials: /credentials
buckets:
- bucket: gs://training-data
read_only: true
lazy: true
- bucket: gs://results
The sections you are most likely to change:
| Section | Controls |
|---|---|
general |
Default cloud, worker threads |
cloud |
Providers, credentials, buckets, change notifications, event log |
disk |
Where data, cache and logs live (default /data) |
cache |
Read cache configuration |
indexing |
Indexing workers and consistency-check policy |
filesystem |
Default owner and group for files written by other tools |
protocol.nfs |
NFS protocol gateway |
protocol.smb |
SMB protocol gateway |
logging |
What gets logged, and in which format |
Configuration settings lists the important settings in each section.
Printing the effective configuration
Every server logs its effective configuration at startup, after all four layers have been applied. Check the log first when a setting doesn’t seem to take effect.
Configuration in scale-out deployments
In a scale-out deployment you configure the config server. It distributes
the configuration to every other server, which fetches it at startup and
receives updates while running. Settings that describe one particular process,
such as its IP addresses, go in the local section, which is never
distributed.
Cloud credentials are the exception: the config server never distributes them. Every server that talks to a bucket needs its own credentials, through workload identity, a mounted file or environment variables.
After changing the configuration in Kubernetes, restart the config server:
kubectl -n bfs rollout restart deploy/config