Buckets
Name the buckets BlixtFS serves, and control how each one is indexed, updated and replicated.
Bucket specs
On the command line and in environment variables, a bucket is named with a bucket spec: a scheme, the bucket name, and optional attributes after a colon.
scheme://bucket-name[:attribute,attribute=value,...]
docker run ... blixtfs/standard:2.9.0 \
s3://my-bucket \
gs://training-data:readonly,lazy \
r2://media-archive:rescan=6h
Schemes
| Scheme | Provider | Directory under the mount root |
|---|---|---|
s3:// |
Amazon S3 | aws/ |
gs:// |
Google Cloud Storage | gcp/ |
azure:// |
Azure Blob Storage | azure/ |
oci:// |
Oracle OCI Object Storage | oracle/ |
minio:// |
MinIO | minio/ |
r2:// |
Cloudflare R2 | cloudflare/ |
cw:// |
CoreWeave AI Object Storage | coreweave/ |
A bucket name without a scheme belongs to the cloud named by --default_cloud.
Attributes
| Attribute | Meaning | Available |
|---|---|---|
readonly |
Never write to this bucket. See Read-only buckets. | 2.8 or earlier |
ro |
Short for readonly. |
2.9 |
anonymous |
Read without credentials, for public datasets. Implies readonly. |
2.9 |
lazy |
Serve immediately; index each directory on first access. | 2.8 or earlier |
pubsub, nopubsub |
Turn change notifications on or off for this bucket. | 2.8 or earlier |
topic=NAME |
Subscribe to an existing notification topic instead of creating one (GCS and S3). | 2.9 |
rescan=DURATION |
Relist the bucket on a schedule, for example rescan=1h. rescan=never turns it off. |
2.9 |
read_replication=N |
Number of cache copies of each chunk (scale-out). | 2.8 or earlier |
write_replication=N |
Number of write-tier copies of unuploaded data (scale-out). | 2.8 or earlier |
The readonly attribute has existed for several releases, but the guarantee
that BlixtFS never writes to the bucket in any way is complete from 2.9. See
Read-only buckets.
BlixtFS refuses to start if a bucket spec contains an attribute it doesn’t recognise, so a typo can never silently leave a bucket writable or unindexed. (Releases before 2.9 logged a warning and ignored it.)
Buckets in the configuration file
In YAML, each bucket is an entry under its provider, and attributes become keys:
cloud:
aws:
enabled: true
buckets:
- bucket: s3://my-bucket
region: us-east-1
- bucket: s3://shared-data
region: eu-west-1
read_only: true
rescan: 1h
topic: arn:aws:sns:eu-west-1:123456789012:owner-events
gcp:
enabled: true
project: my-project
buckets:
- bucket: gs://training-data
lazy: true
- bucket: gs://public-dataset
anonymous: true
S3 bucket entries need the bucket’s actual region. To serve every bucket the
credentials can see, set all_buckets: true on the provider instead of listing
them.
Environment variables
Buckets can also be listed in BUCKETS, or per provider in AWS_BUCKETS,
GCP_BUCKETS, AZURE_BUCKETS, ORACLE_BUCKETS, MINIO_BUCKETS,
CLOUDFLARE_BUCKETS and COREWEAVE_BUCKETS. Separate several buckets with
spaces. Attributes work the same way as on the command line.
Creating buckets from the filesystem
Where your credentials allow it, creating a directory directly under a cloud’s directory in the mount root creates a new bucket:
mkdir /mnt/blixt/gcp/new-bucket
This doesn’t work for Cloudflare R2 or CoreWeave: create those buckets in the provider’s console.
Indexing large buckets
The first time BlixtFS serves a bucket, it lists every object into the index before serving it. For buckets with millions of objects that can take a long time. Two ways to start faster:
- Mark the bucket
lazy. It is served immediately. Each directory is listed the first time someone opens it, and a background crawler indexes the rest. - Set
indexing.lazy: trueto make every bucket lazy.
Keeping up with outside changes
By default BlixtFS subscribes to the bucket’s change notifications (see
Change notifications). For buckets that
can’t deliver them, or where you prefer not to, use nopubsub together with
rescan=. The periodic consistency check
also catches anything missed.