Read-only buckets
Serve a bucket that BlixtFS must never change, such as a public dataset or a bucket you only have read access to.
BlixtFS can serve a bucket it must not change: someone else’s public dataset, a colleague’s bucket you have been given read access to, or your own data that a particular deployment has no business writing to.
The guarantee
A read-only bucket is never written by BlixtFS. Not its objects, not their metadata, not directory markers, not scratch state, not the pubsub object log.
That is stronger than “writes from clients are refused”, and the difference matters. Serving a bucket normally means BlixtFS writes to it as a side effect of reading it: it stamps its own metadata onto objects as it indexes them, and it writes a zero-byte marker object for every directory it discovers. On a read-only bucket it does neither. Everything BlixtFS needs to know lives in its own database instead.
The one exception is change notifications, and only when you ask for them explicitly; see Keeping the index current.
When to use it
Three situations, which need different settings.
You must not change the data, but your credentials could. This is the common one: a production bucket mounted for analysis, or a colleague’s bucket. Declare it read-only and BlixtFS will not touch it.
bfs s3://shared-data:ro
Your credentials cannot write. Read-only credentials, or — the common
case — a bucket in another project or account that has granted your
principal read access. Declaring :ro is still the clearest thing to do, but
BlixtFS also works it out for itself; see
Automatic detection and
Buckets in another project or account.
The data is public and you have no credentials at all. A public dataset.
bfs s3://noaa-ghcn-pds:anonymous
Declaring it
On the command line, or in a *_BUCKET/*_BUCKETS environment variable,
attributes follow the bucket name after a colon:
| Attribute | Meaning |
|---|---|
readonly, ro |
Never write to this bucket. |
anonymous |
Read with no credentials at all. Implies readonly. |
rescan=DURATION |
How often to re-list the bucket. rescan=never means never. |
pubsub, nopubsub |
Turn change notifications on or off for this bucket. |
lazy |
Serve before the index is complete. |
bfs gs://gcp-public-data-landsat:anonymous,rescan=never
An unknown attribute stops BlixtFS from starting rather than being ignored, so a typo cannot quietly leave a bucket writable.
In YAML the same things are keys, and a :suffix is refused:
cloud:
aws:
buckets:
- bucket: s3://shared-data
project: my-aws-account
region: us-east-1
read_only: true
- bucket: s3://noaa-ghcn-pds
project: my-aws-account
region: us-east-1
anonymous: true
rescan: never
To serve everything read-only, regardless of bucket, use --file_readonly.
Automatic detection
At startup BlixtFS checks whether it can actually write to each bucket that is
not already declared read-only. It creates one small object under its own
.__bfs__/ prefix and deletes it again. A bucket declared :ro is never
probed, so a bucket you said not to touch is not touched.
It checks each bucket, not the credentials. A credential that can write in your own project says nothing about a bucket in someone else’s, which is exactly the case that must not be got wrong.
If the probe fails with a permission error, the bucket is served read-only and a warning names it:
Bucket s3://shared-data: the configured credentials cannot write to it
(ReadOnly), so it is served read-only. Add ":ro" to the bucket to say so
deliberately and silence this.
--cloud_write_probe=falseturns the probe off. The cost is that credentials that cannot write are only discovered when a write fails.--strict_bucket_accessmakes the situation fatal instead: BlixtFS refuses to start rather than serve a bucket read-only that you did not say was read-only.
Buckets in another project or account
A bucket someone else owns, shared with your principal for reading, is different from a public one: you still authenticate as yourself, and the grant lives on their bucket.
Whether BlixtFS can serve it comes down to whether the bucket can be named independently of the account it belongs to:
| Cloud | Another project / account? |
|---|---|
| AWS S3 | Yes. Bucket names are global and the client is chosen by region, so a bucket shared by bucket policy or an account-to-account grant works. Give the region in the config entry: BlixtFS does not look it up across accounts. |
| Google Cloud Storage | Yes. Bucket names are global and object access never goes through your project. --gcp_project stays your project; name the other project’s bucket in the bucket list. |
| Azure Blob Storage | No. Every container is addressed through the one configured storage_account, so a container in a different account cannot be named. Point the deployment at that account instead. |
| Oracle | No. Buckets are addressed through the one configured namespace. |
| Cloudflare R2 | No. Buckets are addressed through the one account’s endpoint. |
| MinIO | No. One endpoint per deployment. |
| CoreWeave | No. One endpoint per deployment. |
On AWS and GCS, note also:
- Discovery does not cross the boundary.
all_bucketslists only your own account or project, so a shared bucket must be named explicitly. - Reading the bucket’s own attributes is a separate permission from reading its objects, and a cross-account grant often omits it. BlixtFS falls back to listing a single object, so a grant of just read access on the objects is enough.
- Requester-pays buckets are not supported. BlixtFS sends no billing project or request-payer header, so a requester-pays bucket is refused by the cloud.
- Live notifications need the owner’s cooperation, since setting them up writes to their bucket’s configuration. A shared bucket is normally served with rescans instead; see Keeping the index current.
- On AWS, objects encrypted with a KMS key in the other account also need
kms:Decrypton that key.
Public buckets
anonymous sends no credentials at all, which is not the same as sending
credentials that happen to have no rights: on most clouds a signed request from
an identity with no business in that bucket is rejected, so a public bucket
can only be read unsigned.
| Cloud | Anonymous support |
|---|---|
| AWS S3 | Yes. |
| Google Cloud Storage | Yes. |
| Azure Blob Storage | Yes, if the container’s public access level is container. “Blob” level allows reading a blob by name but not listing, and BlixtFS must list to index. |
| Oracle, Cloudflare R2, CoreWeave | Not supported. R2 serves public content over r2.dev and custom domains rather than the S3 API. |
A deployment whose buckets are all anonymous needs no credentials for that cloud at all, and the account-level startup checks are skipped.
Note that reading a bucket’s own attributes (storage.buckets.get on GCS) is a
separate permission from reading its objects, and public datasets usually grant
only the latter. BlixtFS falls back to listing a single object to decide whether
it can serve the bucket.
Directories and permissions
A bucket has no directories, only keys. Normally BlixtFS papers over that by writing a zero-byte marker object for each directory. On a read-only bucket it cannot, so directories exist only in BlixtFS’s index:
- A directory is implied by the keys underneath it.
a/b/filegives youa/anda/b/, and both behave as ordinary directories through every protocol. - When the last object under an implied directory is deleted from outside BlixtFS, the directory stops existing in the bucket. A consistency check notices and removes it from the index too, within two sweeps.
- A plain
folder/marker written by another tool is treated as the directory it is.
Objects in a read-only bucket carry none of BlixtFS’s metadata, so there is no
stored mode, owner or group. Files are shown with the server’s default mode and
ownership, with the write bits cleared: files as 0444, directories as
0555. ls -l therefore tells the truth, and a well-behaved program checks
the mode rather than discovering EROFS halfway through a job.
The defaults themselves come from the usual settings
(filesystem.default_file_mode, default_user_id, default_group_id). Only
what a client is shown changes: the stored mode is left alone, so making the
bucket writable later restores the ordinary modes with no migration.
Keeping the index current
If nothing else ever changes the bucket, nothing needs to happen. If something does, BlixtFS has to hear about it one of two ways.
Live notifications. A read-only bucket gets them like any other bucket:
read_only says how the bucket should be served, not what your credentials
can do. Turn them off for a bucket whose credentials cannot configure them:
bfs gs://shared-data:ro,nopubsub
Setting notifications up writes to the bucket’s configuration, not to its objects, so it does not break the read-only promise — but it does need rights an operator of someone else’s bucket often lacks, and the attempt is reported if it fails. The bucket keeps serving either way.
An anonymous bucket is the exception: with no credentials at all there is
nothing to create a topic, queue or rule with, so notifications are off and
cannot be turned on.
On AWS and MinIO, configuring notifications on a read-only bucket is skipped entirely, and the skip is deliberate: their notification API replaces the bucket’s entire notification configuration rather than adding to it, so configuring it would silently delete the rules the bucket’s owner set up. BlixtFS reports the skip and falls back to rescans.
Subscribing to the owner’s topic. If the bucket’s owner already publishes its events, BlixtFS can subscribe to that stream instead of configuring anything on the bucket. The subscription belongs to you, in your own account, and the bucket is not touched — so this works on a bucket you may only read, and it is the only way to get live updates from one on S3.
bfs s3://shared-data:ro,pubsub,topic=arn:aws:sns:us-east-1:210987654321/their-events
bfs gs://shared-data:ro,pubsub,topic=projects/their-project/topics/their-events
| Cloud | Topic to name | The owner must grant |
|---|---|---|
| AWS S3 | An SNS topic ARN | sns:Subscribe on their topic, for your principal. |
| Google Cloud Storage | projects/PROJECT/topics/TOPIC |
pubsub.topics.attachSubscription on their topic. |
Naming a topic on any other cloud fails at startup rather than being ignored: a bucket quietly waiting for its next rescan looks exactly like one that is up to date, until it is not. Those clouds use rescans.
Rescans. Otherwise, changes are picked up the next time BlixtFS re-lists the bucket. The interval is per bucket:
bfs s3://shared-data:ro,rescan=15m
Without rescan=, the server-wide indexing.fsck_sleep_seconds (one hour by
default) applies. For a dataset that never changes, rescan=never skips the
re-listing entirely — worth doing, since listing a hundred-million-object
bucket every hour costs real money to learn nothing.
You can always trigger a check by hand:
setfattr -n bfs.fsck_request -v 1 /mnt/my-bucket/some/dir
To see which of these applies to a bucket, read the bfs.freshness attribute:
$ getfattr -n bfs.freshness /mnt/s3/shared-data
bfs.freshness="rescan every 15m0s: changes made outside BlixtFS are picked up
by the next check, not live"
What clients see
Every write is refused, whichever protocol the client uses: creating files
and directories, writing, opening for write, deleting, renaming, chmod,
chown, truncate and setting extended attributes.
The error a client sees depends on the protocol, because most refusals never
reach BlixtFS. Files are presented with their write bits cleared (0444, and
0555 for directories), so the client’s own kernel, or the protocol server in
front of BlixtFS, usually rejects the write first:
| Protocol | What the client sees |
|---|---|
| FUSE | EROFS (“read-only file system”), and the mount itself reports ro |
| NFS | EACCES for opening to write, truncating and setting user.* attributes; EPERM for chmod, chown and utimes |
| SMB | EACCES (access denied) for everything. chown is EINVAL |
| SFTP | EPERM. SFTP has no read-only status code, so the original error is lost. A failed mkdir may appear as EEXIST and a failed rmdir as ENOTEMPTY |
| WebDAV | The client caches the write locally and fails when it uploads, usually with EINVAL. The file disappears from the view once the client discards its copy |
Owning the file doesn’t help: the presented mode has no write bits for anyone.
Don’t rely on a particular errno to detect a read-only bucket. If your application needs to know, read the bucket’s attributes, or check whether the directory is writable before you start.
Consistency checks on a read-only bucket
fsck still runs, and still repairs BlixtFS’s own records: missing rows,
stale entries, wrong link counts, directories the bucket no longer implies.
Repairs that exist only to change the bucket — writing an object’s metadata back, deleting an object the index thinks is gone — are skipped, and say so:
PermissionMismatch: skipped. The bucket is read-only, and repairing this
would write to it.
This holds regardless of --indexing_fix_issues. A read-only bucket is not
written even when every repair is enabled.
Permissions needed
For a read-only bucket, with no notifications:
| Cloud | Permissions |
|---|---|
| AWS S3 | s3:ListBucket, s3:GetObject on the bucket. |
| Google Cloud Storage | roles/storage.objectViewer on the bucket. |
| Azure Blob Storage | Storage Blob Data Reader on the container. |
| Oracle | OBJECT_READ on the bucket. |
| Cloudflare R2 | An Object Read token scoped to the bucket. |
| CoreWeave | Read access to the bucket. Notifications are never available. |
Adding pubsub to a read-only bucket needs the same notification permissions
as a writable one, which are listed in the deployment guide.