Data model
How BlixtFS stores files as objects, and what it keeps in the metadata index.
Files are ordinary objects
Each regular file is stored as one object, under the file’s path relative to
the bucket root. The filesystem has one directory per cloud provider, and one
directory per bucket inside it. A file at /mnt/blixt/aws/my-bucket/reports/2026/q3.pdf
is the object reports/2026/q3.pdf in the S3 bucket my-bucket. The object’s contents are exactly
the file’s bytes: no chunking, no encryption layer and no custom headers are
added to the data.
Directories are stored as zero-byte directory markers, objects whose key
ends in /, such as reports/2026/. A directory that exists only because
objects are stored under it (for example, one created by another tool) is
still shown. For Read-only buckets, directory objects are only stored in
the local metadata index, not in the bucket.
BlixtFS stores a few metadata items on each object: file type, permission, etc. Storing these key attributes in the object store re-inforces the object store as the “single source of truth”. A BlixtFS metadata index can be re-built from any object store. The metadata entries are standard object metadata, prefixed with “bfs_”. The entries don’t affect other tools reading the object. Object metadata is never updated in Read-only buckets.
Because the layout is plain, you can:
- read and write the bucket with your provider’s tools while BlixtFS is serving it
- point BlixtFS at a bucket that already holds data, and see it immediately
- stop using BlixtFS without migrating anything
The metadata index
The metadata index is stored in-memory and persisted to disk.
It records every file and directory: name, parent, type, size and other attributes,
and which version of the object the file corresponds to.
Directory listings and lookups are answered from the index, which is what makes
metadata-heavy workloads such as ls -R, find and builds fast.
The index also holds BlixtFS’s work queues: pending uploads, pending change notifications and pending consistency checks. Pending work survives a restart.
Object versions
Every version of an object has a version: a number that increases each time the object is replaced. Google Cloud Storage provides one natively. For the other providers BlixtFS derives one from object attributes.
Versions let BlixtFS tell old information from new. A cached chunk belongs to one version, so a newer object can never be served from stale cache. A change notification describing an older version than the index already holds is ignored.
Lazy indexing
For buckets with millions of objects, a full listing before serving can take a
long time. A bucket marked lazy is served straight away. Each directory is
listed from the bucket the first time someone opens it, and a background
crawler indexes the rest.