Reads and writes
What happens between an application calling read() or write() and bytes moving to or from the bucket.
Consistency model
With a single BlixtFS deployment, both metadata and data are strongly consistent, as specified and supported by the file protocol semantics.
Across multiple (single node) BlixtFS deployments, consistency follows the semantics of the object storage. When an object has been succesfully uploaded, other BlixtFS deployments can be notified of the upload.
Enterprise and High performance Edition is recommended for workloads that require strong consistency. Please contact us for more details.
Reading
BlixtFS carefully tracks and orchestrates all data for an file. Multiple readers and writers are supported, within the limits and semantics of each file protocol. A read is served from the first layer that has the data:
- Write log(s). Data that has been written but not yet uploaded to object store.
- Read cache(s). Data read from object store, kept on local disk or replicated in read cache servers.
- Remote object in bucket. BlixtFS reads large amounts of data, possibly in parallel, to fill the read cache quickly.
Cached data carry a checksum and can be compressed. Since all data belong to one specific object version, a file changed in the bucket is never answered from stale cache.
Writing
- BlixtFS persists write data to write logs on local disk. In scale-out deployments, write logs are replicated to multiple servers.
- The write request returns as soon as the data is durable in the log. No data is sent to the object store yet.
- When the file is closed, or explicitly committed, BlixtFS schedules an upload. The upload process is persisted, so it survives server restarts, etc.
- A background worker uploads the complete file as one object (using multipart upload as appropriate), then records the new object version in the index.
Other file clients see the new contents as soon as the write returns, because reads check the write log first.
Why whole-object uploads
Object stores replace objects, they don’t patch them. Uploading on close means one upload per file rather than one per write, and the object in the bucket is always a complete, consistent file. Uploads are throttled per object. BlixtFS optimises object uploads to stay below the upload rate limits of cloud providers.
Truncate and small changes
Changing a single byte, truncating a file or appending to it produces a new version of the whole object. For very large files that are modified in place often, keep this in mind when you size network bandwidth.
Local disk management
Write logs and the read caches are stored on local disk. It is a requirement to use a separate volume to the container’s root filesystem to avoid filling it up. See Configuration.
We recommend fast dedicated storage, sized to your working set plus the largest expected burst of un-uploaded writes. BlixtFS stops accepting writes before the disk fills completely. See Performance tuning.