Performance tuning
The settings and deployment choices that most affect BlixtFS throughput and latency.
BlixtFS’s defaults suit most deployments. When you tune, measure before and after each change with your own workload. The dashboards show where time goes.
Put the servers close to the bucket
The biggest single factor is the network path to the object store. Run BlixtFS in the same cloud region as the bucket. Cross-region and on-premises-to-cloud deployments work, but every uncached read pays the round trip.
Give it fast local disk
The chunk cache and write log live in /data. Use a dedicated SSD or NVMe
volume, not the container’s root filesystem or a network disk, and make it
large enough to hold the working set. A cache too small for the working set
reads the same data from the bucket again and again.
| Setting | Default | Effect |
|---|---|---|
cache.stop_write_percent |
95 | Disk use at which new writes are refused until uploads free space |
cleaner.target_percent |
90 | Disk use the cache cleaner shrinks back to |
cache.enable_compression |
true | Compress cached chunks; turn off for data that doesn’t compress, such as video |
Metadata
Cold lookups and listings are answered from the database, so database latency is file-operation latency.
| Setting | Default | Effect |
|---|---|---|
database.cache_size_mb |
512 | In-memory metadata cache per file server |
database.preload_files |
10000 | Files loaded into the cache at startup |
database.connections |
25 | Database connections per server |
For an external database, keep it in the same zone as the file servers.
Parallelism
| Setting | Default | Effect |
|---|---|---|
general.workers |
10 | Worker threads per bucket |
indexing.workers |
32 | Parallel listing during indexing and consistency checks |
nfs.threads |
512 | NFS server worker threads |
Large buckets
- Mark buckets with millions of objects
lazy, so they are served before indexing finishes. - Consistency checks list the bucket. On very large buckets, lengthen
indexing.fsck_sleep_secondsif checks run back to back, and rely on change notifications between them.
Scale out
When one server is saturated, a High Performance deployment spreads metadata across several file servers and data across cache and write tiers. See Scale-out.