Blixt Documentation v2.9

What a consistency check does

A consistency check (fsck) lists part of the bucket and compares it with the metadata index. It finds, and usually repairs:

  • objects in the bucket that the index doesn’t know about (added outside BlixtFS)
  • files in the index whose object is gone (deleted outside BlixtFS)
  • files whose size or version differs from the object
  • missing directories, broken link counts and similar inconsistencies

Most of the time change notifications keep the index current and the check finds nothing. It is the safety net for events that were missed, and the main way buckets without notifications stay current.

When checks run

  • Periodically, over every bucket. The default interval is one hour (indexing.fsck_sleep_seconds: 3600).
  • On request, for a single directory or a whole bucket (see below).

Checks run one at a time from a queue, because each one lists the bucket and costs cloud requests.

Requesting a check

Set the bfs.fsck_request extended attribute on the directory you want checked:

# Check a directory and everything below it
setfattr -n bfs.fsck_request -v 1 /mnt/blixt/aws/my-bucket/reports

# Check a single directory level
setfattr -n bfs.fsck_request -v shallow /mnt/blixt/aws/my-bucket/reports

# Check a whole bucket
setfattr -n bfs.fsck_request -v 1 /mnt/blixt/aws/my-bucket

A request is queued, not run immediately. You can only request a check of a path you can reach through your mount. Requested checks never write to the bucket. They only bring the index up to date.

Asking twice is safe. A request that a queued check already covers is folded into it, and a request made during the cooldown below is queued to start when the cooldown ends, rather than being refused.

To stop users requesting checks, set indexing.user_fsck: false. To limit how often the same directory can be checked, set indexing.manual_fsck_cooldown_seconds (default 300).

Write requests to bfs.fsck_request, and read the status from bfs.fsck. Writing to bfs.fsck also works and means the same thing, but NFS clients serve back the value they wrote, so a script that writes bfs.fsck and reads it straight back sees its own request instead of the status.

Reading the status

getfattr -n bfs.fsck /mnt/blixt/aws/my-bucket/reports
getfattr -n bfs.fsck_state /mnt/blixt/aws/my-bucket/reports

A directory reports the most specific check that covers it: one of the directory itself, of a parent, or of the whole bucket.

Attribute Value
bfs.fsck A one-line summary, or no recent check.
bfs.fsck_state pending, running, waiting, done or failed.
bfs.fsck_duration_seconds Time elapsed so far, or the total once finished.
bfs.fsck_issues Faults found and fixed, as found/fixed.
bfs.fsck_changes What the check altered, as added/updated/removed.
bfs.fsck_not_before When a queued check may start, while it waits out the cooldown.
bfs.fsck_error Why the check failed. Present only when failed.

Read bfs.fsck_changes as well as bfs.fsck_issues. They answer different questions: a check that indexed three objects someone uploaded directly found no faults (the index wasn’t wrong, it was behind), so it reports no issues but three changes. The issue counts alone can’t tell you whether your view of a directory was stale.

bfs.freshness tells you how a bucket learns about outside changes at all: live notifications, a periodic rescan, or neither. It is the first thing to check when a change took an hour to appear.

Repair policy

Repairs are split by what they touch:

Setting Default Allows
indexing.fix_database_issues true Repairs to BlixtFS’s own index, including reconciling deletions made outside BlixtFS.
indexing.fix_grace_minutes 60 Files changed within this many minutes are left alone, since an upload in progress can look like a deletion.
indexing.fix_issues false Repairs that write to or delete from the bucket, such as creating missing directory markers or removing orphaned objects.

Leave fix_issues off unless you’ve reviewed what a check reports. It acts on your data based on the index the check is verifying. In scale-out deployments, set these on the config server.

Some issues are never repaired automatically, because there is no safe choice: for example, a file and a directory with the same name. They are reported in the log for you to resolve.

Read-only buckets

A check never writes to a read-only bucket, whatever fix_issues says. Repairs that would need to write to the bucket are skipped.