Monitoring
Collect BlixtFS metrics with Prometheus, view them in Grafana, and find the logs.
Metrics endpoints and the built-in dashboards require a license that includes Monitoring (Enterprise or High Performance).
Built-in Prometheus and Grafana
The blixtfs/enterprise image runs Prometheus and Grafana inside the
container, already configured to scrape every BlixtFS component.
The monitoring dashboards are available on port 3000:
docker run ... -p 3000:3000 blixtfs/enterprise:latest ...
Open Grafana at http://<host>:3000. It comes with these dashboards:
| Dashboard | Shows |
|---|---|
| Overview | Health and throughput of the whole deployment |
| Cloud | Requests, bytes, errors and retries per cloud provider |
| Database | Query rates, latency and connection use |
| NFS | NFS operations, latency and errors |
| SMB | SMB operations, latency and errors |
| Client | Operations as seen by the file-protocol clients |
In Kubernetes deployments, the monitoring is accessed as a service.
Scraping with your own Prometheus instance
The Blixt Prometheus instance runs on port 9090.
Each component serves Prometheus metrics at /metrics on its own port:
| Component | Port |
|---|---|
| Config server | 4721 |
| File server | 4722 |
| Cache server | 4723 |
| Replication | 4724 |
| Write server | 4725 |
| Updater | 4726 |
| Protocol clients | 4727 |
| Indexing | 4728 |
| Notifier | 4732 |
| NFS gateway | 9587 |
| SMB gateway | 9922 |
In Kubernetes, set serviceMonitor.enabled=true to create a Prometheus
Operator ServiceMonitor.
Some BlixtFS metrics are native histograms. Enable
scrape_native_histograms: true in your Prometheus scrape configuration to
collect them.
Restrict who can reach the metrics endpoints with monitoring.allowed_cidrs.
Metrics worth alerting on
- Cloud request errors and retries, by provider: credential or quota problems
- Database connection use close to the limit: see Troubleshooting
- Consistency checks finding issues repeatedly:
bfs_fsck_issues_total - Local disk use close to
cache.stop_write_percent(95% by default), at which point BlixtFS stops accepting writes
Logs
Each server writes logs to /data/log (in the container’s data directory) and
to standard output, so docker logs and kubectl logs show them too.
| Setting | Default | Meaning |
|---|---|---|
--log_json |
off | Write JSON log lines for log collectors |
--log_fsck |
off | Log consistency-check details |
--log_all |
off | Log everything (very verbose; for troubleshooting only) |
cleaner.log_max_size_mb |
100 | Rotate log files at this size |
cleaner.log_max_files |
5 | Keep this many rotated files |
Other logging.* settings turn on detail for one area each, such as cloud,
database, nfs or pubsub. Only configuration logging is on by default.