Storage Architected for AI™
- Introduction to Quobyte
- The Quobyte Design Principles
- The Quobyte Architecture
- The Quobyte File System
- State-of-the-art Security for Storage
- Quobyte's Storage-as-a-Service Design
- The File Query Engine
- Data Management Services
- Non-disruptive Operations and Automated Maintenance
- Automation and Monitoring
- Integrations
Quobyte: A Fault-tolerant, High-performance Parallel File System Designed for Simplified Operations of Large-scale Storage Infrastructures
Quobyte is the most technologically advanced parallel distributed file system and runs on any commodity Linux x86 or ARM server. It turns a cluster of servers into a linear scalable, high-performance, and reliable storage system. Quobyte combines the simplicity of software storage with the exceptional performance of a parallel file system.
While Quobyte is a parallel distributed POSIX compatible file system at the core, it is a complete data platform combining ultra-high performance NVMe and cost-effective HDD storage with advanced data management, analytics, multi-tenancy, and security features. A Quobyte cluster can serve the most demanding applications, such as AI/machine learning, video rendering, or MPI applications, as well as traditional enterprise workloads like databases or home directories.
This whitepaper provides an in-depth description of Quobyte's design principles, how they translate to unique benefits and enable scalable operations, as well as a detailed discussion of Quobyte's technology "under the hood". If your focus is on installing and managing a Quobyte cluster, reading this white paper is not a requirement.
We start by describing Quobyte's architecture and its benefits, with a section on the design principles of Quobyte. This section illustrates how the composition of a unique approach to fault tolerance and scalability results in a product that is easy to manage at scale while delivering maximum performance. The second part dives into the hardware details, the role of individual Quobyte services, and how they interact and communicate with each other and the client. This is followed by an in-depth discussion of the fault-tolerance and data protection mechanisms used in Quobyte. Finally, we show how everything combines to build a highly performant and flexible file system with a broad range of Data Services and features.
The Quobyte Design Principles
Quobyte was designed with the same principles and technologies that fuel the massive infrastructures of hyperscalers like Google: A distributed system with a shared-nothing architecture running on regular commodity servers connected by a standard IP network.
Not only does this type of architecture provide massive scalability, but it also delivers the simplicity of operations of the software and hardware that is required to manage clusters in the exabyte scale with small teams. These design principles make Quobyte one of the most advanced storage systems in the market, both in terms of technology and unique, scalable operations.
Fault tolerance in Software for Unparalleled Reliability and Operational Simplicity
Quobyte is built with the assumption that "anything can fail at any time": Hardware failure and unavailability are taken care of by the software automatically. This includes issues like failed drives, crashed servers, unavailable racks, network partitions, packet loss, etc.
As a 100% software solution, Quobyte's fault tolerance is entirely implemented in software and does not rely on brittle and complex hardware redundancy like dual controllers, heartbeat cables, or RAID. In addition to not needing specialized hardware, the more important benefit is that Quobyte provides fault tolerance across machines that can be physically distributed - for example, across racks, clusters, or even data centers. This means that a Quobyte cluster can tolerate the loss of a full rack of storage servers, a whole cluster, or even an entire data center without any downtime or data loss.
An often overlooked – but far more important – benefit of Quobyte's fault tolerance in software is that it has been designed so that unavailable drives or servers are part of normal operations and do not cause any disruption to users and applications. With this unparalleled reliability and resiliency, Quobyte can offer truly non-disruptive operations and exceptional uptime. Maintenance windows are a thing of the past with Quobyte. The following operations are examples of what administrators and field techs in the data center can do without disrupting users:
- Rolling updates
These are done automatically by Quobyte or with simple Ansible scripts. Unlike solutions that sell container-based updates as "almost non-disruptive," Quobyte rolling updates mean zero interruption or downtime. - Hardware maintenance at will
Admins or field techs can restart a Quobyte server (or even a full rack or larger failure domain), shut it down for a few hours for hardware fixes, re-cabling or relocation, or restart Quobyte services. A simple power down is all that's needed. - Simplified Maintenance Processes
With Quobyte, hardware operators and storage admins do not need to coordinate their work with each other. Hardware can be replaced, rebooted, moved, recabled etc. at any time without the storage admins having to do anything. This decoupling of processes greatly improves productivity and is key to operating large scale storage infrastructure with hyperscaler efficiency and cost effectiveness. - Add or remove servers and/or disks
Adding servers takes seconds, and the additional capacity and performance are instantly available. Servers can also be removed at any time with one click. Quobyte moves data to other nodes transparently, enabling administrators even to do a full hardware refresh without any downtime. - "Letting the hardware rot"
In a Quobyte cluster, failed disks or a failed server are non-events. The software takes care of recreating redundancy on other machines. As long as there is enough spare capacity in the cluster, admins can simply wait until enough hardware has failed and then send it back to the vendor on a monthly swap. Rebuilds are fully declustered, meaning that lost redundancy is re-created on all disks and nodes of the cluster. This means that rebuilding traffic is just background noise and the healing capabilities scale with the size of the cluster.
The type of software-based fault tolerance Quobyte uses enables hyperscalers to run the world's largest data infrastructures with small teams and unparalleled cost efficiency. By not relying on hardware redundancy, Quobyte takes the "human" out of the loop even when failures happen. In addition to the reduced administrative burden, Quobyte enables low-touch storage clusters e.g., for edge computing.
No Specialized Hardware: Commodity Servers for Maximum Simplicity and Lowest TCO
Quobyte has been designed from the ground up to run on any x86 or ARM server without any custom hardware. This required a completely different approach than what has been done in storage until today: Quobyte is built around the assumption that the platform – including hardware, firmware and sometimes even the kernel – "cannot be trusted" and that the software needs to take care of extensive monitoring and redundancy, e.g., with end-to-end checksums, response time monitoring in software and automatic ongoing health checks for TCP connections.
Unlike traditional storage, Quobyte enables the use of the simplest hardware configurations: an x86 or ARM processor, RAM, a network card, NVMe devices, or an HBA with SATA/SAS drives. Simpler hardware not only simplifies day-to-day operations for admins and field techs but also reduces the cost and TCO of a Quobyte storage cluster.
There is no need for custom-built controllers, extra networks, cables, or switches, which are found in most traditional storage based on the dated concept of controllers and disk shelves or SAN backends (FC, NVMe-over-Fabric). In addition, administrators don't have to learn new skills to operate a Quobyte cluster; it's just standard Linux servers.
As a hardware-agnostic storage solution, customers can pick the hardware that fits their needs:
- From the vendor you are familiar with or get the best discounts
- Mix hardware from multiple vendors in the same Quobyte cluster to minimize dependency on a single vendor
- Standardize on a few server configs for all use cases in your data center to streamline operations and reduce operational cost
- Pick servers that suit your use case and requirements, like 4-node-in-2U servers for the edge with NVMe or high-density HDD servers for cold storage
- Mix NVMe and HDD for high-performance, low-latency and cost-effective storage in one cluster
- Re-use existing servers with Quobyte to protect your investment from the vendor you are familiar with or get the best discounts
- Combine different configurations, like all-flash and HDD servers, and different models in the same cluster, e.g., newer CPUs or larger drives. Quobyte can handle heterogeneous hardware configurations, allowing you to expand your cluster down the road with newer hardware from the vendor you are familiar with or get the best discounts
- Deploy anywhere®: Your own data center, co-locations, public cloud VMs, or bare metal machines, or create hybrid Quobyte clusters that combine on-prem and the public clouds.
- Use the provisioning, management and monitoring tools admins are already used to
For simplicity, Quobyte provides a list of validated recommended hardware platforms with high-performance all-flash servers and high-density HDD configurations from a number of OEMs.
Simplicity Scales: Quobyte Makes Operations at Exabyte Scale Easy
Scalability shouldn't be viewed only from the technical point of "Can the system scale?". The operational scalability is equally important: Can a small team manage a growing system, or does the scalability come at the cost of more manpower? In short, complexity gets exponentially worse at scale - simplicity scales. This is why we made simplicity a key concern when designing Quobyte: Software that is easy to install and operate and doesn't require more human attention as you scale to exabytes:
Quobyte is standard user-land software that any Linux admin knows how to install and operate:
- Quobyte clients do not require kernel modules. Forget the hassle of compiling modules by hand and painful kernel updates.
- Quobyte services are regular systemd services running as non-privileged users
- Quobyte clients and servers run on standard, unmodified Linux distributions. Quobyte provides packages for a number of distros, including RHEL, RockyLinux, Debian, Ubuntu, Oracle Linux, and SLES (check the latest list here)
- Services and clients communicate over standard IP networks and can leverage RDMA where available — no need for complex setups, multicast, or specialized networking drivers.
- As a fully integrated product, Quobyte doesn't require installing and maintaining any extra services or databases.
As a result, Quobyte installations only take a few minutes with the automated installer or via simple Ansible scripts for more control.
Another important concept contributing to the overall simplicity of running Quobyte clusters is the "anything can fail" approach to fault tolerance. Together with extensive monitoring and a highly automated system, all typical failures are handled automatically and do not require attention from an administrator. A Quobyte cluster keeps itself healthy and optimized as long as there is enough capacity to recreate redundancy.
In addition, Quobyte forgives operator errors, which are a frequent source of unplanned downtime. Did you accidentally switch off a server? Pull the wrong networking cable? Forget which drive goes back into which slot? Put a drive in the wrong server? Misconfigured tiering or data placement? The system automatically handles all of these situations.
Empowering administrators is crucial for successfully running large-scale systems, and understanding and anticipating the behavior of a system you operate is key. That's why we embrace the concept of an "open system" where administrators can "look under the hood" instead of working with a black box. From monitoring of internals, human-readable RPC traces to understandable on-disk formats.
Just as importantly, Quobyte is built around deterministic algorithms, like explicit data placement based on policies and does not use unpredictable hashing algorithms. Administrators can rely on predictable behavior and easily understand and predict what will happen when they perform routine tasks like removing a server, changing placement configurations, or adding nodes, etc.
Flexibility: Storage Evolution at Your Fingertips
The advanced fault tolerance in Quobyte yields yet another advantage: Total flexibility.
The ability to transparently move files, i.e., without any downtime or interruption even on the file being moved, allows administrators to add and remove drives and servers and move files and volumes arbitrarily in the cluster without downtime. With this foundation, managers, and administrators can change and adapt a Quobyte cluster to new requirements in record time:
- Add more servers or drives when more performance or capacity is required. This includes the option to scale the NVMe or HDD tier independently.
- Refresh the hardware: Add new servers and remove the old ones. Since there aren't any forced rebalances, this can be done over time with zero haste or impact on users.
- Add dedicated metadata servers: Do you need fast metadata? Adding dedicated servers for metadata is just a matter of a few clicks.
- Change placement configuration, such as shifting resources to different volumes, users, or applications or moving applications and users to dedicated hardware for perfect performance isolation.
- Add temporary cloud resources: When you run out of resources, you can always add temporary VMs or bare metal machines and remove them from the cluster when the peak usage is over.
Scale With Your Ambition™: The Performance of True Scale-out
Quobyte's architecture is built around a decentralized approach where decisions and fault tolerance are implemented locally and without centralized components where possible - sometimes referred to as horizontal scalability. The result is an architecture with linear performance scaling: Doubling the number of servers will double the performance and capacity without diminishing returns at scale. Our design allows a single Quobyte cluster to scale well beyond 1000s of servers and 100,000s of clients concurrently accessing the storage.
In addition to the scalable fault-tolerance and architecture, we have also avoided typical scalability bottlenecks in our architecture – both for theoretical as well as practical scalability:
- No protocol bottlenecks like NFS gateways that severely limit practical scalability due to problems like cache consistency coordination, which limits the number of feasible gateways, or lack of load balancing, which causes performance bottlenecks.
- Explicit policy-based data placement instead of consistent hashing, which severely limits practical cluster sizes and changes to the cluster due to the forced rebalances.
- Centralized reference counting, which is required for deduplication.
- Block allocation on a local level, which doesn't involve other clients or Metadata Services during file appends.
- A single-layer architecture that doesn't suffer from the complexity of multi-layered architectures that have to scale both frontend and NVMe-over-Fabric/FibreChannel backends. A Quobyte cluster can simply be scaled by adding more servers and/or drives — no need to adjust gateway nodes, backend switches, or disk shelves.
Another benefit of linear performance scaling is that Quobyte clusters can achieve exceptional and "unlimited" performance by aggregating many cost-effective mid-range servers. This avoids the high cost of scale-up controllers or gateway nodes found in many legacy storage products.
Quobyte's Security DNA
Since legacy protocols like NFS don't restrict Quobyte, we were able to integrate a range of innovative security features end-to-end.
On the clients and servers, Quobyte avoids the security risk of kernel modules and runs as an unprivileged user. The cluster has no dependency on external or cloud-based services and can easily run in an air-gapped data center.
End-to-end data encryption means that data is never decrypted until a Quobyte client delivers it to the operating system. This yields far stronger protection than NFS-based solutions, where gateway nodes process data in cleartext. In addition to data encryption, Quobyte also offers TLS-secured communication for all communication or across specific networks.
Access to a Quobyte cluster can be restricted in a number of ways, including the traditional IP filters used in storage. For environments where networks are untrusted or for additional security, Quobyte offers X.509 certificates and user authentication tokens.
For isolation, Quobyte offers multi-tenancy with strong isolation down to the hardware level when required. Tenants can be completely isolated from each other in terms of storage, management, and LDAP/AD domains. Quobyte comes with architecture, security, and features that enable storage-as-a-service.
The Quobyte Architecture
On a fundamental level, Quobyte’s architecture is pretty simple: there are Quobyte Clients communicating with Quobyte Services to fulfill application IO requests. Both – the clients and services – run on regular x86 or ARM servers communicating over an IP network such as Ethernet. The Quobyte services run on Linux. Quobyte clients can run on Linux, WIndows, and macOS.
In this section, we’ll refine this picture, starting with the hardware layer and then walking up the stack: Quobyte Services and their roles, the Quobyte Clients, and other ways to access the storage as a user or application. Next, we show how all clients and services communicate using the high-performance Quobyte RPC protocol. We finish this section with an in-depth review of data protection and redundancy in Quobyte, which is the core of its unique fault-tolerant architecture.
Hardware Layer: Exceptional Performance and Reliability on Standard Servers
When we talk about a "Quobyte server," we mean an x86/ARM machine with NVMes and/or SATA/SAS drives attached. It provides storage for the file system and runs services. "Quobyte clients" refer to any type of x86/ARM computer—servers or workstations—that runs the native Quobyte client, which provides the file system to applications and users.
As discussed in the Quobyte Design Principles section, Quobyte recommends keeping the hardware as simple as possible for a number of reasons: Simpler hardware has fewer ways and components that can fail, fixing things in hardware (firmware updates) takes exponentially longer than software, and simpler hardware is just more cost-effective.
The components required in a production-grade Quobyte server are very straightforward:
- x86 or ARM server CPU: ideally single socket, but dual socket is supported as well Higher clock speed reduces latencies, and Quobyte services can take advantage of a high core count
- RAM: A minimum of 196 GB of RAM; Quobyte services can use any free memory for read-caching
- NVMe devices and/or an HBA for SATA/SAS drives: Quobyte does not require journaling devices, caches, or RAID controllers. SATA/SAS devices must be used in JBOD mode, where a physical drive corresponds to a separate block device.
- A single network: Quobyte doesn't require a backend network; a single IP network is sufficient. Quobyte can use RDMA via RoCE, InfiniBand, or omni-path when available. Multiple networks - when available - and routed networks are supported.
- Linux: Any standard Linux distribution that's currently supported by Quobyte
No other components, such as disk shelves, SANs, external controllers, or storage backend networks like NVMe-over-Fabric/FibreChannel, are required.
The Quobyte Services
Quobyte services are regular user-land processes that run as a non-privileged user on Linux servers. They can be managed just like any other systemd service. Three core services (Registry, Metadata, and Data) have many instances running across all Quobyte servers in a cluster.
The default model is running all three service types on each Quobyte server, as it simplifies management and scalability. However, administrators can also run each Quobyte service type on dedicated servers to reduce latencies and for further performance isolation. We speak of dedicated Quobyte Metadata servers or Quobyte Data servers in that case. Thanks to Quobyte's fault tolerance, all services can be live-migrated with their data to another server. So, changing from the model with all services on the same machine to dedicated Quobyte Metadata servers can be done later without downtime.
In addition to the core services, Quobyte has additional services like the webconsole (UI), api, NFS and S3 that provide additional functionality, but aren't required for the operations of the file system.
In addition to running Quobyte services on a physical server, it is possible to run user workloads on the same machines. This "hyper-converged" model has some advantages, e.g., when space is an issue in edge computing, in Kubernetes clusters, or when processes can benefit from reading from local replicas (on the same machine) like Hadoop. Please remember that performance isolation must be implemented carefully, and some aspects of a CPU can't be isolated properly, like memory bandwidth.
No local configuration: Quobyte services need exactly one bit of information to find and join the cluster: A DNS name pointing to all registries; see further down. No additional local configuration is required. As a result, services can easily be installed, disks can be cloned, or machines can be booted via PXE.
Drives are fully managed by Quobyte Services. Unformatted (empty) drives are automatically detected and can be formatted and added to the cluster with one click. Once a drive becomes part of a Quobyte cluster – a "Quobyte device" – it gets a unique identifier and a device label.
A Quobyte drive can be inserted into any Quobyte server. The services automatically detect and mount the drive. There is no need to configure fstab or LVM. This also means that putting a drive back in the "wrong" slot or "wrong" server is not an issue. The drive and its data become immediately available to the cluster, and Quobyte will even move data to meet any failure domain constraints! On the cloud, this translates to the ability to attach a cloud block device that contains a Quobyte device to any VM running the Quobyte services.
Each Quobyte device has one or more Device Tags for management and data placement. Each tag is a simple string. Some tags are predefined and automatically assigned to a device, e.g., "ssd" or "hdd" based on the media type. Administrators can define new tags, e.g., to create additional tiers for slower flash or archival HDDs. Together with placement rules that address these tags, this creates dynamic pools of devices. Tags and placement rules on a device can be changed at any time, the system will migrate data transparently in the background in this case.
Quobyte Registry Service
The Registry Service is the entry point to the cluster for all services and Quobyte clients. It knows the "location" (such as IP endpoints) of services, devices, and clients in the cluster (lookup service), as well as the file system volumes and configured policies.
This service is built around a replicated key-value store and is self-managing and automatically redundant. Each cluster should have five instances of the Registry Service running to ensure automatic failover in case of permanent hardware failure.
To join a cluster, a Quobyte service or client only needs to know the IP addresses of hosts where Registry Service instances are running. This list can be provided as a DNS record with multiple IPs for convenience. This is the only configuration services and clients need to work.
The information stored in the Registry Service changes fairly infrequently and is cached very effectively in the Quobyte Services and Clients. This means the Registry Service does not get many requests and can easily be co-located with other Quobyte services on the same server.
In addition to providing "location" information, the Registry Service also collects monitoring and performance data from all components for alerting, monitoring, and running tasks that keep the cluster healthy and optimized.
Quobyte Metadata Service
Quobyte splits the file system into metadata and data and has separate services responsible for handling each. Quobyte's metadata includes file name, directory tree structures, file size, permissions, ACLs, or xattrs. Most importantly, the Metadata Service stores the location of each file. This location is a fairly small list of device references that hold the actual file data (content), but not individually allocated blocks of a file. This is an important difference to systems that include per-file block allocations as metadata. By delegating the block allocation to the local devices, Quobyte achieves far superior scalability and performance. In general, Metadata Services are no longer part of the IO path once a file has been opened.
The Metadata Services store the metadata for each volume in a separate key-value store that is replicated (see quorum replication) across three Metadata Services (and hence also stored on three separate physical drives and servers). At any time, one of the services is the primary for a volume, and two others act as backups. In case of unavailability or failure, one of the backups takes over instantly. The same mechanism is also used to transparently move volumes between Metadata Services without downtime or disruptions.
The volume databases and primary roles are distributed across all Metadata Services in the cluster. Additional Metadata Services - either on shared servers or dedicated metadata servers - can be added at any time to improve metadata performance and increase the capacity to store more and larger volumes. There is no upper limit to the number of metadata servers in a Quobyte cluster.
Metadata is stored in a highly efficient key-value store (KVS) based on LSM trees, which Quobyte developed specifically for high-performance and low-latency file system workloads. In addition to being optimized for typical file system operations with billions of files, the KVS has been designed to handle highly concurrent scans of files, e.g., for the Quobyte File Query engine or maintenance operations. Making the Metadata Services capable of running complex queries over billions of files across all servers in parallel within minutes.
For reliability, the key-value store is also tightly coupled with the synchronous replication used for redundancy and reliability. A single transaction log is used for persistence, recovery, replication, and backup catchups. In addition, all metadata data is checksum-protected at rest and in transit.
Besides handling and storing metadata and enforcing file access permissions, the Metadata Service has another crucial role: It evaluates policies to decide which devices to use for a new file. The Metadata Service makes sure that files are stored on the desired media type (NVMe, HDD), across the selected failure domains (across machines, racks, etc.), with the proper data protection (EC vs. replication), and many other policies.
Quobyte Data Service
Once a file has been opened, the operations that interact with file data (contents of a file) are performed directly on Data Services. As Quobyte clients know all Data Services and devices with the file contents, they do not need to interact with the metadata service anymore.
The Data Service has three fundamental responsibilities:
- Handle IO operations like read, write, fsync, or truncate and validating checksums of data read/written
- Manage file locking: File locks are handled by the Data Services storing a file for scalable locking support. This includes file locks for erasure-coded files.
- File replication: For files protected with synchronous replication, the Data Services do peer-to-peer primary election and quorum data replication to ensure consistency.
- Transparent file movement: Files that need to be moved to a different drive or server are migrated using the replication mechanisms between Data Services. This movement does not cause any IO interruptions to applications, ensures redundancy is never reduced, and can even migrate erasure-coded files while in use.
Quobyte API Service
Quobyte is API-first, meaning all functionality for controlling a Quobyte cluster is available via the API. The command-line tool and the Quobyte Webconsole UI both use the API to talk to the Quobyte cluster.
Admins can run as many instances of this service as needed on the same servers as the core services or on separate servers.
The Quobyte API is implemented as a JSON REST API via HTTP(S). The documentation provides a full list of API calls. Users can authenticate with their username/password or Access Key and Secret Key credentials. The API service supports multi-tenancy and access control based on the user's role.
Quobyte Webconsole Service
This service provides a web-based graphical user interface. Like the API, admins can run as many instances of this service as they like across Quobyte servers or on dedicated servers.
The webconsole offers all the options to configure a Quobyte cluster, including networks, LDAP/ActiveDirectory backends, security configurations, and a visual editor for the policy engine. In addition, the webconsole provides comprehensive monitoring, performance analytics, and alerting dashboards.
Multi-tenancy is fully implemented across all layers in Quobyte, including in the Quobyte Webconsole. Similarly, support for different user roles is built-in. While admins can see the entire system and all configurations and alerts, there are more restrictive roles available:
- Tenant Admins (read-write)
Can see their namespace (volumes, quotas, clients) and create and manage volumes and quotas. - Regular Users (read-write)
Have a limited view with a dashboard for their usage, quotas, and can manage their S3 and API credentials. - Hardware Operators (read-write)
Have only access to the device and services list. They can drain devices and control drive LEDs. - Filesystem Admins (read-write and read-only)
Can manage volumes and quotas of all tenants and policy rules, but do not have access to the hardware and services part of the system. - Superuser (read-write and read-only)
Can do everything, but like all other management roles can’t access files directly
Quobyte Rapid RPC Protocol: Designed for Massive Performance and Reliability
Quobyte's Rapid RPC Protocol has been designed for modern computers and distributed systems storage. Instead of relying on the dated NFS protocol, it provides the essentials for modern, fast storage for workloads like AI/ML and HPC: parallel n-to-m communication for metadata and data, aggregated multiple TCP connections for higher performance, end-to-end hardware-accelerated checksums, and an RPC design for fast parsing on modern CPUs.
End-to-end Checksums protect each data block (usually 4k to align with the x86/ARM page size) from when a Quobyte client gets the data from the OS through its entire lifetime. The checksum is validated before the block is persisted to find in-transit corruptions. The checksums are then stored on disk (and protected with additional LRE checks). On read, they are validated again to ensure that the data the application reads is exactly what was written - even when data is corrupted on disk, over the network, or in the networking stack of the OS.
RPC Message Checksums protect the actual RPC header and messages to detect corruptions in transit. Bitflips in messages have led to massive outages in the past. Quobyte protects you from these situations by detecting corrupted messages before they cause harm.
Direct Full Mesh Communication. Each Quobyte client directly talks to the Metadata Services responsible for the volumes it needs to interact with. Once a client opens a file, it communicates directly with one or multiple Data Services (striping) to read/write a file, aggregating the performance of 15 or more servers in a single file. There are no bottlenecks or additional network latencies like NFS gateways in the path, and all clients communicate directly with all Quobyte servers (n-to-m communication). The performance of the Quobyte cluster is only limited by the (networking) hardware - and the speed of the devices and CPUs in the servers.
Separate connections and backpressure. Quobyte uses separate TCP connections for each data device and volume. Once a device is busy, backpressure is transmitted to the client to slow down IO from applications instead of having queues build up everywhere, resulting in wasted RAM. Unlike protocols like NFS, Quobyte can ensure that, e.g. metadata operations or IO to other devices is not affected by a heavily utilized device. This is why Quobyte can provide graceful degradation under heavy load where all clients (inside the same QoS class) are throttled fairly instead of the client locking, a typical phenomenon with NFS.
Fairness and Quality of Service (QoS). Thanks to individual connections and end-to-end control over the protocol, Quobyte can provide fairness inside the quality of service classes between files, small and large IO on the same file, and metadata operations. Fairness and QoS are implemented end-to-end in the Quobyte client, the RPC protocol, the Quobyte services, and IO scheduling. Quobyte services also use the QoS to prioritize urgent traffic like rebuilds and deprioritize background tasks like rebalances.
Better performance with multiple TCP connections. The Quobyte client uses multiple TCP connections for each device to maximize throughput over TCP.
TCP and RDMA, no multicast. Quobyte uses TCP by default and supports multiple networks for clients and servers, including routed networks. There is no need for complex networking setups that limit usage, such as multicast. When available, Quobyte can take advantage of RDMA (over RoCE, infiniband, and OmniPath) between clients and servers. This switch is transparent, and clusters can be run in a mixed mode where only a subset of clients and servers have RDMA enabled.
Works with a wide range of Network Setups. Quobyte can run over a single network for "frontend" (client-to-server) and "backend" (server-to-server) traffic. This is the recommended setup to reduce complexity and cost. However, Quobyte can use multiple networks, e.g. one or more separate backend networks. Similarly, Quobyte supports multiple non-routed fronted networks on the servers, e.g., one for the compute/GPU cluster and another for access from workstations. Clients can access servers over routed networks since there aren't any special networking requirements.
The Quobyte Native Clients
Native clients are the counterparts for Quobyte services: They run on the machines where users and applications use the file system. The client is responsible for creating the "file system semantics" and translating file system operations into calls to the various Quobyte services. It communicates directly with all Quobyte services via the Quobyte Rapid RPC protocol and is available for Linux, Windows, and macOS.
While native clients hold local state, such as cached data or the locations of volumes and services, a disconnected or crashed client does not impact the cluster outside of the client machine. Locks that applications on the client hold will time out automatically, allowing automatic failover of applications to another host within seconds.
No Local Configuration
Like the Quobyte services, the native clients are designed for ease of use and scalable operations. This is why they don't need local configuration outside of the cluster's registry endpoint DNS record. Deploying the native client across a large fleet of workstations, compute nodes, and GPU servers is trivial.
Client-side configurations, such as metadata cache TTLs, data cache behavior, and other settings, are controlled through the policy engine and automatically updated and applied to all clients when changed.
The built-in automounter makes all accessible volumes available as subdirectories. There is no need to mount volumes individually or configure services like autofs. The native clients adapt the visible (and accessible) volumes based on the user's tenant membership. For convenience, the volumes can be arranged hierarchically, e.g., in groups in subdirectories like /quobyte/home/user, where every user has their own volume.
Orchestrated Performance and Protocol Optimizations
A major advantage of a native client with a custom RPC protocol is that Quobyte clients can take advantage of performance optimizations unavailable in NFS. Since the Quobyte client runs in user space, optimizations, and new features can be tested and released much faster. A number of sophisticated optimizations in the client result in an overall performance increase vs. traditional protocols:
Parallel IO aggregates the performance of many drives and servers, even inside a single file.
The Adaptive Prefetcher loads data in parallel from the drives and adapts to the speed of the applications and the backpressure from the drives. This enables the prefetcher to optimize throughput with low latency when access patterns change. It also adapts to multiple streams in the same file for applications that have multi-threaded read patterns like modern video codecs. In combination with the auto redundancy, the adaptive prefetcher can also mask the latency of hard drives.
File prefetching automatically loads files in the background based on file name patterns. This masks latencies and eliminates jitter when accessing lots of small-ish files in a sequence, e.g., in one-file-per-frame video format playback (DPX) or AI/deep learning training.
Automatic caching detects access patterns per file in cooperation with the prefetcher. For applications that exhibit random IO, the client automatically disables caching and switches to direct IO to minimize latency. Admins have full control over caching behavior with the policy engine. For maximum compatibility, the Quobyte native clients support NFS cache coherency by default (close-to-open/unlock-to-lock consistency), but these can be relaxed with the policy engine.
Deferred writebacks allow applications to trade strict NFS semantics for faster performance for sequential small file workloads, such as git or compiling source code. Caches are flushed in the background when a file is closed, so the application doesn't have to wait for writes to be acknowledged by the Data Services. Administrators can configure an upper bound in the lag for the writebacks. Once that limit is reached, applications will be throttled.
RDMA is used automatically when available. If allowed by configuration, and both the client and the service it communicates with have RDMA enabled, the client automatically switches from TCP to RDMA for lower latency and lower CPU utilization. The native client can operate in a mixed mode where some services have RDMA while others can only be reached via TCP.
Reliability
Another major advantage of the Quobyte client over traditional storage protocols like NFS is that the Quobyte client understands data redundancy and can automatically fail over to another replica or restore lost data blocks from EC parity data. Failovers are part of the protocol and do not rely on kludges like virtual IPs or multicast.
This failover is completely transparent to applications: All operations are idempotent so that retries do not cause errors to the application, e.g., due to a file create executed twice. Even when disconnected for hours or days, applications will simply continue to work when clients can reach services again. This behavior is configurable with the policy engine to ensure that long-running batch jobs are not killed by temporary unavailability. At the same time, interactive workloads can be set to receive IO errors instead of hanging. In addition, the built-in client-side monitoring alerts admins to applications waiting for IO for prolonged periods.
Due to this orchestration between client and protocol, Quobyte doesn't need error-prone and slow IP-based failover mechanisms that often cause issues when the network isn't configured optimally.
In addition, creating the resilient file-system view in the client instead of NFS gateways also greatly improves scalability: NFS clients were designed when storage was a single server. NFS-based distributed solutions must create a consistent (single copy serializable) view on the gateways. This cache consistency comes at a high price of coordination between gateway nodes that access the same directories and files, severely limiting the practical limit of NFS gateways on a single file system.
Built-in Monitoring of IO and Applications
The native clients have a wealth of information available: They collect IO statistics before and after caches to show real workload patterns vs. what goes over the wire. Clients make this information available on a per-file level, helping admins to understand workloads and quickly identify problems. For each file, the Quobyte client shows
- throughput
- IOPS
- maximum IO depth
- current operations in flight
- detected access patterns, and detected changes in the pattern
- prefetcher horizon
- file locks held and those the application is waiting for
The native clients also aggregate this information for each process and make the top consumers of IOPS, throughput, and metadata operations available. This data is aggregated for overall top activity in the cluster.
In addition to performance monitoring, the Quobyte clients also provide metrics for the alerting system to inform admins of issues before users and applications see them. This includes alerts for applications waiting for IO for too long, being stuck waiting for file locks, or because quotas ran out, as well as inefficient or undefined IO behavior like random IO on erasure coded files or concurrent unaligned file appends.
The Quobyte Native Client for Linux
The native client for Linux comes as a simple RPM or deb package, or as a container image, and runs entirely in user space via FUSE. The native client does not require a custom kernel module, runs on any version of the Linux kernel (version 3.10 or newer), and can be updated easily. It delivers high performance (depending on the hardware) with up to 10GB/s single stream read performance over RDMA.
Users and applications can access the file system from /quobyte (or any other directory, which can be specified when mounting), and volumes are shown as subdirectories.
The client can be used as a systemd service (recommended), via fstab, or manually.

The Quobyte Native Client for Windows
The native Windows driver has the same benefits as the Linux driver when it comes to performance, monitoring, and automatic fault tolerance. It avoids the bottlenecks and hassle of accessing storage via the SMB protocol. The native Windows client is ideal for applications and appliances that require high performance, like video editing, visualizations, or automated electron microscopes. Single-stream performance over TCP varies with the hardware, but a bandwidth of 4GB/s (single-stream sequential read) can easily be achieved.
Quobyte can be accessed in three ways with the native client for Windows:
- As a drive: A single volume can be mounted as a Windows drive, e.g. "Q:" This mode is ideal for applications that cannot handle UNC paths.
- UNC path: All volumes are accessible as network volumes on the local machine
- Inside a folder: Like Linux/macOS, Quobyte can be mounted into an existing folder on a local drive. Volumes are accessible as subfolders.
The Windows client translates Linux permissions and ACLs to and from Windows access control.

The Quobyte Native Client for MacOS
The macOS client runs in user space using macFUSE, similar to the Linux client. Quobyte can be mounted in a subdirectory with a single volume or all volumes as subdirectories.
The macOS client supports extended attributes, colored files/folders, and access control.

The S3 Gateway Service
Quobyte's S3 Gateways provide access to the file system through the S3 object storage interface. It is based on a custom S3 protocol implementation by Quobyte and is tightly integrated with the file system layer. As a result, the S3 Gateway offers unified access to the file system: Users can seamlessly access files and objects and vice versa.
Like the file system, the S3 Gateways are designed for low latency and high throughput. Object PUT and GET calls are directly translated to IO operations on the Quobyte Data Services that hold the file/object data for minimal latency. Placement for files created via an object upload is handled just like normal files by the policy engine and can take advantage of Flash for lower latency.
The S3 Gateway Service is stateless, and each part of a multipart upload can be directed to a different S3 Gateway. Load balancing many clients and HTTP connections across many S3 Gateways is easy to configure and operate. This service can run on the same physical server as other Quobyte Services (Data, Metadata) or run on separate machines or VMs, e.g. to improve security when sharing data with the "outside world".
Quobyte's S3 Gateways provide a tightly integrated and fully transparent translation layer between the file system and S3:
- Objects are mapped onto files. Uploaded objects are stored as regular files on the file system. The object name is translated into a path; for example, an object named something/foo/bar.txt will be stored as a file bar.txt in /something/foo. If necessary, the S3 Gateway creates the directories in the path. It even supports directory objects.
- Files are accessible as objects. No matter if a file was created via any file system driver or S3, the file can be accessed as an object and a file simultaneously. It doesn't matter whether the file was created before a volume was made accessible via S3 or after.
- Existing file system volumes can be exported as buckets. This gives admins control over which volumes can be accessed via S3.
- New buckets can be translated into new volumes or subdirectories on an S3 bucket volume, and the way a new bucket is stored can be configured.
- S3 user metadata is stored in the file's extended attributes (xattrs). This means that S3 user metadata is visible in the file system layer, and xattrs set on a file are visible as user metadata via S3. In addition, user metadata can be queried via the Quobyte File Query Engine, regardless of how it was attached to the file.
- Object versions are mapped onto file versions, which can be accessed from both the file system and the S3 interface.
- Object lock (immutability) is supported as well and translated to Quobyte's file immutability policies.
Users authenticate to the S3 Gateway service with the well-known credentials: an Access Key and a Secret Key. In Quobyte, each key is associated with a user with a uid and a primary group. When the S3 Gateway service executes the file IO, it uses these credentials. This allows for unified access control between the file system and S3:
- All S3 operations are translated to file system IO by the user identified by the S3 credentials
- File system access control—like permissions and ACLs—is enforced for S3 access as well. Admins do not have to manage separate access control for S3 and the file system.
- Unlike traditional S3, Quobyte supports fine granular access control inside a volume, e.g., permissions for individual users and groups on subdirectories.
- ACLs can be displayed and changed through the S3 ACL operations.
The NFS Gateway Service
Like all Quobyte services, the NFS service is a standard systemd service that runs as user quobyte (no root privileges, not based on the kernel nfsd) and doesn't require local configuration.
Installing and running the service on many machines is straightforward and can easily be automated. The service can run in parallel with other Quobyte services or on dedicated hosts.
The NFS Gateway Service supports NFS versions 3 and 4.1 simultaneously. When using NFS version 4, Quobyte supports advanced features such as file locking, NFS4 ACLs, Kerberos authentication, xattrs and in-transit encryption.
Unfortunately, the NFS protocol has no fault tolerance or load-balancing capabilities. Unlike the Quobyte Rapid RPC protocol, NFS requires virtual IPs for fault tolerance. The NFS gateways automatically acquire one or more virtual IPs from the configured pool. Other NFS gateways will automatically take over the virtual IPs if a gateway is disconnected, crashes, or is stopped.
The Quobyte NFS Gateways support multi-tenancy based on configured IP networks. Machines will be assigned to one or more tenants based on their matches with the IP networks configured for them. In addition, access can be restricted to specific IP networks on a volume basis.
Other Connectors
- SAMBA
Quobyte offers Windows connectivity via our native driver and SAMBA. Quobyte supports access control via POSIX ACLs, file locking, and fault tolerance with cttb on Quobyte. - Hadoop/HDFS
The native HDFS driver loads the Quobyte client (libquobyte) into the application process and supports the full HDFS feature set. Since file IO operations do not go through the kernel, users can provide certificates or Access Keys for authentication. - MPI-IO
Quobyte's MPI-IO driver uses libquobyte directly in user space. All file system operations are directly translated into RPC calls to the Quobyte services. The file system IO does not go through the kernel, for security users can provide certificates or Access Keys.
Seamless Data Protection
Quobyte implements data protection fully in the software layer and does not rely on any kind of hardware redundancy. Instead, its design follows the "don't trust the hardware" approach to ensure that Quobyte is truly hardware agnostic.
Quobyte's data protection has two pillars: Data redundancy ensures data is stored consistently across multiple machines to protect against hardware failures, outages, and unavailability. The second pillar is for detecting issues across all layers: end-to-end checksums and hardware monitoring.
End-to-End Checksums
One of the major benefits of the Quobyte native client and the Quobyte Rapid RPC protocol is that it enables the use of true end-to-end checksums. The Quobyte native clients compute the CRC32 checksum for each 4kB data block as soon as they receive it from the operating system. The checksum is hardware accelerated on all modern CPUs (iSCSI polynomials), and computation is almost "free" since the data is hot in the CPU L1/2/3 caches. The block size can be configured but defaults to 4kB, the page cache size on x86/ARM processors.
The checksum stays with the data block for its entire lifetime. The Data Services validate the data and checksum after they receive it over the network. If data is corrupted in transit - outside of the control of the Quobyte services - admins will be alerted to an issue in their networking stack. Here, Quobyte provides the necessary protection that TCP doesn't offer for 64kB+ packets due to the CRC16 being unable to protect more than 64kB. In addition, Quobyte's end-to-end checksums protect against difficult-to-find issues like buggy network drivers or defective ethernet switches.
On the Data Service, the end-to-end checksum is stored on disk with the data itself and protected together with other on-disk metadata using a LRC checksum.
On read, the checksum is verified after reading from disk by the Data Services and by the client before delivering the data to the kernel and the application. If data corruption is detected, the data is read from a good replica or restored from EC parity. Drives that corrupt too many blocks are automatically removed from the cluster.
Software Monitoring
While Quobyte takes a different approach than most products with its "anything can fail" design, it closely monitors the hardware and network to ensure low latency and overall cluster health. This includes the aforementioned end-to-end checksums but also:
Device Watchdogs.Quobyte does not rely on IO timeouts from the kernel or drive firmware. Drives that don't respond in a timely manner will automatically be removed from the cluster to ensure consistent low latency and performance. As a result, Quobyte does not require e.g., SAS drives or specific firmware versions on the drives.
RPC Checksums. Like end-to-end Data Checksums, all RPC messages are protected with CRC32 checksums to ensure that issues in the networking stack are detected.
Metadata Checksums. All metadata stored on disk - transaction logs and checkpoints - is protected by checksums.
TCP connection monitoring. Quobyte doesn't rely on the kernel or the networking layer to check the liveness of TCP connections. They are proactively monitored and dead connections are terminated, and admins receive timely alerts.
SMART. Quobyte uses SMART as an additional source to detect failing drives that impact performance negatively.
Kernel monitoring. Quobyte regularly scans errors from the Linux kernel and modules to detect hardware issues.
Split-brain Safe Synchronous Replication
The key value store used by the Quobyte Registry and the Metadata Service is protected using Quobyte's patented synchronous quorum replication mechanism. Files can be protected using the same replication mechanism optimized for file IO or erasure coding (see next chapter).
The patented quorum replication consists of two core mechanisms:
- The peer-to-peer primary election
When necessary, one of the services that holds a copy of the data becomes primary
through an election mechanism executed by the services with a copy among each other
(peer-to-peer) without needing an extra locking service. This peer-to-peer primary
election is highly scalable and avoids typical bottlenecks. In addition, because it is
lightweight and scalable, Quobyte is able to have a fine granular primary role, e.g.,
for each file, instead of having to group files into replication groups or other
containers.
If a primary fails, is shut down, or disconnected, the remaining replicas instantly
re-elect a new primary that takes over.
- The quorum primary/backup replication
All IO operations are handled by the primary, which sends the updates in parallel to its
local disk and the remote backups. As soon as all available backups and the primary –
but at least a majority – have acknowledged the write, the primary sends the
acknowledgement back to the client and application.
If a primary fails, a new primary is elected, automatically bringing the file to a
consistent state to serve application IO.
Together, these two mechanisms provide the fault tolerance that makes Quobyte so robust using quorums and the strong consistency that creates the illusion of a single copy of the data as required by, e.g., databases. The patented synchronous quorum replication in Quobyte yields a number of advantages over traditional redundancy mechanisms:
- It is implemented entirely in software over IP networks, with no need for RAID, dual controllers, or heartbeat signals
- While primarily designed for low-latency IO and replication inside a cluster, it can tolerate latencies up to 50ms. This means that Quobyte can be deployed in geographically distributed clusters yet provide strong consistency.
- The majority voting (quorum) protects against the well-known issues of data mirroring, like inconsistent copies or stalled systems due to split-brain
- This strong data protection enables the anything-can-fail design since losing a replica is not an emergency situation. There are still two copies of the data available, and rebuilds can be done without having to go into panic mode.
- Quobyte's services protect the data stored on local drives - a so-called shared nothing architecture. There is no need for a storage SAN backend, like NVMe-over-Fabric, and the complexity of coordinating these and having an additional hardware and networking layer.
- The peer-to-peer primary election is key to Quobyte's linear scaling to 1000s of servers as it avoids centralized lock services.
A key feature of Quobyte's replication is that replicas can be added or removed from a file/key-value-store while the file/key-value-store can still be modified. This enables Quobyte to migrate data and support, e.g., removing a disk/server or a hardware refresh without downtime. Replica movements are transparent to applications and users.
The Quobyte key-value-store, used in the Registry and the Metadata Service, tightly couples the replication with the key-value-store's transaction logs. Since Quobyte develops both components in-house, the key-value-store replication benefits from not requiring an additional log. In addition, primaries and backups hold the latest state locally (in RAM and on disk) so that a backup can take over immediately from the primary without lengthy replay operations.
The Quobyte Data Service uses replication with a separate primary for each file. The replication operates on a 4kB block size with in-place updates for low-latency random IO operations. The in-place updates are coupled with the checksum for atomicity and avoid the problems of write-log-based systems, such as write amplification, log purging, and inconsistent performance when logs and caches run full. Similarly, reads are in place and don't have to be retrieved from a write log or caching device.
Sequential IO is batched into larger write operations to ensure maximum performance for throughput workloads on NVMe and HDD.
For small reads, Quobyte clients can do direct quorum reads across replicas round robin. This allows Quobyte to deliver highly accessed small files with high performance from three or more servers in parallel.
Online Erasure Coding
For file data, Quobyte offers online erasure coding (EC) as an alternative to replication. File content is erasure coded on the Quobyte native client (or the S3/NFS gateways) using a patented high-performance consistency protocol.
Erasure-coded files are automatically striped for performance, i.e., an 8+3 erasure-coded file is stored across 11 drives and servers. The performance of these drives/servers is aggregated for performance. The client computes the erasure coding parity blocks for each row using hardware-accelerated Reed-Solomon codes. The data and parity blocks are then sent in parallel in the background to the storage servers.
When reading data, the client will only read the data blocks requested by the application. If a data block is unavailable or corrupted, the client automatically reads the parity data and restores the missing data online. Applications do not have to wait for a rebuild process to access data.
Since erasure coding requires row consistency, meaning that data blocks and parity in a row must be the same version, Quobyte enforces single-writer semantics on a file. Concurrent reads from many clients are supported. Applications that execute inefficient operations on erasure-coded files are automatically detected by the Quobyte native clients and reported to the administrators.
Files can be recoded from replicated to erasure-coded and vice versa and recoded between different erasure coding schemas. This offline recoding is controlled through policies.
A major issue with erasure-coded storage systems is the live migration of erasure-coded files between storage servers. To overcome this problem and offer non-disruptive node removal and data isolation, Quobyte combines erasure coding with replication: When an erasure-coded file or some stripes of the file need to be moved to another server, the stripe is temporarily replicated in addition to the erasure coding. The new replica is added, and the old replica is removed once all data has been migrated.
Auto Redundancy: Combining Replication and Erasure-coding for Optimal Performance and Space Utilization
The auto redundancy mode combines synchronous replication and erasure coding inside the same file. When enabled, the first 8MB of each file is protected with synchronous replication. As the file grows above 8MB, the rest of the data is protected with erasure coding.
The file content can either be placed entirely on Flash, or with auto placement, the first 8MB of the file will be stored on Flash, while the rest will be stored (with erasure coding) on HDD. The auto redundancy has a number of advantages:
- Small files (<8MB) are stored with replication for better space utilization than EC.
- Small files can be read from all three replicas in parallel for higher performance
- Small files are stored on flash for low latency and to avoid random IO workloads on HDDs
- For large files, the first 8MB are served from Flash with low latency. In parallel, the prefetcher in the Quobyte native client prefetches the erasure-coded data from multiple drives. The data stored on Flash helps to reduce the time-to-first-byte and masks the seek latency of the hard drives.
- The policy can handle small and large files in the same system and optimizes performance and space efficiency.
High-Performance Declustered Rebuilds
When a drive or server in a Quobyte cluster fails, the data is not restored in place, which causes typical performance issues in RAID systems or those with replication groups. Instead, Quobyte creates new replicas or EC stripes on all available drives in the cluster that have free capacity.
During a rebuild, all servers and drives contribute to the rebuild. As a result, the load on the CPU, the IO operations, and network traffic is well distributed across the entire cluster. This yields a number of benefits:
- Rebuild times drop with increasing cluster size, counteracting the increase in failures with more hardware.
- Rebuilds cause very little traffic and IO impact on each machine, meaning that the impact on production traffic is almost nonexistent.
- All existing replicas and EC stripes of a single file contribute to the rebuild, decreasing rebuild times even further.
- Thanks to the automatic failover and online recomputation of missing EC stripes, files and volumes stay fully accessible to applications and users during the rebuilding process.
- There is no need for spare drives or servers. As long as the cluster has enough redundancy (servers and drives) and capacity, all resources are used all the time instead of wasted spare drives.
- When hardware fails, no human interaction is required. There is no need for fast replacement of drives and servers and expensive support like "24-hour on-site replacement" can be avoided. Instead, hardware operators can wait until enough hardware has failed and do batch replacements once a month.
- The "let the hardware rot" approach makes Quobyte ideal for remote locations like co-located data centers or edge deployments.
The Quobyte File System: Flexibility of a Truly Software-defined Storage System
In this section, we look at how all pieces of a Quobyte cluster come together to provide a scalable, fault-tolerant, and policy-driven file system that abstracts from the underlying hardware. This includes an in-depth explanation of how your files are spread across servers and disks and how IO is cached from the application to the actual physical storage devices.
The Logical Storage Layer: What the Users see
Quobyte organizes the logical storage into volumes. Each volume is a separate namespace/file system and belongs to one tenant (see Multi-Tenancy). Volumes are unlimited in size/capacity and thinly provisioned. The policy engine assigns actual storage to a file when it is created.
Volumes can be arranged hierarchically in the global namespace; for example, each user can have a separate volume under home/alice, home/bob, etc.
Quotas limit how much data and files can be stored in a volume by a user, group, or tenant. Quobyte supports quotas on the number of files, the physical storage (including redundancy overhead), and the logical storage (the file size a user sees), both of which can be defined for all tiers or for flash and HDD. Quotas are evaluated hierarchically, i.e., a tenant quota will supersede volume or user quotas. This allows administrators to do oversubscription.
Enforcement of Quotas is done in a scalable manner that does not cause significant performance impact. Users can exceed the quota for up to a few seconds before their IO is suspended.
Volumes can be accessed concurrently via all supported interfaces, including NFS and S3, without limitations. Access control is equally unified: Quobyte supports NFS4-like ACLs, which are enforced for all access and auto-translated to/from Windows, S3, NFS4, POSIX, and macOS. Similarly, user-defined metadata is unified, and extended attributes (xattrs) can be accessed through all supported interfaces and as S3 user-defined metadata.
The Quobyte Policy Engine: Software-defined Storage - Defining the Structure of Your Cluster in Software
The policy engine is at the heart of a Quobyte cluster and the key to the total flexibility of Quobyte. The engine operates on a per-file level, and combined with how Quobyte does data protection on a file level without replication groups, allows the individual placement of every new file. Similarly, existing files can be moved to any set of drives in the cluster when policies are changed. Policies can be changed at any time, and data can be moved non-disruptively to adapt a cluster to changing requirements without any interruptions.
This allows admins very fine granular control, like placing Alice's .mp4 files with erasure coding on the high-density HDD servers. More practical examples include:
- perfect performance isolation of workloads (scratch/home), users, projects, or tenants in hardware
- guaranteed isolation of data from different tenants, e.g., for legal or compliance purposes
- automatic transparent tiering of warm and cold data from flash to HDD
- optimal usage of different flash types like AI/deep learning training data on read-optimized flash
- directing data directly to the proper tier, e.g., important projects directly to Flash, while S3 uploads go to HDD
The Quobyte policy engine is a very powerful tool that allows admins to reconfigure hardware usage and optimize applications, yet it is very easy to use. In line with our approach of "simplicity scales", the Quobyte policy engine…
- offers defaults for automatic fault tolerance and performance enabled on installation.
- has pre-configured policies for common applications and use cases.
- has full UI support for configuring, editing, managing, and evaluating policies with a few clicks. Since Quobyte is API first, all modifications of policies can be done via command line tools or the API.
- is forgiving to operator errors. When an admin misconfigures a policy, changes can be reverted on the fly.
- decouples data movement from policy changes. The data movement can be deferred and done during periods of low traffic like the night or weekend. However, administrators have the option to start the data movement immediately.
Tags are used for data placement in Quobyte. Each drive can have one or more tags. In addition to tags assigned automatically, like NVMe, SSD, or HDD, admins can configure as many tags as they need to create additional pools. These can be used, e.g. to identify HDDs in high-density servers used for cold storage, for multiple flash tiers (read-optimized for AI training data, QLC drives), or to assign drives and servers exclusively to a tenant, workload, volume, etc.
Labels are key/value pairs (strings) that can be assigned to tenants and volumes and used to match policy rules.
Policy Rules
A policy rule in Quobyte combines one or more policies applied to all targets that match the policy rule's scope filters. Policies in this sense include data placement, client cache settings, or redundancy (see further down for a full list of policies). The targets include files, clients, volumes, tenants, or global. Specific targets can be selected with the policy scope filters.
Policy rules are evaluated hierarchically: The first policy rule that matches a file is used. In most cases, a policy rule does not define all possible policies. In this case, the policy engine continues evaluating the remaining undefined policies until it finds a match or will fall back to the built-in defaults.
The order in which the policy rules are evaluated is as follows (from top to bottom). Within a hierarchy level, admins can configure the priority of policies:
- Files This is the most granular scope and can target files based on a large number of filters:
- file name (starts/ends/contains/regex)
- file extension
- owner (username)
- expected object size
- current file size
- last access time
- last modification time
- creation time
- extended attribute on the file, parent directory, or any directory in the path. This filter allows users to interact with the policy engine directly
- Clients Apply policies only to requests (open, file create) from specific clients. Available
scope filters are:
- client type (native, S3, NFS)
- client's IP address (network filter)
- Volume Apply to all files inside a volume. Available scope filters are:
- volume uuid
- regular expression match on the volume name
- labels (exists or value match)
- Tenant Apply to all files and volumes inside a tenant. Available scope filters are:
- volume uuid
- regular expression match on the volume name
- labels (exists or value match)
- Global Policies Apply to all files in the system.
- Built-in Defaults Cannot be changed and will be used if no match was found for a policy.

The actual policies fall roughly into two groups: Creation time policies apply when a file is created. They can only be changed by a recode task that essentially copies the file into a new format. When changes are made, these only apply to newly created files. Runtime policies are applied when a file is opened or pushed in real time to all clients. The following policies are currently supported:
- Redundancy for files Files can be protected with synchronous replication (factor 3 or 5), erasure coding (5+3, 8+3, and 12+4), auto redundancy, or unprotected (for cloud-based setups or scratch)
- File striping Replicated files can be striped across multiple drives for performance. This is implicitly enabled for EC files.
- Tag-based Placement for Files Based on required and forbidden tags, which can be enforced strictly (IO error when no space is available) or relaxed (temporarily use other devices until suitable ones become available)
- Failure domain placement for Files and Volumes Place replicas or file stripes across drives that are in unique failure domains, where the failure domain level can be selected (machine, rack, cluster, data center)
- Locality-placement for Files Similar to Hadoop's data placement policy, Quobyte will try to use a replica in the same machine or place all data stripes of a file in a local machine
- Redundancy for volumes Synchronous replication with three or five replicas
- File retention policies This includes immutability, append-only, no deletion or automatic expiration after configurable time
- Volume Snapshots With configurable snapshot interval and retention period
- Volume Security This includes privileged users (to override root), privileges for regular users, and Windows-specific settings and mapping options
- Client-side Settings such as metadata cache settings, automatic file prefetching, Quality-of-Service, file locking, or fsync behavior
The Quobyte policy engine is highly decentralized and runs on all Metadata Services without locking. It decides only the file layout of a file (see next section), and neither the Metadata Service nor the policy engine is involved in the block allocation of a file, which enables Quobyte's scalability.
With explicit file placement, Quobyte does not need unpredictable, consistent hashing. This ensures that the behavior of a Quobyte cluster is deterministic at all times and avoids the dreaded forced rebalances of hashing-based systems. Adding or removing nodes from a Quobyte cluster doesn't cause any rebalance traffic. Instead, the policy engine uses new drives and servers for new files within seconds. With the flexibility to change placement at any time, fine granular control, and hardware isolation, the Quobyte Policy Engine is a core component of offering Storage-as-a-Service.
Failure Domains for Hyperscale Availability
Failure domains in Quobyte are integrated with the Quobyte Policy Engine to ensure data replication across a number of different failure domain levels inside the same cluster.
Quobyte supports disks, machines, racks, clusters, and data centers as failure domains. Disks and machines are enabled by default. The remaining failure domains can be configured based on hostname (FQDN) or IP networks.
Once configured, failure domains can be used for placement decisions. Administrators can set the minimum failure domain spread for volume replicas, file replicas, and erasure-coded files.
Quobyte can tolerate a latency of up to 50ms between the failure domains when used with synchronous replication. This enables setups with Quobyte running replication between three geographically separated data centers. Such a "geo-stretch-cluster" can tolerate the loss of a full data center without any data loss or downtime.
On a smaller scale, using failure domains at the rack level is highly desirable: A cluster with rack-level replica spread can tolerate the failure of a full rack - unfortunately, this is a very common scenario when power or networking is misconfigured. Larger clusters benefit from the ability to update and restart all servers inside a rack simultaneously to reduce the duration of software updates.
Since Quobyte has a shared-nothing architecture and doesn't require hardware redundancy or heartbeat cables, it is easy to spread clusters across many racks. A typical data center design puts a small number of Quobyte servers in each rack to improve distribution and resiliency.
File Layouts: The mapping of a Logical File onto Physical Drives
When created, each logical file in Quobyte is assigned a file layout by the policy engine. The layout assigned to it depends on the matching policies for redundancy (replicated, ec, auto), failure domain spread, and device tags (required/forbidden).
The file layout enables a Quobyte native client and other Quobyte services to locate the drives containing a particular byte of a file.
A file is split into one or more segments. A segment covers a range of bytes, e.g., from byte 0 to byte 2147483648 (2GB). Each segment can have a different size, but they are typically between 2 and 10GB.
Each file segment has its own configuration for data protection, striping, and drive assignment. This means the first segment could be replicated, while the second is erasure-coded.
For replicated segments, each stripe has three or five drives assigned to it—the replicas of the stripe. Erasure-coded segments are always striped, i.e., when the 5+3 schema is selected, the segment has eight stripes and eight drives assigned to it.
However, the file layout does not include any information on how the blocks on disk are assigned to a replica or stripe. The Data Service handles this locally. Quobyte clients can use the list of segments to determine which Data Service and drive to talk to in order to execute read and write operations. The file layout also includes the replicas so clients can automatically fail over to the next one or restore data from the EC parity blocks.
File layouts support sparse files, and unallocated segments can exist. Similarly, Quobyte supports deallocation (hole punching) and maps these onto the file layout.
A file layout can change even while a file is in use, e.g., when replicas are added or removed. The changes in the file layout are propagated to the clients through a highly scalable "gossip protocol" that avoids centralized coordination.
Caching and Consistency in Quobyte File Systems
Quobyte directs clients' IO operations directly to the devices the file has been assigned to, i.e., the final location of the data. There aren't any caching or write-log devices in between, which means that performance depends entirely on the media type selected for placement and not on some shared caching devices that might run full. This results in predictable and consistent performance.
In addition, Quobyte always writes data persistently to the disk, waiting for the drives' acknowledgment. This results in much stronger data durability guarantees in Quobyte, similar to writing files locally with the O_SYNC flag, and it also ensures that drive caches do not create variance in performance.
With the backpressure implemented in the Quobyte Rapid RPC protocol on a per-device basis, applications on the client can write consistently with the actual performance of a drive or multiple drives aggregated for striped or erasure-coded files.
Data Caching in the Quobyte Native Client
The data cache inside the Quobyte native client is adaptive and automatically detects IO patterns. The client will cache data locally when detecting sequential workloads, including multi-stream sequential IO on the same file. For random IO workloads, like databases and VMs, the clients disable caching and prefetching to reduce latency. The data cache is also disabled when a file is opened with the O_SYNC or O_DIRECT flags.
Writes are batched into larger chunks and then sent to the drives in parallel. This parallel IO allows the aggregation of multiple drives' performance for striped and EC files. For the latter, the client caches entire rows and then computes the parity when the row is flushed.
Sequential reads are also served from the cache, which is filled by an adaptive pre-fetcher that detects one or more sequential streams and then fetches a number of chunks in parallel, depending on the speed of the consumer (application) and how fast the drives deliver the data.
For distributed access to files, Quobyte supports the so-called NSF "close-to-open" and "unlock-to-lock" consistency. Clients flush all data to disk before returning from the file close or unlock call. Similarly, fsync/fdatasync calls trigger a flush of the cache to disk before they return unless this behavior is disabled in the policy engine.
In addition to the caching in the Quobyte client, data can be cached by the kernel's page cache. These are flushed on open if the file has changed since the last access to ensure close-to-open consistency.
Page cache flushing can be disabled for files or volumes with read-only data. While this breaks the NFS consistency guarantees, it allows the kernel on the client machines to serve data locally from the page cache without network communication. This is particularly beneficial for volumes with application executables.
For small file workloads designed for local file systems, like compiling, Jupyter Notebooks, or git, Quobyte offers a deferred writeback mode. When enabled, close calls return immediately, so single-threaded applications do not have to wait for the writes to be flushed to disk. The policy engine can selectively enable deferred writeback.
Metadata Caching on the Quobyte Native Client
File metadata is cached on the client for a configurable period of time. The policy engine can control this time-to-live for positive and negative entries separately. Changes to the time-to-live are propagated to the clients automatically. Metadata caching can be disabled completely on a per-volume or per-client basis.
For directory operations, Quobyte clients support the NFS close-to-open consistency on directory listings.
Automatic File Prefetching on the Quobyte Native Client
The policy engine can be configured to automatically detect files being accessed in a sequence and then prefetch files based on that sequence. This feature is particularly useful for video playback, where each frame is stored as a file (DPX), or machine learning, where masking latencies is crucial.
Caching on the Quobyte Servers
The Quobyte Data service uses the local page cache for file reads. Subsequent accesses to the same file will be served from the page cache. However, writes are not cached in the page cache or the device's volatile caches. This ensures consistent performance during write operations without the bottlenecks of caching devices or unpredictable IO stalls when caches run full. Not relying on write-caching devices also significantly improves read performance since reads can go directly to the data block without requiring a lookup on caching devices, which quickly becomes a bottleneck for read and write operations.
The Quobyte Metadata service extensively uses the page cache for volume data (directories, etc.). This service benefits from large amounts of RAM to speed up read operations like stat and readdir.
Predictable Performance
The Quobyte client and services preserve IO patterns: Large sequential IO from applications is potentially batched and sent to the drives as large sequential IO. This means that applications with this IO pattern, like video streaming, traditional HPC, or big data analytics, can take full advantage of HDD streaming performance. This contrasts some systems that turn sequential IO into random 4k writes across the cluster.
For random IO and small file operations, Quobyte disables caches and does not artificially limit the concurrency with caching devices that pose a bottleneck. Instead, Quobyte preserves concurrency over the network and directly on the devices to maximize IOPS and reduce latencies.
Quality-of-Service (QoS) and Fairness
Quality of service and fairness are implemented in all layers of Quobyte: The native clients, the Rapid RPC layer on both the client and server-side, and inside the servers for queuing and IO scheduling.
For QoS, Quobyte offers four priorities. Three that users can use are high, normal, and low - and background, which is used for routine maintenance tasks that are not critical. The three QoS priorities work as expected: the highest priority "overtakes" requests from the lower priorities. Background priority is implemented so that it only gets executed when drives and RPC connections are idle.
To avoid situations where a single distributed application or user can "monopolize" a cluster and cause slowdowns for all users, Quobyte has implemented fairness for metadata and IO operations:
- IO operations from individual files are scheduled fairly on the RPC and disk level
- Small and large IO operations inside the same file are executed fairly to ensure that large IO does not starve small IO
- Metadata operations are executed fairly on the metadata service based on the user's identity
When a drive, metadata, or the entire cluster receives more requests than the hardware can handle, the Quobyte services will use the backpressure mechanism in the Rapid RPC protocol to slow down the clients. This is done – again – in a fair way to ensure that all clients are still able to execute IO. As a result, Quobyte exhibits "graceful degradation" in such overload situations instead of the typical NFS lockups.
Exceptional Scale-out-Performance: Enabled by Quobyte's Unique Architecture
The performance in Quobyte results from a carefully orchestrated architecture that avoids bottlenecks along the entire path from the application, over the network, down to the physical drives. This section provides a summary of how the individual features and architectural design decisions mentioned in previous sections result in Quobyte's exceptional performance and linear scalability:
Direct communication between Native Clients and Quobyte Servers avoids the typical problems of architectures with gateways, such as NFS-based storage: The gateways introduce an extra network hop to get to the actual data, which increases latency. More problematic is the lack of load balancing - unlike lab setups for benchmarking - client load is never evenly distributed across the gateways, creating bottlenecks when too many clients execute IO simultaneously. In contrast, Quobyte native clients can directly go to the server (or servers) that hold the data, which minimizes latency and avoids any bottlenecks.
Parallel IO inside individual files can aggregate the performance of up to 12 servers and NVMe disks in the same file for throughput of up to 20GB/s and single stream read performance of up to 10GB/s on a single file. With full-mesh communication between all Quobyte clients and servers, demanding applications, like AI training, can easily harness the full performance of the entire hardware (servers, network, disks) when reading just from a handful of files or a large number of files.
Quobyte's modern, high-performance RPC protocol, which uses multiple connections per device from each client, helps maximize performance over TCP and can take advantage of RDMA when available for even lower latency and higher throughput.
True scale-out with linear performance scaling is possible due to Quobyte's decentralized policy engine, peer-to-peer primary coordination, and a protocol without gateways. This includes avoiding scalability bottlenecks such as central lock services or deduplication registries.
No block allocation on the global level: When block allocation is managed globally, i.e., not locally by the servers that manage the device, it quickly becomes a massive performance bottleneck and sometimes even a problem for availability when locking is involved. Quobyte uses a hierarchical model where the Metadata Services only assign a file to several drives, and block allocation happens locally on the device when clients write data.
Optimizations for files, such as reads from multiple replicas and deferred writeback, improve the performance of latency-sensitive small file workloads.
A single-layer shared-nothing architecture provides minimal latency, direct-to-disk access with a single network hop, and avoids the bottlenecks, inconsistent performance, and write amplification of write caches.
State-of-the-art Security for your Storage
Quobyte has a comprehensive set of security features that go well beyond the current limited features in NFS-based storage solutions, like true end-to-end data encryption. The multi-layer approach puts multiple fences around your data to ensure maximum security.
User Credentials
On the management layer, Quobyte supports multiple sources for user identities at the same time: The internal database, Active Directory, LDAP, OpenID Connect (OIDC) and OpenStack keystone. Multiple AD/LDAP backends and domains are supported concurrently. Single sign-on is supported via OpenID Connect.
The mapping of user identities, LDAP group memberships, and LDAP groups to Quobyte roles can be configured. In addition, LDAP/AD data can be augmented with information from the internal database to store additional Quobyte-specific information.
On the file system layer, Quobyte operates with with user names and group names instead if host-local uid/gid, similar to NFS v4. Numeric uids and gids are translated to the string representation on the Quobyte Native Clients (or the NFS Gateways) using the operating system's mechanism, e.g., pam/sssd in Linux (which is typically backed by a directory service like LDAP/AD). The clients are also responsible for translating back to numeric IDs when necessary, e.g., for a stat or readdir operation. Quobyte supports a large number of groups, both from the native clients as well as gateways. The NFS gateway automatically looks up user groups from LDAP/AD and isn't limited to 16 groups for NFSv3.
This approach has a number of benefits compared to the old-style numeric uid/gid usage:
- Users and groups are treated the same across all platforms/operating systems as long as their names are the same
- Each tenant can use their own LDAP/AD domain or other source of user identities; the Quobyte cluster does not need to know or interact with their user identity source. This is a prerequisite for storage-as-a-service
- Metadata services do not need to interact and wait for external user identity sources to make decisions about access control
For S3-based access, each Access Key belongs to a file system user. Information such as tenant and group membership for that user is retrieved from the user identity source(s) configured for the management layer.
All operations on the file system are executed by the S3 Gateway on behalf of the file system user. Normal access control, like permissions and ACLs, are enforced for S3 access as well. Files created from S3 have the proper user and group set as owners.
Access Control Lists
Quobyte unifies the access control from all supported protocols and platforms (operating systems). Instead of managing separate permissions and ACLs for Linux and Windows, Quobyte automatically translates to and from NFS4, POSIX, Windows, S3, and macOS to the internal representation of file and directory ACLs. Administrators only need to manage a single ACL per file/directory/volume, which reduces management overhead and removes the security risk of diverging permissions for different platforms.
Quobyte ACLs support multi-tenancy so that file, volume, or directory access can also be granted to users from other tenants.
File System Authentication
Services and clients can be authenticated via various mechanisms. Authentication, in this sense, means that a service is allowed to join the cluster, and a client is allowed to access the cluster's volumes.
In addition to the traditional machine-level authentication, Quobyte supports per-mount and per-user authentication for environments with untrusted or shared client machines.
IP Filters and Trusted Networks
These mechanisms are the traditional way of restricting access to a storage system and are based on IP networks. The assumption is that the network is trusted, i.e., a bad actor cannot easily put a machine into a network, and the root user of hosts is trusted (as the user identities from the host is trusted).
Trusted IP Networks in Quobyte can be defined as one or more IPv4 or IPv6 networks. Any component (Quobyte service, Gateways, or native Clients) connecting from a trusted network can join the cluster as a service and has access to all tenants and volumes. This overrides IP filters and Access Keys but not X.509 certificates. Instead of Trusted Networks, X.509 certificates can be used to authenticate Quobyte services (see below).
Tenant IP Networks define which clients belong to a tenant based on the IP networks. A client can belong to multiple tenants at the same time.
Volume IP Networks can be used to limit or expand the visibility of a volume. Read-only access for certain networks can be defined as well. In combination with policy rules that match specific clients based on IP networks, additional security settings like root permissions can be configured.
X.509 Certificates
X.509 certificates are a good option when the network or hosts cannot be trusted, e.g., when users have root access on their machines. This authentication mechanism can be combined with the other mechanisms.
Quobyte does not require any specific information in the certificate; it just has to be a valid certificate that can be verified locally using the installed CA certificates. Restrictions can be assigned to a certificate in Quobyte. These restrictions can be changed in real time, and certificates can be revoked. Restrictions include:
- Access to one or more tenants
- Operations (IO and metadata) limited to one or more users and/or groups
- Access is limited to a list of volumes
- Read-only access
- "Root squash"
Access Keys
This authentication mechanism is more granular than the previous two. Quobyte supports user sessions where each user has to authenticate to the system with a valid Access Key and Secret Key (same as S3 credentials). The client then only grants access to the tenant(s) of which the user is a member. This mechanism allows multiple users from different tenants to securely access the Quobyte storage cluster from the same client machine.
In addition to authenticating a user, this mechanism also enforces that all IO issued by applications belonging to the session is mapped onto the user's uid and gid. This is particularly useful in containerized environments where IO can be issued by any uid/gid—including root. This authentication mechanism is automatically used by Quobyte's CSI plugin; see the section Kubernetes for more details.
End-to-End Data Encryption
This is a stronger form than at-rest and in-transit encryption. With end-to-end data encryption, file content (data) is encrypted directly on the machine where the user/application generated the data. The Quobyte Native Client encrypts the data after receiving it from the kernel using symmetric encryption (AES-XTS). This encryption is hardware accelerated and does not incur a measurable performance penalty.
The client then sends only encrypted data over the network, which is stored in encrypted form on the Data Services. It is never decrypted until a user or application reads the data again. In this case, the data is sent to the client in encrypted form and is only decrypted before being handed over to the kernel.
This strong form of end-to-end protection is integrated into the Quobyte Rapid RPC protocol and cannot be provided by systems that use NFS. Quobyte's end-to-end data encryption yields a number of significant benefits:
- With data only ever decrypted on the clients, the network and the storage servers do not have to be trusted. Even system administrators cannot access data when logging into the servers, accessing the local data, or logging network traffic.
- Data is only ever stored in encrypted form on the drives. An attacker cannot restore any relevant data from a drive or server they obtained
- Per-file keys are automatically distributed (via TLS connections) to clients upon file access. Thus, there is no need to deal with self-encrypting drives and unlocking them on boot.
- The encryption works with any drive, regardless of type and firmware.
- AES encryption is hardware accelerated on all modern CPUs, and data is en-/decrypted when the CPU has already moved it, which makes the computation overhead negligible.
In-transit Encryption with TLS
TLS encryption of communication is orthogonal to end-to-end data encryption. Both can be used at the same time or separately.
Quobyte allows TLS to be enabled for specific networks, e.g. for fast cleartext communication inside the GPU cluster and secure communication via TLS for "the outside world". When enabled TLS communication is enforced between clients and servers, as well as between servers.
The TLS support in Quobyte is particularly handy when high-performance data synchronization or transfers are required between clusters. Communication between Quobyte servers is direct all-to-all with many TCP connections for each device, similar to client-to-server communication. This distributes the load for encryption and cryptographic checksums required by TLS among all Quobyte servers, yielding far higher throughput than a VPN gateway can provide.
Quobyte's Storage-as-a-Service Design
Quobyte has been designed from the start as a system to provide Storage-as-a-Service to internal or external customers and users. Users benefit from a cloud-like experience with instant provisioning, and administrators can take advantage of automating all storage-related tasks. To enable Storage-as-a-Service, Quobyte has carefully orchestrated core features:
Multi-Tenancy
Multi-tenancy provides strong and secure isolation of multiple tenants on the same Quobyte cluster. This feature can be used to isolate different departments, research groups, or external customers from each other. To each tenant it looks and feels like they had their own Quobyte cluster.
Tenants are completely isolated on the logical storage level. Each tenant has their own namespace for volumes, user and group identities, and quotas. The tenants can use their own LDAP/AD, which doesn't have to be connected to Quobyte.
Monitoring is available to tenant admins as well, including visibility into all clients connected to a tenant and the top IO metrics.
Multi-tenancy is supported across all access protocols (NFS, S3, Hadoop), Quobyte Native Clients, and management levels — the API and Quobyte webconsole support multi-tenancy out of the box.
Clients are assigned to tenants based on IP networks or one of the other authentication mechanisms. Volumes can also be shared with users outside the tenant with IP access lists, and ACLs support users from other tenants and AD/LDAP domains.
While the logical storage level, i.e., the user-visible storage, is isolated, the hardware usage is shared by all tenants by default. However, the Quobyte Policy Engine allows for easy hardware isolation.
Hardware Isolation
Tenants can easily be isolated in hardware using the Quobyte Policy Engine: Each tenant's data and metadata are stored exclusively on the dedicated drives or servers assigned to them. This creates perfect performance isolation from noisy neighbors for both network, CPU, and IO. Giving tenants performance guarantees is easy when they use dedicated hardware.
For administrators, hardware isolation via policy engine configuration is far easier to manage than dedicated storage systems. Since Quobyte supports non-disruptive changes to policies and data placement, they can easily shift resources between tenants or move tenants to/from a shared hardware pool to dedicated hardware when priorities change. These tasks can be done with a few clicks or automated via the Quobyte API.
Similarly, the policy engine allows per-tenant configurations to optimize volumes for the specific needs of their applications or use cases.
Chargebacks and Billing
Real-time usage information is available for every tenant separated by flash and HDD tiers, which can be retrieved via the API. Similarly, quota settings and usage can also be retrieved for all tenants via the API. This information can be used for detailed usage-based billing.
IT managers can implement multiple service levels for billing, like dedicated hardware, all-flash, or a mix. These service levels can easily be translated to policies and quotas in Quobyte and changed with an API call when customers change the service level.
The per-tenant monitoring includes performance metrics such as bytes read/written or read/write operations and metadata operations. These metrics can be used for additional cloud-like billing per million IO operations.
Oversubscription
Tenants are limited by their tenant quota, which limits the number of volumes, files, and logical or physical storage (see Quotas). While a tenant can create arbitrary user, group, and volume quotas, the tenant quota limits everything inside the tenant.
Administrators can use oversubscription for tenant quotas. When these are oversubscribed, and tenants run on shared hardware, administrators can achieve cloud-like cost efficiency.
The File Query Engine: The Filesystem as a Database for AI data storage and other applications
The Quobyte File Query Engine takes advantage of Quobyte's high-performance key-value store, which is used inside the Metadata Services. File system users and administrators can run SQL-like queries over file metadata - including custom metadata in extended attributes - across volumes or the entire cluster.
The File Query engine runs queries in parallel on all available Metadata Services and can scan and filter billions of files in minutes. The query performance automatically grows with the number of Metadata Services, so the query engine scales in lock-step with metadata growth. Results are streamed back to the consuming application/user and generated at the speed of the application.
Unlike many other storage systems or management software that create a copy in a separate traditional database, Quobyte runs the queries on the live metadata. As a result, Quobyte reduces RAM and CPU usage since no additional index needs to be kept in sync and paged into RAM while also delivering results on the latest data rather than outdated snapshots.
The File Query Engine uses the same file system permissions and ACLs that govern file system access. Regular users can query file metadata and only see results they have permissions for. Administrators do not have to manage separate security and access settings for the Query Engine - they can make this feature available to all users without any configuration.
The File Query Engine supports complex queries over a number of file system metadata that can be broadly classified into three categories:
POSIX File System Metadata includes the typical user-visible metadata such as file name, path, file extension, a/m/c/creation time, file size, number of links, and type.
Internal File System Metadata. This category includes Quobyte internal file metadata, such as file immutability, whether a file is encrypted, how it is protected, which devices it is stored on, what media type the file data resides on, etc.
User-defined Metadata. Users can run queries over extended attributes of a file (including those stored via S3 user-defined metadata). Besides testing for the existence of a key (xattr name), queries over extended attributes are dynamically typed and support both numeric and string comparisons. Users can easily use the file system's extended attributes and the query engine instead of external databases or tiny metadata files, as is often true for AI training data. Not only does this remove the problem of keeping the two in sync, but also improves storage performance and usability. This, together with the massive parallel performance, makes Quobyte the ideal AI data storage solution.
An example of a query to get all JPEG files with a width of 1024 pixels or larger (stored in xattrs) would look like this:
qmgmt query files "name~=\".*(jpeg|jpg)\" AND xattr.width>=1024"
Users can run queries via the qmgmt command-line tool, which communicates directly with the API. They can authenticate using the same credentials they use for the Quobyte Webconsole or their S3 credentials. Administrators do not have to set up extra accounts for regular users to take advantage of this fast query mechanism, which is significantly faster and more flexible than running find commands.
Data Management Services
Quobyte Volume Mirroring: Near-Instantaneous Geo-Replication
Volume mirroring allows a source volume to be replicated to one or more remote volumes. The remote volumes subscribe to the source volume's event stream and pull all changes as fast as possible, in most cases in near-real time.
The data transfer from the source to the remote volume is highly parallelized; the Data Services directly fetch the data from their counterparts on the source cluster. File transfers are scheduled based on the size of the files to ensure maximum performance for small and large files. As a result, Quobyte clusters can transfer thousands of files in parallel across all servers, taking advantage of high bandwidth, such as dedicated links between the data centers.
The remote volumes can have different configurations regarding policies, such as file layout and snapshots. Clusters don't have to be identical in hardware configuration or size.
The remote volumes are read-only with volume mirroring until the mirror is terminated. The Quobyte Data Mover can be used to synchronize volumes that are writable in both clusters.
Quobyte Data Mover: Policy-based data synchronization, recoding, and movement
The Quobyte Data Mover is a highly efficient, parallelized, distributed file movement engine controlled by policies. The data mover is a very versatile service that can be leveraged for a multitude of use cases, including:
- Copying files to another Quobyte cluster without expensive and slow rsync tree walks
- Move files to another Quobyte cluster with potentially different hardware configuration, e.g., an offsite archival cluster with HDD
- Bidirectional volume synchronization, where both volumes can be used for reading and writing, is similar to eventual consistency in object storage. Typical use cases include input or AI training data or home directories
- Recoding of files in the same volume, e.g., to recode large files with a wider erasure coding schema for archival
- Hydrate temporary cloud clusters. When a public cloud is used for burst workloads, the Quobyte cluster on the cloud can pull the data and signal once staging is done via the API.
The data mover runs across all Metadata and Data Services and executes the file movement operations where the data is located. Like volume mirroring, the data mover schedules large and small files concurrently to ensure maximum performance despite high-latency links across geographically distributed clusters.
When recoding files or writing files to a remote cluster with different protection, the data mover computes strong checksums on read and verifies the data separately after write.
The selection of files to copy/recode/move is done with policy filters based on criteria such as file age, last access time, size, name, or subdirectory.
Data mover tasks are regular maintenance tasks that can be started, canceled, and monitored via the Quobyte Webconsole or API. The latter is particularly useful for automation, e.g., when staging data on cloud clusters.
External Tiering to Object Storage
Quobyte supports transparent internal tiering between any number of tiers, including cost-effective high-density HDD servers. However, in some circumstances, external tiering can be necessary, such as on the cloud, where block storage pricing is artificially inflated, or when an off-site copy on a separate storage system is desired.
Quobyte leverages the Data Mover for tiering and copying to external object storage, which can be S3-compatible or Azure Blobstore. Files can be tiered out and pulled back in based on policies. The policy filters include whole volumes or subdirectories, filenames or extensions, age, size, last access, etc.
When a file is tiered out, it stays as a so-called stub on Quobyte. It appears as a regular file when listing directories or doing stat calls. However, accessing data is either denied or, when allowed by policy, the Quobyte Native Client directly fetches the file data from the object storage for reading. Files are only tiered in explicitly when a data mover task is started with a policy filter that matches the file (and potentially a large number of files).
The data mover tasks to tier files in/out and can be started by admins, and based on permissions, also by regular users. Multiple object storage backends can be used in the same volume, e.g., a regular S3 bucket and another for archival storage, such as AWS Glacier.
In addition to tiering, the Quobyte Data Mover can also create an independent copy on S3. In this case, each object gets the full file path and name as the key. The copy on S3 can be used and accessed without Quobyte, making this feature particularly useful as a backup and ransomware protection.
Hydrate from Object Storage
Similar to copying volumes and files into an external object storage, Quobyte can copy data from an object storage (S3 or blob-store compatible) into a Quobyte volume.
The copy functionality supports prefix matches and multiple buckets and external storage sources, allowing admins and users to quickly create high-performance Quobyte volumes from existing data on object storage. Typical use-cases for this includes re-creating a Quobyte volume from data stored in an object store, or to create volumes for AI training with custom subsets of object data.
Volume Reports
Volume reports provide detailed information on file age, size, last access, and media types used, among other things. They are generated in real time using the Quobyte File Query Engine and give users, administrators, and management a graphical overview of their data.
Volume reports can be generated for individual volumes, tenants, or the entire cluster.
Snapshots
Snapshots can be generated per volume, either manually by admins or automatically by the policy engine. The schedule and retention period can be configured as a policy at the volume, tenant, or global level.
Snapshots are accessible by users through a .snapshot directory on all platforms and interfaces. Access to the files in a snapshot is governed by the ACLs that we set on the files and directories at the time of the snapshot.
File Retention: Immutable Files and Automatic Expiration
Through the policy engine, files can automatically get a retention policy after having been created and closed by the application. The mechanism supports a number of modes that can be combined:
- Immutable: The file cannot be changed in any way but can be deleted
- Retain until: The file cannot be deleted but can be modified
- Delete after: The file will be deleted after the expiration time
- May extend: The retention period can be extended - but not changed otherwise or removed
- May shorten: The retention period can be shortened - but not changed otherwise or removed
The duration for the immutability settings can be controlled by the policy engine based on a default retention period or from the user/application by setting the timestamp via an extended attribute. For automatic expiration, the file will be deleted after the retention period/timestamp.
Immutability properties cannot be changed after they are activated on a file, not even by a file system administrator. This provides proper WORM semantics and is also an effective protection against ransomware.
Fast Delete
Files can be deleted more efficiently than a "rm -r" using a file deletion task. Like the data mover, this task can delete all files matching a filter in a volume or subdirectory. Admins or regular users can start file deletion tasks.
File System Event Stream
The file system event stream publishes all file system events such as file creation, deletion, access, or metadata change via Kafka.
Non-disruptive Operations and Automated Maintenance
One of the core principles introduced in the first section of this whitepaper is non-disruptive operations coupled with highly automated self-healing and maintenance by the Quobyte cluster. These two are the keys to scalable operations, where small teams can run massive storage infrastructures. This section will look at how this works on a more technical level.
The ability to handle all kinds of changes to the hardware, software, and configuration non-disruptively is mostly a result of how Quobyte's architecture is designed around the "anything can fail" principle and the way the quorum replication allows automatic failover and replica additions/removals without interruption.
A Quobyte cluster has multiple components that keep the system healthy and optimized:
The Quobyte Health Manager
The Health Manager is a distributed and fault-tolerant service that constantly monitors the cluster in terms of device and server health, network and service availability, data redundancy, and other parameters. When necessary, it will start maintenance tasks to re-protect data when a drive or server is lost, run rebalances, or perform automatic updates, among other things. The health manager's behavior is controlled through policies that administrators can adjust to their needs and operational procedures.
The Health Manager is also responsible for scheduling maintenance tasks and prioritizing them to ensure critical tasks like re-protection can be completed as quickly as possible. When scheduling non-critical tasks, the health manager also limits the number of tasks running per volume or device to avoid excessive IO and CPU usage.
Administrators can also decide to run non-critical routine maintenance tasks during certain periods of the day, such as data rebalances or scrubs, at night or on weekends only. Tasks that are limited to certain times will be paused if they don't finish within the time window and automatically resumed during the next window.
Maintenance Tasks
Most operations that operate on metadata and/or data are implemented as Maintenance Tasks in Quobyte. These tasks are run distributed across the services that have the data, e.g., the Metadata or Data Services, and take advantage of being highly parallelized.
Tasks can be automatically started by the Health Manager (see next section) or administrators. Similarly, administrators can cancel, pause, and resume all tasks.
All tasks provide progress reports in the Webconsole or command line. Task progress and completion can also be checked via the API for easy automation.
Hardware Maintenance at Will
Administrators and hardware operators have the ability to switch off any Quobyte server at any time without having to "notify" the Quobyte system beforehand. This activity will not impact users and applications as long as the cluster is healthy.
Administrators can configure the time window for how long a server can be offline as a health manager policy. Since Quobyte uses redundancy mechanisms where losing one server/drive/failure domain is not a problem, this period can be multiple hours since at least two more copies/stripes are available. Once the drive(s) and server(s) return online, a maintenance task will be run to bring the local replicas/EC stripes up to the latest state.
If a server is offline for too long, for whatever reason, a regenerate task will be started to re-protect the affected files by adding a new replica or regenerating data/parity stripes. The cluster will automatically return to a healthy state, even if an administrator forgets to return a server online after maintenance. Similarly, Quobyte forgives operator mistakes such as putting drives back in the wrong slot or server.
When failure domains are configured and used for replica placement, administrators can even shut down an entire rack of Quobyte servers simultaneously for maintenance.
Rolling Software Update
The Health Manager can run automated updates of the Quobyte software on the servers. Updates are only done when the cluster is in a healthy state. When enabled and a new release is available, the health manager updates one server at a time and waits for the service to be fully restarted and caught up before going to the next server. Updates are done using the system's package manager, e.g., dpkg or yum. Administrators can disable the automatic update policy and do them manually, pin releases using the package manager, or run more sophisticated updates, such as per-rack rolling updates, with Ansible playbooks.
Live Migration: Hardware Refreshes without Downtime
Unlike appliance-based storage solutions, Quobyte enables administrators to swap out the complete hardware without any downtime, e.g. to use more efficient new servers or to achieve higher performance.
A complete (or partial) hardware swap can be done with a few simple steps:
- Rebalances are temporarily disabled in the health manager. This step isn't required but avoids superfluous data movement through rebalances.
- The new servers are added to the cluster. They are used for new files immediately.
- The old servers are drained. Quobyte transparently moves files to the new servers. In addition, new files will not be created on servers in drain mode.
Once the drain tasks have been completed, the old servers can be removed from the system.
Similarly, only certain roles can be removed from servers, e.g., to introduce dedicated metadata servers for lower latency.
Optional Rebalancing
Data and volume rebalancing aren't a requirement in a Quobyte cluster but are recommended to keep it well balanced. Rebalancing is desired in most cases when adding new hardware unless the goal is a hardware swap.
The health manager has policies for data and metadata (volume) rebalancing that allow administrators to define how balanced the cluster should be and when rebalances are executed.
When rebalance maintenance tasks are running, they run with background priority and cause minimal impact on HDD and no impact on NVMe performance.
Automation and Monitoring
Quobyte has been designed with the Infrastructure-as-Code principle in mind: highly automated and scripted installation, updates, configuration changes, day-to-day operations, and highly effective distributed monitoring.
API First
Quobyte's Webconsole and command-line tool qmgmt both use the API to communicate with the Quobyte cluster without any hidden operations or side channels. The JSON REST API can automate every task that administrators or users do via Quobyte's UI or CLI.
Like the Webconsole or CLI, the Quobyte API is ready for multi-tenancy.
Users can authenticate using their LDAP/AD username and password or S3 access and secret keys.
Zero-Configuration Services
As mentioned in the section on Quobyte Services, they do not require any local configuration besides the location of the Registry Services. This also includes the drives, which are automatically mounted by the services - no need to configure fstab.
As a result, Quobyte services can be deployed easily on any Linux system with simple tools or Ansible playbooks. Options like cloning system drives or PXE-booting Quobyte servers are also possible.
Ansible playbooks for deploying and updating Quobyte clients and services are available from the Quobyte github account.
Real-time Analytics and Performance Monitoring
Quobyte takes advantage of its native clients running on the machines where the applications are running: The native client collects IO metrics per file and aggregates it by application. These IO metrics include the IO executed by the application and the IO that is sent over the wire. The top IO activity is exported by each client and aggregated by the cluster.
In the Quobyte webconsole, admins get an overview of the top 100 consumers of storage and metadata IO on a per-process or Slurm/PBS job level. Administrators and tenant admins can drill down into per-client metrics and even analyze the IO on a single-file level to analyze issues or understand application IO patterns.




Quobyte also provides proactive alerting based on client-side observations to notify admins of performance issues before users notice. Examples include applications that exhibit inefficient IO patterns, such as random overwrites on erasure-coded files or concurrently appending to a single file. The system will also alert admins to applications and clients whose IO is stuck because devices are too slow, network communication is slow/unreliable, or the user is out of quota.
In addition to performance monitoring, Quobyte provides real-time usage statistics by tenant, user, group, and volume for file count and capacity usage by tier and external storage.

Monitoring with Prometheus and Grafana
Long-term and more detailed performance monitoring is provided through Prometheus exporters and implemented in all Quobyte Services, Native Clients, and Gateways. All components export a wide variety of performance and internal metrics that can be used to check the system's health and overall performance.
For simplicity, Quobyte registries implement the Consul API for automatic service discovery. Prometheus—or any other monitoring system that understands Prometheus data formats—just needs to be pointed to the cluster's registry, and it will automatically detect and scrape all services and clients. This list is dynamically updated when clients or services leave or join the cluster.
Quobyte provides a Grafana dashboard that delivers crucial metrics like percentiles for IO and metadata operations—both measured on the clients and services—as well as long-term trends.

Integrations
Quobyte and Kubernetes (k8s)
There are two independent ways in which Quobyte can be used for Kubernetes. On the one hand, Quobyte volumes can be used by applications running on Kubernetes using persistent volumes. On the other hand, Quobyte Services can run on Kubernetes, providing volumes to the Kubernetes cluster and clients outside of Kubernetes.
These two options can be used independently, i.e., an external Quobyte cluster can provide storage to Kubernetes applications and vice versa, or together, where the applications and Quobyte Services run side-by-side on the Kubernetes cluster.
Providing Storage to Kubernetes: The Quobyte CSI Plugin and the Quobyte Native Client
The Quobyte CSI plugin manages persistent volumes (PVs) and maps persistent volume claims (PVCs) onto Quobyte volumes. It is available as a container image (see Quobyte container images) and can be installed easily with the Quobyte helm chart.
In addition to the CSI plugin, the Quobyte Native Client must run on each Kubernetes worker node. Similar to the CSI Plugin, the native client can be deployed and updated across the entire cluster using the Quobyte client helm chart. The native client runs inside a container and exports the mount via a bidirectional volume to the host operating system namespace.
When the Quobyte CSI Plugin fulfills a PVC request, it maps the volume from the local Quobyte client mount point into the pod's container. If a new volume is requested, the CSI plugin creates it and assigns it a quota if it has a size.
Applications can take advantage of Quobyte's scale-out performance. Quobyte-backed PVs are RWX (read-write-many) and can be accessed concurrently from thousands of pods. This makes Quobyte persistent volumes particularly useful for applications like AI training that require large amounts of data and files to be accessible with high performance from thousands of GPU worker nodes.
Multiple StorageClasses are supported: Each StorageClass can be configured to attach one or more labels (key-value pairs) to the corresponding Quobyte volumes. Administrators can then configure policies in Quobyte to match the labels, e.g., create different volume types such as "all-fash" or "archival".
Kubernetes namespace can be mapped onto Quobyte tenants to ensure proper isolation between tenants inside and outside the Kubernetes cluster.
The CSI plugin supports volume snapshots, and a pod-killer is provided to terminate pods when updating the client image
Providing Storage to Kubernetes: Secure Access Control on Shared Volumes
The Quobyte CSI plugin enforces user authentication and strict mapping of uid/gids to ensure proper access control on shared file systems. When a user submits a PVC, they must provide a valid Access/Secret Key pair as credentials. The CSI plugin uses these to authenticate the user with the Quobyte cluster.
The Quobyte native client then maps all IO coming from the containers/pods accessing the PV onto the user's real uid/gid, regardless of what uid/gid is sent from the container. This solves a typical security gap when using shared file systems with containers: Inside the container, the user has full control over user identities, including root. When executing file system operations these numeric uids/gids are passed on without any translation or filtering, essentially circumventing any file system access control.
Running Quobyte Services on Kubernetes
Quobyte services run in privileged containers and automatically detect and mount Quobyte devices. This is very similar to the systemd version of the Quobyte services, which also run in cgroups for resource isolation.
The services are available as a single container image on quay.io and can be deployed as a StatefulSet with the Quobyte Helm chart. Each service type, including the gateways, webconsole, and API, can be started from the same container image.
There are a few caveats to consider when running Quobyte Services on a Kubernetes cluster:
- Guaranteed Resources. Unless you want to assign dedicated worker nodes to Quobyte, you must ensure that resources aren't oversubscribed on shared worker nodes. In particular, sufficient RAM must be available to the Quobyte Services so they do not get killed by the OOM killer when other applications consume too much RAM. For an in-depth discussion about running Quobyte Services on shared hosts, see this article on our blog.
- DNS records. Quobyte clusters on Kubernetes use the internal DNS to resolve registry records. When accessing the cluster from hosts outside of Kubernetes, the k8s internal DNS records must be resolvable outside the cluster
Running Quobyte on GKS, EKS, and other public clouds
When running on VMs, the Quobyte Services require persistent block storage volumes instead of physical drives. Persistent volume claims are included in the StatefulSets for the Registry, Metadata, and Data Services. Please refer to our tutorials on how to run Quobyte on the Google Cloud Kubernetes Engine (GKE) and the Amazon AWS Elastic Kubernetes Service (EKS).
Quobyte for OpenStack
Quobyte offers drivers for Nova, Cinder, and Glance. For a detailed discussion of the Integration, please see the Quobyte whitepaper on OpenStack.
