中文

Extending Big Data Protection: Scutech DBackup Adds Backup and Recovery for Apache Doris

01 Apache Doris: A Real-Time Analytics Engine for the Big Data Era

Apache Doris is a high-performance real-time analytical database built on a massively parallel processing (MPP) architecture. With sub-second query response across massive datasets, it handles both highly concurrent point queries and high-throughput complex analytics, making it a leading choice for enterprises building real-time analytics capabilities at scale.

Typical Use Cases

01

Real-Time Analytics

Real-time reports, live dashboards, user behavior analysis, A/B testing platforms, and log search and analytics—delivering sub-second interactive experiences directly to end users.

02

Lakehouse Analytics

Unified data warehouse development and accelerated federated queries across data lakes, including direct queries on open table formats such as Iceberg, Delta Lake, and Hudi. One engine supports both real-time and offline analytics.

03

Hybrid Search and Analytics

Integrates full-text search, vector search, and structured analytics to power RAG applications, semantic search, and real-time agent decision-making across the AI data stack.

From internet services to the real economy, Doris is used by industry leaders including Baidu, Xiaomi, BYD, Ford, JD.com, Kuaishou, Meituan, Luckin Coffee, and miHoYo. As their businesses scale, Doris clusters often grow from terabytes to petabytes—a defining characteristic of big data environments. The rising value of data, together with the increasing impact of site failures, is driving organizations to place greater emphasis on data security and business continuity.

Conventional database backup solutions, however, are often ill-suited to the architecture of big data platforms: distributed storage spans many nodes, data volumes are enormous, backup windows are tight, and storage costs are high. As the last line of defense for data security, a backup system must integrate seamlessly with Doris’s distributed architecture and provide fast, precise, and reliable backup and recovery.

DBackup already protects the Hadoop ecosystem, including Hive and HBase, as well as the standalone OLAP engine ClickHouse. Support for Doris extends that protection portfolio into real-time big data analytics.

Apache Doris real-time analytics, lakehouse analytics, and hybrid search workloads scaling from terabytes to petabytes

02 Limitations and Challenges of Native Doris Backup and Restore

Apache Doris provides snapshot-based BACKUP and RESTORE statements for protecting databases, tables, or partitions in a remote storage system. In the large-scale environments where Doris is most often deployed, however, massive data volumes, numerous nodes, and stringent continuity requirements magnify the limitations of the native approach:

01

Single-Task Concurrency Restricts Throughput

Doris allows only one backup or restore job to run at a time within the same database. At scale, this constraint serializes operations, prevents full use of cluster parallelism, and lengthens both backup and recovery windows.

02

Limited Storage Options Drive Up Costs

The native approach uploads backup data to remote object storage such as Amazon S3 or Alibaba Cloud OSS. As backup data accumulates, capacity consumption and storage costs continue to rise, while built-in data-reduction capabilities such as deduplication and compression are absent.

03

Basic Management with Limited Automation

Native statements do not provide job scheduling, recurring execution, or alert notifications. DBAs must write scripts, configure scheduled tasks, and track execution status manually. This is inefficient and increases the risk that a missed or failed backup will go undetected.

04

Complex Cross-Cluster Restore Weakens DR Readiness

Although native restore supports recovery to another cluster, the process involves multiple manual steps, including Repository configuration, snapshot selection, and replica-count settings. Without centralized management or automated recovery orchestration, rapid and accurate service restoration during a disaster-recovery exercise or failover is difficult.

03 DBackup for Doris: Lower Costs with Deduplication, Faster Protection with Multiple Channels

Doris compute-storage integrated cluster topology with Frontend and Backend nodes, a backup server, and a distributed deduplication cluster

Built on the Doris snapshot mechanism and deeply integrated with native BACKUP and RESTORE statements, Scutech DBackup delivers comprehensive backup and recovery protection for enterprise Doris platforms. Its lightweight, agentless architecture protects the entire cluster without requiring an agent on every production node. Backup data can be written flexibly to an object storage pool or an EOBS (Extended Object Storage) deduplicated storage pool, balancing scalable capacity with optimized cost.

01

Broad Version Coverage

DBackup backup and recovery for Doris supports major Apache Doris releases, including versions 2.0, 2.1, 3.0, 3.1, 4.0, and 4.1. From established releases to the latest mainstream deployments, organizations can bring their Doris environments under protection without disrupting the production architecture.

02

Deduplicated Storage Cuts Cost and Improves Efficiency

In addition to third-party object storage, DBackup provides a self-developed, S3-compatible deduplicated storage pool. Backup data is automatically compressed and deduplicated as it is written, eliminating substantial redundancy and reducing storage cost. The pool also supports distributed deployment: data is spread across independent nodes, I/O load is balanced, and protection is provided through multiple replicas and erasure coding. Automatic failover on node faults combines high throughput with high reliability, creating a robust storage foundation for big data backup.

03

Granular, Policy-Based Protection

Backup and recovery can be performed at database, table, or partition level. Frequently changing partitions can be protected more often to approximate incremental backup behavior, while an entire database can receive full protection. The same granularity is available during recovery, reducing elapsed time and resource consumption for both backup and restore.

04

Multi-Channel Acceleration for Higher Throughput

DBackup transfers data across multiple channels in parallel to accelerate backup and recovery. Combined with Doris’s native Tablet-level parallel snapshot mechanism, each Backend (BE) node uploads or downloads the data shards it owns at the same time. The result is true distributed parallel backup and recovery capable of protecting large datasets within the required window.

05

Automated Scheduling Simplifies Operations

DBackup supports fully automated backup plans ranging from immediate execution to recurring schedules, including one-time, manual, hourly, daily, weekly, and monthly policies. Jobs are triggered, executed, and recorded automatically, reducing manual intervention. Operations teams can monitor every backup and recovery job from a unified management console, significantly lowering administrative complexity.

06

Cross-Cluster Restore Accelerates Disaster Recovery

Backup data can be restored to the original Doris cluster or to another cluster. During hardware failure, version upgrades, or data migration, services can be recovered rapidly from a backup set to reduce downtime. Using the available snapshot name and timestamp, the backup host selects the appropriate recovery point and submits the RESTORE statement. Frontend (FE) nodes then assign recovery tasks to the relevant BE nodes, which download snapshot files in parallel and load them locally for efficient recovery.

07

Storage Pool Replication Adds Another Layer of Protection

Storage pool replication provides redundant protection across pools. Backup data can be copied automatically to an off-site storage pool, creating an additional layer of resilience against a catastrophic failure at a single site and further strengthening data security.

Scutech DBackup architecture for multi-channel backup, automated scheduling, original-cluster and cross-cluster recovery, and off-site pool replication

With sub-second query performance and unified real-time analytics, Apache Doris is becoming a core engine in the big data architectures of a growing number of enterprises. Yet petabyte-scale datasets, distributed multi-node deployment, and highly concurrent real-time ingestion make traditional database backup methods inadequate. Purpose-built backup and recovery for big data is therefore essential.

Adding Doris backup and recovery closes another important gap in DBackup’s big data portfolio, which already protects Hive, HBase, HDFS, and ClickHouse. Big data platforms far exceed conventional relational databases in scale, architectural complexity, and operational demands. DBackup’s core capabilities—agentless deployment, EOBS deduplication, multi-channel acceleration, granular protection, automated scheduling, cross-cluster recovery, and storage pool replication—demonstrate that it is well equipped to meet these challenges.

From the Hadoop ecosystem to ClickHouse and Doris, Scutech DBackup continues to expand its support for big data platforms, delivering professional-grade data protection so enterprises can innovate with confidence and without fear of data loss.

Contact