Apache Druid is an open-source analytics database for real-time queries on streaming and batch data. It is designed for sub-second analytics and the project says it can execute OLAP queries in milliseconds across high-cardinality datasets with billions to trillions of rows. Native Apache Kafka and Amazon Kinesis integrations support low-latency ingestion and query-on-arrival. Ingested data is organized into compressed, indexed columns, including time indexes and dictionary and bitmap indexes. Users can query with Druid SQL or JSON-over-HTTP native queries; joins are supported during ingestion and at query time. Its web console can load data, manage datasources and tasks, show server status and segments, and run queries. Extensions connect to storage, databases, and formats including S3, HDFS, Azure, PostgreSQL, Avro, ORC, and Parquet. Druid is free under Apache License 2.0 and can be self-hosted on Linux, macOS, and other Unix-like systems; Windows is not supported. The local quickstart requires at least 6 GiB of RAM and Java 17. Security features are disabled by default, so production deployments need TLS, authentication, and authorization configured.
Who it is for
Druid suits teams building analytics workloads that need low-latency queries over streaming or batch data, including high-concurrency and ad hoc exploration use cases. It requires a supported Unix-like environment and production security configuration.
What is good
- Free and licensed under Apache License 2.0.
- Supports streaming ingestion through Kafka and Kinesis.
- Offers Druid SQL and JSON-over-HTTP queries.
- Web console can manage data and run queries.
- Extensions cover storage, databases, and file formats.
What to know first
- Windows is not supported.
- Local quickstart requires at least 6 GiB RAM and Java 17.
- Security features are disabled by default.
- Not commonly used for full-text search over text logs.
Freedom251 review
Apache Druid: the full review
Apache Druid is a self-hosted option for real-time analytics across streaming and batch data, with SQL access, a web console, and broad extension support. Plan for the local system requirements and configure security controls before production use.
Overview
Apache Druid is an open-source database for teams that need fast analytical answers from large, changing datasets. It is strongest in user-facing analytics and interactive exploration, where fresh data and high query volume matter. Its distributed design offers room to scale, but it calls for operational planning rather than a plug-and-play database.
Druid targets millisecond OLAP queries across datasets ranging from billions to trillions of rows, and applications handling hundreds to 100,000 queries per second. Those are stated capabilities, not universal performance guarantees; deployment and workload matter. The project’s latest stable release is 37.0.0, released May 8, 2026, under the Apache License, Version 2.0.
Key features
Streaming and batch analytics
Native Apache Kafka and Amazon Kinesis integrations support low-latency ingestion and query-on-arrival, including ingestion at millions of events per second with guaranteed consistency. Druid also handles batch data, so it can serve workloads that need both incoming events and larger historical datasets. Its source coverage includes streaming systems, object stores, databases, and files.
Storage and query design
On ingestion, Druid turns data into compressed columnar storage with time indexes, dictionary encoding, and bitmap indexes. That structure is designed to accelerate analytical queries, particularly on high-cardinality data. Users can query through Druid SQL or native JSON-over-HTTP requests; joins are supported both during ingestion and at query time.
Scaling, operations, and integrations
Ingestion, query, and orchestration components are loosely coupled, with deep storage supporting scale-up and scale-out. Continuous backup, automated recovery, and multi-node replication provide a foundation for durability and availability, while extensions connect Druid to systems and formats such as S3, HDFS, Google Cloud Storage, Azure, Avro, ORC, Parquet, MySQL, and PostgreSQL.
The web console can load data, manage datasources and tasks, show server status and segments, and run SQL and native queries. That gives operators a central place for common tasks, though it does not remove the need to manage the underlying deployment.
Pricing
Apache Druid: 0.00 USD per free. The open-source analytics database is downloadable for self-hosting and includes real-time ingestion and SQL access. The project and documentation are licensed under Apache License, Version 2.0. This is a self-managed option rather than a managed service plan; deployment and infrastructure remain with the user.
Platforms
Druid supports API, Linux, macOS, self-hosted, and web environments. The quickstart supports Linux, Mac OS X, and other Unix-like systems, but not Windows. A local quickstart needs at least 6 GiB of RAM and Java 17. Druid is also designed for commodity hardware and cloud environments including AWS, GCP, and Azure.
Security features are disabled by default, so production operators need to configure TLS, authentication, and authorization. Documented authenticator extensions include HTTP Basic authentication, LDAP, and Kerberos. The project directs users to Slack and GitHub for community help and names Cloudera, Datumo, Deep.BI, Imply, and Rill Data as commercial support providers.
Who it's for
Druid is a strong fit for teams building user-facing analytics, high-concurrency low-latency query systems, or tools that need immediate visibility into streaming data. It also suits ad hoc exploration over large datasets and organizations that want to scale ingestion and queries across loosely coupled components.
It is less suitable for someone seeking a general-purpose database or full-text search over text logs. The project notes that Druid is not commonly used for full-text log search, though it can ingest and analyze semi-structured data such as JSON. Teams without the capacity to configure security and operate a distributed self-hosted system should weigh that burden before choosing it.
Pros and cons
- Low-latency analytics at scale: Druid is designed for millisecond OLAP queries on very large, high-cardinality datasets and high query concurrency.
- Fresh streaming data: Kafka and Kinesis integrations support query-on-arrival alongside batch workloads.
- Flexible data connectivity: Extensions cover major storage systems, databases, streaming sources, and file formats.
- Self-hosting demands operational work: The distributed architecture and local system requirements call for infrastructure planning and administration.
- Security needs deliberate setup: TLS, authentication, and authorization must be configured before production because security features are off by default.
- Not a typical full-text log search tool: Teams focused on text search should consider a different fit.
Alternatives
Apache Beam is another free, open-source option if you want a programming model for data processing; execution costs depend on the runner and infrastructure you select.
Apache Storm is a free, open-source alternative for readers comparing streaming-oriented projects.
Feldera may suit teams seeking a SQL-focused open-source edition with all connectors and datasets larger than memory, within its single-node, single-container setup.
Timeplus offers the free, Apache 2.0-licensed Timeplus Proton as a lightweight streaming SQL engine distributed as a single binary under 500MB.
Apache Spark is a free distributed data analytics engine if that broader engine model better matches your workload.
Confluent Cloud is worth considering if you prefer serverless, fully managed Kafka clusters and connectors, with a free Basic plan starting at $0/Month.
Materialize offers a self-managed Community License with up to 24GiB memory and 48GiB disk, or a Cloud Capacity plan priced at 1.50.
Apache Flink is a free Apache License v2 alternative for teams evaluating an open-source stream-processing project.
Browse Streaming Analytics Software, OLAP Databases, Columnar Databases, OLAP Software, and Database Software for more options.
Verdict
Choose Apache Druid when your priority is fast, concurrent analytics on fresh streaming data and large batch datasets, and you can operate a self-hosted distributed system. Its combination of real-time ingestion, SQL, indexing, and extensibility is compelling for that workload. Look elsewhere for full-text log search or when you need security configured out of the box or want to avoid deployment and operations work.
Apache Druid plans and pricing
All plansCompared on database software
- Real-time ingestion
- Yes





