HPCC Systems is a free, open-source platform for enterprise big-data processing and data-lake management. Its architecture pairs ECL, a declarative and modular programming language, with Thor for data refinement and Roxie for data delivery. The ECL compiler optimizes programs for parallel processing and compiles them into C++. Thor handles ingestion, transformation, linking, and indexing, while Roxie serves concurrent queries. Clusters can range from two computers to more than a thousand commodity-hardware nodes. The cloud-native platform runs on Kubernetes and supports Azure Kubernetes Service and Amazon Elastic Kubernetes Service. Its storage plane supports AWS S3, Azure Blob Storage, AWS Elastic Block Store, Azure Files, and Azure Disks. Security capabilities include end-to-end encryption, service meshes, OAuth 2.0 with Azure AD support, JWT, and configurable managers for data and services. The Machine Learning Library provides ECL-accessible algorithms for building and testing prediction models. Integrations include Spark, Kafka, Couchbase, Redis, Memcached, Pentaho, R, JDBC, Java APIs, and Tableau data connectors. It is Apache 2.0 licensed and self-hosted.
Who it is for
HPCC Systems suits organizations with enterprise big-data and data-lake needs that can operate a self-hosted platform. It is also relevant to teams using ECL, Kubernetes, or its listed integrations and query interfaces.
What is good
- Apache 2.0 licensed and free to use.
- Clusters can scale beyond a thousand nodes.
- Supports Azure Kubernetes Service and Amazon EKS.
- Includes ECL-accessible machine-learning algorithms.
- Offers XML, HTTP, SOAP, and REST query interfaces.
What to know first
- The platform is self-hosted.
- A supported Linux system is required for a single-node cluster.
- ECL IDE is available for Windows.
- Apple OSX has client tools only.
Freedom251 review
HPCC Systems: the full review
HPCC Systems combines parallel data processing, data-lake management, and a broad set of deployment, security, and integration capabilities. It is free under Apache 2.0, but deployment is self-hosted and a Linux system is required for a single-node cluster.
Overview
HPCC Systems is a self-hosted platform for teams building and operating large-scale data lakes. It suits organizations with Linux and cluster-operations capacity; its breadth is compelling, but it is not a managed shortcut to big-data processing.
The open-source architecture pairs data refinement with query delivery, so teams can prepare mixed-schema data and serve it from the same platform. The Apache 2.0 license keeps software costs at zero, while leaving deployment and operation to the user.
For other data-lake options, browse Data Lake Software.
Key features
ECL and parallel processing
ECL is a declarative, modular language; its compiler optimizes programs for parallel processing and compiles them into C++. That model is suited to teams comfortable expressing data workflows in code, rather than buyers looking for a no-code environment.
Refinement and query delivery
Thor handles ingestion, transformation, linking and indexing, while Roxie serves high-performance concurrent queries. Keeping preparation and serving in distinct engines gives data teams a path from raw inputs to queryable outputs, though it also means operating a multi-component platform.
Cluster and cloud infrastructure
Clusters scale from two computers to more than a thousand commodity-hardware nodes. The cloud-native platform runs on Kubernetes, including Azure Kubernetes Service and Amazon Elastic Kubernetes Service. Its storage plane supports AWS S3, Azure Blob Storage, AWS Elastic Block Store, Azure Files and Azure Disks. These options suit teams that need deployment flexibility, but the self-hosted model still calls for infrastructure ownership.
Security and governance
Cloud-native security includes end-to-end encryption, Linkerd or Istio service meshes, OAuth 2.0 with Azure AD support, and JWT. Configurable security managers can protect landing zones, file scopes, recordset data, workunit execution and ESP services. Metadata catalog and governance controls are also included, making the platform relevant where access boundaries and data organization matter.
Machine learning and integrations
The Machine Learning Library exposes ECL-accessible algorithms for building and testing qualitative or quantitative prediction models. Official integrations include Spark, Kafka, Couchbase, Redis, Memcached, Pentaho, R, JDBC, Java APIs and Tableau data connectors. ESP exposes ECL queries through XML, HTTP, SOAP and REST, and deployed queries can be called with REST/JSON. This range helps connect existing data and application workflows without making HPCC Systems a fit for teams seeking a narrowly scoped tool.
Pricing
Open source: 0.00 USD per free. The plan is Apache 2.0 licensed and self-hosted. It removes a software license charge, not the work of providing Linux systems, cluster infrastructure and operations. No paid tier is needed to use the platform.
Platforms
HPCC Systems supports API, extension, Linux, macOS, self-hosted, web and Windows. A current supported Linux system is required for a single-node cluster; ECL IDE is available for Windows, while Apple OSX support is limited to client tools. In practice, the platform's listed platform breadth does not remove Linux as a requirement for running a cluster.
Deployment is hybrid, with both ingestion modes, both query interfaces and data sharing supported. Public Stack Overflow community support, free online tutorials and free online training give teams self-service learning and community channels rather than changing the self-hosted operating burden.
Who it's for
HPCC Systems is a strong fit for enterprises that need to refine and serve large, mixed-schema datasets, want control over deployment, and have people able to build ECL workflows and run clusters. Its scaling range, cloud storage choices, security controls and integration set address substantial data-lake workloads.
It is a weaker choice for small teams that want a managed service, avoid Linux administration or do not want to work in a declarative programming language. The zero-price license cannot offset the cost in infrastructure and operational expertise.
Pros and cons
- Pros: Thor and Roxie divide data preparation and concurrent query serving, covering two important stages of a data lake in one platform.
- Pros: ECL compiles optimized parallel workloads to C++, and clusters can grow beyond a thousand commodity-hardware nodes.
- Pros: Kubernetes deployment, several cloud storage targets, configurable security managers and a broad integration roster provide meaningful architectural flexibility.
- Cons: Self-hosting and the Linux cluster requirement put infrastructure setup and ongoing operations on the user.
- Cons: ECL's code-centered approach is less suitable for teams that need to avoid specialized programming.
Alternatives
AWS Lake Formation is worth considering for readers who prefer a paid platform with a free permissions plan and use integrated services such as Amazon S3, whose standard usage rates still apply.
HPE Ezmeral Data Fabric offers a free trial and reseller-dependent pricing, including a Canada store starting price of C$500.00; consider it when evaluating a commercially priced data-fabric option.
lakeFS is freemium, with an enterprise plan by sales contact and managed cloud or self-managed deployment choices, including on-premises, own-cloud and air-gapped environments.
Tencent Cloud Application Performance Management is a freemium web option with developer packages that specify trace-storage periods and Agent*Hours.
Unilake is another free option; its maker says it is licensed under AGPL 3.0 and EUPL and can run fully isolated in an environment of the user's choice.
Apache Hadoop HDFS is a free, open-source alternative developed by the Apache Hadoop project under the Apache Software Foundation.
Parseable is a freemium option with web, Windows and Linux platforms.
Cloudera Data Lake Service is a paid option with pricing by sales contact and no free plan.
Verdict
Choose HPCC Systems if your organization needs a scalable, code-driven data-lake platform and can run its own Linux and Kubernetes infrastructure. Its strongest case is the combination of parallel refinement, high-concurrency query delivery and extensive deployment and integration controls; look elsewhere if managed operations or a gentler programming model matter more than that control.
HPCC Systems plans and pricing
All plansCompared on data lake software
- Free plan
- Yes
- Deployment model
- hybrid
- Ingestion modes
- both
- Metadata catalog
- Yes
- Governance controls
- Yes
- Query interface
- both
- Data sharing
- Yes




