dstack
dstack is an open-source orchestration layer for AI workloads across GPU clouds, Kubernetes, virtual machines, and bare-metal clusters. It provisions infrastructure and schedules jobs, with auto-scaling, port forwarding, and ingress. YAML configurations define fleets, development environments, tasks, services, presets, and volumes. Supported accelerators include NVIDIA, AMD, TPU, and Tenstorrent; documented backends include AWS, Azure, GCP, Kubernetes, GPU cloud providers, remote SSH hosts, and experimental Slurm. Tasks can use frameworks such as accelerate, torchrun, Ray, and Spark. The CLI and HTTP API provide ways to manage resources and connect integrations, and the server can run on a laptop or another environment able to access the clusters in use. Services can expose inference endpoints through gateways with HTTPS, custom domains, auto-scaling, and rate limits. The self-hosted open-source plan is free. Server data and project secrets are stored in plaintext by default unless administrators configure AES-256-GCM encryption. TPU support is limited to single-host instances of up to eight cores.
Who it is for
dstack suits teams managing AI workloads across cloud, Kubernetes, virtual-machine, or bare-metal infrastructure. It may fit users who want YAML-configured workloads and a CLI or HTTP API, provided administrators can address the default plaintext storage setting.
What is good
- Open-source self-hosted plan is free
- Supports NVIDIA, AMD, TPU, and Tenstorrent accelerators
- CLI and HTTP API for resource management
- Inference gateways support HTTPS and rate limits
What to know first
- Server data and secrets are plaintext by default
- TPU support is limited to single-host instances of up to eight cores
- Slurm backend is experimental
Freedom251 review
dstack: the full review
dstack offers a broad orchestration surface for AI workloads, from job scheduling to inference endpoints. Its plaintext-by-default storage and single-host TPU limit are important operational considerations.
Overview
dstack is open-source orchestration software for AI workloads across cloud and on-premises compute. It best suits teams managing mixed accelerator infrastructure that want one YAML-based workflow for provisioning and scheduling. Its breadth is compelling, but plaintext-by-default storage and limited TPU support deserve careful consideration.
The server can run on a laptop or another environment that can reach the clusters and cloud resources in use. Teams can manage fleets, development environments, tasks, services, presets and volumes through the CLI or HTTP API. This is infrastructure software, not a general-purpose desktop or mobile app.
Key features
Provisioning and workload scheduling
dstack provisions infrastructure and schedules workloads across GPU clouds, Kubernetes, virtual machines and bare-metal clusters. It supports both job-style and service workloads, with auto-scaling, port forwarding and ingress. That makes it useful for teams that need to coordinate varied compute in one system; the absence of quota controls is a drawback where administrators need hard resource limits.
Accelerators and frameworks
Out-of-the-box support covers NVIDIA, AMD, TPU and Tenstorrent accelerators. The task guidance names Accelerate, torchrun, Ray and Spark for distributed work, while the maker describes compatibility with any hardware, open-source tools and frameworks. TPU users should note the narrower practical limit: dstack supports single-host instances only, up to eight cores per instance.
Inference services and integrations
Services can be published as inference endpoints through gateways that support HTTPS, custom domains, auto-scaling and rate limits. Documented backends include AWS, Azure, GCP, Kubernetes, several GPU cloud providers and remote SSH hosts; Slurm support is experimental. GPU utilization metrics are supported. The CLI handles resource management, while the HTTP API covers functionality not exposed in the CLI and direct server integrations.
Security and support
Server data and project-scoped secrets are stored in plaintext by default. Administrators can configure AES-256-GCM encryption for stored data, an important step for teams whose workloads or credentials require stronger at-rest protection. Project admins manage secrets. Support channels are GitHub issue reports and the dstack Discord server.
Pricing
dstack OSS
The dstack OSS plan is 0.00 USD per free and provides the open-source orchestration stack for self-hosting. It is the natural fit for teams able to operate the server and their own infrastructure; self-hosting means they take responsibility for deployment and security configuration.
dstack Sky GPU Marketplace
dstack Sky GPU Marketplace has custom pricing: GPU usage is pay-as-you-go, charged per GPU-hour, with rates varying by GPU and provider. Compute is on-demand or spot, and usage is funded with prepaid credits; resource prices appear in the console before provisioning. This suits teams seeking marketplace GPU capacity without a fixed plan price, but variable rates and prepaid billing call for checking costs before launching workloads.
dstack Sky does not currently charge for BYOC mode. dstack Factory is the commercial extension, adding advanced multi-tenancy, usage metering, billing automation and optimized inference presets for frontier open models; it has custom pricing.
Platforms
dstack supports API, Linux, macOS, self-hosted, web and Windows. Its hybrid deployment model allows the server to run wherever it can reach the cloud and on-premises infrastructure being managed.
Who it's for
dstack is best for AI teams orchestrating workloads across heterogeneous accelerators and infrastructure, especially those that want YAML-defined jobs and services, distributed framework support and inference endpoints. It is less suitable for organizations that require built-in quotas, need multi-host TPU deployments, or cannot accept plaintext storage defaults without first configuring server encryption.
Pros and cons
- Broad infrastructure reach: Cloud, Kubernetes, VMs and bare-metal clusters fit a single orchestration layer.
- Wide accelerator coverage: NVIDIA, AMD, TPU and Tenstorrent support suits teams with mixed hardware.
- Useful endpoint controls: HTTPS, custom domains, auto-scaling and rate limits support deployed inference services.
- No quota controls: Teams needing enforced resource ceilings must address that requirement elsewhere.
- Plaintext defaults: Server data and secrets need encryption configuration when at-rest protection is required.
- Single-host TPU ceiling: The eight-core maximum per TPU instance rules out larger multi-host TPU deployments.
Alternatives
For a wider set of GPU cluster management options, browse GPU Cluster Management Software.
- HTCondor is another free option for teams comparing open-source cluster management tools.
- ClearML is worth considering if a freemium product and a self-hosted version described as 100% open source on GitHub better fit the team.
- GPUStack offers a free, open-source GPU cluster manager for teams focused on that role.
- Backend.AI may suit home users who want an open-source option for their own hardware, or users who prefer a free trial.
- Koordinator is another free option for teams comparing infrastructure software.
- NVIDIA ShadowPlay is a free Windows option with a supported-GPU requirement, rather than a comparable cross-environment orchestration choice.
- OpenPBS may fit teams seeking an open-source edition under AGPL 3.0, with community forum support that carries no guarantees.
- HAMi is a free, open-source GPU virtualization middleware option for AI workloads on Kubernetes.
Verdict
Choose dstack if your team needs a flexible orchestration layer spanning cloud and on-premises AI compute, with support for several accelerator families and inference services. Its central advantage is the breadth of workloads and infrastructure in one workflow. Look elsewhere, or plan additional controls, if quota enforcement, encrypted-by-default storage or multi-host TPU support is essential.
dstack plans and pricing
All plansCompared on GPU cluster management software
- Free plan
- Yes
- Deployment model
- hybrid
- Workload scheduling
- both
- Kubernetes support
- Yes
- Quota controls
- No
- GPU utilization metrics
- Yes
- Cloud GPU support
- Yes








