Provider-neutral Kubernetes Dynamic Resource Allocation observability, simulation, and diagnostics for GPU and accelerator workloads.
Model virtual device pools, inspect ResourceClaim and ResourceSlice state, and explain why allocations succeed or fail—without requiring physical accelerator hardware or a specific cloud provider.
Documentation · Quickstart · Architecture · Compatibility · Operations · Contributing · Governance · Security · Support
Kubernetes Dynamic Resource Allocation (DRA) offers fine-grained, driver-controlled accelerator sharing. However, developing and debugging DRA configurations presents a major challenge:
- Hardware Scarcity: Acquiring and configuring dedicated accelerator nodes (e.g., NVIDIA H100 GPUs) for test environments is costly and slow.
- Observability Gap: Native Kubernetes scheduling logs make it difficult to visualize why a resource claim failed to bind to a pod.
DRAForge bridges this gap by providing an evidence-based diagnostics toolkit, a dynamic virtual device simulator, a terminal user interface (TUI), and a real-time interactive relationship graph dashboard.
- Virtual Device Pools: Simulate arbitrary hardware profiles (e.g. GPUs, FPGAs, High-Speed NICs) on worker nodes using custom attributes and capacities.
- Diagnostics Doctor: Honest, non-mocked configuration analysis (e.g. API availability, version compatibility, ResourceSlice consistency checks).
- Explain Engine: Real-time evaluation of selectors, capacity bounds, and node affinity to pinpoint why claims are pending.
- Bubble Tea TUI: Professional terminal-based monitor for dynamic pool capacities.
- Interactive Graph Dashboard: Real-time SVG visualization of relationships between Pods, Claims, Devices, and Pools.
graph TD
subgraph Kubernetes Cluster
Server[DRAForge Server] <--> WebSPA[Vite + React SPA Dashboard]
Controller[DRAForge Controller] <--> SimulatedDevicePool[SimulatedDevicePool CRD]
Plugin[Node Plugin DaemonSet] --> ResourceSlice[ResourceSlice Spec]
APIServer[Kubernetes API Server] <--> Server
APIServer <--> Controller
APIServer <--> Plugin
end
CLI[DRAForge CLI] <--> APIServer
DRAForge components use Kubernetes APIs and Helm contracts rather than provider-specific services. They can run on compatible managed Kubernetes offerings, self-managed clusters, on-premises environments, and local test clusters. Provider-specific Terraform or registry assets in this repository are optional showcases.
go install github.com/oaslananka/draforge/cmd/draforge@latestOr clone and build locally:
git clone https://github.com/oaslananka/draforge.git
cd draforge
task build # Builds all three binaries into bin/
./bin/draforge versionBinaries are also available as pre-built archives from the GitHub Releases page.
Install Docker, kubectl, Helm, jq, curl, Go, Node.js, pnpm, Task, and the kind version pinned in tests/install-e2e/kubernetes-versions.json. Then create the same complete local stack used by the pull-request install gate and keep it for exploration:
DRAFORGE_INSTALL_E2E_KEEP_CLUSTER=1 task e2e:install-kindAccess the dashboard:
kubectl port-forward svc/draforge-server -n draforge-system 8080:8080Build the local CLI/TUI when needed:
task build
./bin/draforge doctor
./bin/draforge tuiDestroy the disposable cluster when finished:
kind delete cluster --name draforge-install-e2eThis provider-specific path is an optional demonstration of DRAForge on one managed Kubernetes service. It is not required for installation, testing, or releases.
⚠️ Billable Resources: This task provisions a live DOKS cluster and DOCR registry on your DigitalOcean account and incurs cloud costs. Always runtask demo:downwhen finished to destroy all billable resources.
To deploy the short-lived DOKS showcase:
task demo:upThis task explicitly applies the non-production values-showcase-docr.yaml profile, which enables unauthenticated HTTP exposure. It audits resource limits, provisions infrastructure, builds images remotely, installs the Helm release, and outputs the demo URL. Do not use this profile with sensitive clusters or as a production deployment.
A normal Helm install creates only internal ClusterIP services. Production public access requires TLS and an operator-managed OIDC/identity-aware proxy; see the installation guide.
The node plugin defaults to isolated demo CDI output. Host-integrated kubelet CDI output is an explicit nodePlugin.outputMode=node opt-in and fails closed when Kubernetes or the host CDI directory is unavailable.
To tear down the showcase and clean all billable resources:
task demo:down# Run all unit tests (fast, no race)
go test ./...
# or: task test:unit
# Run unit tests with race detector and coverage
go test -race -coverprofile=coverage.out ./...
# or: task test:race
# Run all Go vet checks
go vet ./...
# or: task vet
# Run frontend unit and integration tests
pnpm --dir web test
# or: task web:test
# Run full Go CI suite (unit + race)
task testThe tagged smoke tests run against any compatible Kubernetes cluster that serves the required DRA APIs and are gated by DRAFORGE_E2E=1. The Portable Kubernetes E2E workflow validates this path on credential-free kind for relevant pull requests and can use a protected short-lived kubeconfig for an external cluster:
DRAFORGE_E2E=1 go test -tags=e2e ./tests/e2e/... -vSee docs/release.md for the full release process.
task release:local
# or: goreleaser release --snapshot --clean --skip=docker,sbom,signNEXT_VERSION="${NEXT_VERSION:?set NEXT_VERSION to the intended SemVer, for example 0.3.1}"
release_tag="v${NEXT_VERSION}"
git tag -a "$release_tag" -m "DRAForge $release_tag"
RELEASE_TAG="$release_tag" RELEASE_MAIN_REF=main bash scripts/verify-release-tag.sh
git push origin "$release_tag"Pushing the annotated tag starts the protected release workflow. Published v* tags are immutable and are never moved or reused.
| Command | Description | Example |
|---|---|---|
draforge version |
Print binary version and commit details | draforge version |
draforge discover |
Lists active DRA pools, devices, and resource claims | draforge discover -o json |
draforge claims |
Summarizes claims status in a formatted table | draforge claims |
draforge graph |
Generates relationship graph in DOT or Mermaid formats | draforge graph -o mermaid |
draforge explain |
Troubleshoots why a ResourceClaim is pending | draforge explain my-pending-claim |
draforge doctor |
Executes cluster and driver diagnostics checks | draforge doctor |
draforge tui |
Launches the terminal-based monitor console | draforge tui |
draforge serve |
Runs HTTP API server and Vite dashboard frontend | draforge serve |
| Resource | Description |
|---|---|
| Contributing Guide | Setup, testing, PR expectations, and what not to commit |
| Security Policy | Vulnerability reporting and security model |
| Installation Guide | Demo and production installation profiles |
| Operations Guide | Day-two operations, upgrades, metrics, logs, and cleanup |
| Troubleshooting Guide | Diagnostics for common issues and pending claims |
| Maintainer Checklist | Internal review and release processes |
| Release Process | Snapshot and tagged release workflow |
| Dashboard Guide | Web dashboard setup and usage |
| Simulator Scenarios | Scenario authoring and fault injection |
| Kubernetes DRA Compatibility | DRA API support matrix |
| Maintainers | Current project maintainers |
| Governance | Decision-making and roles |
| Support | How to get help |
DRAForge exposes a read-only dashboard API, but Helm does not expose it externally by default. Read endpoints can reveal operational cluster metadata, so production public access must use TLS and an operator-managed identity-aware proxy. Cluster modifications remain CLI-only and authenticate with the administrator's local kubeconfig (see ADR-0009).
For Kubernetes DRA API support and known limitations, see Kubernetes DRA Compatibility.
This project is licensed under the Apache License, Version 2.0. See LICENSE for the full license text.