Local-first, zero data retention
Everything runs on your machine. Datasets, prompts, schemas, and logs are never sent to ExaScale. Local AI via Ollama by default; remote telemetry is off by default.
DataRig Lite is free for your desktop: describe a pipeline in plain English, review exactly what it builds, then approve — governed Iceberg tables, orchestrated dbt models, and live dashboards, without your data ever leaving your laptop.
Part of the ExaScale DataRig platform: start free on your desktop, scale to managed Cloud, or deploy in your own VPC — same pipelines, same governance, zero rewrites.
local.user.orders (Iceberg · Polaris)revenue_daily + freshness & row-count testsorders_daily, scheduled 06:00Everything you need, packaged in one install. Nothing hand-wired.
Proprietary platforms charge you rent to write pipelines yourself. Hand-rolled open source costs you a weekend a month. DataRig removes both problems.
Everything runs on your machine. Datasets, prompts, schemas, and logs are never sent to ExaScale. Local AI via Ollama by default; remote telemetry is off by default.
Every generated artifact — ingestion specs, DAGs, dbt models, SQL, dashboards — is validated and hash-bound. Nothing executes until you approve the exact plan.
Your tables are Apache Iceberg v2 behind an open REST catalog. Leave anytime and your data comes with you in standard formats any engine can read.
SQL explorer, orchestration, dbt docs, dashboards, CDC connectors — pre-connected and one click away. You never assemble seven projects by hand.
DataRig profiles your sources, infers types, and drafts plans grounded in registered adapters — a tool name in a prompt is never mistaken for executability.
DataRig Lite is free forever — complete, and local-first, not a crippled trial. Paid managed editions come later, for teams that outgrow their laptop.
You stay in control: DataRig proposes the exact plan, you approve the exact hash, your machine runs it.
Drop in CSV, JSON, or Parquet files — or connect PostgreSQL / SQL Server for continuous CDC. All local.
Say what you want: “load this into Iceberg,” “add dbt tests,” “build a revenue dashboard.”
Inspect source analysis, plan steps, and generated files. Approve binds to an exact content hash.
Watch live Spark logs, staged progress, and status chips — with error-specific recovery if anything fails.
Browse your Iceberg catalog tree, inspect schemas, and run read-only SQL — right beside the chat. AI-suggested queries load here for review; they never auto-run.
SELECT order_date, SUM(amount) AS revenue FROM local.user.orders GROUP BY order_date ORDER BY order_date DESC;
| order_date | revenue |
|---|---|
| 2026-08-25 | 18,204.50 |
| 2026-08-24 | 16,930.00 |
| 2026-08-23 | 21,417.75 |
| 2026-08-22 | 14,882.10 |
Specialist consoles are packaged, wired, and one click away — with the DataRig console as your single home base.
Sources → Iceberg → Airflow → dbt → Superset. Each stage is a governed skill with its own validated contract.
Land CSV, JSON, JSONL, and Parquet files — or HTTPS APIs — into governed Iceberg tables. Full load today; XML profiling is review-ready.
Capture changes from PostgreSQL and SQL Server with Kafka + Debezium. Snapshot, updates, deletes, restart recovery — verified end to end.
dbt models generated with tests and docs, executed by Airflow against Trino — every artifact reviewable before it runs.
Airflow DAGs created, scheduled, paused, retried, and refreshed in place. Run-scoped docs and live status inside the console.
Freshness checks, row-count parity, and null/duplicate tests proposed as part of the plan — not bolted on afterwards.
Superset datasets, charts, and dashboards provisioned from your governed business layer, queried through Trino.
The pipelines you build in Lite are the same artifacts every edition runs: Iceberg tables, Airflow DAGs, tested dbt models, Superset dashboards. When you scale, your work scales with you — no rewrites, no migration project.
Individual engineers, pilots, local-first teams
Teams that outgrew the laptop but not their budget
Regulated orgs that need sovereign control
Pipeline definitions carry forward across every edition. Your data stays in open Apache Iceberg formats everywhere.
ExaScale is the bridge when your data
doesn't live in one wall.
Planned: read-only federation to Snowflake, Redshift, Databricks, and BigQuery catalogs — join lakehouse and warehouse data without ETL sprawl, behind the same review-gate stack.
Planned: import existing semantic models and metadata from the platforms you already pay for — your definitions travel with you instead of being rebuilt per vendor.
Federation arrives read-only first. Any future write-back adapter lands behind explicit, hash-bound approvals — never silent synchronization.
Designed, not yet scheduled — join the waitlist to shape the priority.
Stop hand-wiring load scripts and cron jobs. Describe the pipeline, review the plan, ship the dashboard — all locally.
Let your team self-serve pipelines safely. Review gates mean nothing mutates data without an approved plan.
A governed lakehouse on a laptop, with zero data egress and open Iceberg formats — no vendor bill, no lock-in.
Your team pays for AI assistants in two platforms and still can’t join the data. DataRig is planned to bridge warehouses and lakehouse — read-only first, always gated.
DataRig Lite is the complete local product — free forever. Paid editions exist for teams that want managed operations, not extra features on your desktop.
Install DataRig Lite free and ship your first governed pipeline today.