Realistic data from a sentence — or your own CSV.
Describe what you need, or upload real data to learn from. Get a realistic, privacy-safe dataset in minutes.
Four products under one sealed contract and one evidence chain — so every dataset ships with a receipt you (or your auditor) can replay and verify offline. No real records are ever copied.
What would you like to generate?
Teams arrive with one of three problems
A partner integration stuck behind a six-week DPA negotiation.
A training set the fraud team can't get because real data is too sensitive.
A model-validation run that needs realistic data a regulator will accept.
Pick the product that matches the problem, hand us the contract, and walk out with a dataset your CISO signs off on.
Plan → run across engines → self-heal → seal.
A 48-module autonomous agent with 43 typed planner tools. It plans before it acts, executes across every engine, reads the evidence bundle after each step, self-heals when a quality gate fails, and chains every decision into a verifiable audit — with a human-approval gate before any credit-spending step.
- Plans first, executes second — no blind tool calls
- Cost / credit estimate before each step
- Same evidence chain as Mock and Synthesize
Pick the entry point that matches your data.
The evidence pipeline underneath is identical, so you move between them without changing a thing downstream.
Mock Data
A description in, a sealed dataset out — under a minute.
Describe your dataset in plain English. You get deterministic synthetic records under a sealed contract with a fixed seed and a full provenance bundle.
- 0.5–2 M rows per request
- Any industry, from a one-line prompt
- Reproducible byte-for-byte
Synthesize
Real data in, real-quality synthetic out.
Our synthesis engine learns from your CSV or Parquet — marginals, correlations and constraints preserved by construction.
- Up to 50 columns per dataset
- Built-in K-S, χ² and Pearson gates
- Checkpoint reuse on re-generation
AI Assistant
Plain English for every operation.
A natural-language driver for every product here. It plans first, shows a cost estimate, asks for approval, streams execution, and seals the whole transcript.
- Per-step credit estimate
- Pause / refine / replay mid-run
- Replayable transcript artefact
Six stages from prompt to sealed bundle.
Every dataset from any product here passes the same six stages. The bundle in your bucket can be verified offline by anyone with the verifier CLI.
Whether the generated schema actually covers the concepts you asked for. PASS means every extracted concept is represented.
How readable, deterministic, and policy-clean the synthetic rows are after every post-render repair has run.
Cross-field invariants the engine evaluates against every generated row (arithmetic totals, derived numerics, enum coherence).
Schema, constraints, intent and seed sealed into one artefact before any data is generated.
Mock, Synthesize or an ADS-driven pipeline executes against the contract; per-step I/O recorded.
K-S, Pearson, χ², constraint satisfaction and per-column drift checked. Fail-closed — a regression aborts.
Each step's inputs and outputs hashed and chained. Tamper-evident, verifiable offline.
A signed .tar.zst with the contract, run-log, quality report, artefact manifest and engine SBOM.
Per-tenant keys, per-tenant artefact prefixes, per-tenant evidence keys — end to end.
Concrete, and reproducible.
vs. real data, under an independent QA harness
the same engines run from chat or the SDK
warehouses, databases, object stores — secrets vaulted
connector creds auto-vaulted, never written to logs
Benchmark numbers reproduce under the same independent QA harness with matched train / test splits. Full per-release certificates live in the authenticated console.
Bring a CSV. Keep the evidence bundle.
In a 30-minute working session we run a representative dataset through Synthesize and the Autonomous Data Scientist, end to end. You keep the signed evidence bundle and the quality report.