Synthetic data you can prove.
RadMah generates realistic data for AI, software testing, healthcare and industrial systems — for the many situations where production data is sensitive, restricted, or simply does not exist. Every run is validated, and arrives with evidence describing what was produced and how.
From a spreadsheet to a substation — with receipts.
Describe
prompt, schema or source
Generate
engine or simulator
Validate
every row, every rule
Evidence
sealed, and yours to re-check
- sealed contract
- constraint report
- artifact hashes
- determinism report
blake3 9f2c 41ab 7d05 c318 … e6b1
Real data is often the data you cannot use.
Teams need realistic data to train models, validate software and test operational systems. The data that would be most useful is frequently the data they are least able to touch: it carries personal information, it is contractually restricted, it is expensive to capture, or it describes an event nobody wants to reproduce in the real world — a plant failure, an intrusion, a rare clinical presentation.
RadMah exists to generate the data those teams need, and to make the result verifiable — so the dataset is not just plausible, but accompanied by a record of how it was produced and what was checked.
Four stages, and a record of all of them.
The same path whether the output is a table of invoices, a FHIR bundle, or a week of substation telemetry.
- 01
Describe or connect
Say what you need in plain English, upload a file, or point at a connector. A price is shown before anything runs.
- 02
Generate or simulate
The request is compiled into a sealed contract and routed to the right engine — tabular, healthcare, or an industrial simulator.
- 03
Validate
Rows are checked against the rules you declared and the invariants the contract implies. A contradiction is refused, not papered over.
- 04
Verify
The run seals an evidence record — contract, constraint report, artifact hashes, determinism report — that you can re-check yourself.
Three things it does. One contract, one evidence chain.
A simulated plant run can feed a synthesis job; the agent can drive either. They are capabilities of one system, not separate tools you integrate yourself.
Generate
Create realistic data from a prompt, a schema, or a dataset you already have.
Simulate
Produce operational data and scenarios that no ordinary dataset contains.
Verify and orchestrate
Check what was produced, prove how it was produced, and drive all of it from your own code.
Generation is half the job. The other half is the record.
A quality score asks you to trust the vendor that produced it. RadMah seals each run into an evidence bundle instead: the contract that defined the job, the constraint report, an index of every artifact with its own hash, and the determinism report. A verifier ships with the SDK and checks the bundle offline — no API key, no call back to us.
- Re-check the bundle without trusting our runtime
- The same contract, seed and engine version reproduce byte-identical data
- Tamper-evident: edit any artifact and the hash comparison fails
No identifier-shape values detected; dataset clean.
- Inspected rows
- 100
- Inspected attributes
- 15
- Repairs applied
- 0
- Scanner
- v1-row-level
Every run is measured against your own data.
A vendor benchmark tells you how an engine behaved on somebody else’s data. So RadMah measures yours: each synthesis run scores the delivered rows against your source — distributional similarity, downstream utility, privacy indicators and exact-copy rate — and seals the result into that job’s evidence bundle, next to the contract and seed that produced it. Fidelity depends far more on the size and shape of a table than on any headline figure, which is why we publish the measurement rather than a number.
What a run measuresData that only exists when a plant is running.
Most synthetic data stops at the table. RadMah extends the same contract and the same evidence chain into operational environments: a water-treatment station simulated from process physics, emitting telemetry on the OT protocols your equipment actually speaks, with packet captures included — and labelled attack scenarios you could never safely run against production.
- Six OT protocols, driven by process models rather than canned CSV files
- 67 plant templates, with a CI gate that fails the build if the registry drifts
- ATT&CK for ICS scenarios with per-event ground truth
A request as ordinary as
“a water-treatment pump station, sixty seconds, Modbus and OPC-UA, eight signals, with an over-pressure alarm”
produces a simulated run with the pressures, flows and alarm behaviour a process model implies — not a random walk dressed up as telemetry.
Four teams, one problem.
They need realistic data, and they need to be able to say where it came from.
Data and ML teams
Training and evaluation data when the production tables are restricted or too small.
Software and data engineering
Deterministic fixtures that make a test suite diffable instead of flaky.
OT, industrial and security teams
Plant telemetry and labelled attack scenarios without touching a live process.
Regulated and healthcare teams
Realistic clinical structures, and a record of exactly how each dataset was produced.
Run it in our cloud, or in yours.
RadMah, Inc. is a Delaware C corporation and a wholly owned subsidiary of ITLOX, Inc..
We run it. You pick the residency region your data stays in.
An offline container bundle, a licence verified on your own host, and no call back to RadMah at runtime.
Per-tenant keys, storage prefixes and row-level filtering, exercised by gates in CI.
Bundles verify offline, so an audit never depends on our availability.
Bring a dataset, or just describe one.
Run a job end to end and keep what comes out of it — the data, the quality report, and the sealed evidence bundle.