Trust is the product.

Synthetic data is only worth buying when the real data is out of reach — and only worth trusting when you can check it. Plexoria is built around that second half.

Why synthetic, and why us

For domains like fraud, the real data is legally locked away — it can't be shared or resold. A realistic, labelled, legally-shareable synthetic substitute is genuinely scarce. But the market is also full of free, undocumented generators. So we don't compete on being cheap or being big. We compete on being trustworthy and deep.

Four commitments, on every dataset

Interpretable, not anonymised. Named, meaningful columns instead of opaque PCA components. You can engineer features and actually reason about the data.

Labelled hard cases. We don't ship one undifferentiated blob. Attack patterns are labelled as typologies — including covert ones that deliberately hug the legitimate distribution, so you can stress-test a detector against the cases it really misses.

Proven, not asserted. A reproducible baseline benchmark ships in the bundle. Quality is a number you can re-run yourself, not a marketing claim.

Signed provenance. A SHA-256 manifest and an Ed25519 signature travel with every bundle, and the key fingerprint is published out-of-band. You can prove a dataset is intact and really from us. See how →

Honest by construction

No real data ever enters our pipeline — the generator is rule-based and seeded, so there is no personal data to leak and nothing to re-license. Every datasheet states plainly what a dataset is and isn't. Fraud detection is adversarial and non-stationary; no static dataset — synthetic or real — is ever "current". Our datasets are a benchmark for developing and stress-testing methods, not a substitute for production decisions. We'd rather under-claim and be believed.

A dataset is generated, not extracted — so we can regenerate it, reweight the difficulty, or add new typologies on demand. A frozen real extract can't. That is the long-term advantage of doing this well.

One pipeline, many datasets

Fraud is the first line. The same orchestration, quality gates, signing, and governance carry over to the next — so every dataset in the catalogue meets the same bar and wears the same signature. Consistency is the brand.