Live demos

See it work. Right now.

Four production demos that run against the real Bhala API. Real model, real sub-second latency, and a timestamped audit record on every call, retrievable by receipt id — an unkeyed SHA-256 you can recompute offline, not a signature.

How to read the scores below. They run from 0.5 to 1.0. 0.5 means the model is guessing at random — it has learned nothing from the text. 1.0 means it is right every single time. There are only two possible answers, so random guessing gets you 0.5; anything above that is real information the model picked up. (Technical name: AUC.)

Counterfactual

Counterfactual fairness audit

Paste any decision document. The race signal the model picked up is projected out — subtracted from the model's reading of the text — and the stripped version is scored as a comparison control; erasure is not removal. Scored across six domains as an internal index with no legal threshold. A disparate-impact ratio is a pool statistic and cannot be computed from one document. Intersectional analysis — overlapping traits read together in one shared space, not one axis at a time. Content hash you can recompute offline.

Six domains · sub-second · internal index, no legal threshold

Open demo
28-axis

Bias audit (28 protected axes)

Score any text on 28 BBQ + StereoSet + CrowS-Pairs + WinoBias axes. Per-axis drift, global Bias Score Index (BSI), top contributing axis. Content hash you can recompute offline.

28 axes · one embedding space

Open demo
0.9339 on HateCheck

Hate-speech classifier

Nine per-category detectors (women, LGBTQ+, jews, muslims, black, asian, latino, disabled, migrants). Scores 0.9339 on the HateCheck benchmark it was fit on — right about 14 times out of 15 — and tells 'I hate X' apart from 'Saying I hate X is bigoted'. That number is benchmark-specific: on other English hate corpora it falls to 0.47-0.62, at or barely above guessing at random.

9 categories · 0.9339 on HateCheck — right ~14 times out of 15 (benchmark-specific)

Open demo
RAG guardrail

Hallucination / grounding check

Given a model output and the source documents it was supposed to be grounded in, score how strongly the output is supported. Drop-in guardrail for RAG pipelines.

Q&A · summarization · dialog · content hash

Open demo
Production beta

Want production API access?

The demo endpoints above are rate-limited. For production-shape API keys (no rate limit, full audit-log access, SLAs), join the beta.

Join the waitlist