STATUS 2026-07-28: contains claims superseded by the resume decision + race-direction runs. Do not send externally until rebuilt. See docs/BENCHMARKS.md.

Press & Media

Press kit and media resources.

Founder bio, technical claims summary, and downloadable assets for journalists, researchers, and analysts covering Bhala’s work on programmable embeddings.

Technical claims

What Bhala has built, stated precisely.

Each claim below is documented on the benchmarks page with methodology, dataset references, and reproducibility notes. Citations appreciated.

How to read the scores below. Where a claim prints a score between 0.5 and 1.0, that is a detection score (technical name: AUC). 0.5 means the model is guessing at random — it has learned nothing from the record. 1.0 means it is right every single time. There are only two possible answers, so random guessing gets you 0.5; anything above that is real information the model picked up. Plain percentages such as 73.2% accuracy or a 95–100% transfer rate are ordinary percentages and do not sit on this scale.

An encoder pretrained on a single language transfers across 17

The Bhala encoder is pretrained on isiZulu alone. On MASSIVE 60-way intent classification, using the field's strictest representation-quality test, it reaches 73.2% on Swahili (above GPT-4o zero-shot at 70.6%), 72.5% on Korean, 69.7% on Hindi, and 66.5% on Amharic — none of these languages were in the training corpus. No labels, no fine-tuning: only a small monolingual sample is used to adapt the input per language. Results clear 38–43× over random across all four languages. To our knowledge, no other published model satisfies this exact test condition (zero target-language training) on typologically distant languages. Source: Mhlambi 2026, “Structure is all you need.”

Operator transfer at 95–100% on data the operator was never fitted on

A sentiment direction estimated from contrastive pairs in one language applies zero-shot to others without re-estimation. Tested pairs to date: Zulu→Swahili (100%), Zulu→Xhosa (95%). Intent operators (book→cancel, alarm→calendar) transfer at 100% on tested pairs. These rates were not measured against a norm-matched random direction, and a norm-matched random direction is a floor to beat, not a zero (Rogue Scalpel, arXiv 2509.22067) — read them as an upper bound. Each application writes a timestamped record to an append-only log you can retrieve by receipt id; the id is an unkeyed SHA-256 you can recompute offline, so it attests the record's contents but not its origin — not a signature. Cross-family transfer of the same operator class is in progress.

Architectural inductive bias accounts for ~80 percentage points of accuracy

A standard pipeline approaches random performance on this task. Bhala's architecture closes that gap to production-grade accuracy. The contribution of each architectural component is documented in the technical paper, available on request.

Detection is the product — the 100% strict-flip claim is retracted (2026-07-28)

Detection is the product: finding the proxies for race in a record. A proxy is something on the record that ISN'T race — but gives it away anyway. A ZIP code is not race. But if a neighbourhood is 96% Black, writing that ZIP down tells a computer nearly as much as writing the person's race would. Same for the university: an HBCU is not a race, but it is a very good clue. Nobody has to intend this; the model works it out on its own. Erasure, by contrast, is not removal: erasing three of the internal directions that carry the clue leaves race readable at 0.947-0.993 on the same model, 0.949-0.998 on a channel — one field of the resume, such as surname or ZIP — that was never used in the fit, and 0.996-1.000 cross-channel, meaning you scrub the clue out of one field and race still reads off a different field (4 models x 600 real resumes). Removing race from ONE channel took 12 directions, and left surname readable at only 0.535 by a simple probe (a small classifier trained to guess the attribute from the model's internals) and 0.508 by a stronger one — both at the 0.5 guessing floor — while ZIP 0.9936, school 0.9958 and first name 0.986 went untouched. The edit is cheap (qualification 0.4677 -> 0.4672) but it moved top-10 selection share beyond a random-direction control on 1 of 4 models. The earlier claim — 100% strict-flip on 28 protected dimensions across BBQ, StereoSet, CrowS-Pairs and WinoBias, verified by an independently-trained classifier — is withdrawn: the verification is circular in the same way "erase → 0.500 (guessing at random)" is circular, because the direction is fit on a channel and the checking probe is trained on that same channel, and the lexical-swap control has not been run. The multi-layer probe behind those axes covered 10,064 sentence pairs (500 per axis) at 16 probed layers, not 15,966 pairs at all 60 layers. What holds is detection: ZIP 0.984 (right about 49 times out of 50), HBCU 0.983 (about 49 times out of 50), name 0.945 (about 19 times out of 20), gender 0.979 (about 24 times out of 25) across 4 models, with 12/12 ZCTAs confirmed against 2020 Census DHC P8, statutory HBCU status, and a perfect 1.000 rank match against the census Black population share — the higher that share, the higher the model scores the ZIP. Those four detection scores were re-measured 2026-07-29 at full resume length — 8,000 characters / 2,048 tokens, 90.8% of all text, 500 matched pairs per axis, same four backbones. The earlier figures (ZIP 0.991 / HBCU 0.991 / gender 0.986 / name 0.964) were taken at 420 characters / 160 tokens, roughly 7% of a median resume, which left the injected demographic marker at about 17x its natural share of the input. Every number fell, and the name fell hardest: the gap between the proxies and the explicit marker widened from 0.026 to 0.039, so the proxies carry more of the signal than the truncated run showed, not less. The resume corpus is 41% software/IT and 23% finance, so these results are scoped to white-collar tech and finance screening, not to hiring in general.

Hate steers, per model and per method — and the same direction is a jailbreak

Hate steers, per model and per method: K-steering at layer 10 gives +0.188 [0.146, 0.231] on Qwen2.5-7B against a random-direction floor of +0.0081, and steers Mistral-7B-Instruct-v0.2. NULL on Mistral-7B-v0.1 base (baseline 0.2981, headroom present) and Gemma-2-9b (baseline 0.0549, floor-limited). The simpler rank-1 diff-of-means method — one single direction, taken as the difference between two group averages — scores +0.030 and fails. The same direction inverted is a jailbreak vector -- dual-use, never a safety dial. A norm-matched random direction is a floor to beat, not a zero (Rogue Scalpel, arXiv 2509.22067).

Hate-speech production: 11 corpora, HateCheck 0.90, TweetEval-hate 0.77

Bhala is evaluated jointly across 11 hate-speech corpora — HateCheck, CONAN, Civil Comments, Berkeley MHS, SBIC, DynaHate, TweetEval-hate, TweetEval-offensive, HateXplain, Stormfront, Hate-Speech-18 — for ~134K labeled examples. We publish every per-corpus number, not just the strong ones: CONAN 0.93, Civil Comments 0.92, Berkeley MHS 0.91, HateCheck 0.90, SBIC 0.88, Hate-Speech-18 0.84, TweetEval-hate 0.77, DynaHate 0.73, TweetEval-offensive 0.71, TweetEval-irony 0.58, MLMA 0.52. HateCheck 0.90, on the same 0.5-to-1.0 scale, matches or beats HateBERT (0.85–0.88) and HateXplain BERT (0.83), which were fully fine-tuned on hate data; TweetEval-hate 0.77 is Twitter performance without any Twitter pretraining. The weakest result, MLMA (multilingual EN/FR/AR) at 0.52, is barely above the 0.5 guessing floor — a single headline benchmark score hides that, which is why we report the full table. Live in production since 2026-05-02.

CPU-deployable: <50ms inference, no GPU

Bhala runs on consumer CPU with sub-50ms single-query latency. No GPU required for inference. Designed for edge deployment in low-resource settings, but the same model serves the hosted API.

Founder

Sabelo Mhlambi

Founder, Bhala AI.

Sabelo Mhlambi is the founder of Bhala AI. His background spans NLP engineering and AI ethics: software engineering at Mapbox and Iris.tv, and fellowships at the Berkman Klein Center for Internet & Society (Harvard), the Carr Center for Human Rights Policy (Harvard Kennedy School), Stanford PACS, and TechCongress. He is the founder of Bantucracy, a research initiative on Ubuntu ethics and AI.

Bhala is the commercial vehicle for a decade of work at the intersection of low-resource language technology, decolonial AI, and representation theory. The company's central technical claim — that a sensitive concept's influence can be located inside a model's learned representations, on models Bhala did not train, and read out with a content hash an auditor can recompute offline — emerged from that lineage. The removal half of that claim has been retracted: erasure is not removal, the attribute stays readable after the edit, and no proof-of-removal is offered.

Assets

Logos and photos.

Right-click to save, or contact press@mail.bhala.ai for additional formats.

Paper & reproducibility

Methodology, code, weights.

Preprint forthcoming

Structure is all you need

Mhlambi · April 2026 · Bhala AI

Bhala’s architecture and training objective, with component-level ablations and cross-family transfer results across 17 MASSIVE target languages. Argues that linguistic inductive bias carries information currently paid for in scale, and that the parameter count required for a given capability is a function of the task, not of a scaling law. Preprint link added here on release.

Benchmarks page

Per-benchmark numbers, evaluation protocol, dataset references, and links to the audit harness.

View benchmarks

Live demo

Apply sentiment, intent, and bias operators to your own text. Each call returns a timestamped audit record retrievable by receipt id; the id is an unkeyed SHA-256 you can recompute offline, not a signature.

Open demo

Press inquiries

For interviews, technical briefings, or independent reproduction access, write to press@mail.bhala.ai. We respond within two business days.