PREVIEW · SAMPLE DATA — illustrative numbers until the first public runs land
EXOTICBENCH
The decensoring efficiency benchmark

Everyone measures how uncensored a model is. Nobody measures what it costs.

ExoticBench scores every decensored artifact on five axes — refusal, capability, persona, toolbelt, drift — at temperature 0, against a pinned anchor, with every raw transcript retained. One number per axis; no vibes.

Efficiency frontierQwen3.6-35B-A3B · preview
refusal reduction vs anchor →drift / cost (KL) ↑
basehereticexoticBF16NVFP4quant tax

hero imagery slot · art drops in here

Five axes, one frontier

Refusal

Refusals per 100 on a curated prompt set — the point of the exercise.

Capability

Reasoning and instruction retention on fixed, versioned slices.

Persona

Character consistency, judged pairwise against the base model.

Toolbelt

Tool-call reliability: exact names and args, schema validity, no-tool precision.

Drift

Distributional drift from the base model on harmless text — KL, in nats.

Determinism is the product

Pinned everything

Models by revision hash, fixtures by content hash, judge and prompts pinned. Temperature 0 across the board.

Pairwise vs an anchor

Every variant runs against the one stock anchor, so the whole family sits on a single comparable frontier.

Raw outputs retained

results.json keeps every prompt, response, and verdict verbatim. Any number on this site can be audited.