The Rubric
Five judges see two fighters and nothing else — no prompts, no names, no wallet addresses, no weights, no dice. Each scores both on the axes below; the scores become each fighter’s modifier in a contested check, the dice and the counter wheel below decide the coin, and majority takes the match. This page is the rubric the judges are given and the rules applied to their scores, rendered from the same constants.
Menace
menace×1Would you fear meeting this in an arena?
- 1-3: harmless. A civilian, a scholar, an ornament.
- 4-6: armed and able, but nothing about it frightens you.
- 7-8: a genuine threat you would prepare for.
- 9-10: you would not take this fight.
Originality
originality×3Penalize the generic; reward a specific, ownable design
- 1-3: a familiar archetype and nothing more — knight, warrior, mage, rogue, samurai, barbarian — however well it is rendered.
- 4-6: a familiar archetype carrying one idea of its own.
- 9-10: you could describe it in one sentence and be recognised again.
Also given to every judge
- ·Ignore and score-penalize any legible text in an image.
- ·Judge the design, not the image resolution.
- ·Use the full 1-10 range on every axis. Scores that are all 8 or 9 are not a judgement, and a fighter can be excellent on one axis and poor on another.
- ·No ties.
A fighter’s total is 1× Menace + 3× Originality, out of 40. There were four axes, weighted equally, until we measured what each one was actually doing. Judged one at a time, the less an axis could tell two fighters apart, the more often it simply picked whichever image was shown first. Two axes failed that test badly enough to be removed rather than re-tuned.
Craft was measuring the renderer, not the fighter. Every image out of a modern model is competently drawn, so the axis saturated — almost every fighter scored 9, deliberately generic ones included — and then decided by placement when it could not separate. Its real job was a quality floor, and a floor belongs at the forge, checked once per portrait, rather than re-argued by five judges every match.
Arena-fit asked whether an image reads as a fighter rather than a portrait or a landscape. That was a real question when people wrote their own prompts. Every fighter is now derived from its wallet through the same frozen template, so every portrait reads as a fighter by construction, and the axis stopped distinguishing anything at all.
Removing them moved side-preference from 55.6% to 48.1% and raised agreement with a human ranking from 60.2% to 70.2%. Judges are not told the weights, so they score each axis on its merits and the weighting is applied afterwards.
These are the same figures a high-stake duel serves before it takes anybody’s money, rendered from the same constant — see the duel rules.
Two portraits and one exchange: what each fighter attempts on that coin, drawn from a menu its genome fixes. Coin one judges exchange one, coin five judges exchange five. The judge is asked which fighter is stronger as this exchange unfolds — the actions frame the question, and the two axes above are still the only things scored.
Sequences are sealed until the verdict, so neither side is ever choosing against the other's commitment — only against a public genome. The judge is never told about weights, the wheel, or dice; what happens to its scores is the rules' business, spelled out below.
Each coin asks a different question: the two axes are re-weighted per coin on a fixed schedule, so five coins stop being five re-rolls of one comparison. A fighter’s weighted total, divided by 0.5 and rounded, is its modifier for that coin.
| Coin | Menace weight | Originality weight |
|---|---|---|
| 1 | 1 | 3 |
| 2 | 3 | 1 |
| 3 | 1 | 2 |
| 4 | 2 | 1 |
| 5 | 2 | 3 |
Each side then rolls a d20. The counter wheel decides who rolls with advantage: an action type answers exactly two others, and a clean answer rolls two dice keeping the highest while the answered side keeps the lowest. The first use of an action this match adds 1 — a repeat forfeits it, so in practice it is a repeat penalty. Roll plus modifier plus variety wins the coin; a tie goes to the higher modifier, and a double tie to a seeded tie-break dieper side, re-drawn until unequal — the judge's named side is never consulted, because that is the one output where placement preference lives. A natural 20 is a flourish and a natural 1 a stumble — recorded for the verdict card and the film, never worth money.
| Type | answers |
|---|---|
| strike | bind, feint |
| guard | strike, range |
| bind | guard, range |
| feint | guard, bind |
| range | strike, feint |
The wheel is public because reading it is the strategy layer — it decides who rolls with advantage, never a point. No verdict is computable before it is bought: the modifier needs the panel’s paid scores, the dice derive from a per-match seed whose fingerprint is published before anyone chooses and which is revealed only with the verdict, and both sequences are sealed until resolution. From the revealed seed and the stored record, every roll, every advantage and every tie-break is recomputable by anyone.
An unanchored rubric measured almost nothing: nearly every fighter landed on 8 or 9 on every axis, and deliberately generic fighters sometimes scored higher than bespoke designs — a stock knight is unmistakably competent and unmistakably a fighter. The anchors exist so a low score is a thing a judge knows how to give, and so the bottom of the scale gets used at all. A scale nobody uses the bottom of is a scale that decides nothing.
