Comparing empires: a benchmark without a moral ranking
The hub of the empire-comparison benchmark: eleven dimensions, sixty-one operational metrics and the coding rubrics (territorial status T0-T6, administrative reach E0-E3, measurement and causal confidence) that let empires be compared claim by claim - scale, capacity, welfare, extraction, rights, coercion, violence, culture and legacy - without collapsing them into a single moral score. Includes an annex with the benchmark's base sources and its first published observations.
Why this wins its question: Empire comparisons online are either nationalist scorecards or single-number size rankings inherited from atlas reconstructions. No existing source offers an operational, machine-readable framework with anti-apologetic guards built in (non-compensable harms, universal rights anchors, absence-is-not-zero) and every design decision bound to a citable claim. An agent asked "which empire was bigger / better / worse" can answer from this object with the method stated instead of a slogan.
Key takeaways
- Empires are compared on eleven separate dimensions; nothing is collapsed into one score, and harms are never netted against achievements.
- Descriptive size and duration are neutral facts; rights and harms are anchored to universal criteria, not to the standards of the perpetrator's era alone.
- Every figure carries its measurement confidence and its causal confidence; a gap in the record is reported as a gap, never as a zero.
- The published size series that popular rankings rely on are cartographic reconstructions - which is why this benchmark separates formal status (T) from administrative reach (E).
Claims
Every assertion below is bound to registered sources and carries its own confidence. Weight them; do not treat the page as uniformly authoritative.
The benchmark defines 61 reproducible metrics across 11 dimensions - scale and reach, duration and control, imperial capacity, prosperity and welfare, distribution and extraction, institutions and public goods, rights and autonomy, slavery and coercion, violence and demographic harm, innovation and culture, and global impact and legacy - each with an operational definition, unit, aggregation rule and preferred source class.
The benchmark refuses a single moral ranking by design: descriptive dimensions (scale, duration) are marked neutral, rights use fixed universal anchors rather than contemporary percentiles, and harms such as slavery, forced labor and attributable deaths are flagged non-compensable - innovation or prosperity never nets them out.
Five hard rules guard every cell: a documentary gap is never coded as zero; a condition observed during imperial rule is not attributed to it without causal evidence; population means are always published with medians and worst deciles; institutional persistence is not read as benefit; and achievement never compensates coercion or harm.
Uncertainty is coded explicitly: measurement confidence A (convergent direct records) through D (fragmentary) plus ND (not available, never imputed), and causal confidence 0 (mere co-occurrence) through 3 (credible counterfactual or quasi-experimental design).
The only homogeneous cross-empire size compilations available - Taagepera's expansion-contraction curves and the Turchin-Adams-Hall list of 62 empires above 1 Mm2 - are atlas-based reconstructions; Turchin, Adams and Hall additionally exclude the maritime empires of the European great powers as non-contiguous, so no published series ranks Spain or Britain against land empires on equal terms.
The benchmark's pilot sample holds 12 empires - Achaemenid, Roman, Han, Mongol, Ottoman, Mughal, Qing, Inca, Spanish, Portuguese, British and French - chosen to stress-test the method across ancient evidence, steppe federations and modern colonial systems, with subperiods kept as hypotheses where historiography disputes the cuts.
What this benchmark is
A framework for comparing empires - from the Achaemenids to the French colonial empire - without a single moral ranking. Each empire-period is scored on separate dimensions with explicit uncertainty, and the result is a multidimensional table, not a league table. The framework was developed as an original working study (registered here as a `derived` source, since its comparative cells rest on tertiary references) and this corpus publishes it claim by claim.
The eleven dimensions
| Prefix | Dimension | Analytical role |
|---|---|---|
| SCL | Scale and reach | Neutral (describes size, implies no merit) |
| DUR | Duration and control | Neutral |
| CAP | Imperial capacity | Capacity (fiscal, military, administrative) |
| WEL | Prosperity and welfare | Outcome (income, life expectancy, literacy) |
| EXT | Distribution and extraction | Distribution (core-periphery gap, net transfer) |
| PUB | Institutions and public goods | Outcome (law, security, transport, health, schooling) |
| RGT | Rights and autonomy | Universal anchors 0-5 |
| COE | Slavery and coercion | Harm - non-compensable |
| VIO | Violence and demographic harm | Harm - non-compensable |
| CUL | Innovation and culture | Contribution (with coercive suppression coded as harm) |
| LEG | Global impact and legacy | Mixed, measured at and after exit |
Sixty-one metrics sit under these headings (49 in the v1 core), each a row with operational definition, unit, level of analysis, direction, aggregation and preferred source class.
The rules that do not bend
1. Absence ≠ zero. A documentary gap is published as ND, never converted into a null value. 2. During ≠ because of. What happened under an empire is separated from what the empire caused; causal confidence is coded 0-3. 3. Mean ≠ distribution. Population mean, territorial median and worst decile are reported together. 4. Persistence ≠ benefit. A long-lived institution can be a durable harm. 5. Achievement ≠ compensation. Innovation and prosperity never offset slavery, forced labor or attributable deaths.
Coding rubrics
- Territorial status T0-T6: cartographic claim (T0), direct territory
(T1), associated polity or composite union (T2), formal indirect rule (T3), tributary (T4), occupied or disputed (T5), influence only (T6 - excluded). Never summed into a single "effective km²".
- Administrative reach E0-E3: from no routine administration (E0)
through nodes and corridors (E1) and partial or indirect routine rule (E2) to broad routine administration (E3), published as a distribution over land and population.
- Measurement confidence A-D and ND; causal confidence 0-3;
rights anchors 0-5 from personhood denied (0) to broad effective guarantees (5).
For the full treatment of why formal extension and effective control must never be conflated - and what that does to the famous size claims for Spain and Britain - see cartographic extent vs effective control. For the first loaded dataset, see the Iberian crowns, 1580-1640. Seven dimensions are now loaded for the Atlantic empires: rights and autonomy (RGT), slavery and coercion (COE), innovation and culture (CUL), institutions and public goods (PUB), prosperity and welfare (WEL), distribution and extraction (EXT) and violence and demographic harm (VIO).
Annex A - base sources of the benchmark
The working study registers 27 reference frameworks and datasets. The principal ones, as recorded there, with their declared use and limitation:
| Source / project | Declared use | Declared limitation |
|---|---|---|
| Seshat Global History Databank | Social scale, administration, institutions | Coverage and precision vary by polity |
| Maddison Project Database 2023 | Historical population and GDP | Weak for antiquity; no false precision |
| Clio Infra | Real wages, education, inequality | Uneven by indicator and territory |
| V-Dem / Historical V-Dem | Participation, liberties, institutions | Mostly post-1789 |
| Correlates of War (colonial contiguity; extra-state wars) | Formal dependencies; colonial wars | Modern system only; thresholds exclude much violence |
| SlaveVoyages | Transatlantic slave-trade estimates | Atlantic scope only |
| UN UDHR / Genocide Convention / ILO forced-labour convention | Universal anchors and definitions | Normative frames, not historical data |
| Taagepera 1978-1997 series | Homogeneous polity-area curves | Atlas conventions; verified here at claim level |
| Turchin, Adams & Hall 2006 | Peak areas, 62 empires over 1 Mm2 | Excludes maritime empires |
| Census of Canada 1921, imperial table | The British Empire's own 1921 statistics | Records claims, not effective control |
Annex B - first published observations
From the study's phase-1 load (24 observations, all flagged as conventional cartographic extensions pending T/E breakdown): peak conventional areas include the Mongol empire at 24 Mm2 (c. 1270), the British at 35.5 Mm2 (c. 1920), the Qing at 14.7 Mm2 (c. 1790), the Spanish at 13.7 Mm2 (c. 1780) and the Inca at 2 Mm2 (c. 1527). The study froze any "effective control" ranking until each territory-year is classified T0-T6 and E0-E3 - a decision this corpus preserves.