Skip to content
NCBNational
Capability
Benchmark
Portuguรชs

Method

Capability is measured apart from wealth

The benchmark asks whether a country's capacity to act is distinct from its income. Every choice is open to challenge.

Why this exists

If the claim is right, countries with similar incomes will have different capability profiles. If not, the dimensions mostly track income.

  • All 52 countries set the scale. A frame built around one country would only describe that country.
  • A high score is not a recipe. Mechanisms depend on local conditions. The shape says where to look; context says what to build.
  • Capability changes below the national level, in groups small enough to act. A country score is a coarse proxy for their conditions.
  • Treat this as a measuring instrument. It tests whether an attempt to raise a capability worked. Confidence, gaps and revisions sit beside each score.
  • The diagnostics tests whether dimensions collapse into income. The limits page records known failures.

This measures capability, not wealth

The benchmark measures capacity to anticipate, coordinate, learn, adapt and build under uncertainty. Wealth, quality of life and popularity are outside its scope.

DimensionQuestion
AnticipationHow capable is the country of identifying and preparing for emerging change?
AgencyHow able are individuals and organizations to turn an intention into action?
CoordinationHow effectively can independent actors organize around shared objectives?
TrustHow much cooperation is possible beyond immediate personal networks?
LearningHow effectively does the country acquire, distribute, and update knowledge?
ExperimentationHow easily can new approaches be attempted, tested, abandoned, and improved?
AdaptabilityHow effectively can the system respond when circumstances change?
BuildingHow capable is the country of turning plans and knowledge into functioning systems?
Shared PurposeTo what extent can people imagine themselves as participants in a common project?

How an indicator becomes a score

  1. Take the latest comparable value for each country and indicator, with its source and year.
  2. Apply the declared transform: per million people, log, or distance from a target.
  3. Winsorize extreme outliers with Tukey fences at three interquartile ranges.
  4. Normalize to 0 through 100 against the frame set by all 52 countries. Reverse indicators where lower is better.
  5. Average the available indicators inside a dimension with equal weights.
  6. Compute confidence separately as coverage ร— recency ร— source quality.

Missing indicators lower coverage and drop out of the mean. Nothing is imputed. Equal weighting keeps v0 easy to challenge.

The registry has 69 indicators: 35 with data, 26 gaps and 8 retired rows. Gaps have no comparable dataset; retired rows have a rejected one. Both lower confidence and define the collection agenda. 1 values come from reproducible source adapters and 2 from published tables, with retrieval dates stored.

National scores have a second reading

The benchmark compares countries at the national level, while selected destination pages show constituent data as corroboration or context.

The comparison layer reads observations with geometry=national. That keeps every country on the same unit of comparison and means a state or province cannot silently move a national score.

A destination page may also show published values for constituent units. Each fixture declares whether those values aggregate to the national figure, stand independently, or provide context only. The rule stays beside the values, and the source remains visible. These rows corroborate, question or explain a national result; they do not create per-state capability scores.

Every indicator states what it measures

The dataset labels each indicator as C, I, O or P so the classification can be checked.

  • C, direct capability measure. Measures the thing itself. Days to register a company measures how hard it is to start a business.
  • I, capability input. Measures something that supports the capability. Research spending is an input. It does not prove a country reads the future well.
  • O, downstream outcome. Measures a result that usually follows from the capability. Patents can show experimentation, but also defensive filing. Korea files at volume, so its patent number says less than it looks.
  • P, perception proxy. Records what people or experts say, not what they did. The Worldwide Governance Indicators aggregate expert opinion. Seven were retired because they tracked income per head more than capability.

Source quality affects confidence only

Each tier affects confidence. Delphi estimates have the least weight.

TierWeight
official statistical1.00
international organization0.95
academic survey0.85
composite index0.70
expert panel0.50
llm delphi0.30

A tier says who published a number. The sources page lists the publisher, database and request.

Thin evidence appears on the chart

Confidence never enters the score. Thin evidence gets a dashed edge and hollow point. The gap widens as confidence falls.

  • Coverage is the share of a dimension's indicators that have a value.
  • Recency decays after two grace years, over a twelve-year window, to a floor of 0.1.
  • Source quality is the mean tier weight of the values that are present.
  • The product stays below 1 in practice, so the bands reflect real values.

Momentum uses only matching indicators

Momentum shows score change over time on the current frame. Only indicators observed at both ends count.

  • Historical values use today's frame, so the change reflects the country.
  • The same indicators are used at both ends. A new indicator cannot create movement.
  • The basket may be smaller than the dimension, so the trend level can differ from the score. Its size is printed beside the trend.
  • Ten-year and twenty-year spans are published. A missing span shows how far the data reaches.
  • Each indicator has its own line back to 1990 where data exists. Nothing is carried forward or filled in.
  • Each point carries the published value, normalized value and source tier.
  • Each run compares its data with the previous file and logs restated, added or dropped values.
  • Values more than five years old do not count for a year. Historical values outside the frame clamp to 0 or 100, and the clamp is recorded.
  • Adoption indicators often rise for every country. Compare each change with the median before calling it progress.

Documented deliveries sit outside the score

Evidence records describe work that a gap indicator cannot measure. They do not change scores.

  • Each record carries a published number, reference period, source and retrieval date.
  • Each record states what the case does not show.
  • Records never affect scores or confidence.
  • A gap becomes scorable when a comparable series covers at least two countries.

A panel reviews what the data misses

Each panelist has a fixed stance. The panel interprets the source-backed evidence and reviews the indicators.

  • Round 1: each panelist scores dimensions with thin source coverage from the evidence brief and its knowledge.
  • Round 2: each panelist sees the anonymized round-1 scores and rationales, then revises or defends its scores.
  • We keep the median and interquartile range. A range above 25 points is unresolved disagreement.
  • Panel estimates stay in their own file and never enter the indicator score or confidence. The blended view uses one only when no indicator is observed.
  • The panel also rates each indicator's class, validity, wealth-proxy risk and redundancy.
  • The Delphi page shows the current run and its provenance. The active run is a working session, not a panel.

The country set defines the scale

All 52 countries set each indicator's fences and endpoints. The frame stays fixed within a version. Adding a country rebases it and requires a major version bump.

CountryWhy it is included
BrazilPrimary reference case; large, diverse upper-middle-income democracy
United StatesHigh innovation and agency; large-scale institutional complexity
NetherlandsStrong institutions, coordination and social trust
SwitzerlandHighly decentralized but unusually coordinated system
SingaporeHigh-capacity, highly coordinated small state
South KoreaRapid development, technology adoption and execution capacity
EstoniaSmall state known for digital institutional experimentation
IndiaLarge, diverse emerging economy with significant bottom-up capability
ChileLatin American comparison with relatively strong institutions
South AfricaUnequal, institutionally complex middle-income comparison case
MexicoSecond largest Latin American economy; deep manufacturing base tied to North America
ArgentinaStrong research and human capital against repeated macroeconomic rupture
ColombiaLarge economy rebuilding state capacity after prolonged internal conflict
PeruSustained growth with persistent institutional instability and high informality
UruguaySmall state with the strongest institutional trust in the region
Costa RicaSmall state that moved into high-value manufacturing and services without an extractive base
GermanyLarge manufacturing economy coordinated through federal states and industry associations
FranceCentralized state with a long tradition of directing industrial policy
United KingdomServices and finance concentration with weak recent productivity growth
SpainSouthern European comparison with strong infrastructure delivery and high unemployment
PolandPost-socialist convergence case that rebuilt institutions and industry together
SwedenHigh-trust Nordic state with an unusual mix of large firms and startups
FinlandSmall state with an institutionalized foresight function and strong measured learning
IrelandSmall open economy whose output figures are distorted by foreign direct investment
CanadaResource-rich federal democracy with persistent productivity questions
AustraliaResource exporter far from its markets, with high administrative capacity
JapanAging high-capability manufacturer testing whether execution survives demographic decline
ChinaState-directed development at continental scale, the clearest contrast to the rest of the set
IndonesiaLarge archipelago state coordinating across extreme geographic dispersion
VietnamFast industrial catch-up on a low income base
PhilippinesServices export and remittance economy with weak industrial depth
MalaysiaMiddle-income manufacturer testing the move into higher-value production
ThailandThe middle-income trap as a case: strong assembly, thin innovation, aging fast
TurkeyIndustrial middle power with repeated macroeconomic instability
IsraelSmall state with the highest venture density in the world and deep civil divisions
United Arab EmiratesState-led diversification away from oil, executed quickly and from the top
NigeriaLargest African economy, with capability concentrated outside the state
KenyaEast African digital finance leader, where a private rail reached population scale
RwandaSmall state with a strong delivery reputation and a narrow political base
EthiopiaLarge low-income state attempting state-led industrialization under conflict
BoliviaResource-dependent landlocked state with strong social movements and weak formal institutions
ParaguayLandlocked agro-exporter with a small state and fast recent growth
EcuadorDollarized oil exporter cycling through repeated institutional redesigns
VenezuelaState collapse case; the sparse recent data is itself the finding
PanamaServices and logistics hub built around a single asset it operates well
GuatemalaLargest Central American economy with a chronically underfunded state
HondurasLow-capacity state where remittances stand in for absent institutions
El SalvadorSmall state undergoing a centralized security-led institutional rebuild
NicaraguaAuthoritarian consolidation case with thinning independent statistics
Dominican RepublicFast-growing tourism and services economy with weak public delivery
CubaState-run system outside most international statistical programs, so coverage is thin by design
HaitiState breakdown case; shows what the frame floor looks like

The assumptions are public

The decision log records each choice and what evidence would overturn it.

  • 0 and 100 are the weakest and strongest values among the 52 countries. They are not a sample of the world, and a low score is not a percentage of capability.
  • Scores use only the latest observation. Trends use a matched basket against today's frame. Nothing is back-filled or imputed.
  • With 52 countries, diagnostics are hints, not established results.
  • Doing Business series are frozen at 2019 and are marked down by the recency term.
  • Retiring the perception composites left Coordination, Trust and Shared Purpose with one or two indicators each. The limits page carries the detail.
  • Political uniformity is never treated as a capability.

How to argue with any of this and what would make each decision fall.