Method
Capability is measured apart from wealth
The benchmark asks whether a country's capacity to act is distinct from its income. Every choice is open to challenge.
Why this exists
If the claim is right, countries with similar incomes will have different capability profiles. If not, the dimensions mostly track income.
- All 52 countries set the scale. A frame built around one country would only describe that country.
- A high score is not a recipe. Mechanisms depend on local conditions. The shape says where to look; context says what to build.
- Capability changes below the national level, in groups small enough to act. A country score is a coarse proxy for their conditions.
- Treat this as a measuring instrument. It tests whether an attempt to raise a capability worked. Confidence, gaps and revisions sit beside each score.
- The diagnostics tests whether dimensions collapse into income. The limits page records known failures.
This measures capability, not wealth
The benchmark measures capacity to anticipate, coordinate, learn, adapt and build under uncertainty. Wealth, quality of life and popularity are outside its scope.
| Dimension | Question |
|---|---|
| Anticipation | How capable is the country of identifying and preparing for emerging change? |
| Agency | How able are individuals and organizations to turn an intention into action? |
| Coordination | How effectively can independent actors organize around shared objectives? |
| Trust | How much cooperation is possible beyond immediate personal networks? |
| Learning | How effectively does the country acquire, distribute, and update knowledge? |
| Experimentation | How easily can new approaches be attempted, tested, abandoned, and improved? |
| Adaptability | How effectively can the system respond when circumstances change? |
| Building | How capable is the country of turning plans and knowledge into functioning systems? |
| Shared Purpose | To what extent can people imagine themselves as participants in a common project? |
How an indicator becomes a score
- Take the latest comparable value for each country and indicator, with its source and year.
- Apply the declared transform: per million people, log, or distance from a target.
- Winsorize extreme outliers with Tukey fences at three interquartile ranges.
- Normalize to 0 through 100 against the frame set by all 52 countries. Reverse indicators where lower is better.
- Average the available indicators inside a dimension with equal weights.
- Compute confidence separately as coverage ร recency ร source quality.
Missing indicators lower coverage and drop out of the mean. Nothing is imputed. Equal weighting keeps v0 easy to challenge.
The registry has 69 indicators: 35 with data, 26 gaps and 8 retired rows. Gaps have no comparable dataset; retired rows have a rejected one. Both lower confidence and define the collection agenda. 1 values come from reproducible source adapters and 2 from published tables, with retrieval dates stored.
National scores have a second reading
The benchmark compares countries at the national level, while selected destination pages show constituent data as corroboration or context.
The comparison layer reads observations with geometry=national. That keeps every country on the same unit of comparison and means a state or province cannot silently move a national score.
A destination page may also show published values for constituent units. Each fixture declares whether those values aggregate to the national figure, stand independently, or provide context only. The rule stays beside the values, and the source remains visible. These rows corroborate, question or explain a national result; they do not create per-state capability scores.
Every indicator states what it measures
The dataset labels each indicator as C, I, O or P so the classification can be checked.
- C, direct capability measure. Measures the thing itself. Days to register a company measures how hard it is to start a business.
- I, capability input. Measures something that supports the capability. Research spending is an input. It does not prove a country reads the future well.
- O, downstream outcome. Measures a result that usually follows from the capability. Patents can show experimentation, but also defensive filing. Korea files at volume, so its patent number says less than it looks.
- P, perception proxy. Records what people or experts say, not what they did. The Worldwide Governance Indicators aggregate expert opinion. Seven were retired because they tracked income per head more than capability.
Source quality affects confidence only
Each tier affects confidence. Delphi estimates have the least weight.
| Tier | Weight |
|---|---|
| official statistical | 1.00 |
| international organization | 0.95 |
| academic survey | 0.85 |
| composite index | 0.70 |
| expert panel | 0.50 |
| llm delphi | 0.30 |
A tier says who published a number. The sources page lists the publisher, database and request.
Thin evidence appears on the chart
Confidence never enters the score. Thin evidence gets a dashed edge and hollow point. The gap widens as confidence falls.
- Coverage is the share of a dimension's indicators that have a value.
- Recency decays after two grace years, over a twelve-year window, to a floor of 0.1.
- Source quality is the mean tier weight of the values that are present.
- The product stays below 1 in practice, so the bands reflect real values.
Momentum uses only matching indicators
Momentum shows score change over time on the current frame. Only indicators observed at both ends count.
- Historical values use today's frame, so the change reflects the country.
- The same indicators are used at both ends. A new indicator cannot create movement.
- The basket may be smaller than the dimension, so the trend level can differ from the score. Its size is printed beside the trend.
- Ten-year and twenty-year spans are published. A missing span shows how far the data reaches.
- Each indicator has its own line back to 1990 where data exists. Nothing is carried forward or filled in.
- Each point carries the published value, normalized value and source tier.
- Each run compares its data with the previous file and logs restated, added or dropped values.
- Values more than five years old do not count for a year. Historical values outside the frame clamp to 0 or 100, and the clamp is recorded.
- Adoption indicators often rise for every country. Compare each change with the median before calling it progress.
Documented deliveries sit outside the score
Evidence records describe work that a gap indicator cannot measure. They do not change scores.
- Each record carries a published number, reference period, source and retrieval date.
- Each record states what the case does not show.
- Records never affect scores or confidence.
- A gap becomes scorable when a comparable series covers at least two countries.
A panel reviews what the data misses
Each panelist has a fixed stance. The panel interprets the source-backed evidence and reviews the indicators.
- Round 1: each panelist scores dimensions with thin source coverage from the evidence brief and its knowledge.
- Round 2: each panelist sees the anonymized round-1 scores and rationales, then revises or defends its scores.
- We keep the median and interquartile range. A range above 25 points is unresolved disagreement.
- Panel estimates stay in their own file and never enter the indicator score or confidence. The blended view uses one only when no indicator is observed.
- The panel also rates each indicator's class, validity, wealth-proxy risk and redundancy.
- The Delphi page shows the current run and its provenance. The active run is a working session, not a panel.
The country set defines the scale
All 52 countries set each indicator's fences and endpoints. The frame stays fixed within a version. Adding a country rebases it and requires a major version bump.
| Country | Why it is included |
|---|---|
| Brazil | Primary reference case; large, diverse upper-middle-income democracy |
| United States | High innovation and agency; large-scale institutional complexity |
| Netherlands | Strong institutions, coordination and social trust |
| Switzerland | Highly decentralized but unusually coordinated system |
| Singapore | High-capacity, highly coordinated small state |
| South Korea | Rapid development, technology adoption and execution capacity |
| Estonia | Small state known for digital institutional experimentation |
| India | Large, diverse emerging economy with significant bottom-up capability |
| Chile | Latin American comparison with relatively strong institutions |
| South Africa | Unequal, institutionally complex middle-income comparison case |
| Mexico | Second largest Latin American economy; deep manufacturing base tied to North America |
| Argentina | Strong research and human capital against repeated macroeconomic rupture |
| Colombia | Large economy rebuilding state capacity after prolonged internal conflict |
| Peru | Sustained growth with persistent institutional instability and high informality |
| Uruguay | Small state with the strongest institutional trust in the region |
| Costa Rica | Small state that moved into high-value manufacturing and services without an extractive base |
| Germany | Large manufacturing economy coordinated through federal states and industry associations |
| France | Centralized state with a long tradition of directing industrial policy |
| United Kingdom | Services and finance concentration with weak recent productivity growth |
| Spain | Southern European comparison with strong infrastructure delivery and high unemployment |
| Poland | Post-socialist convergence case that rebuilt institutions and industry together |
| Sweden | High-trust Nordic state with an unusual mix of large firms and startups |
| Finland | Small state with an institutionalized foresight function and strong measured learning |
| Ireland | Small open economy whose output figures are distorted by foreign direct investment |
| Canada | Resource-rich federal democracy with persistent productivity questions |
| Australia | Resource exporter far from its markets, with high administrative capacity |
| Japan | Aging high-capability manufacturer testing whether execution survives demographic decline |
| China | State-directed development at continental scale, the clearest contrast to the rest of the set |
| Indonesia | Large archipelago state coordinating across extreme geographic dispersion |
| Vietnam | Fast industrial catch-up on a low income base |
| Philippines | Services export and remittance economy with weak industrial depth |
| Malaysia | Middle-income manufacturer testing the move into higher-value production |
| Thailand | The middle-income trap as a case: strong assembly, thin innovation, aging fast |
| Turkey | Industrial middle power with repeated macroeconomic instability |
| Israel | Small state with the highest venture density in the world and deep civil divisions |
| United Arab Emirates | State-led diversification away from oil, executed quickly and from the top |
| Nigeria | Largest African economy, with capability concentrated outside the state |
| Kenya | East African digital finance leader, where a private rail reached population scale |
| Rwanda | Small state with a strong delivery reputation and a narrow political base |
| Ethiopia | Large low-income state attempting state-led industrialization under conflict |
| Bolivia | Resource-dependent landlocked state with strong social movements and weak formal institutions |
| Paraguay | Landlocked agro-exporter with a small state and fast recent growth |
| Ecuador | Dollarized oil exporter cycling through repeated institutional redesigns |
| Venezuela | State collapse case; the sparse recent data is itself the finding |
| Panama | Services and logistics hub built around a single asset it operates well |
| Guatemala | Largest Central American economy with a chronically underfunded state |
| Honduras | Low-capacity state where remittances stand in for absent institutions |
| El Salvador | Small state undergoing a centralized security-led institutional rebuild |
| Nicaragua | Authoritarian consolidation case with thinning independent statistics |
| Dominican Republic | Fast-growing tourism and services economy with weak public delivery |
| Cuba | State-run system outside most international statistical programs, so coverage is thin by design |
| Haiti | State breakdown case; shows what the frame floor looks like |
The assumptions are public
The decision log records each choice and what evidence would overturn it.
- 0 and 100 are the weakest and strongest values among the 52 countries. They are not a sample of the world, and a low score is not a percentage of capability.
- Scores use only the latest observation. Trends use a matched basket against today's frame. Nothing is back-filled or imputed.
- With 52 countries, diagnostics are hints, not established results.
- Doing Business series are frozen at 2019 and are marked down by the recency term.
- Retiring the perception composites left Coordination, Trust and Shared Purpose with one or two indicators each. The limits page carries the detail.
- Political uniformity is never treated as a capability.
How to argue with any of this and what would make each decision fall.