Glossary
The glossary explains the terms
This page explains the terms used by the benchmark. If a letter, band or dashed line is unclear, start here.
The four letters beside every indicator
Each indicator has a class: C, I, O or P. The registry keeps the label so it can be checked.
- Cdirect capability measure
- Measures the thing itself. Days to register a company measures how hard it is to start a business.
- Icapability input
- Measures something that supports the capability. Research spending is an input. It does not prove a country reads the future well.
- Odownstream outcome
- Measures a result that usually follows from the capability. Patents can show experimentation, but also defensive filing. Korea files at volume, so its patent number says less than it looks.
- Pperception proxy
- Records what people or experts say, not what they did. The Worldwide Governance Indicators aggregate expert opinion. Seven were retired because they tracked income per head more than capability.
What is being measured
- Capability
What a country is able to do, separately from how rich it is.
The benchmark asks what a country can do: anticipate change, coordinate, learn, adapt and build under uncertainty. Wealth and capability are related, but the design keeps them separate enough to compare.
- Dimension
One of the nine capabilities, each scored on its own.
There are nine dimensions. Each has its own question and 0 to 100 score. Scores stay separate, so countries with the same average can have different shapes.
Building asks whether a country turns plans into working systems. Trust asks how much cooperation is possible beyond people who already know each other.
- Indicator
One published statistic used as evidence for one dimension.
Each dimension uses published indicators with a source and year. The score averages indicators with data; raw values stay visible for checking.
- Indicator family
A group of indicators inside one dimension that answer the same question.
Some dimensions ask two questions under one name. Trust asks whether people rely on strangers and whether they rely on institutions, and an indicator belongs to one of those two families. The family changes nothing about the score, which stays the equal-weight average of whatever is observed. It exists so the diagnostics can report which family the evidence came from, because several readings of one question are not several independent signals.
Trust holds a social family and an institutional family. The current release observes one row in each family where it scores, while court performance remains a gap.
- Measurement class
Whether an indicator measures the capability, an input, a result, or an opinion.
The registry labels every indicator C, I, O or P. C measures the capability itself; I measures an input; O measures a result; P records a perception. The benchmark prefers C and I. P was retired after it tracked income too closely.
Time to register a company is C. Research spending is I. Patents are O. An expert survey about government quality is P.
How a number is made
- Score
A position from 0 to 100 inside a fixed comparison frame.
A dimension score runs from 0 to 100. Zero is weakest and 100 strongest in this frame. A 10 is near the floor, not 10 percent of a capability.
- Score band
Four named ranges a score falls in: weak, below middle, above middle, strong.
Each score falls into one of four bands. The labels are relative to this frame. Check confidence before interpreting a weak score.
- Comparison frame
The countries whose values fix the ends of every scale.
All countries set each indicator’s endpoints. The frame stays fixed within a version, so score changes reflect data. Adding a country changes the frame and restates scores.
- Frame rebase
A new scale after the country set changes.
Adding a country can move the endpoints and restate published numbers. The dataset gets a major version bump, the benchmark is rescored, and old and new numbers cannot be compared. The change is announced.
- Normalization
Turning a raw value into a 0 to 100 position, reversing where lower is better.
Raw values use different units. Normalization turns each into a 0 to 100 position within its indicator frame. Lower-is-better indicators are reversed so higher always means better.
- Distance from target
A transform that rewards values close to a defined target.
Some measures are best near a target. Budget execution scores distance from 100, so the closest country ranks highest.
A budget execution value of 95 is five percentage points from the approved budget. A value of 130 is thirty points away.
- Winsorizing
Pulling extreme outliers back to a boundary so one country cannot stretch the scale.
An extreme value can compress every other country into a narrow band. Winsorizing clips values beyond three interquartile ranges before building the scale. It is used sparingly.
- Out of frame
A value beyond the ends of the scale, so its score is clamped and flagged.
Current values sit inside the frame by construction. Historical or late-arriving values can fall outside it, clamp to 0 or 100, and get flagged. The flag shows where information was lost.
- Ingest route
How a value gets into the dataset: from an API, a published table, or nowhere yet.
Every indicator declares one of five routes. World Bank values come from its API. Values from a reproducible source adapter are fetched or parsed by code tied to a named release. Values from published tables keep the retrieval date. A gap has no comparable dataset. A retired row has a rejected dataset. Both lower confidence.
Generalised interpersonal trust is parsed by the Joint EVS/WVS adapter from the publisher-weighted A165 results table. GEM indicators remain entered by hand from published tables.
How good the evidence is
- Confidence
How well a score is evidenced, reported beside it and never inside it.
Confidence is coverage × recency × source quality, from 0 to 1. It describes the evidence, not the score. The same score can have very different confidence.
Coordination for every country currently sits at 0.08 confidence, because one indicator of seven has data and it stopped in 2019.
- Coverage, recency, source quality
The three parts of confidence.
Coverage is the share of indicators with a value. Recency declines after two grace years over a twelve-year window. Source quality is the average source tier. The three multiply.
- Confidence band
Four named ranges: very thin, thin, usable, good.
Confidence has four bands. Very thin means the score rests on one or two indicators and should not be quoted alone. Good means most indicators are present, recent and official. Thin evidence appears as a dashed radar edge and hollow point.
- Source tier
Who published a number, ranked from national statistics office down to a model panel.
Every value carries a source type, such as a statistical agency, international organization, survey or model panel. The tier affects confidence, not the score, and shows when a line mixes sources.
- Wealth proxy
An indicator that mostly restates income per head.
Each indicator is correlated with log GDP per capita. Above 0.7, it is flagged as a wealth proxy and removed in a sensitivity test. The panel gets the same test.
What is missing
- Gap
An indicator the model asks for that no comparable dataset covers.
A gap stays in the registry, lowers confidence and appears in the collection agenda. Removing it would make the numbers look better without adding evidence.
Cost and schedule performance of major public projects is a gap. It is probably the single best measure of execution and no comparable international dataset exists.
- Retired indicator
A dataset this project rejected, with the reason recorded.
A retired indicator stays in the registry, is not fetched or scored, and lowers coverage like a gap. The reason remains available for challenge.
- Known artefact
A place where the model produces a number that is wrong about the world.
Artefacts are measurement failures recorded with severity, evidence and a possible fix. The viewer publishes them on the limits page. Read them before quoting a score.
Coordination, Trust and Shared Purpose currently rest on one or two indicators each, so their scores move enough to mislead.
How things change over time
- Momentum
How much a dimension moved over ten or twenty years, on the same ruler.
History uses today's frame, so score change reflects the country. Ten-year and twenty-year spans answer different questions. Values up to five years old can count at a span end, and clamped values are reported.
Brazil gained 26.2 points on Agency over ten years, against a median of 11.4 across all 40 countries, with two of the four basket indicators clamped at the frame edge.
- Matched basket
Only the indicators present at both ends of a span are used for a trend.
A trend uses only indicators observed at both ends of the span. The basket can be smaller than the full dimension, so its level can differ from the score.
- Indicator line
One indicator's own history, as far back as its data goes.
An indicator is comparable with itself, so its line reaches back to 1990 where data exists. Each point carries the published, normalized and source-tier values. Nothing is filled in or carried forward.
What sits beside the score
- Geometry
The spatial level an observation describes, such as a country or state.
National geometry is the comparison layer. State, province, region and municipality geometries describe constituent units and can appear on a destination page without entering the national score.
- Reconciliation rule
The declared relationship between a national value and its constituent values.
An aggregate can be read as a national quantity assembled from its parts. An independent value should be compared alongside the national value. A context-only value adds detail without claiming that the two layers can be combined.
- Behavioral check
A published series shown next to a dimension and left out of it.
Some series measure something real about a capability and still fail the tests this benchmark applies before a number is scored, usually because they mostly track national income. A check is fetched and published like an indicator and then excluded from the scale, the average, the indicator count and the confidence. The reason it is not scored travels with the number, so a reader can weigh the evidence without the benchmark asserting it.
Bribery incidence asks whether a firm was itself asked for a bribe. It reads on trust and it also tracks income, so Trust publishes it beside the score and never inside it.
- Evidence record
A documented case for something the indicators cannot measure. It does not change the score.
An evidence record describes a country doing something the current indicators miss. It includes a published number, the period it covers, its source, limits and delivery status. It does not affect scores or confidence. If a comparable series later covers at least two countries, the gap can become an indicator.
Brazil’s records run from Embrapa in 1973 to Pix in 2020. The immunisation programme is recorded as operating below its peak: 99 percent coverage in 2003, 91 percent in 2024.
- Institutional capability network
A sourced map of which organisations hold capability and how they constrain or support one another.
A country-specific network maps public institutions and selected outside organizations through sourced links such as funding, regulation, audit, appointment, training and delivery. It guides investigation and never affects scores or confidence.
Brazil’s first network links the federal backbone to a São Paulo pilot, including the BNDES, Finep, Enap, the STF, the STJ, FAPESP, state universities and the municipality of São Paulo.
- Delphi panel
Language models with fixed analytical stances, interpreting what the data misses.
A panel of models with different analytical stances reviews thin or questionable dimensions from the source-backed evidence brief and audits the indicators. In round two, panelists see anonymized reasoning from round one. Panel estimates stay separate from indicator scores, observations and confidence.
- Provenance
How a panel run was produced, recorded on the run file itself.
A run records whether it is a gateway panel, working session, human panel or mock run. Mock runs exercise the pipeline and are never country evidence. The label lives in the file.
- Dissent
Panel disagreement above the reporting threshold.
The panel keeps a median and interquartile range. A cell is unresolved when its middle half spans more than a quarter of the scale. Stable disagreement is a result.
- Capability agenda
The scores turned into a list of things to do, computed from the data.
A country document generated from scored output. Low scores with usable evidence are items to raise. Thin evidence becomes an item to measure first. The rest are holds. Declared gaps form the measurement agenda. The agenda regenerates with the data.
The Brazil agenda lists Building, at 9.4 with usable confidence, as the first dimension to raise, and Coordination, at confidence 0.08, as a dimension to measure before managing.
- Interpretation layer
One language's rendering of the ground data. The numbers never translate.
The ground layer keeps English ids, registry definitions and JSON output. Lexicons translate vocabulary and document strings from that data. They cannot change numbers. Missing translations fall back to registry English.
BRA.json is the ground record. BRA.en.md and BRA.pt-BR.md render it through two lexicons.
- Blended score
The published fallback: the indicator score, or the panel estimate when no indicator evidence exists.
Every dimension carries a blendedScore and blendedFrom label. The value is the indicator score when the dimension clears its coverage floor, the panel estimate only when no indicator is observed, or none. It is never a mix.
A dimension with one observed indicator remains unmeasured and does not fall back to Delphi. The fallback is reserved for a dimension with no observed indicators.
- Revision log
The append-only record of what each ingest restated, added or dropped.
Each ingest compares its data with the previous file and appends restated, added or dropped values. The log records when a published number changes.