Challenge this
The benchmark is built to be argued with
Every decision names the evidence that would overturn it. Known failures sit beside the scores they affect. Bring a series, a case or an objection.
Public disputes stay attached to the number
0 disputes in the ledger. New submissions await maintainer review.
No disputes have been submitted yet. Use Challenge beside a score to file the first one.
Start with the known failures
10 artefacts are open: places where the model produces a number that misdescribes the world. Two have high severity.
| Id | Severity | What it is |
|---|---|---|
| A3 | high | Coordination and Trust remain weakly separable from wealth |
| A12 | high | Coordination and Trust are scored on thin evidence |
| A8 | structural | Every correlation here is a hint, not a result |
| A10 | structural | The frame is 52 countries wide, and they are not the world |
| A1 | medium | Experimentation is not measured, it is inferred from patents |
| A2 | medium | Per-capita normalisation flattens India |
| A7 | medium | Learning overstates Brazil and understates Korea, Estonia and Singapore |
| A9 | medium | Coordination reads far too low for small, competent states |
| A11 | medium | Building measures industrial output, and reads as delivery capacity |
| A6 | low | Doing Business indicators are frozen at 2019 |
The limits page carries the full entry, the test behind it and a possible fix.
Each decision says what could overturn it
The decision log lists each challenge clause, newest first. New evidence adds a superseding entry.
| Id | The choice | What would overturn it |
|---|---|---|
| D60 | A series that cannot be scored is published beside the score, with the reason attached | Evidence that readers treat a check as a score, which would mean publishing it beside the dimension does the harm the exclusion was meant to avoid, and the row should be retired instead. Or a change in the series that breaks its correlation with income, which would make it an indicator and move it to the other registry. |
| D59 | The data publishes a quoting contract for automated readers, and no MCP server | Evidence that agents are a real consumer rather than an assumed one: referrer or user-agent data showing automated fetches of data/out, or a citation of an NCB number in a generated document that drops the confidence. Either would justify the enforcing surface. The other trigger is the Delphi layer becoming interactive, so a caller wants to run a scored what-if against the frame instead of reading a fixed file. That is a tool call, not a document, and it is the point where MCP earns its keep. It belongs as a route handler in apps/web reading @ncb/core and data/out the way the pages do, so that no third copy of the model exists. |
| D58 | The country's wiring is a system matrix, and its ramp is fixed rather than fitted | A country whose map fills so few cells that the matrix reads as empty rather than sparse, which would mean the surface needs a coverage threshold before it renders. Or relation weights in the schema, which would make a count the wrong quantity to put in a cell. |
| D57 | Trust is two families, and it publishes nothing until both are answered | A harmonised social-trust series and a comparable institutional-performance series that together clear the floor, whose combined dimension survives the D42 wealth-attribution test. That closes the decision by satisfying it. The decision is wrong instead if the two families turn out to correlate above the redundancy threshold once both are measured, which would mean Trust asks one question after all and the split encodes a distinction the world does not make. |
| D56 | The institution map reads as a directory and a relation ledger, never as a drawn network | A country map whose relation set does not sort cleanly into the four families, which would mean the families encode Brazil rather than the model. Or evidence from readers that the whole-network shape is the question they arrive with, which would make the matrix the page's first surface rather than its missing one. |
| D55 | Budget execution is the first API-backed Coordination replacement | Cross-country delivery evidence showing that budget execution fidelity has little relationship to the ability of independent actors to complete shared objectives, or a better comparable indicator that measures joint delivery directly and survives the same wealth-attribution test. |
| D54 | An institution map explains capability and never scores it | A cross-government official register that publishes the same organisations and capability relationships with stable identifiers, or a cross-country construction whose network measures survive differences in documentation coverage. The first would replace the curated skeleton. The second would justify considering a network measure as evidence, but it would still need a separate decision before entering any score. |
| D53 | The radar answers a pointer, and it has no resting number | Evidence that readers want the numbers printed on the chart itself. Nine labels on a 260 unit square collide at the sizes this chart is drawn at, which is why they are not there, but a larger chart on the country page could carry them and would make the readout redundant. |
| D52 | A candidate series is tested before it becomes a registry row | A second ingester. The probe speaks World Bank, and the series that would fill Coordination and Trust are the ones the World Bank does not publish, so its long-term job is to be generalised or retired. |
| D51 | Latin America is covered whole, not sampled | Evidence that thin-coverage countries distort an indicator's Tukey fences enough to change well-evidenced countries' readings, or a regional dataset that covers the twelve better than the World Bank does. |
| D50 | The falsification conditions are an index, not a second document | A decision entry that needs more than one falsification condition, or an artefact whose severity is a judgment rather than a word. Both would mean the structure belongs in data rather than in prose, and the entries would move to a file the documents render from instead of the other way round. |
| D49 | The fetch describes itself once, and the viewer prints that description | A second ingester. The World Bank shape is generalised in sources.ts only as far as one publisher needs, and an ILOSTAT or OECD adapter would want a per-publisher description rather than a World Bank one with other publishers listed beside it. |
| D48 | The panel column takes the same wealth test as the indicators | A gateway run whose panel column holds under 0.70 on the dimensions the indicators cannot measure. That is the evidence that would let a panel estimate carry a dimension, and this diagnostic is how it gets checked. |
| D47 | Every country sets the frame it is scored against | Absolute anchoring per indicator, which would remove the country set from the scale entirely and make a score comparable across versions. It needs a defensible floor and ceiling for each of the 33 scored indicators, hand-set and defended one at a time, and it trades a frame that describes this set for one that describes a claim about the world. |
| D46 | A case study is an address, and the list of them is a filter | The corpus growing past the point where shipping it to the browser is reasonable, which makes the index a server-rendered list reading the same query parameters. Or evidence that readers never use the record pages, which would mean the anchor was enough and the cost was the two extra clicks to reach a mechanism. |
| D45 | A dimension with fewer than two observed indicators publishes no score | Replacement indicators landing in Coordination and Trust, which is what A12 asks for and would make the floor moot for them. Or evidence that readers treat an empty axis as a zero anyway, which would mean the withholding needs words on the chart and not only in the table. |
| D44 | The homicide rate is retired, because it was the wealth signal it was hired to remove | A country set wide enough that homicide stops tracking income, which would mean the correlation was an artefact of ten reference countries. Or a behavioural trust measure landing beside it, which would let the dimension carry homicide without homicide carrying the dimension. |
| D43 | The country page opens on the agenda, and the split has one home | A country page that reads better with the shape first and the agenda second, which would move the block under the radar. Or a hold list long enough to deserve a column of its own. |
| D42 | A candidate is judged on what it does to its dimension, not on its own correlation | Evidence that a high delta is routinely the right answer, which would make the diagnostic noise. Or a dimension-level wealth test derived from a partial correlation rather than a leave-one-out, which would be the stronger statistic if the country set ever grows enough to support it. |
| D41 | The pipeline stops discarding its own warnings, and the viewer carries them | A run where the exemplar exclusion empties most dimensions, which would argue for annotating clamped exemplars instead of excluding them. A persistent ingest failure record, which would replace the silent carry-forward. |
| D40 | A date is metadata, and a document a reader is told to open is a link | A viewer page that renders the limits document itself, which would give the viewer a local target for {limits} while the markdown keeps the repository link. Or a versioned documentation site, which would replace the branch in docHref with a release. |
| D38 | The homepage is global, and one country is a layer on top of it | Use showing readers cannot enter the grid without a worked example. A decision to make this viewer a country-specific product, which would put the focal case back and make the global grid the secondary surface. |
| D37 | The output directory describes itself | A consumer that standard tooling cannot serve from the Data Package, which would argue for a real API. Or the schemas drifting from the files in practice, which would mean generation is wired to the wrong place and validation of the emitted output should be added to bench validate. |
| D36 | A record states where the delivery stands, and a claim can carry two numbers | Records that repeatedly need a third number, or statuses that keep landing in the wrong bin — either means the lifecycle deserves a dated series, not two stamps, and the slot design should be replaced rather than extended. |
| D35 | The agenda is computed, and language is an interpretation layer | A reader study or field use showing the raise cutoff misleads at 50, which would justify deriving it from the score distribution instead. Records gaining translated fields, which would remove the mixed-language cost. A lexicon whose translation drifts from the registry meaning, which would justify review rules for lexicon changes rather than plain PRs. |
| D34 | Countries are identified by ISO 3166-1 alpha-3, recorded after the fact | A primary data source that keys on something else and outweighs the World Bank in the registry. That would justify an internal ID with per-source mappings, and this entry should be superseded when it happens. |
| D33 | Evidence records get an inclusion rule before they get more records | A corpus that satisfies every test and still reads as advocacy — that would mean selection bias lives somewhere the rule does not reach, and the rule needs to move from authoring discipline to independent review. Or a demonstrated need to document sub-national or non-state deliveries, which test three currently excludes. |
| D32 | Evidence is drawn as a gradient, and every chart is a control | Nothing foreseeable. If the gradient reads as noise at small sizes, the icon-labelled radars can fall back to the threshold. |
| D31 | A record carries its mechanism, and patterns get their own page | Enough records to compare mechanisms rather than list them, at which point the page becomes a query rather than a list. |
| D30 | Every number opens onto the field it sits in | Nothing foreseeable for the panel itself. A reader who wants the underlying distribution rather than the normalized one needs a raw axis with a log option, which is a further step. |
| D29 | The eighth dimension keeps one name, and the axes carry their marks | Evidence that the grid is unreadable without words, which would mean going back to labels and making the cards larger. |
| D28 | Icons are copied in, one per concept, never alone | A need for many more icons, at which point installing the package beats maintaining a copied set. |
| D27 | Forty countries, and output split one file per country | A frame rebase, which would be a versioned event with its own decision, or a move to absolute anchoring per indicator. |
| D26 | Every term is defined once, in plain language, in the model | Nothing foreseeable. A second surface that needs the same definitions, such as a printed report or an API, reads the same file. |
| D25 | Every point carries its provenance, and every run records what moved | A move to per-country output files, which would change where the series lives but not what it has to carry. |
| D24 | Two spans for a dimension, and a full line for every indicator | Enough indicator history to compute a full-dimension basket at both ends, which would make the matched basket unnecessary and both spans directly comparable to the headline score. |
| D23 | The perception layer is retired, and the cost is visible | Observable replacements: court throughput and case clearance, budget execution rates, cross-agency programme delivery, voter turnout, volunteering rates, civic participation. Each one that lands raises the coverage this decision knocked down. If none land, the honest conclusion is that Coordination and Trust cannot be measured with public data, and they should be reported as unmeasured rather than scored. |
| D22 | Momentum is measured on one ruler and a matched basket | Enough indicator history to score a full basket at both ends, which would let momentum use the whole dimension rather than a subset. A move to absolute anchoring would also change what a fixed ruler means and this decision would need restating. |
| D21 | GEM is wired by hand, venture capital stays a gap | An inspectable venture capital or business R&D series that covers the reference set, or a GEM licensing change that stops the data being usable this way. |
| D20 | Documented deliveries are recorded as evidence and never scored | A comparable delivery series covering the reference set, at which point large_project_delivery leaves gap, the records become source notes on a scored indicator, and this decision is superseded. |
| D19 | Extended countries get no visual marking | A case where the distinction changes how a number should be read, most likely a country clamping at 0 or 100. Flag outOfFrame on that cell rather than reinstating a badge on the country. |
| D18 | One display for every 0 to 100 score | Evidence that readers misread band edges as real differences, or a move away from a frame-relative scale. |
| D17 | Confidence bands are fixed thresholds, and not a red-to-green scale | Evidence about how readers actually act on the bands, or a change to the confidence formula that shifts its range. |
| D16 | The normalization frame is pinned to the ten reference countries | A sustained pattern of outOfFrame cells, or a decision to move to absolute anchoring per indicator. Either way, rebasing is a versioned event: bump a frame version, re-publish, and say plainly that the old numbers are not comparable. |
| D15 | The World Bank is the only wired ingestion source in v0 | Writing the next adapter. Each one is independent work. |
| D14 | Provenance is stored, never inferred | Nothing. |
| D13 | Panel diversity comes from vendors, not from model size | Evidence that stance dominates model, in which case several stances on one model would be as good and simpler to reason about. |
| D12 | Panel disagreement is recorded, not averaged away | Evidence that panel dissent is noise rather than signal — for instance if dissent does not correlate with low coverage across several runs. |
| D11 | Delphi output never enters the capability score | Nothing at v0. Any future blending must be a new, explicit, named field, never a change to score. |
| D10 | Inspectability is a hard filter on sources | A source opening its microdata, or an explicit decision to accept composite indices with a recorded quality penalty via source.tier. |
| D9 | Gap indicators stay in the registry | Nothing. Do not delete gaps to make numbers look better. |
| D8 | Only the most recent observation, no trends | The cross-section holding up. Trend is the obvious v1 extension and the data is already fetched from 2000 onward. |
| D7 | Winsorize with Tukey fences at k = 3 | Evidence that a specific indicator's distribution needs a transform rather than a clip. Prefer adding a transform to the registry over lowering k. |
| D6 | Dimensions are scored at any coverage above zero | A published deliverable. Before anything is published, either introduce a coverage floor below which a dimension reports null, or mark low-coverage cells visually in every output. Open, unresolved. |
| D5 | Missing data is dropped, never imputed | Nothing. But see D6 — this is why a floor may be needed. |
| D4 | Confidence is reported beside the score, never inside it | Nothing we can foresee. This is close to load-bearing. |
| D3 | Equal weights inside a dimension | Delphi construct-validity ratings that are stable across several real panels. Weight by panel-rated validity only once the panel itself has been shown to agree. |
| D2 | Normalisation is relative to the country set, not to an absolute frontier | A move to a large enough country set that absolute anchoring becomes possible, or an explicit decision to anchor against fixed reference values per indicator. |
| D1 | Nine dimensions, no headline ranking | Evidence that the nine dimensions are so correlated that the shape carries no information beyond a single factor. Watch diagnostics.dimensionPairs: if most pairs sit above r ≈ 0.9 on a larger country set, the dimensional structure is not earning its keep. |
The decision log holds 59 entries with each choice, reason and cost.
Ways to object
Use the repository so the argument remains visible after it is settled.
- Dispute a decision. Name the decision id and evidence in an issue. A new entry supersedes a decision; the old one stays visible.
- Fill a gap. Point to a published series covering at least two countries with comparable definitions, an open URL, publisher, reference period and method. National statistical sources are welcome.
- File an evidence record. Document a national delivery with one published number and a statement of what it does not show. Records never enter a score, and one in five must document erosion or collapse.
- Add a language. Add a lexicon data file. The English ground layer lets readers check the translation against its source.
CONTRIBUTING.md has the rules for each, docs/EVIDENCE.md has the inclusion test for records, and docs/WHY.md states the claim under test. Objections go to the issue tracker.
How to cite it
Quote the dataset version with every score. Adding a country changes the frame and restates all scores.
Envisioning (2026). NCB, the National Capability Benchmark, dataset 4.4.0. https://github.com/envisioning/national-capability-benchmarkThe code is MIT, in LICENSE. Data keeps the terms of its publishers, listed in NOTICE.md . Keep the attribution when redistributing a number.
This is a prototype. Read the limits and the confidence beside a score; a thin dimension rests on one or two indicators and cannot carry an argument alone.