Skip to content
NCBNational
Capability
Benchmark

Challenge this

The benchmark is built to be argued with

Every decision names the evidence that would overturn it. Known failures sit beside the scores they affect. Bring a series, a case or an objection.

Public disputes stay attached to the number

0 disputes in the ledger. New submissions await maintainer review.

No disputes have been submitted yet. Use Challenge beside a score to file the first one.

Start with the known failures

10 artefacts are open: places where the model produces a number that misdescribes the world. Two have high severity.

IdSeverityWhat it is
A3highCoordination and Trust remain weakly separable from wealth
A12highCoordination and Trust are scored on thin evidence
A8structuralEvery correlation here is a hint, not a result
A10structuralThe frame is 52 countries wide, and they are not the world
A1mediumExperimentation is not measured, it is inferred from patents
A2mediumPer-capita normalisation flattens India
A7mediumLearning overstates Brazil and understates Korea, Estonia and Singapore
A9mediumCoordination reads far too low for small, competent states
A11mediumBuilding measures industrial output, and reads as delivery capacity
A6lowDoing Business indicators are frozen at 2019

The limits page carries the full entry, the test behind it and a possible fix.

Each decision says what could overturn it

The decision log lists each challenge clause, newest first. New evidence adds a superseding entry.

IdThe choiceWhat would overturn it
D60A series that cannot be scored is published beside the score, with the reason attachedEvidence that readers treat a check as a score, which would mean publishing it beside the dimension does the harm the exclusion was meant to avoid, and the row should be retired instead. Or a change in the series that breaks its correlation with income, which would make it an indicator and move it to the other registry.
D59The data publishes a quoting contract for automated readers, and no MCP serverEvidence that agents are a real consumer rather than an assumed one: referrer or user-agent data showing automated fetches of data/out, or a citation of an NCB number in a generated document that drops the confidence. Either would justify the enforcing surface. The other trigger is the Delphi layer becoming interactive, so a caller wants to run a scored what-if against the frame instead of reading a fixed file. That is a tool call, not a document, and it is the point where MCP earns its keep. It belongs as a route handler in apps/web reading @ncb/core and data/out the way the pages do, so that no third copy of the model exists.
D58The country's wiring is a system matrix, and its ramp is fixed rather than fittedA country whose map fills so few cells that the matrix reads as empty rather than sparse, which would mean the surface needs a coverage threshold before it renders. Or relation weights in the schema, which would make a count the wrong quantity to put in a cell.
D57Trust is two families, and it publishes nothing until both are answeredA harmonised social-trust series and a comparable institutional-performance series that together clear the floor, whose combined dimension survives the D42 wealth-attribution test. That closes the decision by satisfying it. The decision is wrong instead if the two families turn out to correlate above the redundancy threshold once both are measured, which would mean Trust asks one question after all and the split encodes a distinction the world does not make.
D56The institution map reads as a directory and a relation ledger, never as a drawn networkA country map whose relation set does not sort cleanly into the four families, which would mean the families encode Brazil rather than the model. Or evidence from readers that the whole-network shape is the question they arrive with, which would make the matrix the page's first surface rather than its missing one.
D55Budget execution is the first API-backed Coordination replacementCross-country delivery evidence showing that budget execution fidelity has little relationship to the ability of independent actors to complete shared objectives, or a better comparable indicator that measures joint delivery directly and survives the same wealth-attribution test.
D54An institution map explains capability and never scores itA cross-government official register that publishes the same organisations and capability relationships with stable identifiers, or a cross-country construction whose network measures survive differences in documentation coverage. The first would replace the curated skeleton. The second would justify considering a network measure as evidence, but it would still need a separate decision before entering any score.
D53The radar answers a pointer, and it has no resting numberEvidence that readers want the numbers printed on the chart itself. Nine labels on a 260 unit square collide at the sizes this chart is drawn at, which is why they are not there, but a larger chart on the country page could carry them and would make the readout redundant.
D52A candidate series is tested before it becomes a registry rowA second ingester. The probe speaks World Bank, and the series that would fill Coordination and Trust are the ones the World Bank does not publish, so its long-term job is to be generalised or retired.
D51Latin America is covered whole, not sampledEvidence that thin-coverage countries distort an indicator's Tukey fences enough to change well-evidenced countries' readings, or a regional dataset that covers the twelve better than the World Bank does.
D50The falsification conditions are an index, not a second documentA decision entry that needs more than one falsification condition, or an artefact whose severity is a judgment rather than a word. Both would mean the structure belongs in data rather than in prose, and the entries would move to a file the documents render from instead of the other way round.
D49The fetch describes itself once, and the viewer prints that descriptionA second ingester. The World Bank shape is generalised in sources.ts only as far as one publisher needs, and an ILOSTAT or OECD adapter would want a per-publisher description rather than a World Bank one with other publishers listed beside it.
D48The panel column takes the same wealth test as the indicatorsA gateway run whose panel column holds under 0.70 on the dimensions the indicators cannot measure. That is the evidence that would let a panel estimate carry a dimension, and this diagnostic is how it gets checked.
D47Every country sets the frame it is scored againstAbsolute anchoring per indicator, which would remove the country set from the scale entirely and make a score comparable across versions. It needs a defensible floor and ceiling for each of the 33 scored indicators, hand-set and defended one at a time, and it trades a frame that describes this set for one that describes a claim about the world.
D46A case study is an address, and the list of them is a filterThe corpus growing past the point where shipping it to the browser is reasonable, which makes the index a server-rendered list reading the same query parameters. Or evidence that readers never use the record pages, which would mean the anchor was enough and the cost was the two extra clicks to reach a mechanism.
D45A dimension with fewer than two observed indicators publishes no scoreReplacement indicators landing in Coordination and Trust, which is what A12 asks for and would make the floor moot for them. Or evidence that readers treat an empty axis as a zero anyway, which would mean the withholding needs words on the chart and not only in the table.
D44The homicide rate is retired, because it was the wealth signal it was hired to removeA country set wide enough that homicide stops tracking income, which would mean the correlation was an artefact of ten reference countries. Or a behavioural trust measure landing beside it, which would let the dimension carry homicide without homicide carrying the dimension.
D43The country page opens on the agenda, and the split has one homeA country page that reads better with the shape first and the agenda second, which would move the block under the radar. Or a hold list long enough to deserve a column of its own.
D42A candidate is judged on what it does to its dimension, not on its own correlationEvidence that a high delta is routinely the right answer, which would make the diagnostic noise. Or a dimension-level wealth test derived from a partial correlation rather than a leave-one-out, which would be the stronger statistic if the country set ever grows enough to support it.
D41The pipeline stops discarding its own warnings, and the viewer carries themA run where the exemplar exclusion empties most dimensions, which would argue for annotating clamped exemplars instead of excluding them. A persistent ingest failure record, which would replace the silent carry-forward.
D40A date is metadata, and a document a reader is told to open is a linkA viewer page that renders the limits document itself, which would give the viewer a local target for {limits} while the markdown keeps the repository link. Or a versioned documentation site, which would replace the branch in docHref with a release.
D38The homepage is global, and one country is a layer on top of itUse showing readers cannot enter the grid without a worked example. A decision to make this viewer a country-specific product, which would put the focal case back and make the global grid the secondary surface.
D37The output directory describes itselfA consumer that standard tooling cannot serve from the Data Package, which would argue for a real API. Or the schemas drifting from the files in practice, which would mean generation is wired to the wrong place and validation of the emitted output should be added to bench validate.
D36A record states where the delivery stands, and a claim can carry two numbersRecords that repeatedly need a third number, or statuses that keep landing in the wrong bin — either means the lifecycle deserves a dated series, not two stamps, and the slot design should be replaced rather than extended.
D35The agenda is computed, and language is an interpretation layerA reader study or field use showing the raise cutoff misleads at 50, which would justify deriving it from the score distribution instead. Records gaining translated fields, which would remove the mixed-language cost. A lexicon whose translation drifts from the registry meaning, which would justify review rules for lexicon changes rather than plain PRs.
D34Countries are identified by ISO 3166-1 alpha-3, recorded after the factA primary data source that keys on something else and outweighs the World Bank in the registry. That would justify an internal ID with per-source mappings, and this entry should be superseded when it happens.
D33Evidence records get an inclusion rule before they get more recordsA corpus that satisfies every test and still reads as advocacy — that would mean selection bias lives somewhere the rule does not reach, and the rule needs to move from authoring discipline to independent review. Or a demonstrated need to document sub-national or non-state deliveries, which test three currently excludes.
D32Evidence is drawn as a gradient, and every chart is a controlNothing foreseeable. If the gradient reads as noise at small sizes, the icon-labelled radars can fall back to the threshold.
D31A record carries its mechanism, and patterns get their own pageEnough records to compare mechanisms rather than list them, at which point the page becomes a query rather than a list.
D30Every number opens onto the field it sits inNothing foreseeable for the panel itself. A reader who wants the underlying distribution rather than the normalized one needs a raw axis with a log option, which is a further step.
D29The eighth dimension keeps one name, and the axes carry their marksEvidence that the grid is unreadable without words, which would mean going back to labels and making the cards larger.
D28Icons are copied in, one per concept, never aloneA need for many more icons, at which point installing the package beats maintaining a copied set.
D27Forty countries, and output split one file per countryA frame rebase, which would be a versioned event with its own decision, or a move to absolute anchoring per indicator.
D26Every term is defined once, in plain language, in the modelNothing foreseeable. A second surface that needs the same definitions, such as a printed report or an API, reads the same file.
D25Every point carries its provenance, and every run records what movedA move to per-country output files, which would change where the series lives but not what it has to carry.
D24Two spans for a dimension, and a full line for every indicatorEnough indicator history to compute a full-dimension basket at both ends, which would make the matched basket unnecessary and both spans directly comparable to the headline score.
D23The perception layer is retired, and the cost is visibleObservable replacements: court throughput and case clearance, budget execution rates, cross-agency programme delivery, voter turnout, volunteering rates, civic participation. Each one that lands raises the coverage this decision knocked down. If none land, the honest conclusion is that Coordination and Trust cannot be measured with public data, and they should be reported as unmeasured rather than scored.
D22Momentum is measured on one ruler and a matched basketEnough indicator history to score a full basket at both ends, which would let momentum use the whole dimension rather than a subset. A move to absolute anchoring would also change what a fixed ruler means and this decision would need restating.
D21GEM is wired by hand, venture capital stays a gapAn inspectable venture capital or business R&D series that covers the reference set, or a GEM licensing change that stops the data being usable this way.
D20Documented deliveries are recorded as evidence and never scoredA comparable delivery series covering the reference set, at which point large_project_delivery leaves gap, the records become source notes on a scored indicator, and this decision is superseded.
D19Extended countries get no visual markingA case where the distinction changes how a number should be read, most likely a country clamping at 0 or 100. Flag outOfFrame on that cell rather than reinstating a badge on the country.
D18One display for every 0 to 100 scoreEvidence that readers misread band edges as real differences, or a move away from a frame-relative scale.
D17Confidence bands are fixed thresholds, and not a red-to-green scaleEvidence about how readers actually act on the bands, or a change to the confidence formula that shifts its range.
D16The normalization frame is pinned to the ten reference countriesA sustained pattern of outOfFrame cells, or a decision to move to absolute anchoring per indicator. Either way, rebasing is a versioned event: bump a frame version, re-publish, and say plainly that the old numbers are not comparable.
D15The World Bank is the only wired ingestion source in v0Writing the next adapter. Each one is independent work.
D14Provenance is stored, never inferredNothing.
D13Panel diversity comes from vendors, not from model sizeEvidence that stance dominates model, in which case several stances on one model would be as good and simpler to reason about.
D12Panel disagreement is recorded, not averaged awayEvidence that panel dissent is noise rather than signal — for instance if dissent does not correlate with low coverage across several runs.
D11Delphi output never enters the capability scoreNothing at v0. Any future blending must be a new, explicit, named field, never a change to score.
D10Inspectability is a hard filter on sourcesA source opening its microdata, or an explicit decision to accept composite indices with a recorded quality penalty via source.tier.
D9Gap indicators stay in the registryNothing. Do not delete gaps to make numbers look better.
D8Only the most recent observation, no trendsThe cross-section holding up. Trend is the obvious v1 extension and the data is already fetched from 2000 onward.
D7Winsorize with Tukey fences at k = 3Evidence that a specific indicator's distribution needs a transform rather than a clip. Prefer adding a transform to the registry over lowering k.
D6Dimensions are scored at any coverage above zeroA published deliverable. Before anything is published, either introduce a coverage floor below which a dimension reports null, or mark low-coverage cells visually in every output. Open, unresolved.
D5Missing data is dropped, never imputedNothing. But see D6 — this is why a floor may be needed.
D4Confidence is reported beside the score, never inside itNothing we can foresee. This is close to load-bearing.
D3Equal weights inside a dimensionDelphi construct-validity ratings that are stable across several real panels. Weight by panel-rated validity only once the panel itself has been shown to agree.
D2Normalisation is relative to the country set, not to an absolute frontierA move to a large enough country set that absolute anchoring becomes possible, or an explicit decision to anchor against fixed reference values per indicator.
D1Nine dimensions, no headline rankingEvidence that the nine dimensions are so correlated that the shape carries no information beyond a single factor. Watch diagnostics.dimensionPairs: if most pairs sit above r ≈ 0.9 on a larger country set, the dimensional structure is not earning its keep.

The decision log holds 59 entries with each choice, reason and cost.

Ways to object

Use the repository so the argument remains visible after it is settled.

  • Dispute a decision. Name the decision id and evidence in an issue. A new entry supersedes a decision; the old one stays visible.
  • Fill a gap. Point to a published series covering at least two countries with comparable definitions, an open URL, publisher, reference period and method. National statistical sources are welcome.
  • File an evidence record. Document a national delivery with one published number and a statement of what it does not show. Records never enter a score, and one in five must document erosion or collapse.
  • Add a language. Add a lexicon data file. The English ground layer lets readers check the translation against its source.

CONTRIBUTING.md has the rules for each, docs/EVIDENCE.md has the inclusion test for records, and docs/WHY.md states the claim under test. Objections go to the issue tracker.

How to cite it

Quote the dataset version with every score. Adding a country changes the frame and restates all scores.

Envisioning (2026). NCB, the National Capability Benchmark, dataset 4.4.0. https://github.com/envisioning/national-capability-benchmark

The code is MIT, in LICENSE. Data keeps the terms of its publishers, listed in NOTICE.md . Keep the attribution when redistributing a number.

This is a prototype. Read the limits and the confidence beside a score; a thin dimension rests on one or two indicators and cannot carry an argument alone.