Skip to content
NCBNational Capability Benchmark

Cheap intelligence makes capability the bottleneck

Intelligence is getting easier to reach. What stays scarce is the ability to choose a direction, coordinate people and build what the technology makes possible. That ability is what this benchmark tries to measure.

What changed

Agents lower the cost of execution, AI speeds up science, and robotics carries both into physical production.

A capability that once took decades to build can now arrive in months, through tools and knowledge developed somewhere else. Knowing how a thing works matters less when more people can reach the means to make it.

The bottleneck moves to choosing a useful direction, coordinating the people who will act, trusting their work, deploying what they build, maintaining it and scaling it. Those are institutional and social capacities. The benchmark already measures them, and falling execution costs make them matter more.

Where the claim holds, and where it fails

If a country's ability to act is separate from its wealth, the nine capabilities will not simply follow income. Here is how far that holds in the current data.

Countries can be rich without being equally able to anticipate change, coordinate, learn, experiment, adapt, build or hold a shared purpose. The benchmark tests whether those capacities can be observed on their own. If the claim holds, two countries at the same income have different capability shapes, and a country can raise a dimension before it gets richer. If it fails, the dimensions track GDP per head and the benchmark is an income table with extra steps.

The answer today is split. Five of the nine sit below the wealth-tracking line and four sit at or above it.

  1. Shared Purpose0.4546 countries
  2. Coordination0.4744 countries
  3. Experimentation0.6249 countries
  4. Trust0.6335 countries
  5. Building0.6450 countries
  6. Learning0.7250 countries
  7. Adaptability0.8250 countries
  8. Agency0.8550 countries
  9. Anticipation0.8750 countries

Absolute correlation with log GDP per capita, on a 0 to 1 axis. The line marks 0.7, the point at or above which the diagnostics call a dimension wealth tracking. Country counts differ because a dimension below the coverage floor publishes no score.

Shared Purpose moves least with income, at 0.45. Anticipation moves most, at 0.87, which is close enough to income that the project treats it as a known failure and says so beside the number.

The diagnostics run this test on every release and also rescore the model with the wealth-correlated indicators removed. The limits record which dimensions do not survive that removal. The method explains how a statistic becomes a score, and docs/WHY.md states the full argument with the evidence that would sink it.

A correlation with income is not proof that income causes the capability. It marks a dimension whose current indicators cannot separate the two, which is a data problem the project publishes rather than hides.

Leverage: what multiplies a capability

The second layer covers the resources through which a capability gets multiplied.

Singapore and Brazil can reach the same frontier model. The systems around that model decide what happens next. A country with capability and no leverage leaves its ideas undeployed. The layer would measure the resources:

  • AI access and compute
  • data, connectivity and energy
  • capital and technical talent
  • robotics and advanced manufacturing
  • scientific capacity and institutions that deploy emerging technologies

Velocity: how fast a capability moves

The third layer follows the rate at which a capability changes, because under fast change the level alone is misleading.

A country can hold a strong capability and still lose ground if its institutions cannot learn at the speed of the technology around them. The layer would measure:

  • AI adoption and skill acquisition
  • time from idea to prototype to deployment
  • institutional experimentation and regulatory adaptation
  • diffusion of new practices and business creation
  • scientific adoption curves

Both layers are computed already, as offline fixtures the viewer does not publish. Neither reaches a score, a confidence or an agenda. Decision D65 sets what each one has to pass first: a settled method, a six-month review record and a review from outside the maintainers. Until then the foundation is the published layer and the other two are research.

Brazil is the first field case

One country tests whether the frame says anything a national statistics office does not already say.

Brazil is where the thesis meets a real institutional setting, through Envisioning's Brasil Capaz work on raising individual and collective capability. The benchmark gives that work a destination: a shape to move and a set of indicators that says which parts are measured and which are guessed.

Brazil gets no special treatment in the model. What it has is a reading of its own: the Brazilian layer carries the institution map, the subnational spread and the agenda in Portuguese, beside the English profile. The programme argument belongs there, where its audience is.

What the benchmark is for

It gives a public-sector leader, a journalist and a citizen one question and one body of evidence.

The benchmark exists so that different readers can ask what a country can actually do and check the answer against its sources. It sits inside Envisioning's wider work: signals detect exponential change, the benchmark measures whether a society can absorb it, and the capability programmes and institutional design labs work on raising the number.

Read the decision record before quoting a result. This page is the frame for the work and no substitute for its sources or its limits.