Back to insights
Climate AssuranceNovember 2025Research note

Measurement Architecture for Carbon-Market Integrity

A technical account of what a carbon credit represents, where integrity fails, and how baselines, monitoring, verification, registries, and claims must connect.

Institutional analysis1,165 wordsBy Ram Labs ResearchEvidence reviewed 20 August 2026
Principal finding

A registry can prevent duplicate serial numbers, but it cannot prove that a baseline is credible, a reduction is additional, or a removal will persist. Integrity emerges only when the physical measurement, counterfactual, verification, authorization, transfer, retirement, and public claim are traceable as one evidence chain.

~29% global emissions under direct carbon pricing

World Bank 2026 estimate across emissions trading systems and carbon taxes, not a measure of credit-market coverage or policy stringency.

Evidence[1]
87 implemented direct-pricing policies

World Bank count in 2026; instruments differ substantially in scope, exemptions, allocation, and price.

Evidence[1]
$107bn public revenue in 2025

World Bank estimate for emissions trading systems and carbon taxes, expressed in US dollars; not carbon-credit transaction value.

Evidence[1][2]
~1bn tCO2e unretired credit pool in 2024

World Bank 2025 estimate; unretired supply is heterogeneous and does not by itself establish low quality or future usability.

Evidence[3]

Separate carbon pricing from carbon crediting

The phrase carbon market covers instruments with different mechanics. A carbon tax fixes a price; an emissions trading system fixes a cap and allows allowance trading; a crediting mechanism issues units against an estimated baseline for a project or programme. The World Bank reports that direct carbon pricing covered about 29% of global greenhouse-gas emissions through 87 implemented policies in 2026 and mobilized $107 billion for public budgets in 2025. Those figures describe taxes and trading systems. They do not establish the environmental quality of individual credits. Conflating the categories makes both policy analysis and procurement weaker.

A credit is not a physical tonne stored in a database. It is a quantified claim that emissions were lower, or removals higher, than under a specified counterfactual, subject to a methodology, monitoring period, uncertainty treatment, verification process, and use rule. The counterfactual cannot be observed directly. This is the central epistemic difficulty: registries can track issued units, but the causal claim originates in a model of what would otherwise have happened. Serious assessment therefore starts with methodology and evidence, not with a project photograph, rating badge, or token history.

Evidence[1][2][3]

Build a unit-level chain of custody

A decision-grade credit record should connect the activity identifier, host jurisdiction, methodology and version, baseline vintage, monitoring report, raw or reconciled measurements, uncertainty deductions, verifier opinion, issuance batch, serial number, ownership transfers, authorization status, corresponding adjustment where applicable, retirement, and final claim. The UNFCCC Article 6.4 registry assigns unique identifiers and records activity, host Party, vintage, holdings, transfers, and use. That is necessary infrastructure for traceability and avoidance of unit duplication. It is not a substitute for evaluating the underlying mitigation result.

The ledger should be append-only at the event level, with corrected records linked to rather than silently replacing prior states. Documents need cryptographic hashes, stable versions, responsible parties, and timestamps. Public disclosure should expose enough evidence to reproduce issuance while protecting narrowly defined confidential information. A buyer should be able to answer four separate questions: did the physical activity occur; was the credited difference conservatively quantified; was the unit authorized and transferred without double counting; and is the buyer's proposed claim consistent with the unit's status and applicable rules?

Evidence[4][5][7]

Additionality is a structured counterfactual test

Additionality asks whether the credited mitigation would occur without the incentive or enabling conditions created by the mechanism. It is not proven by showing that an activity is desirable, expensive, or uncommon. Evidence may include regulatory surplus, investment analysis, barrier analysis, common-practice assessment, and evidence that credit revenue changes a real decision. Each test has failure modes: selective financial assumptions, rapidly changing technology costs, unenforced regulation, or a project definition chosen to make common practice appear rare. The UNFCCC Article 6.4 Supervisory Body now maintains a formal standard for demonstrating additionality in mechanism methodologies.

The baseline must also update when policy, technology, commodity prices, or practice changes. A static baseline can over-credit an activity that becomes normal during a long crediting period. Good design specifies renewal frequency, data hierarchy, conservative defaults, and conditions that suspend issuance. For sectoral programmes, baseline governance should be insulated from the project developer and accompanied by sensitivity analysis. The analytical output is not simply 'additional' or 'not additional'; it should show which assumptions dominate credited volume and how issuance changes under plausible alternatives.

Evidence[5][6]

Measurement needs uncertainty, leakage, and reversal accounting

Monitoring plans should begin with a mass-balance or causal diagram that identifies sources, sinks, boundaries, instruments, sampling frequency, missing-data rules, and quality controls. Meter accuracy alone is insufficient when model parameters, spatial sampling, activity data, or attribution dominate uncertainty. Report gross effect, leakage, uncertainty deduction, and net credited result separately. The UNFCCC's leakage standard matters because an activity can displace emissions beyond its project boundary. A local reduction is not an atmospheric benefit if production or land-use pressure simply moves elsewhere.

Removal credits add a time dimension. Stored carbon may be released through fire, harvest, equipment failure, insolvency, or policy change. Durability should be described as a monitored risk distribution over a defined period, not compressed into the word permanent. Buffer pools and replacement obligations can mutualize some risk, but they require correlated-risk analysis: a regional drought can affect many projects simultaneously. Engineered and biological storage have different failure modes and verification methods. They should not be compared only by a single price per tonne without disclosing duration, reversal liability, and monitoring obligations.

Evidence[5][6][7]

Verification is a risk-based evidence process

ISO 14064-3 specifies principles and requirements for validating and verifying greenhouse-gas statements and is programme-neutral. In practice, an assurance plan should identify intended users, materiality, level of assurance, competence requirements, conflicts, sampling logic, site-visit rationale, data-system controls, and the treatment of detected misstatements. Remote sensing can expand coverage, but it does not eliminate the need to validate calibration, detection limits, attribution, ground truth, and data availability. Independence is organizational as well as personal: fees, repeat engagements, and consultancy relationships should be visible.

Verification should test the claim that matters, not merely reconcile a spreadsheet to a methodology template. That requires tracing samples back to instruments and operational records, rerunning key calculations, challenging baseline choices, testing controls around manual adjustments, and examining adverse evidence. Findings should distinguish corrected errors, unresolved limitations, and scope exclusions. A clean opinion within a narrow scope must not be presented as confirmation of broad project quality. Where uncertainty or data gaps are material, conservative deductions or delayed issuance are more credible than false numerical precision.

Evidence[7][5][6]

Procure evidence, then decide what can be claimed

The global pool of unretired credits approached one billion tonnes in 2024, according to the World Bank, but aggregate supply reveals little about fitness for a particular use. Procurement should begin with a use case: compliance, contribution finance, internal carbon pricing, results-based payment, or a claim against residual emissions. Screen legal eligibility separately from environmental quality. Then score baseline risk, additionality evidence, quantification uncertainty, leakage, durability, safeguards, host authorization, verification quality, registry state, and claim compatibility. Record exclusion reasons so portfolio decisions remain auditable.

The most defensible public communication preserves these distinctions. Report gross emissions and reductions inside the organization's own inventory before describing external credits. Name programme, methodology, vintage, host, quantity, retirement identifier, and the exact role of the units. Avoid language implying that a credit erases a physical emission or delivers certainty beyond the evidence. Carbon markets can direct finance and support cooperation, but trust cannot be declared into existence. It is produced by conservative quantification, interoperable records, competent independent challenge, and claims whose scope matches what was actually verified.

Evidence[1][3][4][7]
Research boundary

Scope and limitations

Carbon-pricing counts, coverage, revenues, issuances, and unretired supply change annually and depend on the World Bank's instrument boundaries and conversion methods. Article 6.4 rules and infrastructure are still evolving. No single public metric establishes credit quality, and the framework here does not rate a named programme or project. Project-level conclusions require the applicable methodology, monitoring reports, verification statements, registry records, host authorization, and current claims guidance.

Evidence base

References

Source review: 20 August 2026. Quantitative values retain their original definitions, periods, and boundaries.

  1. 01
    State and Trends of Carbon Pricing 2026

    World Bank · 2026

    www.worldbank.org
  2. 02
    State and Trends of Carbon Pricing 2026: Full Report

    World Bank · 2026

    documents1.worldbank.org
  3. 03
    State and Trends of Carbon Pricing 2025

    World Bank · 2025

    openknowledge.worldbank.org
  4. 04
    Article 6.4 Mechanism Registry

    United Nations Framework Convention on Climate Change · 2026

    unfccc.int
  5. 05
    Standard: Demonstration of additionality in mechanism methodologies (v.02.0)

    UNFCCC Article 6.4 Supervisory Body · 2026

    unfccc.int
  6. 06
    Standard: Addressing leakage in mechanism methodologies (v.01.0)

    UNFCCC Article 6.4 Supervisory Body · 2025

    unfccc.int
  7. 07
    ISO 14064-3:2019: Verification and validation of greenhouse gas statements

    International Organization for Standardization · 2019

    www.iso.org