The Wealth Delta Tax: Phase One

Author

K. Ogata

Published

September 20, 2026

Keywords

Wealth Delta Tax, wealth taxation, political economy, tax politics, institutional design, political durability, democratic legitimacy, fiscal legitimacy, cooperative taxation, political coalition formation, policy stability, tax reform, implementation politics

Version: 1.01  |  Date: 20 Sep 2026  |  Word count: 5,880 (excl. front matter)

Author Disclosure

Portions of the drafting, editing, literature organisation, and structural review of this paper were assisted by publicly available large language models, including Anthropic’s Claude and OpenAI’s ChatGPT. These tools were used as aids to the author’s research and writing process; the substantive arguments, analysis, interpretations, and conclusions are the author’s own.

This work received no external funding, sponsorship, or other financial support. The author is solely responsible for the content of the paper and for any errors that remain.

Revision History

Revision Date Details
0.01 3 August 2026 First Draft
1.00 15 August 2026 Published to website
1.01 20 September 2026 Crosslinks added: §6.3 Agrawal cross-base externality footer note now cites (BEHAV.A §D) alongside (BEHAV §9.2); §7 conclusion extended with forward pointer to (FAL) as the falsification counterpart to the Phase One empirical cluster agenda

Abstract

The companion paper series has produced a complete mechanism design. This paper draws the line between what that design establishes and what only implementation can answer.

Seven empirical clusters are identified — coherent groups of questions that are unanswerable at the design stage because the data they require does not yet exist. For each cluster, the paper states the working assumption currently in force (the best available estimate in the absence of Phase One evidence) and a prospective evaluation design specifying what Phase One data would constitute a genuine test of that assumption. The clusters cover: behavioural response to the cooperative architecture, where the compliance decision at WDT concentration levels is made within professional intermediary networks rather than by individual taxpayers directly; avoidance under imperfect compliance, mapped against the nine-shape behavioural taxonomy in BEHAV; migration and the cross-base fiscal externality, where the Agrawal et al. (2025) finding of approximately six times the direct revenue loss in parallel taxes is the most consequential figure Phase One must test; valuation route and assessment window adoption, which is load-bearing for Governing Council calibration of the flexibility levy; administrative-layer intervention effects, where the taxpayer history record is the highest-priority sub-item; OBR independence under the sustained pressures of acting as WDT mandate-guardian; and the conditions under which the corporate instrument could eventually displace corporate income tax.

The paper does not make design decisions, resolve any open question, or supply the behavioural magnitudes the revenue model requires. Its evaluation designs are written to be useful independently of whether Phase One happens on any particular timeline: they specify precisely what would close each cluster. A Phase One that confirms working assumptions and one that requires the design to be revised are equally valid outcomes. What would not be valid is discovering after implementation that no measurement framework was in place to distinguish the two.

Glossary

Cross-base externality: The fiscal loss in parallel taxes (principally income tax and value added tax) that results from wealth-tax-driven emigration; identified in the Phase One context by Agrawal et al. (2025) at approximately six times the direct wealth-tax revenue loss.

Desk research: Analysis conducted without access to data generated by the operation of the mechanism itself; the limit of what the companion paper series can establish.

Empirical boundary: The point at which a design question cannot be answered by further analytical work and requires evidence from a live system.

Institutionally mediated compliance: The compliance posture adopted not by a taxpayer directly but by the professional network (advisers, family offices, trustees, corporate structures) through which ultra-high-net-worth individuals conduct their financial affairs.

Phase One: The first operating period of the WDT, characterised by a high exemption threshold, a small taxpayer population, and a primary objective of building valuation infrastructure and generating behavioural evidence before the system scales.

Phase One empirical question: A question that is unanswerable at the design stage because the data it requires does not yet exist; it describes how people will behave under a mechanism that has never operated.

Working assumption: A position adopted in the absence of Phase One evidence, on the basis of the best available adjacent evidence and first-principles reasoning, that Phase One is designed to test and, where necessary, correct.

1. Introduction

Analytical work runs out at a specific point: when the question requires data that does not yet exist because the mechanism has never operated.

The companion paper series has produced a complete design: a progressive, accrual-basis wealth tax with symmetric loss refunds, a pre-funded Sovereign Wealth Fund, and a governance architecture derived from the mechanism’s own transactions. What it has not produced, and cannot produce, is evidence of how the taxable population will respond once the mechanism operates. This paper draws that line.

Seven empirical clusters are identified: coherent groups of related questions unanswerable at the design stage. For each, the paper states what the design currently assumes in the absence of evidence and what a Phase One evaluation framework would need to measure to test it. The clusters cover behavioural response to the cooperative architecture, avoidance under imperfect compliance, migration and the cross-base fiscal externality, valuation route and assessment window adoption, administrative-layer intervention effects, OBR independence, and conditions for the corporate instrument displacing corporate income tax.

This is not a gap register. The consolidated open-questions register tracks every unresolved item with its assigned resolution path. This paper specifies why a particular class of question has the character it does, what the design assumes in lieu of evidence, and what Phase One data would constitute a genuine answer. The (PHASE1 §5) evaluation designs are specifications of what would close each cluster — useful independently of whether Phase One happens on any particular timeline.

2. The Boundary of Desk Research

The companion paper series has done what desk research can do. The behavioural robustness paper sets out a taxonomy of five friction types and nine behavioural shapes. The position closure paper specifies how each of the four closure event types settles. The rates paper models revenue against seventy-three historical starting years. The literature review characterises what adjacent empirical evidence exists and where it stops being informative. None of this resolves the empirical clusters identified here, because none of it constitutes evidence about how the WDT’s specific population will respond to the WDT’s specific design. The adjacent evidence supports the directional claims but does not supply the magnitudes, and for several clusters does not even supply confident direction.

The questions in the seven clusters are open because the data they require is a byproduct of the mechanism operating, and the mechanism has not operated — not because something has been left underspecified. The rates paper is explicit: its revenue figures are pre-behavioural starting points. The same qualification applies to the behavioural robustness paper’s account of the membrane interventions and the valuation paper’s account of route adoption. (PHASE1 §3) collects those working assumptions and states them precisely; (PHASE1 §5) specifies what Phase One data would constitute a test of each.

The seven clusters are Phase One empirical questions specifically: questions requiring implementation data that Phase One is positioned to answer. Items in the consolidated register requiring formal modelling extensions (the Domar-Musgrave extension, the welfare comparison, the macroeconomic modelling questions) or jurisdiction-specific legal analysis (exit and bankruptcy closure design, corporate levy interaction with CIT during transition) are out of scope here. They are noted in (PHASE1 §6) and their resolution paths are in the register.

3. The Working Assumptions Currently in Force

The design operates on a working assumption for each of the seven clusters. These are best available estimates in the absence of Phase One evidence, stated precisely enough that Phase One data can confirm, revise, or contradict them. The purpose is not to entrench them but to make the test explicit.

Cooperative architecture (#11, #17 residual). The reciprocal features of the design (symmetric loss refunds, governance participation, visible SWF entitlement) alter the compliance calculus of the taxable population in a favourable direction relative to a purely extractive system, and the effect is large enough to be detectable in Phase One administrative data. The basis is the cooperative compliance literature’s finding that procedural fairness improves compliance outcomes, combined with the first-principles argument that reciprocal treatment changes the rational calculus of the population best positioned to resist the system. The qualification — developed in (LR.A §3) and BEHAV — is that the compliance decision at WDT concentration levels is made within institutional networks rather than by individual taxpayers directly, so the cooperative architecture’s effect is adviser-mediated rather than direct. Whether that mediation preserves, amplifies, or attenuates the effect is unknown.

Avoidance under imperfect compliance (#13). Avoidance will concentrate in the shapes the design has identified as structurally available (principally asset restructuring, timing manipulation, and cross-border asset migration without personal exit), and shapes involving personal exit will be bounded by the structural closure mechanics in CLOSE. The nine-shape taxonomy in BEHAV is the analytical frame (BEHAV §4); Phase One is its empirical test.

Migration and cross-base externality (#15). WDT-driven emigration during Phase One will generate income tax and VAT losses materially larger than the direct wealth-tax revenue loss, consistent with the Agrawal et al. (2025) finding of approximately six times the direct loss in comparable jurisdictions. The design does not assume this externality away. It assumes the bridging facility, the re-entry rule, and membrane investment will reduce emigration rates relative to adversarial exit-tax regimes, and that Phase One’s high threshold and small population limit aggregate exposure. What the design does not assume is that the Agrawal et al. multiplier applies to the WDT population specifically; that is the Phase One measurement question.

Valuation route and assessment window adoption (#28). The distribution of taxpayer elections across the four routes and seven window lengths will be sufficiently dispersed that the Governing Council has meaningful calibration data within the first two assessment cycles. Route adoption is expected to track asset composition, and default-long window elections are expected to be common in the early years before the assessment window premium is calibrated. Both expectations are first-principles reasoning from the design’s own incentive structure; neither has empirical support.

Administrative intervention effects (BEHAV-1). Each of the six interventions specified in BEHAV produces a detectable positive effect on the relevant dimension of taxpayer engagement: annual entitlement statements increase refund-claim rates; provisional refund notification reduces disputed assessments; the classification register with safe harbour reduces compliance friction for borderline assets; the legible SWF annual number increases public salience; the default-long window reduces first-year administrative burden; the taxpayer history record produces the claimed effects on psychological engagement and procedural-fairness perception. The last is the highest-priority sub-item — the most novel intervention with the least adjacent precedent — and what remains after the governance paper settled the structural protection question is entirely empirical.

OBR independence (#22). The Office for Budget Responsibility meets the structural properties (GOV §6.3) specifies for a mandate-guardian institution (long-tenure appointment, transparent methodology, published forecasts, institutional separation from the executive), and those properties were adequately demonstrated during the 2022 mini-budget period and following the December 2025 chair resignation. This assumption is held with less confidence than the others. The OBR is an existing institution with a short and partially stressed track record; Phase One generates the data that makes a fair assessment possible.

Corporate instrument transition (#32). Phase One will produce attribution data (tranche-three share of corporate ownership over time, attribution test pass and fail rates by entity type, non-reconciliation rates by wealth band) sufficient to frame the long-run CIT displacement question as a policy decision. The design does not assume the answer; it assumes Phase One will make the question answerable.

4. The Empirical Clusters

4.1 Behavioural Response to Cooperative Architecture

Open question #11 asks whether the cooperative architecture (symmetric loss refunds, governance participation, visible SWF entitlement) shifts the distribution of behavioural shapes in a favourable direction relative to an otherwise identical extractive design. The adjacent evidence is supportive but indirect.

The cooperative compliance literature Gangl et al. (2015) supports the general claim that perceived fairness improves compliance outcomes, but models the compliance decision as an individual one. At the wealth levels where WDT revenue concentrates, the compliance decision is made within an institutional network of advisers, family offices, trustees, and corporate structures; the individual taxpayer’s attitudes are a secondary input into a process shaped primarily by professional norms and intermediary compliance culture. (LR.A §3) develops this characterisation, drawing on Klepper & Nagin (1989), Sakurai & Braithwaite (2003), and the OECD Cooperative Compliance programme. (LR.A §3.1) confirms it as a genuine literature gap: direct compliance evidence for cooperative design effects among professionally-mediated, ultra-high-net-worth taxpayers does not exist.

The measurement problem Phase One inherits is organisational, not individual. The question is whether the cooperative architecture alters how professional intermediaries advise their clients — whether the reciprocal framing shifts the risk calculus of the advisory networks through which compliance decisions are actually made. The #17 residual is precisely this: not the general cooperative compliance claim, which has directional support, but compliance evidence among this population under this design. The measurement design in (PHASE1 §5.1) specifies what that evidence would require.

4.2 Avoidance Under Imperfect Compliance

Open question #13 concerns the empirical extent and form of avoidance under imperfect compliance. The nine-shape taxonomy in BEHAV classifies the forms avoidance can take, but it cannot supply the distribution: how much of the taxable population will concentrate in which shapes, or how that distribution will shift across assessment cycles as the membrane matures.

The shapes generating the most measurement-relevant signal are Shapes 4, 5, and 6. Shape 4 (avoidance through restructuring) is the most directly observable, because restructuring that shifts assets across routes or reduces declared delta requires corporate and legal changes visible in administrative data. Shape 5 (deferral and timing manipulation) is the shape the valuation architecture is most directly designed to contain through the delta-based self-correction mechanism; its Phase One prevalence will be the primary test of whether that mechanism performs as designed. Shape 6 (cross-border asset migration without personal exit) sits at the intersection of this cluster and (PHASE1 §4.3); the cross-base externality Agrawal et al. (2025) identify may partly reflect asset migration rather than personal exit alone.

Shapes 7, 8, and 9 (personal exit, partial exit, and active resistance) are the most consequential at significant scale. The working assumption is that they will be bounded by the structural closure mechanics in CLOSE and membrane quality in BEHAV; whether that holds is tested jointly by this cluster and (PHASE1 §4.3). The measurement design in (PHASE1 §5.4) specifies what administrative data would allow the Governing Council to track the shape distribution across cycles.

4.3 Migration and the Cross-Base Externality

Three structural positions are settled and are not what Phase One is testing. The no-punitive-exit-taxation position (that departure is a legitimate termination event, not an avoidance act requiring a coercive fiscal response) is specified in CLOSE. The bridging facility, operating on the symmetric bond structure in (GOV.B §E.3), decouples physical departure from asset settlement. The re-entry rule closes the strategic cycling loophole and provides a structural incentive for long-term participants to return. None of this constitutes evidence that these mechanics work at the scale Phase One will test.

The measurement problem that remains is quantitative. Agrawal et al. (2025) find that wealth-tax-driven migration generates income tax and VAT losses approximately six times larger than the direct wealth-tax revenue loss in comparable jurisdictions. This cross-base multiplier is the central figure this cluster needs to test. The Agrawal et al. finding is the best available estimate, but it is drawn from subnational wealth taxes in Spain where the tax was not accrual-based, the symmetric refund mechanism was absent, and the no-punitive-exit-taxation position was not in force. Whether those differences compress or expand the multiplier for the WDT population is not determinable from the existing evidence.

During Phase One, income tax and VAT operate in parallel with the WDT, making the cross-base externality a live fiscal exposure rather than an abstract long-run concern. In Phase Two, where WDT revenue displaces those taxes, the externality disappears structurally. CLOSE characterises this honestly; RATES excludes behavioural migration responses from its central revenue claim for the same reason. The (PHASE1 §5.3) measurement design accordingly carries the largest fiscal implications of any in the seven clusters. It also addresses the Shape 6 ambiguity: Phase One administrative data will not cleanly separate asset-migration effects from personal-exit effects in the aggregate externality figure, and (PHASE1 §5.3) specifies how to decompose them.

4.4 Valuation Route and Assessment Window Adoption

Open question #28 (behavioural adoption rates across the four valuation routes and seven window lengths) is the most directly administrative of the clusters, but the data it generates is load-bearing for Governing Council calibration in ways that make early measurement essential.

The assessment window premium illustrates why. RATES excludes the premium from its reference revenue model because calibrating it requires Phase One election data that does not yet exist. The premium has two components: a deferral charge anchored to sovereign borrowing cost, which can be set from first principles, and a flexibility levy calibrated against observed election patterns, which cannot. Phase One election data converts that arbitrary parameter into a data-driven one.

Route adoption matters for a related but distinct reason. The four routes differ substantially in administrative burden, deterrence properties, and revenue timing. Routes A and B generate revenue on the standard annual cycle; Route C generates revenue through the must-transfer mechanism’s continuous self-correction; Route D defers all taxation to realisation, meaning Phase One revenue from Route D taxpayers will be lower than the cohort model suggests until the first realisation events occur. The distribution across routes therefore affects both the aggregate revenue figure and its timing.

The default-long window election creates an interaction with (PHASE1 §4.5): if it produces widespread uptake of longer windows, the Governing Council’s calibration problem for the flexibility levy changes in character. The (PHASE1 §5.4) measurement design tracks not only raw election distributions but whether the default is driving them, because the policy implications of default-induced inertia and revealed preference differ.

4.5 Administrative Intervention Effects

BEHAV proposes six administrative-layer interventions: an annual entitlement statement; provisional refund notification in loss years; an asset classification register with safe harbour; a legible SWF annual number; a default-long assessment window election; and a taxpayer history record. None changes the mechanism’s economics. They change how the mechanism is experienced, and BEHAV treats that as a substantive institutional claim. Whether the claim holds is what BEHAV-1 asks.

The evaluation logic differs across the six by friction type. The annual entitlement statement and provisional refund notification address visibility friction: the testable prediction is that making refund entitlement visible in real time increases claim rates and reduces unclaimed refunds, trackable directly from administrative data. The classification register addresses compliance friction: safe harbour uptake should correlate with reduced dispute rates for borderline asset classes, trackable through tribunal and audit outcome data. The legible SWF annual number addresses legitimacy friction (the gap between abstract entitlement and felt stake) and is the hardest to measure from administrative data alone; it is primarily a survey question, addressed in (PHASE1 §5.5).

The default-long window election’s evaluation logic within this cluster is distinct from (PHASE1 §4.4): the question is not what election distribution it produces but whether it reduces first-year administrative burden and error rates, which the assessment record can track separately from the distribution itself.

The taxpayer history record is the highest-priority sub-item. Nothing structurally similar exists in the professional insurance or cooperative compliance literatures, so adjacent evidence is least informative here. The structural protection question is settled: GOV specifies the record as a mandatory Administrator output with distributed archiving at designated institutions. What remains is entirely empirical: whether making the record visible to taxpayers in real time produces the claimed effects on psychological engagement and procedural-fairness perception. The record’s value accumulates over cycles rather than being visible from the first year, so the (PHASE1 §5.5) measurement design specifies a minimum observation window and the markers that would distinguish a functioning intervention from an administrative output taxpayers do not consult.

4.6 OBR Independence — Empirical Assessment

(GOV §6.3) specifies the structural properties a mandate-guardian institution needs: long-tenure appointment, transparent methodology, published forecasts, and institutional separation from the executive. Those properties define the benchmark for #22. Whether the Office for Budget Responsibility meets them is a question about an existing institution with a short and partially stressed track record; no fair assessment is possible without implementation data.

The two stress events the UK jurisdiction has produced are informative but not conclusive. The 2022 mini-budget period tested OBR independence in one mode: the then-Chancellor presenting a fiscal event without an OBR forecast and the institutional consequences that followed. The December 2025 chair resignation tested it in another: whether the conditions of leadership transition are consistent with the separation-from-executive criterion. Each event illuminates one dimension of the structural properties. Neither speaks to how the institution would behave as WDT mandate-guardian specifically, under the sustained pressures that role creates: published forecasts of a novel tax base, actuarial accountability for the SWF refund reserve, and the political visibility of certifying Phase One milestone achievement.

This cluster has a feature distinguishing it from the others. The other six generate data that flows to the Governing Council as calibration inputs. This one generates data the Governing Council uses to evaluate an institution sitting alongside it rather than beneath it. The body best placed to evaluate OBR independence is not the Governing Council itself, which has a direct interest in the OBR’s outputs. The (PHASE1 §5.6) measurement design therefore relies on the published record (forecast history, milestone certification audit trail, and the visible pattern of Governing Council responses) rather than internal assessment by an interested party.

4.7 Corporate Instrument Transition

The corporate instrument cluster contains one item: #32, whether the corporate instrument should eventually displace corporate income tax outright. The attribution test question (previously BEHAV-3) is settled: the binary test is the steady-state design, and any Phase One on-ramp for intermediaries demonstrating improving attribution capability is a Governing Council calibration parameter, not a design question.

Whether the corporate instrument could eventually render CIT redundant depends on whether it achieves sufficient coverage of the economic activity CIT currently reaches, which in turn depends on Phase One attribution rates across the three ownership tranches and reconciliation failure patterns across wealth bands. The Phase One data bearing on #32 is the attribution data the instrument generates in normal operation: tranche-three share of corporate ownership over time (indicating whether permanently unattributed shareholdings are declining), attribution test pass and fail rates by entity type (indicating which intermediary categories are structurally resistant), and non-reconciliation rates by wealth band.

None of this will answer #32 directly. It will indicate whether the conditions under which the displacement question becomes answerable are being met. The (PHASE1 §5.7) measurement design specifies the observation schedule and threshold conditions that would bring #32 from a speculative long-run question into the range of a policy decision the Governing Council could reasonably begin to frame.

5. What Phase One Would Need to Measure

The evaluation designs below are prospective specifications: what measurement would close each cluster, written independently of any assumed implementation timeline. Where (BEHAV §10) specifies which interventions are appropriate in which phase, this section specifies what monitoring each intervention requires.

5.1 Measuring Behavioural Response to the Cooperative Architecture

The measurement problem for #11 and #17 is organisational, not individual. Standard compliance surveys are inadequate because the compliance decision at WDT concentration levels is made within professional networks. The evaluation design has two components.

The administrative data component requires three time series tracked across every assessment cycle: the distribution of declared behavioural shapes (inferred from route elections, window elections, restructuring events, and departure notifications); the refund claim rate (proportion of loss-year taxpayers filing complete refund claims, as a share of those entitled); and the dispute rate (proportion of assessments resulting in tribunal referral, tracked by asset class and route). A shift toward Shapes 1 and 2 over the first three to five cycles, a declining unclaim rate as the membrane matures, and a declining dispute rate would each constitute directional confirmation of the relevant working assumption.

The professional network component requires a practitioner survey and qualitative interview programme with a representative sample of tax advisory firms active in the WDT population, conducted at minimum at the end of cycles one, three, and five. It should track how advisers characterise the WDT to clients in initial briefings, whether that characterisation shifts over cycles, whether advisers distinguish WDT compliance posture from their general HMRC approach, and whether they report changing client receptiveness to cooperative positions. Variation by firm type and client concentration is the primary indicator of whether adviser norm transmission is occurring.

The minimum observation window is five assessment cycles. Early cycle data will understate the steady-state effect; a meaningful evaluation requires enough cycles to observe whether the trend is present.

5.2 Measuring Avoidance Under Imperfect Compliance

The measurement design for #13 uses administrative data sources already generated by the mechanism’s normal operation. No additional collection infrastructure is required.

The primary data sources are: annual return data (asset declarations, route elections, window elections, declared deltas); restructuring event notifications; departure notifications under the position closure framework; and tribunal and audit outcome data. Together these are sufficient to infer the shape distribution at the population level.

Shape 4 (avoidance through restructuring) is the most directly observable. Restructuring that shifts assets from Route A or B to Route D, or that interposes new legal entities affecting route classification, is reportable and auditable. A sustained increase in Route D elections across cycles without a corresponding change in underlying asset composition is the primary signal of Shape 4 concentration.

Shape 5 (deferral and timing manipulation) is observable through window election data and the delta continuity record. Whether taxpayers elect longer windows only when their asset growth profile makes deferral advantageous, rather than as a blanket avoidance strategy, is testable from election data combined with declared growth trajectories.

Shape 6 (cross-border asset migration without personal exit) is the hardest to measure from domestic data, because migrated assets may no longer generate domestic reporting obligations. It requires third-party reporting from financial institutions and automatic exchange of information under existing international frameworks. The Shape 6 component of the cross-base externality can be estimated as the residual between the total externality measured in (PHASE1 §5.3) and the personal-exit component attributable to Shape 7.

5.3 Measuring the Cross-Base Externality

The measurement design for #15 requires three data streams operating in parallel.

The WDT administrator supplies the departure notification record: every position closure via jurisdictional exit, with declared wealth level, asset composition, route elections, and cumulative lifetime contributions at departure. This establishes the size and composition of the emigrating population, but not what happened to that population’s contributions to parallel tax systems after departure.

HMRC supplies the income tax and VAT revenue record, disaggregated by taxpayer identifier to the extent legally permissible. The objective is to estimate the income tax and VAT revenue that would have been collected from the departing population in the absence of departure, and compare it to what was actually collected in the years following each departure event. The Agrawal et al. (2025) multiplier of approximately six times the direct wealth-tax revenue loss provides the prior expectation; Phase One data replaces it with a jurisdiction-specific estimate.

The Shape 6 component requires a third stream: automatic exchange of information reports covering financial assets held by UK-resident taxpayers in foreign jurisdictions, matched against WDT annual return data to identify systematic discrepancies. This component is distinct from Shape 7 in that the taxpayer remains resident and continues to generate domestic income tax and VAT, but the WDT base and associated investment income may have migrated.

The Governing Council should commission an independent externality assessment at the end of cycles two, four, and six, producing a point estimate of the cross-base multiplier, a confidence interval, and a Shape 6 / Shape 7 decomposition. The trend across assessments — whether the multiplier is declining as the bridging facility and re-entry rule become established — is as informative as the level at any single point.

5.4 Measuring Valuation Route and Assessment Window Adoption

The data for #28 is generated directly by the mechanism’s reporting requirements. No additional infrastructure is needed.

The Administrator should publish three monitoring outputs at the end of each assessment cycle. First, the route election distribution: for each of the four routes, the number of electing taxpayers, total declared wealth, and year-on-year change, disaggregated by asset class to test whether adoption tracks asset composition. Second, the window election distribution: for each of the seven window lengths, the number of elections, associated wealth, and the split between first-year elections and renewals — the separation matters because first-year elections under a default-long structure may reflect administrative inertia rather than preference, while renewals are stronger evidence of revealed preference. Third, at the end of cycles two and four the Governing Council should have sufficient data to calculate the flexibility levy component, compare it to initial calibration assumptions, and publish a revised parameter.

One additional disaggregation is needed for the (PHASE1 §5.5) interaction: window elections should be tracked separately for taxpayers who received the default-long notification and those who actively elected without it. This separates the default effect from the underlying preference distribution — necessary to calibrate the flexibility levy against revealed preference rather than administratively induced behaviour.

5.5 Measuring Administrative Intervention Effects

The six interventions require different measurement approaches by friction type.

The annual entitlement statement and provisional refund notification are both measurable from the assessment record. Relevant metrics: refund claim rate per cycle; average time between provisional notification and claim filing; proportion of refund claims filed without error or amendment. A rising claim rate and shortening filing time over the first three cycles supports the visibility-friction account; a stable or declining rate indicates the notification design needs revision.

The asset classification register and safe harbour are measurable from the dispute record. Relevant metrics: proportion of tribunal referrals involving safe-harbour assets; proportion involving assets on the register without a safe harbour; proportion involving assets outside the register entirely. A declining referral rate for safe-harbour assets relative to unclassified assets supports the compliance-friction account; which asset classes generate the most unresolved disputes indicates where the register requires expansion.

The legible SWF annual number cannot be evaluated from administrative data. The evaluation requires an annual public attitudes survey covering a representative sample of the general population, not only WDT taxpayers, tracking SWF awareness, self-reported sense of personal stake, and attitudes toward WDT legitimacy. The survey should be commissioned independently of the Administrator and published alongside the cycle-end data release.

The default-long window election is evaluated via the (PHASE1 §5.4) disaggregation. The additional metric within this cluster is first-year assessment error rates: whether taxpayers who received the default-long notification have lower error rates than those who actively elected, controlling for wealth level and asset composition.

The taxpayer history record requires a minimum observation window of four assessment cycles, at which point taxpayers have a record long enough to show meaningful variation. Two outcomes should be tracked: consultation rate (whether taxpayers access their record before assessment, measurable through Administrator access logs); and the correspondence between consultation and assessment accuracy (whether those who consult submit assessments with lower amendment rates). The second metric requires a data linkage arrangement between access logs and assessment outcomes that should be established as part of Phase One infrastructure design, not retrospectively.

5.6 Measuring OBR Independence

The measurement design for #22 tracks each of (GOV §6.3)’s four structural properties through the OBR’s own published record, without requiring internal Governing Council assessment.

Long-tenure appointment: the relevant metric is not whether chairs serve full terms, but whether departures cluster around political events (changes of government, Budget cycles, fiscal stress periods) in a pattern inconsistent with term-limit expiry.

Transparent methodology: the proportion of WDT-relevant forecast revisions accompanied by a published methodology note; the lag between revision and publication; and the proportion of forecast inputs derived from publicly reproducible data sources. These are trackable by any independent research body using public outputs alone.

Published forecasts: forecast accuracy for WDT revenue, SRR balance trajectory, and LRR fill timeline, measured against outturns on a rolling window. The Governing Council should commission an independent forecast accuracy assessment at cycle four, covering cycles one through three.

Institutional separation from the executive: the frequency and content of OBR publications contradicting official government positions on WDT revenue adequacy or milestone achievement, and the pattern of Governing Council responses. An OBR that consistently validates government positions without published analytical challenge signals incomplete separation regardless of the formal independence structure.

The Governing Council should establish an independent advisory panel, drawn from academic institutions, international fiscal bodies, and former central bank officials, to produce a published OBR independence assessment at the end of Phase One. The panel’s terms of reference should be fixed in advance, methodology published before the assessment begins, and conclusions transmitted to the Governing Council without prior review by the executive.

5.7 Measuring the Corporate Instrument Transition

The Governing Council should track three data series from the corporate instrument’s normal reporting cycle and specify in advance the threshold conditions that would bring #32 into the range of a policy decision.

The three series: tranche-three share of total corporate ownership by assessed market capitalisation, published per cycle; attribution test pass and fail rates by entity type (DC pension schemes, foreign institutional holders, intermediary chains of varying depth), published annually; and non-reconciliation rates by wealth band, measuring the proportion of corporate levy provisional amounts not claimed against individual assessments within the reconciliation window. Each series should be published in the standard cycle-end data release with a rolling four-cycle average.

The threshold conditions should be specified before the end of cycle two, not retrospectively. The Governing Council should resolve that the CIT displacement question will be placed on the policy agenda only if three conditions are jointly met for the stated number of consecutive cycles: tranche-three share below a specified percentage of total assessed corporate wealth; attribution test pass rates for the two most common intermediary categories above a specified threshold; and non-reconciliation rates stabilised below a specified level across all wealth bands. The specific figures are Governing Council parameters set against cycle-one baseline data.

If #32 reaches the policy agenda, it will require additional analysis Phase One cannot supply: the legal interaction between the corporate instrument and CIT during transition, treatment of CIT loss carryforwards, and revenue implications of full CIT displacement at Phase Two scale. What Phase One supplies is the evidence base that makes that analysis worth doing.

6. Limitations and Further Work

6.1 Formal modelling gaps

No items in this paper. PHASE1 specifies evaluation designs; the formal modelling gaps are identified in LR.A.

6.2 Phase One empirical unknowns

This paper is the structured account of what Phase One is designed to answer. The seven empirical clusters — behavioural response to the cooperative architecture, avoidance under imperfect compliance, migration and cross-base externality, valuation route and window adoption, administrative-layer intervention effects, OBR independence assessment, and corporate instrument transition — each carry a working assumption currently in force and a prospective evaluation design specifying what data is required, what observation window is necessary, and what patterns would constitute confirmation, revision, or contradiction.

6.4 Structural and irreducible limits of the design

The evaluation designs are useful independently of whether Phase One happens on any particular timeline. They specify what would close each cluster; if the mechanism is never implemented, the paper stands as the honest account of what the series has and has not established.

6.5 Governing Council calibration parameters

No items specific to this paper beyond those carried from the companion papers each evaluation design references.

7. Conclusion

The series has established a complete design. The open questions this paper addresses are questions the design cannot answer from the desk, because the answers describe how people will behave under a mechanism that has not operated. A design gap can in principle be closed by further analytical work; these cannot.

The working assumptions in (PHASE1 §3) are provisional by design. An assumption that survives Phase One intact and one that requires revision are equally valid outcomes. What would not be valid is discovering after implementation that no measurement design was in place to distinguish the two; that is what (PHASE1 §3) exists to prevent. The evaluation designs are also useful independently of whether Phase One happens: they specify precisely what would need to be answered differently for a reader who finds the design wanting to change that judgment.

The revenue figures will be the first thing Phase One corrects. They are pre-behavioural, stated as such in (RATES), and the (PHASE1 §5.3) measurement design carries the largest fiscal implications of any in the seven clusters. If the Agrawal et al. (2025) multiplier applies to the WDT population at anything close to its estimated magnitude, the fiscal case for Phase One rests on the membrane investment and structural closure mechanics outperforming a comparable adversarial regime — a claim Phase One is positioned to test.

Phase One is valuable regardless of whether the design is right. A Phase One that finds the cooperative architecture’s effects smaller than assumed, the cross-base externality larger than estimated, or certain membrane interventions ineffective still produces the evidence base from which a revised design could be built.

The empirical clusters in (PHASE1 §4) address questions about how the mechanism will be calibrated and administered. A complementary set of questions — whether the mechanism works at all, stated in their strongest falsifiable form — is addressed in (FAL). That paper identifies the five propositions on which the WDT most depends, specifies what evidence or modelling would count against each, and sets the minimum requirements a hostile model must meet to constitute a serious falsification attempt. The relationship between PHASE1 and FAL is sequencing: FAL identifies what would constitute a decisive result; Phase One produces the first data against which those results can begin to be tested.

References

Agrawal, D. R., Foremny, D., & Martínez-Toledano, C. (2025). Wealth tax mobility and tax coordination. American Economic Journal: Applied Economics, 17(1), 402–430. https://doi.org/10.1257/app.20220615
Gangl, K., Hofmann, E., & Kirchler, E. (2015). Tax authorities’ interaction with taxpayers: A conception of compliance in social dilemmas by power and trust. New Ideas in Psychology, 37, 13–23. https://doi.org/10.1016/j.newideapsych.2014.12.001
Klepper, S., & Nagin, D. (1989). The anatomy of tax evasion. Journal of Law, Economics, and Organization, 5(1), 1–24. https://doi.org/10.1093/oxfordjournals.jleo.a036959
Sakurai, Y., & Braithwaite, V. (2003). Taxpayers’ perceptions of practitioners: Finding one who is effective and does the right thing? Journal of Business Ethics, 46(4), 375–387. https://doi.org/10.1023/A:1025641518700