Pre-analysis planEmpirical asset pricing

The Price of
Divergence

Quantifying alpha decay and tracking error caused by ESG rating provider algorithmic bias

Pre-registered protocol. Methodology only: this study declares its hypotheses, data architecture, tests, and falsification conditions in advance. It contains no results.

A large literature reports positive risk-adjusted returns to ESG-screened portfolios — almost always conditioned on a single rating vendor. This protocol asks whether that alpha survives a change of vendor when the mandate, universe, optimiser, calendar and constraint level are all held fixed. The identifying variation is the choice of rating provider and nothing else.

56 / 38 / 6
Measurement / scope / weights
share of rating divergence
0.38–0.71
Pairwise correlation
across six major providers
39–58%
Implied active share between
two identical ESG mandates
60%
Annual turnover floor at
95% monthly rating retention
The threshold operator, made operable
§1.4 · bivariate normal, q = 0.25
MSCI SUSTAINALYTICS 50% shared holdings
0.54
0.30within observed range0.95
Overlap50.3%
Active share49.7%
Jaccard0.336
Disjoint names124 / 250
Idio. TE floor1.89%
Two equal-weighted top-quartile portfolios drawn from a 1,000-name universe. The tracking-error floor assumes 30% idiosyncratic volatility and factor exposures that net out — the systematic sector component of §3.8 sits on top of it.

0 · Front Matter

0.1 Abstract of the proposed study

A large empirical literature reports a positive risk-adjusted return to portfolios formed on environmental, social and governance (ESG) scores. That literature is almost uniformly conditioned on a single rating vendor. Because vendor scores for the same issuer correlate at levels far below the 0.99 achieved by credit rating agencies: Berg, Kölbel and Rigobon (2022) report pairwise correlations of 0.38 to 0.71 across six major providers , the "ESG portfolio" is not a well-defined object. It is a vendor-specific object.

This study asks whether the reported ESG alpha survives a change of vendor, holding the mandate, the universe, the optimiser, the rebalancing calendar and the constraint level fixed. The identifying variation is the choice of rating provider and nothing else. If a mandate written as "top-quartile ESG" produces materially different portfolios, materially different factor loadings, and a non-zero tracking error but a statistically zero return differential across two reputable providers, then the ESG alpha documented in the literature is not a property of corporate sustainability. It is a property of a proprietary algorithm, and the resulting sector and style tilts are an uncompensated active risk borne by the beneficiary.

0.2 Primary research questions

# Question Primary test object
RQ1 How much of the cross-sectional dispersion between two providers' scores is systematic (a stable function of observable firm characteristics, size, disclosure intensity, sector) versus idiosyncratic (rater noise)?
RQ2 Does the systematic component translate into persistent, non-diversifiable factor and sector exposure at the portfolio level? Factor-block tracking-error decomposition
RQ3 Is the return differential between two identically mandated, differently rated portfolios statistically distinguishable from zero once the Fama–French five-factor model is imposed?
RQ4 What is the annualised basis-point friction, turnover drag plus market impact, attributable specifically to rating migration, as distinct from price drift and index reconstitution? Counterfactual "frozen-ratings" turnover
RQ5 Conditional on RQ3 and RQ4, what is the implied cost, in Sharpe-ratio and information-ratio units, of single-vendor dependence for a fiduciary? Shadow price of the constraint

0.3 Hypotheses

The study is designed to be capable of disproving its own motivating conjecture. Each hypothesis is stated with the null that would falsify the divergence thesis.

  • H1 (Systematic bias). . Rejection establishes that the two vendors load differently on observable characteristics, i.e. that divergence is not white noise.
  • H2 (Portfolio friction). active share between Portfolio A and Portfolio B . The theoretical prior under bivariate normality and the observed correlation range is an active share of 39%–58% (§A.1), so H2 is close to a manipulation check.
  • H3 (No compensation). in the FF5 model. Failure to reject H3 while H2 is rejected is the paper's central result: divergence generates risk without return.
  • H4 (Sector-bet dominance). the systematic (factor + sector) block explains of . Rejection establishes that the divergence is a bet, not diversifiable churn.
  • H5 (Friction dominance). rating-migration-induced turnover drag bp p.a. Rejection quantifies the price of the divergence in the currency fiduciaries actually use.
  • H6 (Alpha attenuation). The single-vendor ESG–return coefficient is attenuated by , where is the inter-rater reliability ratio; under the observed correlation range this implies 29%–62% attenuation (§A.2). Instrumenting one vendor's score with another's isolates the common component . If is materially smaller in magnitude than the rescaling would predict, the single-vendor coefficient was being carried by the vendor-specific bias term rather than by common ESG content, which is the divergence thesis restated at the firm level.

0.4 Contribution relative to the existing literature

  1. The divergence literature (Phase 1) documents disagreement at the score level. This protocol carries it through a constrained optimiser to the portfolio level, which is where fiduciary consequences are realised. The mapping from score dispersion to portfolio dispersion is nonlinear (a threshold operator) and is derived in closed form in §1.4.
  2. It separates divergence into a priced-risk channel (systematic tilt) and a transaction-cost channel (idiosyncratic churn), which the existing literature conflates.
  3. It supplies a falsification test, synthetic rater pairs with matched cross-sectional correlation, that distinguishes a genuine divergence effect from the mechanical consequence of any two imperfectly correlated screens.
  4. It is designed against a point-in-time rating archive, addressing the retroactive-restatement problem documented by Berg, Fabisik and Sautner (ECGI WP; SSRN 3722087), which invalidates most published ESG backtests built on current-vintage vendor files.

0.5 Workflow decomposition

The four phases are specified as separable work packages with explicit hand-off artefacts. Where a multi-agent implementation is used, the boundaries below are the natural assignment; notation is fixed centrally in §A.9 so that outputs remain composable.

Phase Work package Hand-off artefact to the next phase
1 Theory and literature synthesis Structural measurement model ; testable restrictions
2 Data architecture and normalisation Point-in-time panel PANEL_PIT: (PERMNO × month) with , , GICS, returns, characteristics, coverage flags
3 Optimisation and econometrics Portfolio return series , weight matrices , regression tables
4 Friction and institutional impact Turnover decomposition, cost-adjusted alphas, capacity curves, fiduciary framing

Phase 1: Theoretical Framework and Literature Synthesis

1.1 The anchor: Berg, Kölbel and Rigobon (2022)

Citation. Berg, F., Kölbel, J. F., and Rigobon, R. (2022). "Aggregate Confusion: The Divergence of ESG Ratings." Review of Finance, 26(6), 1315–1344. DOI: 10.1093/rof/rfac033.

What the paper establishes. Working with six providers (KLD (discontinued 2017), Sustainalytics, Vigeo Eiris (now Moody's ESG), RobecoSAM (now S&P Global), Asset4 (now Refinitiv/LSEG) and MSCI) the authors map every provider indicator into a common taxonomy of 64 categories and decompose observed rating disagreement into three orthogonalised channels:

Channel Definition Share of divergence
Measurement Providers assign different values to the same firm on the same category 56%
Scope Providers include different sets of categories 38%
Weights Providers apply different weights to a common set of categories 6%

Pairwise correlations across providers range from 0.38 to 0.71. The paper further documents a rater effect: a provider's overall assessment of a firm propagates into its category-level assessments, so measurement error is structured within rater rather than independent across categories.

Why the 56/38/6 split is the load-bearing fact for this study. The three channels have entirely different portfolio consequences, and the existing literature does not distinguish them:

  • Weights divergence (6%) is economically benign. It is a monotone re-mixing of a common information set. Rank orderings are largely preserved and it produces very little set-membership churn.
  • Scope divergence (38%) is a systematic distortion. A provider that scores a category the other omits has, in effect, added a characteristic to the objective function. Because ESG categories map non-uniformly onto sectors (e.g. carbon intensity onto Energy and Utilities; data privacy onto Information Technology; supply-chain labour onto Consumer Discretionary), scope divergence necessarily loads onto GICS sectors. Scope divergence is the sector-bet channel.
  • Measurement divergence (56%) is the dominant channel and is mixed: the rater effect implies a firm-level systematic component (a provider's house view, correlated with disclosure volume and firm size), while the residual is classical noise. Measurement divergence is therefore the source of both the size/quality tilt and the turnover churn, and separating its two parts is the econometric heart of Phase 2.

Methodological extension proposed here. Berg et al. obtain their shares from a sequential substitution of rater-specific elements with common elements. Sequential decompositions are order-dependent: the share attributed to scope depends on whether measurement is harmonised first. This study re-estimates the decomposition using a Shapley value over the three channels, averaging the marginal contribution of each channel across all substitution orders:

(1)

where is the reduction in obtained by harmonising the channels in . The Shapley shares satisfy efficiency, , and are order-invariant. Reporting alongside the published 56/38/6 is a direct, falsifiable replication contribution.

1.2 A structural measurement model of the rating

Let denote the latent, mandate-relevant ESG quality of firm at time , the object every provider claims to measure. Let be a vector of observable firm characteristics (log market capitalisation, disclosure intensity, sector indicators, R&D intensity, leverage, geographic revenue mix, analyst coverage). The rating produced by provider is modelled as:

(2)

Two elements of this specification carry the analysis.

The bias vector . This is the algorithmic fingerprint of the provider, the projection of its scope choices, weighting scheme and rater effect onto observable characteristics. The industry-standard characterisations are directly testable restrictions on :

  • MSCI rewards disclosure capacity and compliance infrastructure. Restriction: and .
  • Sustainalytics gap-fills with sub-industry averages where firm-specific data are absent. Restriction: conditional on sub-industry, the within-industry variance of is compressed relative to , i.e. , with the compression increasing in the firm's missing-data density.

These are hypotheses about vendor behaviour, not assumptions. §2.5 specifies their estimation and §2.6 their falsification.

The divergence decomposition. Define observed divergence in normalised units:

(3)

The two components have qualitatively different portfolio consequences, and the entire economic argument of the paper rests on this separation:

(4)
(5)

Because is a linear function of persistent characteristics, it does not mean-revert and cannot be diversified away by holding more names; it is removed only by neutralising the portfolio to . Because is transient, it is diversified in the cross-section but is not free, it is paid for at every rebalance, in the currency of Phase 4.

1.3 The errors-in-variables problem and its consequence for reported ESG alpha

Suppose the true cross-sectional relation between latent ESG quality and expected excess return is . A researcher who runs the regression on a single vendor's score recovers not but an attenuated coefficient:

(6)

This identity is the reason inter-rater correlation is not merely a curiosity. Under the symmetric-noise benchmark, the observed 0.38–0.71 correlation range is the reliability ratio, implying that published single-vendor ESG coefficients are attenuated toward zero by 29% to 62% (§A.2). Two implications follow, and they cut in opposite directions, which is precisely why the study must test both:

  1. A genuine ESG effect is understated. Correcting for measurement error should magnify a real . This is the logic of the instrumental-variables correction pursued by Berg, Kölbel, Pavlova and Rigobon in their noise-correction work (NBER WP 30562; SSRN 3941514), who instrument one provider's rating with another's.
  2. A spurious ESG effect is manufactured. If what a screen actually selects is rather than , then the "ESG alpha" is the return to the characteristics in (size, profitability, low leverage, high disclosure) repackaged. This is testable: the FF5 loadings of the screened portfolio will absorb it, and will vanish.

The instrument and its exclusion restriction. Using as an instrument for requires . This is not innocuous: both providers ingest the same corporate sustainability reports, CDP submissions and regulatory filings, so a common input error violates exclusion. The protocol therefore requires:

  • Over-identification with a third provider ( from LSEG/Refinitiv, ISS ESG or S&P Global CSA), enabling a Hansen test of the over-identifying restrictions;
  • A Hausman test of OLS against IV;
  • Reporting of the first-stage -statistic against the Stock–Yogo weak-instrument criterion. Note that the first-stage equals , so even the weakest observed reliability implies a strong first stage in a firm-level panel; weak identification binds only in short, small cross-sections such as a single month's regression;
  • An explicit statement, in the limitations section, that a rejected test is evidence of correlated vendor error, itself a finding of interest, since correlated error means the market cannot diversify vendor choice.

1.4 From score dispersion to portfolio dispersion: the threshold operator

Portfolio mandates do not apply the score linearly. They apply a threshold: "hold only issuers in the top quartile." This is the single most consequential and least analysed step in the chain, because a threshold operator amplifies dispersion in the region where it bites.

Let be jointly standard normal with correlation after normalisation (§2.4), and let the eligibility rule be with . The probability that a firm survives both screens is:

(7)

where is the bivariate standard normal CDF. The conditional retention rate, the fraction of Portfolio A's names that Portfolio B also holds, and the Jaccard similarity of the two eligible sets are:

(8)

For two equal-weighted portfolios of equal cardinality, the active share of Cremers and Petajisto (2009) reduces exactly to .

10%20%30%40%50%60%70%0.30.40.50.60.70.80.958.1%39.1% observed range, Berg et al. (2022) provider score correlation ρ
Figure 1The threshold operator amplifies disagreement. Two portfolios written to the identical top-quartile mandate differ in 39%–58% of their holdings across the correlation range that Berg, Kölbel and Rigobon report — an active share associated with genuinely active management, obtained with no active view.

Evaluating this at the correlation range reported by Berg et al. yields the study's ex ante motivation, computed numerically in §A.1:

Overlap Jaccard Implied active share A vs B
0.38 41.9% 0.265 58.1%
0.54 50.3% 0.336 49.7%
0.71 60.9% 0.437 39.1%

Two portfolios written to the identical mandate, over the identical universe, at the identical constraint level, differ in roughly two-fifths to three-fifths of their holdings purely as a function of which vendor's file was loaded. For calibration, an active share of 39%–58% is in the range that Cremers and Petajisto associate with genuinely active management, here obtained with no active view whatsoever. This is the quantity H2 formalises, and it is the reason the study is worth running.

The amplification result. Note that across the relevant range: at the scores agree strongly, yet 39% of holdings differ. The threshold operator converts moderate score disagreement into severe portfolio disagreement, because disagreement is concentrated exactly at the cut-point where the mass of the distribution is greatest. This nonlinearity is the mechanism by which "modest" rating divergence becomes a first-order portfolio problem, and it is the primary theoretical contribution of Phase 1.

1.5 Foundational literature: verified sources

The following are verified against publisher records. Each entry states the specific role it plays in this protocol.

Divergence and its measurement

  1. Berg, Kölbel and Rigobon (2022), Review of Finance 26(6), 1315–1344. DOI: 10.1093/rof/rfac033.: Anchor. Supplies the scope/measurement/weights decomposition and the correlation range that calibrates §1.4.

  2. Christensen, D. M., Serafeim, G., and Sikochi, A. (2022). "Why Is Corporate Virtue in the Eye of the Beholder? The Case of ESG Ratings." The Accounting Review, 97(1), 147–175. DOI: 10.2308/TAR-2019-0506.: Establishes that greater ESG disclosure is associated with greater rating disagreement, not less. This is the direct antecedent for the disclosure term in and a warning against the intuition that mandatory disclosure regimes will resolve divergence mechanically.

  3. Chatterji, A. K., Durand, R., Levine, D. I., and Touboul, S. (2016). "Do Ratings of Firms Converge? Implications for Managers, Investors and Strategy Researchers." Strategic Management Journal, 37(8), 1597–1614. DOI: 10.1002/smj.2407.: The pre-Berg benchmark. Distinguishes theorisation (what a rater means by social performance) from commensurability (how it converts that construct to a number). Maps onto scope and measurement divergence respectively, and provides an independent, earlier-vintage estimate of disagreement, useful for testing whether divergence has narrowed over time.

Asset pricing consequences of disagreement

  1. Gibson Brandon, R., Krueger, P., and Schmidt, P. S. (2021). "ESG Rating Disagreement and Stock Returns." Financial Analysts Journal, 77(4), 104–127. DOI: 10.1080/0015198X.2021.1963186.: The closest existing work to RQ3. Treats disagreement as a firm-level characteristic and relates it to returns. This protocol departs from it in the object of study: Gibson Brandon et al. price disagreement as a characteristic; this study prices the portfolio-construction consequence of disagreement, which is a distinct and, for a fiduciary, more actionable quantity.

  2. Avramov, D., Cheng, S., Lioui, A., and Tarelli, A. (2022). "Sustainable Investing with ESG Rating Uncertainty." Journal of Financial Economics, 145(2), 642–664. DOI: 10.1016/j.jfineco.2021.09.009.: Supplies the equilibrium theory. ESG uncertainty raises the market premium, depresses demand for ESG assets, and raises both CAPM alpha and effective beta. Provides the theoretical prior that divergence is not return-neutral in equilibrium, and thus the alternative hypothesis against which H3's null is tested.

  3. Serafeim, G., and Yoon, A. (2023). "Stock Price Reactions to ESG News: The Role of ESG Ratings and Disagreement." Review of Accounting Studies, 28(3), 1500–1530. DOI: 10.1007/s11142-022-09675-3.: Establishes the information-processing channel: consensus ratings predict future ESG news, but predictive ability degrades in the presence of rater disagreement, and price reactions to ESG news weaken. Motivates the interpretation of a null as evidence that neither vendor holds an informational edge.

Equilibrium pricing of sustainability preferences (for interpreting the sign of any surviving alpha)

  1. Pedersen, L. H., Fitzgibbons, S., and Pomorski, L. (2021). "Responsible Investing: The ESG-Efficient Frontier." Journal of Financial Economics, 142(2), 572–597. DOI: 10.1016/j.jfineco.2020.11.001.: Supplies the constrained-optimisation framework this protocol operationalises in Phase 3, and the language of the ESG-adjusted efficient frontier along which the shadow price is measured.

  2. Pástor, Ľ., Stambaugh, R. F., and Taylor, L. A. (2022). "Dissecting Green Returns." Journal of Financial Economics, 146(2), 403–424. DOI: 10.1016/j.jfineco.2022.07.007.: Critical for interpretation. Realised green outperformance in the sample window is driven by unanticipated shifts in climate preferences, not by higher expected returns; theory implies green stocks should earn lower expected returns. Any positive in-sample must therefore be tested against a sentiment-shock control before it is called alpha.

  3. Bolton, P., and Kacperczyk, M. (2021). "Do Investors Care About Carbon Risk?" Journal of Financial Economics, 142(2), 517–549. DOI: 10.1016/j.jfineco.2021.05.008.: Documents a carbon premium in unscaled emissions, providing a candidate omitted risk factor for the augmented specification in §3.6.

Data integrity, the papers that discipline Phase 2

  1. Aswani, J., Raghunandan, A., and Rajgopal, S. (2024). "Are Carbon Emissions Associated with Stock Returns?" Review of Finance, 28(1), 75–106. DOI: 10.1093/rof/rfad013.: Shows that returns correlate with vendor-estimated emissions but not firm-disclosed emissions, and that vendor estimates track financial fundamentals. This is the sharpest published demonstration that ESG-data "alpha" can be an artefact of vendor modelling, and it is the direct methodological precedent for this study's central conjecture. (Note for completeness: the same issue carries a Comment and a Reply; the exchange should be cited in full, as the disagreement is itself informative about the fragility of vendor-data inference.)

  2. Berg, F., Fabisik, K., and Sautner, Z. "Is History Repeating Itself? The (Un)Predictable Past of ESG Ratings." ECGI Finance Working Paper; SSRN 3722087. Working paper, verify publication status at time of submission.: Documents retroactive restatement of historical ESG ratings by a major provider, such that a backtest run today on current-vintage files does not reproduce what an investor could have known in real time. This is the single most important methodological constraint on Phase 2 and the reason point-in-time archives are non-negotiable.

  3. Berg, F., Kölbel, J. F., Pavlova, A., and Rigobon, R. "ESG Confusion and Stock Returns: Tackling the Problem of Noise." NBER Working Paper 30562; SSRN 3941514. Working paper, verify publication status at time of submission.: The noise-correction methodology underlying §1.3.

Econometric and portfolio-construction methods

  1. Fama, E. F., and French, K. R. (2015). "A Five-Factor Asset Pricing Model." Journal of Financial Economics, 116(1), 1–22. DOI: 10.1016/j.jfineco.2014.10.010.: The benchmark model of Phase 3. Note the authors' own finding that HML becomes redundant in the presence of RMW and CMA; §3.7 addresses the resulting multicollinearity.
  2. Cremers, K. J. M., and Petajisto, A. (2009). "How Active Is Your Fund Manager? A New Measure That Predicts Performance." The Review of Financial Studies, 22(9), 3329–3365. DOI: 10.1093/rfs/hhp057.: Active share, used in §1.4 and §3.8.
  3. Novy-Marx, R., and Velikov, M. (2016). "A Taxonomy of Anomalies and Their Trading Costs." The Review of Financial Studies, 29(1), 104–147. DOI: 10.1093/rfs/hhv063.: The transaction-cost framework of Phase 4, including the buy/hold-spread (banding) result that motivates §4.4.
  4. Harvey, C. R., Liu, Y., and Zhu, H. (2016). "…and the Cross-Section of Expected Returns." The Review of Financial Studies, 29(1), 5–68. DOI: 10.1093/rfs/hhv059.: Multiple-testing discipline; the hurdle applied in §5.3.
  5. Ledoit, O., and Wolf, M. (2008). "Robust Performance Hypothesis Testing with the Sharpe Ratio." Journal of Empirical Finance, 15(5), 850–859.: The studentised bootstrap test for the difference of two Sharpe ratios, which is the correct inference for comparing Portfolio A with Portfolio B. (DOI not independently confirmed in this pass; verify before submission.)

Literature required but not cited here, by type. Consistent with the no-hallucination guardrail, the following are categories of source the protocol requires; specific works must be located and verified by the research team rather than assumed:

  • The canonical GRS joint-alpha test (Gibbons, Ross and Shanken, Econometrica, 1989) and the Fama–MacBeth (1973) cross-sectional procedure, both standard, both to be cited from the primary source.
  • HAC covariance estimation (Newey and West, 1987, and the 1994 automatic-bandwidth extension) and the stationary bootstrap (Politis and Romano, 1994).
  • Covariance shrinkage (Ledoit and Wolf, on linear shrinkage toward structured targets) and weight-constraint regularisation (Jagannathan and Ma, on the equivalence between no-short-sale constraints and shrinkage).
  • Delisting-return bias correction (Shumway, 1997, Journal of Finance), required for CRSP handling in §2.7.
  • Low-frequency spread estimators (Corwin and Schultz high–low estimator; Abdi and Ranaldo close–high–low estimator) as alternatives to TAQ-based effective spreads.
  • Square-root market-impact estimation (Almgren et al.; Frazzini, Israel and Moskowitz on live institutional trading costs) for the calibration of in §4.3.
  • Practitioner divergence studies from index providers and asset managers (e.g. work published in the Journal of Portfolio Management and Financial Analysts Journal on ESG rating dispersion), useful for out-of-academic-sample corroboration but to be treated as non-peer-reviewed evidence.

Phase 2: Data Architecture and Algorithmic Bias Normalisation

2.1 Design principle

Every design choice in this phase is subordinated to one rule: the panel must contain only information an investor could have acted on at the close of month . Divergence research is unusually exposed to violations of this rule, because ESG vendor files are delivered as current-vintage databases in which historical scores have been restated, coverage has been backfilled, and methodologies have been retro-applied. A study that fails here does not produce a biased estimate; it produces an estimate of a strategy that was never investable.

2.2 Asset universe

Primary universe: Russell 1000, point-in-time constituents.

Parameter Specification Rationale
Index Russell 1000, as-of constituent lists from FTSE Russell reconstitution files Broad large/mid-cap coverage; rules-based, transparent, market-cap-ranked reconstitution
Reconstitution handling Annual June reconstitution plus quarterly IPO additions, applied with the effective date, not the announcement date Avoids the well-documented reconstitution-anticipation effect contaminating June–July returns
Sample window 2016-01 through 2026-06 (126 months) See §2.3 for the mandatory sub-period split
Share class Issuer-level aggregation; where multiple classes exist, hold the primary CRSP share class (SHRCD 10/11) and aggregate float ESG ratings are assigned at the issuer level, returns at the security level; a naive merge double-counts dual-class issuers
Exclusions Non-US-domiciled ADRs, REITs where GICS treatment changes mid-sample, closed-end funds, SHRCD outside 10/11 Maintains a homogeneous GICS and accounting treatment

Robustness universe: S&P 500. The S&P 500 is run as a secondary specification, not the primary, for three reasons: committee-based inclusion introduces selection endogeneity correlated with governance quality; the smaller cross-section reduces the number of names per quartile from ~250 to ~125, roughly -inflating the idiosyncratic tracking error; and its large-cap skew mechanically compresses the disclosure-intensity variation that identifies . Where the S&P 500 result differs from the Russell 1000 result, the difference is itself reportable evidence on the size-dependence of algorithmic bias.

Coverage-conditioned estimation universe. All primary tests are run on the dual-covered intersection: firms with a valid, in-force rating from both providers at . The coverage-attrition table (§2.8) is a required exhibit, because conditioning on dual coverage is itself a selection and its magnitude must be visible to the reader.

2.3 The methodology-break problem: and a mandatory sample split

Two structural breaks in vendor methodology fall inside a 2016–2026 window. Neither can be ignored, and neither is adequately handled by a linear time control.

Date Event Consequence
September 2018 Sustainalytics launches the ESG Risk Rating, replacing its legacy best-in-class ESG Rating. The new product is an absolute, unbounded risk score in which lower is better, built from exposure and management components, and explicitly designed to be comparable across industries. The score is not merely rescaled, the construct changes from peer-relative performance to absolute unmanaged risk. Chain-linking across this date is not defensible.
30 May 2024 Morningstar Sustainalytics deploys a methodology enhancement (corporate-governance baseline uplift, two standalone Corporate and Stakeholder Governance Material ESG Issues), live in all client systems by 5 June 2024. The vendor reports that 9% of the coverage universe changed ESG Risk Rating category. A 9% category migration is an exogenous shock to eligible-set membership that is unrelated to firm fundamentals. It must be modelled as an event, not absorbed into the residual.

Required design response.

  1. Primary sample: 2018-10 through 2026-06 (93 months), the period over which the Sustainalytics ESG Risk Rating is methodologically stable in construct. This is the head-to-head window.
  2. Legacy sample: 2016-01 through 2018-08, run separately with the legacy Sustainalytics ESG Rating, reported as a robustness exhibit only. Results are never pooled across the September 2018 boundary.
  3. May-2024 break: an event study, not a nuisance. Because the vendor announced the change and quantified its impact, the induced turnover is a clean natural experiment on methodology-driven, fundamentals-orthogonal rebalancing. Estimate abnormal turnover and abnormal return in the month window around 30 May 2024, and report the friction attributable to a single unilateral vendor decision as a standalone exhibit in Phase 4. This converts a data problem into the protocol's most direct evidence on H5.
  4. MSCI methodology actions must be handled symmetrically. MSCI ESG Ratings are published as a AAA–CCC letter grade underpinned by a continuous industry-adjusted score; both the letter thresholds and the underlying key-issue weights are subject to periodic annual methodology review. The research team must obtain the vendor's dated methodology-change log and construct a corresponding break calendar before estimation. Every break date enters the specification as an indicator interacted with the eligible-set-change dummy.

Consequence for the primary window: 93 months is a short sample for alpha inference. §3.9 reports the resulting minimum detectable effect and prescribes the remedies. This is disclosed at the design stage rather than discovered at the referee stage.

2.4 Normalisation: making incommensurable scales comparable

The two vendors produce scores that differ in direction (higher-is-better vs lower-is-better), support (bounded 0–10 industry-adjusted vs an unbounded-in-principle risk score), construct (peer-relative vs absolute) and granularity (heavily discretised vs quasi-continuous). Any comparison must therefore be invariant to monotone transformation. Three steps, applied in order, at every month , using only the cross-section available at .

Step 1: Sign harmonisation. Sustainalytics' ESG Risk Rating is a risk measure in which lower is better. Define the sign-harmonised score so that higher is better for both vendors:

(9)

This step is the most common silent error in the applied literature. A sign inversion converts the study's central result into its exact negation and will not announce itself in any diagnostic. The protocol requires a hard assertion in code: in every month, with the job failing if the assertion is violated.

Step 2: Rank / normal-score transformation (primary specification). Because the underlying scales are ordinal-incommensurable, the primary normalisation is the van der Waerden normal-score transform using the Blom plotting position, applied cross-sectionally within month:

(10)

where is taken over the dual-covered cross-section at with mid-ranks for ties. This transform is invariant to any strictly increasing reparameterisation of either vendor's scale, which is exactly the invariance required when one construct is peer-relative and the other absolute. It also renders the bivariate-normal geometry of §1.4 an assumption about the copula only, which is testable rather than imposed.

Step 3: Winsorised cross-sectional Z-score (robustness specification). For comparability with the applied literature and to preserve cardinal information about the magnitude of divergence:

(11)

All headline results are reported under both transforms. Divergence between them is diagnostic of tail-driven rather than central-tendency-driven disagreement, and is reported as such.

2.4.1 The industry-neutralisation decision: a design fork that must not be collapsed

Standard practice is to z-score within GICS industry group. For this study, that would destroy the object of interest. The sector bet induced by scope divergence is precisely what RQ2 measures; industry-neutralising the score removes it before it can be estimated. The protocol therefore constructs and carries both layers throughout, and treats their difference as an estimand rather than a nuisance:

(12)

The contrast between the portfolio built on and the portfolio built on identifies the sector-allocation channel by construction: any difference in tracking error, factor loading or alpha between the two is attributable to the sector content of the rating. This is the cleanest available identification of the "uncompensated sector bet" in the study's title, and it requires no additional data.

2.5 Estimating the algorithmic bias vector

Estimate, by month and then pooled, a seemingly-unrelated system in which the same characteristic vector explains each vendor's normalised score:

(13)

The characteristic vector is specified ex ante and frozen before estimation:

  • market capitalisation (the size/attention channel)
  • Disclosure intensity: count of non-missing ESG data points reported, or a vendor-independent disclosure score; plus an indicator for publication of a standalone sustainability report and for third-party assurance
  • GICS sector indicators (10 or 11 sectors, dated per §2.6)
  • R&D / assets, capex / assets, leverage, gross profitability, asset growth (the FF5 characteristic space, so that any style tilt is visible)
  • Foreign sales share, employee count, unionisation-free operational scale proxies
  • Analyst coverage, institutional ownership, index membership indicators
  • Firm age since first CRSP appearance

Test of H1. Stack the system and test equality of the bias vectors by a Wald statistic on the difference, with standard errors clustered two-way by firm and month:

(14)

Individual coefficient contrasts deliver the vendor-behaviour restrictions of §1.2 as directional -tests: and under the "MSCI rewards compliance capacity" characterisation.

Decomposing observed divergence. The fitted system yields the operational split of §1.2:

(15)

This share is the Phase 2 headline statistic and the bridge to Phase 3.

2.6 Detecting industry-average gap-filling

The hypothesis that Sustainalytics substitutes sub-industry averages where firm-specific evidence is absent generates three distinct, independently testable signatures. All three should be run; agreement across them constitutes evidence, any one alone does not.

(a) Variance-ratio test. Under gap-filling, within-sub-industry dispersion is compressed:

(16)

Scale-dependence is removed by applying the test to the normalised rather than raw scores, and by reporting a rank-based Fligner–Killeen statistic alongside the parametric ratio.

(b) Shrinkage-intensity regression. Regress each firm's deviation from its sub-industry mean on its missing-data density (the fraction of the vendor's own indicator set that is unreported by the firm):

(17)

A negative , firms that disclose less sit closer to their peer average, is the signature of gap-filling. The identical regression on provides the vendor contrast; is the directional prediction.

(c) Point-mass / clustering diagnostic. Test for excess probability mass at sub-industry medians using a discontinuity-style density test on (median-centred), reporting excess kurtosis and the mass within of zero against the empirical null generated by permuting industry labels.

2.7 Point-in-time integrity, look-ahead and survivorship

Identifier crosswalk. The linkage must be historical, not current.

Step Rule
Vendor → security Match on ISIN/SEDOL as of the rating date, never on a current-vintage identifier file
CUSIP handling Use CRSP NCUSIP (historical CUSIP), never CUSIP (current). Ticker matching is prohibited outright
CRSP ↔ Compustat CCM link table with LINKTYPE ∈ {LU, LC}, LINKPRIM ∈ {P, C}, and the observation date inside [LINKDT, LINKENDDT]
Issuer ↔ security Map vendor issuer IDs to PERMCO, then select the primary PERMNO within PERMCO
GICS Use dated GICS. The Real Estate sector separation (2016) and the March 2023 GICS revision (retail/consumer-staples reclassification) both fall inside the window; a current-vintage GICS assignment silently rewrites sector history

Rating availability lag. Vendor files carry both a rating as-of date and a publication/dissemination date; these differ, sometimes by weeks. The eligible set at month is formed only from ratings whose dissemination date precedes the formation date:

(18)

A month variant is run as a conservatism check; a material difference between and results indicates the strategy depends on rapid vendor-file ingestion, which is itself a reportable operational finding.

Archival vintages. Where the vendor supplies an as-was archive (monthly snapshot files), that archive is the primary source and the current-vintage file is used only to measure restatement. Define the restatement diagnostic:

(19)

Reporting the distribution of , and re-running the headline test on current-vintage data to show how much the measured alpha changes, is a required exhibit. It converts the Berg–Fabisik–Sautner critique from a caveat into a quantified sensitivity. If an as-was archive cannot be obtained for a vendor, that vendor's results must be labelled as subject to unquantified restatement bias, and the paper must say so in the abstract.

Survivorship and delisting. Portfolio membership is determined by the PIT universe at formation and is not conditioned on survival to the end of the sample. Delisted names are held to the delisting date and the CRSP delisting return DLRET is applied. Where DLRET is missing and the delisting code indicates a performance-related delisting (CRSP codes 500 and 520–584), the standard correction of the delisting-bias literature (Shumway, 1997) is applied; a variant and a "missing-set-to-zero" variant are reported as bounds. Proceeds are reinvested pro rata across the surviving eligible set at the delisting date.

2.8 Missing data

Coverage is missing not at random: the probability of being rated rises with size, index membership, disclosure and analyst attention, the same variables that constitute . Naive imputation is therefore not a neutral repair; it manufactures the bias under study.

Prohibited. Imputing a missing vendor score from that vendor's own industry mean. This is mechanically identical to the gap-filling behaviour being tested in §2.6, and would guarantee a false positive.

Required, in this order:

  1. Primary, complete-case on the dual-covered intersection. Report the attrition table: universe size, MSCI coverage, Sustainalytics coverage, intersection, by year, in both counts and market-cap share. Coverage differentials across vendors are a first-order finding in their own right.
  2. Selection quantification: Heckman-type model. Estimate a first-stage probit for dual coverage and include the inverse Mills ratio in the score regressions of §2.5. The exclusion restriction (a variable affecting coverage but not the score) is the weak link; index-membership timing and analyst-coverage shocks are candidates, and the identifying assumption must be argued explicitly rather than asserted.
  3. Robustness, multiple imputation. MICE with imputations using only PIT-available covariates, combined by Rubin's rules:
(20)
  1. Bounds, worst-case sensitivity. Report Manski-style worst-case bounds on obtained by assigning uncovered firms the most and least favourable feasible score. If the headline conclusion survives the bounds, coverage selection is not driving it; if it does not, the paper must report the result as bounded rather than point-identified.

2.9 Deliverable of Phase 2

A single reproducible panel, PANEL_PIT, keyed on (PERMNO, month-end), containing: sign-harmonised and both-normalisation vendor scores under universe-wide and industry-neutral variants; eligibility flags , ; the decomposed divergence ; dated GICS; the FF5 characteristic set; CRSP returns with delisting adjustment; liquidity fields (ADV, effective spread, quoted depth); coverage and restatement flags; and a per-cell provenance stamp recording vendor file vintage and dissemination date. No downstream analysis may read any file other than PANEL_PIT.

Phase 3: Portfolio Optimisation and Econometric Testing

3.1 Inputs to the optimiser

Covariance matrix. With and an estimation window of 60 months, the sample covariance is singular or near-singular and its extreme eigenvalues are severely biased. Two estimators are carried in parallel and results reported under both:

Linear shrinkage. Shrink the sample covariance toward a structured target (constant-correlation, or a single-index target):

(21)

Factor model. A fundamental factor structure with market, the four FF5 style factors, and GICS sector blocks:

(22)

The factor form is not optional: the sector block is what makes the tracking-error decomposition of §3.8 possible, and it is the mechanism by which "uncompensated sector bet" becomes a measured quantity rather than a claim.

Expected returns. Sharpe maximisation requires , and estimated from realised means is the dominant source of error in mean-variance optimisation. To ensure the study measures the effect of the ESG constraint rather than the effect of estimation noise, three specifications are run and the headline conclusion must hold across all three:

  1. Equilibrium anchor (primary). Reverse-optimised from market weights, , with the implied risk-aversion coefficient. This is the Black–Litterman prior and it makes the unconstrained benchmark coincide with the market, so any deviation is attributable to the constraint.
  2. Characteristic-based. , with the PIT FF5 characteristics and estimated by Fama–MacBeth on data through only.
  3. Historical. 60-month trailing means, reported solely to demonstrate the fragility that motivates (1) and (2).

Two -free control portfolios are run alongside: minimum-variance subject to the same constraints, and equal-weight within the eligible set. If the A-versus-B result is materially different for the -free portfolios, the finding is an artefact of expected-return estimation and must be reported as such.

3.2 Benchmark portfolio : unconstrained mean–variance

The maximum-Sharpe problem is a linear-fractional program. It is solved exactly as a convex quadratic program by the standard homogenisation: introduce with and normalise the excess-return functional to unity.

(23)

Equivalent QP (solved in practice):

(24)

Baseline parameterisation: (5% single-name cap, consistent with UCITS-style diversification and with the concentration limits in most institutional IPS documents), (±10 percentage points of active sector weight) in the risk-controlled variant and in the unconstrained variant. Both sector-bound variants are required: the contrast between them prices the sector bet directly, since a mandate that forbids the sector deviation must express the ESG constraint through name selection instead.

3.3 Portfolio A and Portfolio B: the ESG-constrained programs

Two constraint architectures are in institutional use, they are not equivalent, and the difference between them is informative. Both are run.

Architecture (i), exclusionary eligibility screen. The mandate excludes ineligible names outright:

(25)

is identical with . Nothing else changes: same , same , same caps, same calendar, same solver, same seed. This is the experimental design, a single-factor manipulation.

Architecture (ii), portfolio-level score constraint. The mandate requires the weighted-average portfolio score to clear the threshold, which is how most real ESG mandates and index methodologies are drafted:

(26)

Architecture (ii) is a single linear constraint and therefore admits a shadow price, which is the study's cleanest scalar summary of the cost of the mandate.

The shadow price of the ESG constraint. Attach multiplier to the score constraint. The Karush–Kuhn–Tucker stationarity condition for the interior names is:

(27)

and by the envelope theorem the multiplier is the marginal Sharpe cost of tightening the mandate by one unit of normalised score:

(28)

The single most policy-relevant statistic in the study is the difference : the Sharpe-ratio cost of the choice of vendor, holding the stringency of the ESG mandate fixed. Report its time series, its mean with HAC standard errors, and its dispersion.

3.4 The Divergence portfolio : the study's key test asset

Comparing with confounds two things: the names both vendors like (which cancel) and the names on which they disagree (which do not). To isolate the second, construct a self-financing long–short portfolio from the symmetric difference of the eligible sets:

(29)
(30)

Why this is the right test asset. Both legs consist exclusively of firms that a reputable provider certifies as top-quartile ESG. The portfolio is therefore ESG-neutral by construction under the mandate's own definition of ESG quality, and every unit of its risk and return is attributable to inter-vendor disagreement. Its FF5 alpha answers RQ3 in its sharpest form: is there any compensation, of any kind, for taking one vendor's view over the other's?

A value-weighted variant (weights proportional to float within each leg) and a beta-neutralised variant (leg weights scaled so ex ante) are both reported, since equal weighting introduces a mechanical small-cap tilt that would otherwise be mistaken for a divergence effect.

3.5 Rebalancing calendar and implementation rules

Rule Specification
Frequency Monthly (primary) and quarterly (secondary), both end-of-month formation with next-month returns
Formation lag Signal at per §2.7; returns accumulate over
Buffer / banding Baseline: none (hard 75th-percentile threshold). Variants: enter at the 75th, exit at the 65th and at the 60th percentile
Solver Convex QP, identical warm start and tolerance across A and B; solver status logged per month; any non-converged month excluded identically from all portfolios
Cash Fully invested, ; long-only in the primary specification
Rounding No lot rounding in the primary specification; a round-lot variant reported in Phase 4 costs

The banding variants are not cosmetic. Novy-Marx and Velikov (2016) show that a buy/hold spread is the most effective single lever for reducing the trading costs of a signal-based strategy. If a modest band collapses the divergence-induced turnover without materially changing the portfolio's ESG profile, that is a directly implementable recommendation for asset owners and belongs in the paper's conclusion.

3.6 Asset-pricing regressions

Level specification, each portfolio against the Fama–French five-factor model:

(31)

For the self-financing divergence portfolio the left-hand side is the raw return , since no risk-free rate is invested.

Difference specification, the direct test of H3. Rather than differencing two separately estimated alphas (which understates the covariance of the estimates), estimate the differenced series directly so that the standard error is correct by construction:

(32)
(33)

Sector-augmented specification, the direct test of H4. Append excess returns on GICS sector-mimicking portfolios (long sector , short the cap-weighted universe) to absorb allocation effects. The change in measures how much of the raw active return was a sector bet:

(34)

Augmented models (robustness). FF5 + momentum (UMD) as the six-factor model; the -factor model of Hou, Xue and Zhang as a non-nested alternative; a betting-against-beta and quality-minus-junk augmentation to test whether ESG screens are quality proxies; a carbon factor following the emissions literature; and a climate-sentiment control following Pástor, Stambaugh and Taylor (2022), without which any positive in-sample green alpha is not interpretable as an expected-return effect.

Inference.

(35)

At months this yields (§A.8). Standard errors are reported three ways: OLS, Newey–West HAC, and a stationary block bootstrap with expected block length calibrated to the residual autocorrelation, with the most conservative governing the reported significance.

Joint test across portfolios. For the vector of alphas , apply the Gibbons–Ross–Shanken statistic:

(36)

with the number of test portfolios and the number of factors.

Firm-level cross-sectional specification. Because the portfolio-level test is power-constrained (§3.9), the firm-level panel is run as the higher-power complement, with two-way clustered standard errors and Fama–MacBeth as the cross-sectional alternative:

(37)

Under the divergence hypothesis, and carry the loading and are individually insignificant once is included, the signature of two noisy measurements of a common signal, neither of which dominates.

3.7 Multicollinearity and identification hygiene

The design has three collinearity problems, each of which must be diagnosed rather than assumed away.

  1. and are strongly correlated by construction (). Including both in a firm-level regression inflates variances. Remedy: report the orthogonalised pair, the common component and the divergence , which are uncorrelated when the two scores have equal variance, as they do after normalisation. This reparameterisation is exactly identified and interpretable: is the consensus ESG signal, is the vendor disagreement.
  2. HML is close to redundant in FF5 in the presence of RMW and CMA, a point Fama and French (2015) make themselves. Remedy: report the variance inflation factors and the condition number for every regression; report the FF5 and the FF5-without-HML specifications side by side; never interpret an individual in isolation.
  3. Sector-mimicking portfolios overlap with style factors (Energy with value, Technology with growth and investment). Remedy: orthogonalise against before inclusion, and report the incremental from the sector block via a partial -test rather than reading individual .

3.8 Isolating tracking error and the active sector bet

Definition. Annualised tracking error of A against B, from realised monthly active returns:

(38)

Ex-ante risk decomposition, the test of H4. With active weights and the factor covariance structure of §3.1:

(39)

and, partitioning , the active sector tracking error is the sector block of the systematic term:

(40)
(41)

The cross-term is retained explicitly rather than allocated, since sector and style exposures are not orthogonal in practice and silently dropping it overstates the sector share.

Ex-post return attribution: Brinson decomposition. Treating B as the benchmark and A as the portfolio, with sector weights and within-sector returns :

(42)

Active share and turnover-adjusted overlap. Reported monthly alongside the theoretical benchmark of §1.4, so that realised divergence can be compared with the value predicted by the observed inter-rater correlation:

(43)

The joint test that constitutes the paper's central claim. "Uncompensated" is a joint statement: risk without return. It must be tested jointly rather than as two separate results.

(44)

Because "failure to reject " is not evidence of a zero, the protocol requires an equivalence test rather than a bare null result: a two-one-sided-test (TOST) procedure against an economically meaningful margin (proposed: 50 bp p.a., the order of a typical active-management fee), reporting whether is statistically inside . Simultaneously, a Ledoit–Wolf (2008) studentised bootstrap tests directly, which is the correct inference for the Sharpe comparison and is not obtainable from the alpha regression.

3.9 Statistical power: declared in advance

For an annualised residual volatility and years of monthly data, the minimum detectable annualised alpha at 5% two-sided size and 80% power is:

(45)

where the second factor is the standard-error inflation from estimating factor means (the same quantity that appears in the GRS statistic).

The theoretical idiosyncratic tracking-error floor for two equal-weighted quartile portfolios drawn from a 1000-name universe is computed in §A.4: at and 30% idiosyncratic volatility it is approximately 1.9% p.a. before any systematic component. Combining this with §A.3, over the 7.75-year primary window the study can detect an annualised alpha differential of roughly 2%, but is underpowered against a 50 bp effect, which is precisely the magnitude an asset owner would care about.

This is a design constraint, not a defect, and it is addressed rather than concealed:

  1. Report the MDE prominently in the results section; a null result is reported as "not distinguishable from zero at detectable magnitudes," never as "no effect."
  2. Use the equivalence test of §3.8 so that a null carries positive information.
  3. Estimate factor loadings from daily returns (with Dimson-style lead–lag adjustment for non-synchronous trading) to sharpen , which reduces and therefore the MDE.
  4. Lean on the firm-level panel (§3.6), where the effective sample is rather than , for the primary inference on the divergence coefficient.
  5. Extend the sample backward on the legacy Sustainalytics product and forward as data accrue; report the MDE at each horizon.
  6. Report the cost-side results (Phase 4) as the primary economic magnitude, since turnover drag is estimated with far greater precision than alpha and is measured, not inferred.

Phase 4: Institutional Impact and Friction Modelling

4.1 Turnover: definition and the drift correction

Turnover must be measured against drifted weights, not against last month's target weights; the difference is the passive drift a manager never trades and is a common source of overstated turnover in published backtests.

(46)
(47)

4.2 Decomposing turnover: isolating the rating-migration component

Total turnover has four sources, only one of which is the object of this study. They are separated by a sequence of counterfactual portfolios, each differing from the last in exactly one input:

Component Counterfactual construction Interpretation
Optimiser re-run with all inputs frozen except realised prices Mechanical rebalancing to targets
Add index reconstitution and corporate actions; ratings frozen Universe maintenance, common to any mandate
Add live rating migration; risk and return inputs frozen The divergence cost
Add live Estimation-noise churn in the optimiser
(48)

The frozen-ratings control portfolio is the key device: it holds the vendor's file at its vintage and lets everything else evolve. The difference in turnover between the live and frozen portfolios is attributable to rating migration by construction, and it is estimable with far greater precision than any alpha.

Further split rating-migration turnover by the source of the migration, which maps back to Phase 2:

(49)

is estimated from the May-2024 Sustainalytics event window and any dated MSCI methodology actions. A material is the study's most direct evidence that a single vendor decision imposes a real, involuntary cost on every mandate written against that vendor.

The turnover floor implied by rating persistence. Estimate the monthly quartile transition matrix for each vendor. With the probability of remaining in the top quartile, the expected tenure in the eligible set and the implied minimum one-way turnover are:

(50)

Calibration (§A.6) shows how severe this is: a monthly top-quartile retention of 0.95, which sounds high, implies a 60% annual one-way turnover floor from rating migration alone, before any optimiser or price-drift turnover. This single calculation is likely to be the most quoted result in the paper, and it requires no return data at all.

4.3 The transaction-cost model

Implicit costs are modelled as a spread term plus a concave market-impact term, the standard square-root specification:

(51)
Parameter Source Notes
TAQ effective spread (Lee–Ready signing) where available; Corwin–Schultz high–low or Abdi–Ranaldo close–high–low estimator as the low-frequency fallback Report both; large-cap US equity effective spreads are typically a few basis points, so the estimator choice matters more for the mid-cap tail
Trailing 21-day realised daily volatility from CRSP
Trailing 21-day median dollar volume; NASDAQ volume halved for the double-counting convention in the early sample
Calibrated to the published institutional-trading-cost literature; estimated as a free parameter in a sensitivity Report results across rather than at a single point
Commission schedule assumption, stated explicitly

Aggregate drag. Realised cost in period and the annualised drag:

(52)

Net-of-cost performance. Every headline statistic is re-reported net:

(53)

Calibration. §A.5 tabulates the mapping from turnover and average cost to annual drag. At 60% one-way annual turnover and a 20 bp round-trip-equivalent cost, the drag is 24 bp p.a.; at 120% turnover and 40 bp it is 96 bp p.a. Set against the MDE of §3.9, this is the study's central asymmetry: the friction is measurable with precision, and the alpha that would justify it is not.

4.4 Break-even and capacity

Break-even alpha. The alpha a vendor's signal must generate simply to pay for the turnover it induces:

(54)

where is the one-off turnover required to migrate a portfolio from the A-mandate to the B-mandate, the transition cost of switching vendor, a number no asset owner is currently quoted.

Capacity. Setting the drag equal to an alpha budget and inverting the impact function gives the maximum sustainable participation rate, and hence the AUM at which the strategy self-destructs:

(55)

§A.7 shows the participation ceilings implied by illustrative parameters. The economically important qualitative result is that capacity falls with the square of turnover under the square-root law: doubling divergence-induced turnover cuts sustainable AUM by roughly a factor of four at a fixed alpha budget.

Mitigation levers to be priced, not merely described. Each is run as a full portfolio variant so that the paper delivers actionable numbers rather than advice:

Lever Implementation Reported metric
Banding Enter at 75th, exit at 65th/60th percentile , average portfolio ESG score,
Vendor consensus Screen on , or on the intersection TE vs either single-vendor portfolio; turnover; coverage loss
Sector neutrality Screen on the industry-neutral (§2.4.1) Change in ASTE share; change in
Score-constraint architecture Architecture (ii) instead of (i) ; turnover; active share
Rebalance frequency Quarterly vs monthly Drag; signal decay cost

4.5 Conclusion design: what the paper argues, under each outcome

The protocol pre-commits to the interpretation of each possible result, so that the conclusion is not written to fit the data.

Outcome 1: Divergence thesis supported. is large, the systematic block dominates , is statistically equivalent to zero within a 50 bp margin, and imposes a measurable drag. Conclusion: the "ESG alpha" reported on a single vendor's file is not identified as a property of corporate sustainability. Vendor choice is a live, unrewarded active bet, concentrated in sector allocation, financed by the beneficiary.

Outcome 2: Divergence thesis rejected. is reliably non-zero. Conclusion: one vendor's algorithm contains information the other's does not. The paper then pivots to which characteristics carry the differential, recovered from interacted with returns, and the contribution becomes a decomposition of vendor skill. This is a publishable result, and the design must be equally capable of producing it.

Outcome 3: Indeterminate. The confidence interval on contains both zero and economically material values. Conclusion: report the tracking error and turnover findings, which are precisely estimated, as the paper's contribution, and report the alpha result as bounded. State the MDE and the sample length required to resolve it. This outcome must be reported as readily as the other two.

4.6 Implications for institutional asset owners

The following addresses the governance and fiduciary consequences of the empirical design. It is an analysis of the economics of the decision, not legal advice; the applicable duties differ by jurisdiction, plan type and mandate, and a plan's counsel should be consulted on the legal standard that governs it.

The structure of the fiduciary problem. Prudence standards in most institutional settings are process standards, not outcome standards: a fiduciary is judged on whether the decision-making procedure was reasonable given what was knowable, not on whether the investment performed. That framing is what makes this study's results operative. If two reputable vendors, applied to the same mandate, produce portfolios that differ in 39%–58% of their holdings, then a trustee who selects one vendor has made a consequential active decision. Three propositions follow, in increasing order of strength:

  1. Vendor choice is a material investment decision and should be documented as one. It is currently treated in most investment policy statements as an operational or data-procurement matter, frequently delegated to a manager or a consultant and disclosed in neither the IPS nor the risk report. If the choice moves half the portfolio, it belongs in the same governance tier as the benchmark, the constraint set and the rebalancing policy.
  2. A single-vendor mandate should be accompanied by a documented provider-sensitivity analysis. The natural artefact is a "rating provider sensitivity report", a periodic exhibit reporting the active share, tracking error, sector-active weights and turnover of the mandate as implemented against at least one alternative provider. The methodology of Phase 3 is exactly the specification for such a report, and it is inexpensive once the second vendor's file is licensed. This is a concrete, adoptable recommendation and should be presented as the paper's principal practitioner deliverable.
  3. Under a mean-variance framing, uncompensated tracking error is difficult to justify on pecuniary grounds alone. If while , then relative to a consensus or intersection construction, the single-vendor mandate adds active risk and cost without an expected-return offset. Where a plan's governing framework permits non-pecuniary or beneficiary-preference considerations, that is a separate and legitimate basis for the decision, but it is a different justification, and the two should not be conflated in the record. Where the framework requires pecuniary justification, the burden falls on demonstrating that the chosen vendor's methodology is better aligned with the plan's financial-materiality view, which is an argument about and can be made explicitly with the tools of §2.5.

On black-box dependence specifically. Three features of the vendor relationship documented in Phase 2 are governance issues independent of any return result:

  • Unilateral methodology change. A vendor can move 9% of a coverage universe across rating categories on an announced date, forcing involuntary turnover in every mandate written against it. The asset owner bears the cost and has no contractual recourse in a standard data licence.
  • Retroactive restatement. Where historical scores are rewritten, an owner cannot reconstruct what the mandate would have held at a past date, which undermines the ability to audit past compliance with the mandate. Contracting for as-was archival delivery is a low-cost mitigation and should be a standard licence term.
  • Coverage asymmetry. Vendors cover different universes, so the effective investable set is partly determined by the vendor's commercial coverage decisions rather than by the plan's policy.

The regulatory direction of travel is toward transparency, not standardisation. Regulation (EU) 2024/3005 on ESG rating activities applies from 2 July 2026, with ESMA as the supervisor; providers must be authorised or recognised, must disclose the methodologies, models and key assumptions underlying their ratings, and are subject to conflict-of-interest and record-keeping requirements. Critically, the regulation mandates disclosure of methodology, not convergence of methodology. It should therefore be expected to make observable without making smaller. The practical consequence for an asset owner is that the divergence documented here is likely to persist as a structural feature of the market while becoming, for the first time, systematically auditable, which raises rather than lowers the governance expectation on the owner to examine it. This regulatory window also creates a natural extension: the pre- and post-July-2026 comparison is a candidate difference-in-differences design on whether mandated methodology disclosure narrows measurement divergence, and the panel built in Phase 2 is already the right dataset for it.

5. Robustness, Falsification and Inferential Discipline

5.1 The placebo that the study lives or dies by

Any two imperfectly correlated screens produce non-overlapping portfolios and non-zero tracking error. A finding of large active share is therefore not evidence of an ESG-specific phenomenon unless it is benchmarked against what any pair of correlated signals would produce. The required falsification test:

Synthetic rater placebo. Generate pseudo-score pairs with a cross-sectional correlation matched month-by-month to the realised but with no ESG content, a Gaussian copula over the same marginals, and separately a bootstrap that permutes the vendor's own residual across firms within sub-industry. Run the full Phase 3 pipeline on the synthetic pairs, times, and compare the realised statistics against the synthetic null distribution:

(56)

The interpretation is sharp. If observed active share and turnover sit inside the synthetic null, the divergence effect is the generic consequence of imperfect correlation, real, costly, but not specific to ESG algorithms, and the paper says so. If the observed sector share of tracking error and the systematic share of divergence sit far in the right tail of the null (which is the prediction, since synthetic noise has no characteristic structure by construction), the effect is algorithmic bias rather than sampling variation. This single test is what separates a contribution from a tautology, and it should be run before any other result is looked at.

5.2 Robustness battery

Dimension Variants required
Universe Russell 1000 / S&P 500 / Russell 3000 large-mid
Threshold 75th percentile (primary); 60th, 80th, 90th; and AAA–A vs BBB–CCC letter-grade cut for MSCI
Normalisation Normal-score / winsorised Z / raw-score percentile; universe-wide vs industry-neutral
Weighting Optimised / cap-weighted within eligible set / equal-weighted
Constraint architecture Exclusionary (i) / portfolio-score (ii)
Rebalance Monthly / quarterly; no band / 10 pp band / 15 pp band
Covariance Ledoit–Wolf shrinkage / fundamental factor model / 60-month sample with ridge
Expected returns Equilibrium / characteristic / historical / minimum-variance / equal-weight
Factor model FF5 / FF6 / -factor / FF5 + sector block / FF5 + climate-sentiment control
Delisting Shumway / / zero
Signal lag 1 month / 3 months
Vendor vintage As-was archive / current vintage (to quantify restatement bias)
Sample period Full / pre- and post-COVID / pre- and post-2022 ESG-flow reversal / excluding the May-2024 methodology event
Third vendor Repeat the entire A-vs-B design as A-vs-C and B-vs-C

The A-vs-C and B-vs-C replications matter more than they appear to. A finding that holds only for one vendor pair is a statement about those two firms; a finding that holds across all three pairs is a statement about the ESG-rating industry, which is the claim the paper wants to make.

5.3 Multiple testing

The robustness battery generates a large family of alpha estimates. Reporting the best of them as the result would be exactly the error that Harvey, Liu and Zhu (2016) document across the factor literature. The protocol therefore requires:

  • Pre-specification of the primary specification, in writing, before data extraction. Everything else is labelled secondary.
  • Romano–Wolf stepdown for family-wise error control across the family of alphas, which respects the strong cross-specification dependence that Bonferroni ignores.
  • Benjamini–Hochberg FDR control at reported alongside.
  • The hurdle of Harvey, Liu and Zhu applied to any claim of a new priced effect.
  • A specification curve: every variant's and confidence interval plotted in a single ordered exhibit, so the reader sees the full distribution of results rather than a chosen point. Given the study's null-oriented hypothesis, the specification curve is arguably the single most honest exhibit available and should be a main-text figure.

5.4 Reproducibility

Requirement Specification
Pre-registration Deposit this protocol, with the primary specification and hypotheses fixed, prior to data extraction (OSF or equivalent)
Environment Containerised; all package versions pinned; single fixed random seed recorded for every stochastic procedure
Data provenance Every field in PANEL_PIT carries vendor file name, vintage date and extraction timestamp
Code Deterministic pipeline, raw → PANEL_PIT → portfolios → tables; no manual steps; unit tests on the sign-harmonisation assertion, the PIT lag, and the link-table date bounds
Disclosure Licensed vendor data cannot be redistributed; publish code, the full data dictionary, the coverage-attrition table, and a synthetic-data replication that reproduces every table's structure
Conflicts Disclose all vendor data licences, any vendor funding, and any vendor review of the manuscript prior to submission

5.5 Limitations, stated in advance

  1. Power. The primary window is 93 months. The study cannot resolve alpha differentials below roughly 2% p.a. at the portfolio level (§3.9). This is disclosed in the abstract, not buried.
  2. Two vendors are not the industry. MSCI and Sustainalytics are the two largest, but the A-vs-C and B-vs-C replications are what license any industry-wide claim.
  3. Correlated vendor error. If (likely, given shared corporate-disclosure inputs) the IV correction of §1.3 is inconsistent and the reliability ratio is overstated. The Hansen test detects this but does not repair it.
  4. US large-cap only. Divergence is documented to be larger where disclosure is weaker; the US large-cap result is therefore a lower bound on the global effect, and should be described as such rather than generalised.
  5. Ex-ante costs, not executions. The cost model is calibrated, not observed. Access to real institutional execution data would strengthen Phase 4 materially and should be pursued.
  6. The mandate is a stylisation. Real mandates combine screens, exclusions, engagement policies and benchmark-relative constraints. The single-threshold design is chosen for identification, and the paper should be explicit that it isolates one mechanism rather than replicating any actual product.
  7. Regime dependence. The sample spans an ESG inflow boom and a subsequent reversal. Pástor, Stambaugh and Taylor (2022) imply that realised green returns in such a window reflect preference shocks rather than expected returns; conclusions about expected alpha must be conditioned accordingly.

5.6 Indicative execution sequence

Stage Work Gate before proceeding
1 Pre-registration; vendor licensing including as-was archives; methodology-change log obtained from both vendors Protocol frozen and deposited
2 Build PANEL_PIT; coverage-attrition table; sign-harmonisation and PIT assertions pass Panel validation report signed off
3 Estimate ; test H1; gap-filling diagnostics; Shapley re-decomposition Systematic share of divergence estimated
4 Run the synthetic-rater placebo (§5.1) before looking at any live portfolio result Placebo null distribution in hand
5 Build ; regressions; TE decomposition; equivalence tests Primary specification results
6 Turnover decomposition; cost model; capacity; mitigation-lever variants Phase 4 exhibits
7 Robustness battery; specification curve; multiple-testing corrections Full results set
8 Drafting; practitioner-facing provider-sensitivity report template Submission

Target outlets. Review of Finance and Journal of Financial Economics for the full paper, given the anchor literature's home; Financial Analysts Journal or Journal of Portfolio Management for a practitioner-facing companion built around the turnover floor, the capacity curve and the provider-sensitivity report template, which is the part of this work most likely to change institutional behaviour.

Appendix A: Calibration Tables

All tables in this appendix are computed, not asserted. The generating script is reproduced in Appendix B.

A.1 Eligible-set overlap and implied active share

Bivariate normal scores, top-quartile screen (, ). Active share assumes equal-weighted portfolios of equal cardinality, for which .

Overlap (retention) Jaccard Implied active share
0.30 0.0951 38.0% 0.235 62.0%
0.38 0.1047 41.9% 0.265 58.1%
0.45 0.1136 45.5% 0.294 54.5%
0.54 0.1258 50.3% 0.336 49.7%
0.60 0.1345 53.8% 0.368 46.2%
0.71 0.1522 60.9% 0.437 39.1%
0.80 0.1691 67.6% 0.511 32.4%
0.90 0.1930 77.2% 0.629 22.8%
0.95 0.2098 83.9% 0.723 16.1%

The region of interest is , the range reported by Berg, Kölbel and Rigobon (2022). Note the amplification: even at , far above anything observed between ESG vendors, and still well below the ~0.99 typical of credit ratings , nearly a quarter of holdings differ.

A.2 Attenuation of the single-vendor ESG–return coefficient

Under symmetric classical measurement error and common bias, the inter-rater correlation is the reliability ratio , and OLS on one vendor's score recovers rather than .

Attenuation of IV rescaling of the point estimate, IV standard-error inflation, First-stage
0.38 62.0% 2.63× 2.63× 0.144
0.45 55.0% 2.22× 2.22× 0.203
0.54 46.0% 1.85× 1.85× 0.292
0.60 40.0% 1.67× 1.67× 0.360
0.71 29.0% 1.41× 1.41× 0.504
0.80 20.0% 1.25× 1.25× 0.640

The two middle columns are deliberately identical, and the coincidence is the point. Under classical, symmetric, uncorrelated measurement error in the just-identified case, IV scales the point estimate by and its standard error by the same :

(57)

The correction therefore buys the right magnitude, not additional statistical power. Any paper claiming that noise correction reveals a previously insignificant ESG effect must be departing from this benchmark (through over-identification, heteroskedasticity-based identification, or a violation of the exclusion restriction) and must say which. The last column shows that the first stage is nonetheless informative in a firm-level panel: at , across tens of thousands of firm-months yields a first-stage far above conventional weak-instrument thresholds, so weak identification is a concern only in short, small cross-sections.

A.3 Minimum detectable annualised alpha (5% two-sided, 80% power)

, before the GRS standard-error inflation of §A.8.

(ann.)
1.0% 1.25% 1.01% 0.86% 0.72% 0.63%
2.0% 2.51% 2.01% 1.73% 1.45% 1.25%
3.0% 3.76% 3.02% 2.59% 2.17% 1.88%
4.0% 5.01% 4.03% 3.46% 2.89% 2.51%
6.0% 7.52% 6.04% 5.19% 4.34% 3.76%

is the primary window (2018-10 to 2026-06); is the full 2016–2026 span. The bolded cell is the study's realistic operating point.

A.4 Idiosyncratic tracking-error floor, Portfolio A vs Portfolio B

Equal-weighted top-quartile portfolios from a 1,000-name universe (), factor exposures assumed to net out, so this is a floor to which the systematic component of §3.8 must be added:

(58)
Overlap per side
0.38 41.9% 145 1.36% 2.04% 2.72%
0.45 45.5% 136 1.32% 1.98% 2.64%
0.54 50.3% 124 1.26% 1.89% 2.52%
0.60 53.8% 115 1.21% 1.82% 2.43%
0.71 60.9% 98 1.12% 1.68% 2.24%
0.80 67.6% 81 1.02% 1.53% 2.04%
0.90 77.2% 57 0.85% 1.28% 1.71%

Read against §A.3, this is the study's central power problem stated precisely: the tracking error the design generates (≈2%) is of the same order as the alpha it can detect (≈2%), which is why the equivalence test and the firm-level panel are structural requirements rather than robustness extras.

A.5 Annual cost drag (basis points)

.

One-way turnover bp 10 bp 20 bp 40 bp 60 bp
20% 2 4 8 16 24
40% 4 8 16 32 48
60% 6 12 24 48 72
80% 8 16 32 64 96
120% 12 24 48 96 144
160% 16 32 64 128 192

A.6 Turnover floor implied by rating persistence

Two-state Markov chain on top-quartile membership; is monthly retention.

(monthly) Monthly exit rate Expected tenure Annual one-way turnover floor
0.90 0.100 10.0 months 120.0%
0.93 0.070 14.3 months 84.0%
0.95 0.050 20.0 months 60.0%
0.97 0.030 33.3 months 36.0%
0.98 0.020 50.0 months 24.0%
0.99 0.010 100.0 months 12.0%

A 95% monthly retention rate reads as a stable rating. It implies a 60% annual turnover floor from rating migration alone: 24 bp p.a. at a 20 bp average cost, before any price drift, index reconstitution or optimiser churn.

A.7 Capacity: maximum participation rate before impact exhausts the alpha budget

Illustrative parameters: half-spread 3 bp, , daily volatility 2%, .

Alpha budget One-way turnover Max participation
25 bp 40% 7.98%
25 bp 80% 1.59%
50 bp 40% 35.40%
50 bp 80% 7.98%
100 bp 40% (unconstrained at these parameters)
100 bp 80% 35.40%

The quadratic relationship is the point: at a fixed alpha budget, doubling turnover from 40% to 80% reduces the sustainable participation rate by roughly a factor of four.

A.8 Inference parameters

Newey–West automatic bandwidth, :

60 93 126 150 200 252
3 3 4 4 4 4

GRS standard-error inflation, , from estimating factor means:

(ann.) 0.0 0.4 0.6 0.8 1.0 1.2 1.5
Inflation 1.000 1.077 1.166 1.281 1.414 1.562 1.803

A.9 Notation

Symbol Meaning
Firm, month, GICS sector, GICS sub-industry group
Rating provider; = MSCI, = Sustainalytics, = third provider
Latent, mandate-relevant ESG quality
Raw and sign-harmonised vendor score
Normalised score, universe-wide and industry-neutral
Characteristic vector driving algorithmic bias
Provider 's bias-loading vector (its algorithmic fingerprint)
Provider-specific rater noise
, , Divergence and its systematic / idiosyncratic components
Inter-rater reliability ratio,
Eligible set under provider at
Symmetric-difference sets defining the divergence portfolio
Weights, active weights, drifted pre-rebalance weights
Covariance; factor loadings; factor covariance; idiosyncratic variance
Shadow price of the ESG constraint, in Sharpe units
, , Differential alpha, tracking error, information ratio
Active sector tracking error
, One-way turnover; its rating-migration component
, , Per-share cost function, impact coefficient, impact exponent
, Screen quantile (0.25) and its normal threshold

Appendix B: Calibration script

The tables in Appendix A are generated by the following, which is the canonical reference for every number quoted in the protocol.

import numpy as np
from scipy.stats import norm, multivariate_normal as mvn

def top_q_joint(rho, q=0.25):
    """P(firm lands in the top-q tail of BOTH raters) under a Gaussian copula."""
    z = norm.ppf(1 - q)
    return 1 - 2*norm.cdf(z) + mvn(mean=[0, 0], cov=[[1, rho], [rho, 1]]).cdf([z, z])

# A.1 overlap, Jaccard, active share
for rho in [0.30, 0.38, 0.45, 0.54, 0.60, 0.71, 0.80, 0.90, 0.95]:
    p = top_q_joint(rho)
    overlap, jaccard = p/0.25, p/(0.50 - p)
    print(rho, p, overlap, jaccard, 1 - overlap)

# A.3 minimum detectable annualised alpha
k = norm.ppf(0.975) + norm.ppf(0.80)          # 2.802
mde = lambda sigma, years: k * sigma / np.sqrt(years)

# A.4 idiosyncratic tracking-error floor
N_sel = 250
te_idio = lambda rho, sig: np.sqrt(2*N_sel*(1 - top_q_joint(rho)/0.25))/N_sel * sig

# A.5 annual cost drag, one-way turnover convention
drag_bp = lambda turnover, c_bp: 2 * turnover * c_bp

# A.6 turnover floor from top-quartile retention
to_floor = lambda p11: 12 * (1 - p11)

# A.7 capacity: participation rate at which impact exhausts the alpha budget
def max_participation(alpha_budget, turnover, half_spread=0.0003, kappa=0.5, sigma_d=0.02):
    allowance = alpha_budget/(2*turnover) - half_spread
    return None if allowance <= 0 else (allowance/(kappa*sigma_d))**2
← Research