arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2512.11785v2 [math.PR] 20 Aug 2026

Universal entrywise eigenvector fluctuations in delocalized spiked matrix models and asymptotics of rounded spectral algorithms

Shujing Chen Thanks: Email: schen344@jhu.edu Affiliation: Department of Applied Mathematics & Statistics, Johns Hopkins University    Dmitriy Kunisky Thanks: Email: kunisky@jhu.edu. Affiliation: Department of Applied Mathematics & Statistics, Johns Hopkins University
August 20, 2026
Abstract

We consider the distribution of the top eigenvector v^\widehat{v} of a spiked matrix model of the form H=θvv+WH=\theta vv^{*}+W, in the supercritical regime where HH has an outlier eigenvalue of comparable magnitude to W\|W\|. We show that, if vv is sufficiently delocalized, then the distribution of the individual entries of the projector v^v^\widehat{v}\widehat{v}^{*} (not, we emphasize, merely the inner product |v^,v|2|\langle\widehat{v},v\rangle|^{2}) is universal over a large class of generalized Wigner matrices WW having independent entries, depending only on the first two moments of the distributions of the entries of WW. This complements the observation of Capitaine and Donati-Martin (2021) that these distributions are not universal when vv is instead sufficiently localized. Further, for WW having entrywise variances close to constant and thus resembling a Wigner matrix, we show by comparing to WW drawn from the Gaussian orthogonal or unitary ensembles that averages of entrywise functions of v^v^\widehat{v}\widehat{v}^{*} behave as they would if v^\widehat{v} had Gaussian fluctuations around a suitable multiple of vv. We also establish such results for several possibly dependent spiked matrices, showing that, if such matrices are entrywise uncorrelated, then their leading eigenvectors behave as they would with independent Gaussian fluctuations. We apply these results to spectral algorithms with rounding procedures for synchronization problems over the cyclic and circle groups, obtaining the first precise asymptotic error rates for such algorithms. Using our analysis of multiple spiked matrices, we also show that multi-frequency spectral algorithms using estimates from several matrices often have asymptotic error rate superior to that of naive spectral algorithms using just one matrix.

1 Introduction

We will study a universality phenomenon in spiked matrix models, of which we take the simple rank-one case

H=θvv+W,H=\theta vv^{*}+W,

where θ\theta\in\mathbb{R}, v𝕊n1()nv\in\mathbb{S}^{n-1}(\mathbb{C})\subset\mathbb{C}^{n} (the unit sphere of n\mathbb{C}^{n}), and Whermn×nW\in\mathbb{C}^{n\times n}_{\mathrm{herm}} is a Hermitian “noise” matrix normalized such that W\|W\| is of constant order. From a statistical point of view, HH can be seen as an “observation” of the hidden “signal” vv, from which one seeks to estimate vv. A natural choice of estimator is v^=v1(H)\widehat{v}=v_{1}(H), the unit eigenvector of HH corresponding to the largest eigenvalue of HH, which we denote λ^=λ1(H)\widehat{\lambda}=\lambda_{1}(H). Our general goal will be to understand in a fine-grained way the quality of estimate of vv that we obtain from v^\widehat{v}.

Let us explain some of what is known about such models under the following assumption on WW, which we will also use in the first part of our main results.

Definition 1.1 (Generalized Wigner matrix).

A generalized Wigner matrix with Whermn×nW\in\mathbb{C}^{n\times n}_{\mathrm{herm}} with parameters (γW,ξW)(\gamma_{W},\xi_{W}) is a random matrix having the following properties:

  1. 1.

    The entries on and above the diagonal (Wij)1ijn(W_{ij})_{1\leq i\leq j\leq n} are independent.

  2. 2.

    𝔼Wij=0\mathbb{E}W_{ij}=0 for all i,j[n]i,j\in[n].

  3. 3.

    The magnitudes |Wij||W_{ij}| admit a uniform tail bound of the form

    [n|Wij|t]ξW1exp(tξW)\mathbb{P}[\sqrt{n}\cdot|W_{ij}|\geq t]\leq\xi_{W}^{-1}\exp(-t^{\xi_{W}}) (1.1)

    We will call this condition the rescaled entries n|Wij|\sqrt{n}\cdot|W_{ij}| being ξW\xi_{W}-sub-Weibull (Definition 2.1).

  4. 4.

    The numbers σij2:=𝔼|Wij|2\sigma_{ij}^{2}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\mathbb{E}|W_{ij}|^{2} satisfy

    γW1nσij2γW for ij,i,j[n],\gamma_{W}^{-1}\leq n\cdot\sigma_{ij}^{2}\leq\gamma_{W}\text{ for }i\neq j,\ i,j\in[n], (1.2)

    as well as

    j=1nσij2=1 for all i[n].\sum_{j=1}^{n}\sigma_{ij}^{2}=1\text{ for all }i\in[n]. (1.3)

This class of matrices is a well-studied universality class, sharing many basic behaviors. For instance, as a useful point of reference for the results we discuss next, under these assumptions the empirical distribution of WW almost surely converges weakly to the semicircle distribution supported on [2,2][-2,2], and λ1(W)\lambda_{1}(W) and W\|W\| both converge in probability to 2 (see, e.g., [AGZ09] for an exposition of these classical results, or our Theorem 2.13).

We mostly take vv to be deterministic, and later in our results we will mention the case of vv random, in which case it is random independently of WW (in this situation results for vv deterministic can be applied by conditioning on the value of vv).

The following describes the main features of the phase transition that the behavior of the spectrum of HH undergoes as θ\theta varies, a version for this setting of the Baik–Ben Arous–Péché transition. We sketch how the proof of this version follows from the isotropic local law proved by [BEK+14] in Theorem 2.21 below; this technique has been used by several works such as [KY13b, KY14] in the past and we merely adapt the particular local law on which it is based.

Theorem 1.2.

Let W=W(n)hermn×nW=W^{(n)}\in\mathbb{C}^{n\times n}_{\mathrm{herm}} be a sequence of generalized Wigner matrices with parameters (γW,ξW)(\gamma_{W},\xi_{W}) not depending on nn and let v=v(n)𝕊n1()v=v^{(n)}\in\mathbb{S}^{n-1}(\mathbb{C}). Write H=H(n)=θvv+WH=H^{(n)}=\theta vv^{*}+W, λ^=λ^(n)=λ1(H(n))\widehat{\lambda}=\widehat{\lambda}^{(n)}=\lambda_{1}(H^{(n)}) for the largest eigenvalues of these matrices, and v^=v^(n)\widehat{v}=\widehat{v}^{(n)} for the associated eigenvectors of HH. Then, the following hold, with all convergences in probability as nn\to\infty:

λ^\displaystyle\widehat{\lambda} λ(θ):={2if θ1θ+θ1if θ>1},\displaystyle\to\lambda(\theta)\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left\{\begin{array}[]{ll}2&\text{if }\theta\leq 1\\ \theta+\theta^{-1}&\text{if }\theta>1\end{array}\right\},
|v^,v|2\displaystyle|\langle\widehat{v},v\rangle|^{2} ρ(θ)2:={0if θ11θ2if θ>1}.\displaystyle\to\rho(\theta)^{2}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left\{\begin{array}[]{ll}0&\text{if }\theta\leq 1\\ 1-\theta^{-2}&\text{if }\theta>1\end{array}\right\}.

It will also be useful for us to have a specific notation for

τ(θ)2:=1ρ(θ)2={1if θ1θ2if θ>1},\tau(\theta)^{2}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}1-\rho(\theta)^{2}=\left\{\begin{array}[]{ll}1&\text{if }\theta\leq 1\\ \theta^{-2}&\text{if }\theta>1\end{array}\right\}, (1.8)

the typical value of 1|v^,v|21-|\langle\widehat{v},v\rangle|^{2}, the magnitude of the component of v^\widehat{v} orthogonal to vv.

This gives a precise understanding of the top eigenvalue λ^\widehat{\lambda} and eigenvector v^\widehat{v} of HH to leading order, though for the eigenvector only in terms of its correlation with its noiseless counterpart vv. For both quantities, it is reasonable to ask about the next-order fluctuations: how do a rescaling of λ^λ(θ)\widehat{\lambda}-\lambda(\theta) and v^ρ(θ)v\widehat{v}-\rho(\theta)v behave? (The latter is not quite well-defined for the reason described in Remark 1.5 below, but let us ignore this issue for the moment.)

While it would be natural to conjecture that both of these quantities should also behave universally over a generalized Wigner matrices or some other such large class, in fact this is not the case. In our setting, it is known that non-universality of the fluctuations of both eigenvalues and eigenvectors—that is, a sensitivity to the specific distribution of entries of WW—arises when vv is localized, having some large entries. See, for instance, [CDMF09, CDMF12, RS13, BGGM11, KY13b, KY14] for such results on eigenvalues, [CDM21, BDW21], and [MY22, FFHL22] for similar results about eigenvectors but in different settings (where the signal is of macroscopic rank and the noise is Gaussian in the first case, and where θ=θ(n)\theta=\theta(n)\to\infty in the second case). In the former works, it is also shown that the fluctuations of eigenvalues are universal when vv is sufficiently delocalized.

Figure 1: Universality and non-universality in spiked matrix models. We illustrate a simple example of the contrast between universality for delocalized signals and non-universality for localized signals. Let P=v^v^P=\widehat{v}\widehat{v}^{*} be the projection onto the leading eigenvector of a spiked matrix model of dimension n=2000n=$2000$ with supercritical signal strength θ=1.5\theta=$1.5$. We plot the fluctuations of the diagonal entry P11=|v^1|2P_{11}=|\widehat{v}_{1}|^{2} under its natural scaling in each signal regime. We compare noise matrices whose upper-triangular entries are drawn independently from Gaussian and Rademacher distributions. In the left panel, the signal v=e1v=e_{1} is localized and P11=|v^,v|2P_{11}=|\langle\widehat{v},v\rangle|^{2}; the visibly different laws of n(P11ρ(θ)2)\sqrt{n}(P_{11}-\rho(\theta)^{2}) illustrate the non-universality phenomenon proved in [CDM21]. In the right panel, the signal v=𝟏/nv=\mathbf{1}/\sqrt{n} is delocalized, and the agreement of the laws of n(P11ρ(θ)2/n)n(P_{11}-\rho(\theta)^{2}/n) illustrates the universality of entry statistics proved in Theorem 1.4.

1.1 Summary of contributions

To the best of our knowledge, until now a gap has remained in the above line of work concerning a situation that is particularly important in applications: it is not known that the eigenvector fluctuations—those of v^\widehat{v}are universal provided that vv is delocalized, in any stronger sense beyond just the behavior of the specific inner product v^,v\langle\widehat{v},v\rangle. We illustrate a basic version of this phenomenon in Figure 1.

This is the task that we take up in this paper: we will show such a universality result in a rather strong entrywise sense, in the style of the results for eigenvectors of purely random matrices with no low-rank perturbation of [KY13a]. We treat both the behavior of the individual entries nv^iv^j¯n\cdot\widehat{v}_{i}\overline{\widehat{v}_{j}} and averages of the form

1n2i,j=1nψ(nvivj¯,nv^iv^j¯).\frac{1}{n^{2}}\sum_{i,j=1}^{n}\psi(n\cdot v_{i}\overline{v_{j}},n\cdot\widehat{v}_{i}\overline{\widehat{v}_{j}}). (1.9)

In both cases, we scale by nn since the typical scale of each of these numbers is Θ(1/n)\Theta(1/n). If we view ψ\psi as a measurement of proximity of two complex numbers, the latter gives us access to various ways of viewing the amount of error made by a spectral algorithm by measuring various notions of entrywise distance between vvvv^{*} and v^v^\widehat{v}\widehat{v}^{*}.

In particular, when we have prior information that vv belongs to some special subset of n\mathbb{C}^{n}, we may include rounding functions in the definition of ψ\psi, that encode a procedure estimating vvvv^{*} by first computing v^v^\widehat{v}\widehat{v}^{*} and then “projecting” its value in some suitable entrywise sense to the set in which the entries vivj¯v_{i}\overline{v_{j}} of vvvv^{*} can possibly lie. As a simple example, if we know in advance that v{±1/n}nv\in\{\pm 1/\sqrt{n}\}^{n} and WW is real-valued, then it is sensible to take the entrywise sign sgn(v^iv^j¯)/n\mathrm{sgn}(\widehat{v}_{i}\overline{\widehat{v}_{j}})/n as an estimate of vivj¯v_{i}\overline{v_{j}}. We can then measure the natural notion of error of what fraction of these rounded signs disagree with those of vvvv^{*}, i.e., the quantity:

1n2i,j=1n𝟏{sgn(v^iv^j¯)sgn(vivj¯)},\frac{1}{n^{2}}\sum_{i,j=1}^{n}\mathbf{1}\{\mathrm{sgn}(\widehat{v}_{i}\overline{\widehat{v}_{j}})\neq\mathrm{sgn}(v_{i}\overline{v_{j}})\},

arguably a more natural and meaningful notion of error for this discrete setting than continuous quantities like v^v^vvF2\|\widehat{v}\widehat{v}^{*}-vv^{*}\|_{\mathrm{F}}^{2}. To the best of our knowledge, ours are the first precise asymptotics for the amount of entrywise error made by such algorithms under general non-Gaussian models of the noise WW.

We further extend these results to the case of several spiked matrices H(a)=θv(a)v(a)+W(a)H^{(a)}=\theta v^{(a)}v^{(a)^{*}}+W^{(a)} for a[L]a\in[L]. In this case, the main structural assumption on the noise matrices W(a)W^{(a)} is that the tuples (Wij(1),,Wij(L))(W^{(1)}_{ij},\dots,W^{(L)}_{ij}) are independent over different pairs (i,j)(i,j) (up to Hermitian symmetry); however, we note that within these tuples the same entry of different W(a)W^{(a)} can be dependent or even deterministic functions of one another. This has an important consequence for the special case where these tuples are uncorrelated, because in this situation, one can compare to the case of W(a)W^{(a)} having independent Gaussian entries. Thus, the case of H(a)H^{(a)} being entrywise uncorrelated in fact belongs to the same universality class as a case of H(a)H^{(a)} being independent (as for Gaussian random variables being uncorrelated implies being independent).

We call this phenomenon pseudoindependence of spiked matrix models, and we propose that it has interesting ramifications for the design of spectral algorithms. In particular, if the v(a)v^{(a)} are deterministic functions of one another, then clearly in the case of W(a)W^{(a)} independent one expects to gain more information about v(a)v^{(a)} from more independent matrix observations, and by our analysis the same holds under pseudoindependence. This suggests that multi-spectral algorithms, ones using not just a given spiked matrix HH but several deterministic entrywise transformations thereof, may be superior to direct spectral algorithms in some settings. We illustrate this proposal by showing that multi-spectral algorithms are indeed superior for the problem of synchronization over the cyclic and circle groups.

1.2 Theoretical results

We always use the parameter LL to designate the number of matrices we work with. For the sake of brevity we state our results for the case of general L1L\geq 1, but the reader may find it instructive to read the results with the simplifying choice L=1L=1, to recover the case of a single spiked matrix discussed above.

1.2.1 Universality

Our first main result says that, provided that the v(a)v^{(a)} are sufficiently delocalized and that the W(a)W^{(a)} are sufficiently unstructured, statistics of the noisy top eigenvectors v^(a)\widehat{v}^{(a)} do not depend on the entry distribution of the W(a)W^{(a)}. The specific delocalization assumption we make on the v(a)v^{(a)} is the following:

Assumption 1.3.

For a[L]a\in[L], we assume that v(a)nv^{(a)}\in\mathbb{C}^{n} has v(a)=1\|v^{(a)}\|=1, and that

maxa[L]v(a)Cvn1/2+εv\max_{a\in[L]}\|v^{(a)}\|_{\infty}\leq C_{v}n^{-1/2+\varepsilon_{v}}

for some εv(0,1/20)\varepsilon_{v}\in(0,1/20) and Cv>0C_{v}>0. We call (εv,Cv)(\varepsilon_{v},C_{v}) the parameters of this assumption on v(a)v^{(a)}.

Our first main result is then as follows.

Theorem 1.4 (Universality of entry statistics).

Let θ>1\theta>1 and v(1),,v(L)𝕊n1()v^{(1)},\dots,v^{(L)}\in\mathbb{S}^{n-1}(\mathbb{C}) satisfy Assumption 1.3. Let W(1),,W(L),X(1),,X(L)W^{(1)},\dots,W^{(L)},X^{(1)},\dots,X^{(L)} be generalized Wigner matrices (Definition 1.1) with uniform parameters (γW,ξW)(\gamma_{W},\xi_{W}) and (γX,ξX)(\gamma_{X},\xi_{X}) such that the tuples (Wij(1),,Wij(L))(W^{(1)}_{ij},\dots,W^{(L)}_{ij}) are independent over 1ijn1\leq i\leq j\leq n, and likewise for the (Xij(1),,Xij(L))(X^{(1)}_{ij},\dots,X^{(L)}_{ij}). Letting w(ij)2Lw^{(ij)}\in\mathbb{R}^{2L} contain the real and imaginary parts of Wij(a)W^{(a)}_{ij} for a[L]a\in[L] and likewise x(ij)2Lx^{(ij)}\in\mathbb{R}^{2L} for Xij(a)X^{(a)}_{ij}, suppose that the second moments of these vectors match (recalling that their first moments are fixed to be zero and so automatically match by the assumption of the W(a)W^{(a)} and X(a)X^{(a)} being generalized Wigner matrices): for all i,j[n]i,j\in[n], we have

𝔼w(ij)w(ij)=𝔼x(ij)x(ij).\mathbb{E}w^{(ij)}w^{(ij)^{\top}}=\mathbb{E}x^{(ij)}x^{(ij)^{\top}}.

Let v^(W,a)\widehat{v}^{(W,a)} be the eigenvector associated to the largest eigenvalue of θv(a)v(a)+W(a)\theta v^{(a)}v^{(a)^{*}}+W^{(a)} and v^(X,a)\widehat{v}^{(X,a)} that associated to the largest eigenvalue of θv(a)v(a)+X(a)\theta v^{(a)}v^{(a)^{*}}+X^{(a)}. Let ϕ:L\phi:\mathbb{C}^{L}\to\mathbb{R} be a 𝒞5\mathcal{C}^{5} function such that ϕ\phi and its first five derivatives are bounded. Then, for any ε>0\varepsilon>0,

max1i,jn|𝔼ϕ(nv^i(W,1)v^j(W,1)¯,,nv^i(W,L)v^j(W,L)¯)\displaystyle\max_{1\leq i,j\leq n}\bigg|\mathbb{E}\phi\left(n\cdot\widehat{v}^{(W,1)}_{i}\overline{\widehat{v}^{(W,1)}_{j}},\dots,n\cdot\widehat{v}^{(W,L)}_{i}\overline{\widehat{v}^{(W,L)}_{j}}\right)
𝔼ϕ(nv^i(X,1)v^j(X,1)¯,,nv^i(X,L)v^j(X,L)¯)|Cn1/2+10εv+ε,\displaystyle\hskip 56.9055pt-\mathbb{E}\phi\left(n\cdot\widehat{v}^{(X,1)}_{i}\overline{\widehat{v}^{(X,1)}_{j}},\dots,n\cdot\widehat{v}^{(X,L)}_{i}\overline{\widehat{v}^{(X,L)}_{j}}\right)\bigg|\leq Cn^{-1/2+10\varepsilon_{v}+\varepsilon},

for a constant CC depending only on the parameters L,ε,θ,ϕL,\varepsilon,\theta,\phi in the statement of this result, the parameters (εv,Cv)(\varepsilon_{v},C_{v}) of Assumption 1.3 on vv, and the generalized Wigner matrix parameters (γW,ξW)(\gamma_{W},\xi_{W}) and (γX,ξX)(\gamma_{X},\xi_{X}).

Remark 1.5.

The reason for taking products of pairs of entries of the various v^\widehat{v} above is that v^\widehat{v} itself is only defined up to a global phase (or a global sign flip in the real-valued case), while v^v^\widehat{v}\widehat{v}^{*}, the orthogonal projection to the v^\widehat{v} direction, is a well-defined geometric object; above we formulate our result in terms of its entries.

1.2.2 Gaussian approximation and pseudoindependence

The above is a conceptually interesting result, but one that does not allow concrete calculations of the actual values of entrywise quantities like the above. In the special case where the Wij(a)W_{ij}^{(a)} are uncorrelated for a fixed (i,j)(i,j) over different aa, however, by comparing to special Gaussian distributions of noise matrices we can obtain tractable approximations. We consider which WW can be compared by Theorem 1.4 to the following classical distributions of random matrices:

Definition 1.6 (Gaussian orthogonal and unitary ensembles).

Let n1n\geq 1. We then define the following distributions of random matrices in hermn×n\mathbb{C}^{n\times n}_{\mathrm{herm}}.

  • The Gaussian orthogonal ensemble, denoted GOE(n)\mathrm{GOE}(n) for given dimension nn, is the law of Wsymn×nW\in\mathbb{R}^{n\times n}_{\mathrm{sym}} with WijW_{ij} independent for 1ijn1\leq i\leq j\leq n drawn as

    Wij\displaystyle W_{ij} 𝒩(0,1n) for 1i<jn,\displaystyle\sim\mathcal{N}\left(0,\frac{1}{n}\right)\text{ for }1\leq i<j\leq n,
    Wii\displaystyle W_{ii} 𝒩(0,2n) for 1in.\displaystyle\sim\mathcal{N}\left(0,\frac{2}{n}\right)\text{ for }1\leq i\leq n.

    Equivalently, {nWij:1i<jn}{n/2Wii:1in}\{\sqrt{n}\cdot W_{ij}:1\leq i<j\leq n\}\cup\{\sqrt{n/2}\cdot W_{ii}:1\leq i\leq n\} are i.i.d. with law 𝒩(0,1)\mathcal{N}(0,1).

  • The Gaussian unitary ensemble, denoted GUE(n)\mathrm{GUE}(n) for given dimension nn, is the law of Whermn×nW\in\mathbb{C}^{n\times n}_{\mathrm{herm}} with WijW_{ij} independent for 1ijn1\leq i\leq j\leq n drawn as

    Wij\displaystyle W_{ij} 𝒩(0,1n) for 1i<jn,\displaystyle\sim\mathcal{N}_{\mathbb{C}}\left(0,\frac{1}{n}\right)\text{ for }1\leq i<j\leq n,
    Wii\displaystyle W_{ii} 𝒩(0,1n) for 1in.\displaystyle\sim\mathcal{N}_{\mathbb{R}}\left(0,\frac{1}{n}\right)\text{ for }1\leq i\leq n.

    Here 𝒩(0,σ2)=𝒩(0,σ2)\mathcal{N}_{\mathbb{R}}(0,\sigma^{2})=\mathcal{N}(0,\sigma^{2}) is the ordinary Gaussian measure, while 𝒩(0,σ2)\mathcal{N}_{\mathbb{C}}(0,\sigma^{2}) is the complex Gaussian measure, the law of Z=X+𝒊YZ=X+\bm{i}Y with X,Y𝒩(0,σ2/2)X,Y\sim\mathcal{N}_{\mathbb{R}}(0,\sigma^{2}/2) independently, so that 𝔼|Z|2=σ2\mathbb{E}|Z|^{2}=\sigma^{2}. Equivalently, a GUE matrix WW has {2nReWij:1i<jn}{2nImWij:1i<jn}{nWii:1in}\{\sqrt{2n}\cdot\mathrm{Re}W_{ij}:1\leq i<j\leq n\}\cup\{\sqrt{2n}\cdot\mathrm{Im}W_{ij}:1\leq i<j\leq n\}\cup\{\sqrt{n}\cdot W_{ii}:1\leq i\leq n\} are i.i.d. with law 𝒩(0,1)\mathcal{N}(0,1).

Here and throughout we denote the imaginary unit by 𝒊\bm{i} to avoid conflict with indices named ii. It is easy to check that GOE and GUE matrices are both generalized Wigner matrices, up to a negligible renormalization in the GOE case.

When our matrices are drawn from these distributions, the behavior of v^\widehat{v} can be analyzed quite precisely. That is because XGOE(n)X\sim\mathrm{GOE}(n) and XGUE(n)X\sim\mathrm{GUE}(n) are respectively orthogonally and unitarily invariant, satisfying the property that QXQQXQ^{*} has the same law as XX for any orthogonal or any unitary QQ, respectively. Because of this, the component of v^\widehat{v} that is orthogonal to vv is itself a likewise invariant random vector in this subspace, whose norm by Theorem 1.2 is close to τ(θ)=θ1\tau(\theta)=\theta^{-1}. Thus, setting aside the sign ambiguity in v^\widehat{v} mentioned in Remark 1.5, approximating this component by a Gaussian random vector (which is also orthogonally invariant) of comparable norm, we expect

v^(X)(d)𝒩𝔽(ρ(θ)v,τ(θ)2n(Invv)+σ(θ)2nvv),\widehat{v}(X)\stackrel{{\scriptstyle\text{(d)}}}{{\approx}}\mathcal{N}_{\mathbb{F}}\left(\rho(\theta)v,\frac{\tau(\theta)^{2}}{n}(I_{n}-vv^{*})+\frac{\sigma(\theta)^{2}}{n}vv^{*}\right), (1.10)

for some σ(θ)=Θ(1)\sigma(\theta)=\Theta(1) (the correct value of this constant may also be determined explicitly, but in our case of vv delocalized this second term will not play a role in the relevant asymptotics). Here, 𝔽{,}\mathbb{F}\in\{\mathbb{R},\mathbb{C}\} is the field we are working over depending on whether XX is a real or complex Wigner matrix. Further, since the mean is a constant multiple of vv, we should be able to neglect the contribution of 1n(τ(θ)2+σ(θ)2)vv\frac{1}{n}(\tau(\theta)^{2}+\sigma(\theta)^{2})vv^{*} in the covariance. Similar ideas have appeared before in the literature in [CH13, LCC24], but we give a full derivation of a precise version of such an approximation specific to our setting in Section 4.1.

Through Theorem 1.4, together with such analysis, we expect to obtain more concrete predictions for generalized Wigner matrices whose first two moments match those of either a GOE or GUE matrix. That leads to the following related definition of a more restrictive class of random matrices, in which we also allow some slack in the assumption of exact moment matching that we will see we may deal with in the proof of our next result, and which makes these results more useful for applications.

Definition 1.7 (Weakly Wigner tuples).

Let L1L\geq 1. For 1pqn1\leq p\leq q\leq n, define

w(pq):=(ReWpq(1),ImWpq(1),,ReWpq(L),ImWpq(L))2L.w^{(pq)}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left(\mathrm{Re}W^{(1)}_{pq},\mathrm{Im}W^{(1)}_{pq},\ldots,\mathrm{Re}W^{(L)}_{pq},\mathrm{Im}W^{(L)}_{pq}\right)\in\mathbb{R}^{2L}.

A weakly Wigner tuple 𝐖=(W(1),W(L))\mathbf{W}=(W^{(1)},\ldots W^{(L)}) with parameters (ξW,εW,CW)(\xi_{W},\varepsilon_{W},C_{W}), where each W(a)𝔽an×nW^{(a)}\in\mathbb{F}_{a}^{n\times n} with 𝔽a{,}\mathbb{F}_{a}\in\{\mathbb{R},\mathbb{C}\} is Hermitian, is a tuple of matrices satisfying the following properties:

  1. 1.

    The vectors w(pq)w^{(pq)} over 1pqn1\leq p\leq q\leq n are independent.

  2. 2.

    𝔼Wpq(a)=0\mathbb{E}W^{(a)}_{pq}=0 and thus 𝔼w(pq)=𝟎\mathbb{E}w^{(pq)}=\bm{0} for all p,q[n]p,q\in[n], a[L]a\in[L].

  3. 3.

    The n|Wpq(a)|\sqrt{n}\cdot|W^{(a)}_{pq}| are ξW\xi_{W}-sub-Weibull for all p,q[n]p,q\in[n], a[L]a\in[L].

Define the covariance matrices

Σ(pq)𝐖:=𝔼w(pq)w(pq)2L×2Lsym\Sigma^{(pq)}_{\mathbf{W}}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\mathbb{E}w^{(pq)}w^{(pq)\top}\in\mathbb{R}^{2L\times 2L}_{\mathrm{sym}}

which consist of the blocks

Σ(pq)𝐖,ab:=(𝔼ReWpq(a)ReWpq(b)𝔼ReWpq(a)ImWpq(b)𝔼ImWpq(a)ReWpq(b)𝔼ImWpq(a)ImWpq(b)).\Sigma^{(pq)}_{\mathbf{W},ab}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\begin{pmatrix}\mathbb{E}\mathrm{Re}W^{(a)}_{pq}\mathrm{Re}W^{(b)}_{pq}&\mathbb{E}\mathrm{Re}W^{(a)}_{pq}\mathrm{Im}W^{(b)}_{pq}\\ \mathbb{E}\mathrm{Im}W^{(a)}_{pq}\mathrm{Re}W^{(b)}_{pq}&\mathbb{E}\mathrm{Im}W^{(a)}_{pq}\mathrm{Im}W^{(b)}_{pq}\end{pmatrix}.

We further require the following properties from these covariances:

  1. 4.

    Σ𝐖(pp)CW/n\|\Sigma_{\mathbf{W}}^{(pp)}\|\leq C_{W}/n for all 1pn1\leq p\leq n.

  2. 5.

    Let Σ=(1000)\Sigma_{\mathbb{R}}=\begin{pmatrix}1&0\\ 0&0\end{pmatrix}, Σ:=(1/2001/2)\Sigma_{\mathbb{C}}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\begin{pmatrix}1/2&0\\ 0&1/2\end{pmatrix}, and ΣG𝔽E:=1nDiag(Σ𝔽1,,Σ𝔽L)\Sigma_{\mathrm{G}\mathbb{F}\mathrm{E}}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\frac{1}{n}\mathrm{Diag}(\Sigma_{\mathbb{F}_{1}},\ldots,\Sigma_{\mathbb{F}_{L}}). Then, for each 𝔽a{,}\mathbb{F}_{a}\in\{\mathbb{R},\mathbb{C}\}, a[L]a\in[L],

    Σ𝐖(pq)ΣG𝔽ECWn1+εW, for all 1p<qn.\|\Sigma^{(pq)}_{\mathbf{W}}-\Sigma_{\mathrm{G}\mathbb{F}\mathrm{E}}\|\leq\frac{C_{W}}{n^{1+\varepsilon_{W}}},\text{ for all }1\leq p<q\leq n.

We will repeatedly use the variable 𝔽a{,}\mathbb{F}_{a}\in\{\mathbb{R},\mathbb{C}\} specifying whether the matrices we are working with are real or complex weakly Wigner matrices.

We note that it is actually not quite true that all matrices in weakly Wigner tuples are generalized Wigner matrices, because the exact condition (1.3) need not hold. However, the weakly Wigner tuples assumptions imply that it holds up to an O(nεW)O(n^{-\varepsilon_{W}}) error, which we will see in our arguments suffices for our purposes.

Our second main result formalizes the above sketch of an argument combined with Theorem 1.4. We obtain a precise prediction for averages of entrywise functions of v^v^\widehat{v}\widehat{v}^{*} for v^=v^(W)\widehat{v}=\widehat{v}(W) defined with weakly 𝔽\mathbb{F}-Wigner matrices WW for 𝔽{,}\mathbb{F}\in\{\mathbb{R},\mathbb{C}\} matching what we predicted for GOE and GUE matrices, respectively, based on their special invariance properties. As will be useful in applications, we may also allow these functions to depend on the corresponding entries of vvvv^{*}, allowing us to analyze various notions of the quality of entrywise approximation that v^v^\widehat{v}\widehat{v}^{*} gives to vvvv^{*}, as we suggested above in (1.9).

Theorem 1.8.

Let L1L\geq 1 and θ>1\theta>1, suppose that 𝐖\mathbf{W} is a weakly Wigner tuple (Definition 1.7) with parameters (ξW,εW,CW)(\xi_{W},\varepsilon_{W},C_{W}). Let v(a)𝕊n1(𝔽a)v^{(a)}\in\mathbb{S}^{n-1}(\mathbb{F}_{a}) satisfy Assumption 1.3 uniformly for all a[L]a\in[L]. For each a[L]a\in[L], let v^(W,a)\widehat{v}^{(W,a)} be the eigenvector associated to the largest eigenvalue of θv(a)v(a)+W(a)\theta v^{(a)}v^{(a)*}+W^{(a)}. Let ψ:L×L\psi:\mathbb{C}^{L}\times\mathbb{C}^{L}\to\mathbb{R} be a 𝒞5\mathcal{C}^{5} function such that the values of ψ\psi and its first five derivatives are bounded. Define the associated function Ψ:(hermn×n)L×(hermn×n)L\Psi:(\mathbb{C}^{n\times n}_{\mathrm{herm}})^{L}\times(\mathbb{C}^{n\times n}_{\mathrm{herm}})^{L}\to\mathbb{R} by

Ψ(𝐀,𝐁):=1n2i,j=1nψ((Aij(a))a=1L,(Bij(a))a=1L).\Psi(\mathbf{A},\mathbf{B})\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\frac{1}{n^{2}}\sum_{i,j=1}^{n}\psi\left(\left(A^{(a)}_{ij}\right)_{a=1}^{L},\left(B^{(a)}_{ij}\right)_{a=1}^{L}\right).

Let g(a)𝒩𝔽a(0,In)g^{(a)}\sim\mathcal{N}_{\mathbb{F}_{a}}(0,I_{n}), a[L]a\in[L], be independent standard Gaussian random vectors over the corresponding fields, i.e., having entries gi(a)𝒩𝔽a(0,1)g^{(a)}_{i}\sim\mathcal{N}_{\mathbb{F}_{a}}(0,1) drawn i.i.d. Then, for any ε>0\varepsilon>0,

|𝔼Ψ((nv(a)v(a))a=1L,(nv^(W,a)v^(W,a))a=1L)\displaystyle\Bigg|\mathbb{E}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\widehat{v}^{(W,a)}\widehat{v}^{(W,a)*}\right)_{a=1}^{L}\right)
𝔼Ψ((nv(a)v(a))a=1L,((ρ(θ)nv(a)+τ(θ)g(a))(ρ(θ)nv(a)+τ(θ)g(a)))a=1L)|\displaystyle\hskip 56.9055pt-\mathbb{E}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left((\rho(\theta)\sqrt{n}v^{(a)}+\tau(\theta)g^{(a)})(\rho(\theta)\sqrt{n}v^{(a)}+\tau(\theta)g^{(a)})^{*}\right)_{a=1}^{L}\right)\Bigg|
C(n1/2+10εv+ε+nεW/2+ε),\displaystyle\hskip 113.81102pt\leq C\left(n^{-1/2+10\varepsilon_{v}+\varepsilon}+n^{-\varepsilon_{W}/2+\varepsilon}\right),

where CC is a constant depending only on L,(𝔽a)a=1L,ε,θ,ψL,(\mathbb{F}_{a})_{a=1}^{L},\varepsilon,\theta,\psi in the statement of this result, the parameters (εv,Cv)(\varepsilon_{v},C_{v}) of Assumption 1.3 on v(a)v^{(a)}, and the weakly Wigner tuple parameters (ξW,εW,CW)(\xi_{W},\varepsilon_{W},C_{W}).

Suppose further that v(1),,v(L)v^{(1)},\ldots,v^{(L)} are also drawn at random and independently of 𝐖\mathbf{W} such that n(vi(1),,vi(L))\sqrt{n}(v^{(1)}_{i},\ldots,v^{(L)}_{i}) for i[n]i\in[n] are i.i.d. according to some compactly supported probability measure μ\mu on 𝕊0(𝔽1)××𝕊0(𝔽L)\mathbb{S}^{0}(\mathbb{F}_{1})\times\dots\times\mathbb{S}^{0}(\mathbb{F}_{L}).11 1 Concretely, these entries are allowed to lie in 𝕊0()={1,+1}\mathbb{S}^{0}(\mathbb{R})=\{-1,+1\} or the complex unit circle 𝕊0()=U(1)\mathbb{S}^{0}(\mathbb{C})=U(1). Let (u1,,uL),(v1,,vL)μ(u_{1},\dots,u_{L}),(v_{1},\dots,v_{L})\sim\mu and ga,ha𝒩𝔽a(0,1)g_{a},h_{a}\sim\mathcal{N}_{\mathbb{F}_{a}}(0,1) for a[L]a\in[L] all be independent. Then, for any ε>0\varepsilon>0,

|𝔼Ψ((nv(a)v(a))a=1L,(nv^(W,a)v^(W,a))a=1L)\displaystyle\Bigg|\mathbb{E}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\widehat{v}^{(W,a)}\widehat{v}^{(W,a)*}\right)_{a=1}^{L}\right)
𝔼ψ((uava¯)a=1L,((ρ(θ)ua+τ(θ)ga)(ρ(θ)va+τ(θ)ha)¯)a=1L)|\displaystyle\hskip 56.9055pt-\mathbb{E}\psi\left(\left(u_{a}\overline{v_{a}}\right)_{a=1}^{L},\left((\rho(\theta)u_{a}+\tau(\theta)g_{a})\overline{(\rho(\theta)v_{a}+\tau(\theta)h_{a})}\right)_{a=1}^{L}\right)\Bigg|
C(n1/2+ε+nεW/2+ε),\displaystyle\hskip 113.81102pt\leq C\left(n^{-1/2+\varepsilon}+n^{-\varepsilon_{W}/2+\varepsilon}\right), (1.11)

where CC is a constant depending only on L,(𝔽a)a=1L,ε,θ,ψL,(\mathbb{F}_{a})_{a=1}^{L},\varepsilon,\theta,\psi in the statement of this result, the entrywise distribution of the signal μ\mu, and the weakly Wigner tuple parameters (ξW,εW,CW)(\xi_{W},\varepsilon_{W},C_{W}).

Remark 1.9.

We emphasize the important coordinatewise assumption that, for each a[L]a\in[L], v(a)𝕊n1(𝔽a)v^{(a)}\in\mathbb{S}^{n-1}(\mathbb{F}_{a}), where 𝔽a=\mathbb{F}_{a}=\mathbb{R} if W(a)W^{(a)} is real symmetric and 𝔽a=\mathbb{F}_{a}=\mathbb{C} if W(a)W^{(a)} is complex Hermitian; in particular, we exclude the case of v(a)nnv^{(a)}\in\mathbb{C}^{n}\setminus\mathbb{R}^{n} while W(a)W^{(a)} is real symmetric.

We think of ψ\psi as an entrywise loss function for the estimation of (v(a)v(a))a=1L(v^{(a)}v^{(a)*})_{a=1}^{L} by (v^(a)v^(a))a=1L(\widehat{v}^{(a)}\widehat{v}^{(a)*})_{a=1}^{L}, measuring some notion of distance between vectors in L\mathbb{C}^{L}. Ψ\Psi then averages this loss over the entries of a pair of LL-tuples of matrices. The final bound (1.11) then describes a “LL-letter formula” for the expectation of any such loss: though it is a complicated expectation, it only involves O(L)O(L) many random variables while giving precise estimates for large nn. In particular, if we have an asymptotic sequence of v(a)=v(n,a)𝕊n1(𝔽a)v^{(a)}=v^{(n,a)}\in\mathbb{S}^{n-1}(\mathbb{F}_{a}), a[L]a\in[L], drawn as above for some fixed joint entrywise distribution μ\mu on 𝔽1××𝔽L\mathbb{F}_{1}\times\dots\times\mathbb{F}_{L} not depending on nn, a sequence of weakly Wigner tuples 𝐖=𝐖(n)=(W(n,1),,W(n,L))\mathbf{W}=\mathbf{W}^{(n)}=(W^{(n,1)},\dots,W^{(n,L)}) satisfying the definition with any parameters (ξW,εW,CW)(\xi_{W},\varepsilon_{W},C_{W}) not depending on nn, and v^(a)=v^(n,a)\widehat{v}^{(a)}=\widehat{v}^{(n,a)} the top eigenvectors of θv(n,a)v(n,a)+W(n,a)\theta v^{(n,a)}v^{(n,a)^{*}}+W^{(n,a)}, then the above implies the exact asymptotic result

limn𝔼Ψ((nv(a)v(a))a=1L,(nv^(a)v^(a))a=1L)\displaystyle\lim_{n\to\infty}\mathbb{E}\Psi\left(\left(n\cdot v^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\cdot\widehat{v}^{(a)}\widehat{v}^{(a)*}\right)_{a=1}^{L}\right) (1.12)
=𝔼ψ((uava¯)a=1L,((ρ(θ)ua+τ(θ)ga)(ρ(θ)va+τ(θ)ha)¯)a=1L).\displaystyle=\mathbb{E}\psi\left(\left(u_{a}\overline{v_{a}}\right)_{a=1}^{L},\left(\left(\rho(\theta)u_{a}+\tau(\theta)g_{a}\right)\overline{\left(\rho(\theta)v_{a}+\tau(\theta)h_{a}\right)}\right)_{a=1}^{L}\right).

Thus, we characterize the asymptotic expectation of any entrywise measurement of the loss achieved by such a spectral algorithm as nn\to\infty by an expectation of fixed dimension on the right-hand side, over just the LL-tuples (va)a=1L(v_{a})_{a=1}^{L}, (ua)a=1L(u_{a})_{a=1}^{L}, (ga)a=1L(g_{a})_{a=1}^{L}, and (ha)a=1L(h_{a})_{a=1}^{L}, which is complicated to evaluate in closed form for most ψ\psi but can easily be estimated numerically by straightforward Monte Carlo integration methods.

Remark 1.10 (Pseudoindependence).

We emphasize again that dependence across aa is allowed within both tuples: the signals v(1),,v(L)v^{(1)},\ldots,v^{(L)} may be dependent or deterministic functions of one another, and, for each (i,j)(i,j), the noise entries Wij(1),,Wij(L)W^{(1)}_{ij},\ldots,W^{(L)}_{ij} may likewise be dependent or deterministic functions of one another. In this case, the result may be viewed as describing the pseudoindependent asymptotic behavior of such matrices even though in reality they are highly dependent.

1.3 Applications to spectral and multi-spectral algorithms

We give applications to the following setting where spectral algorithms have been used with entrywise rounding in the literature [Sin11, CSC12, RG20, GZ19, CT22]. Let GG be a compact Lie group, and write Haar(G)\mathrm{Haar}(G) for its Haar measure. We suppose that we draw x1,,xni.i.d.Haar(G)x_{1},\dots,x_{n}\stackrel{{\scriptstyle\mathrm{i.i.d.}}}{{\sim}}\mathrm{Haar}(G), from which we can form the matrix of pairwise differences:

Mij:=xixj1,M_{ij}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}x_{i}x_{j}^{-1},

a matrix MGn×nM\in G^{n\times n} that is GG-Hermitian in the sense that Mji=Mij1M_{ji}=M_{ij}^{-1}. We observe a noisy version of this matrix,

Yij{Mijwith probability p,Haar(G)with probability 1p}Y_{ij}\sim\left\{\begin{array}[]{ll}M_{ij}&\text{with probability }p,\\ \mathrm{Haar}(G)&\text{with probability }1-p\end{array}\right\} (1.13)

for each 1i<jn1\leq i<j\leq n, and set Yii=eY_{ii}=e the group identity22 2 The diagonal entries will not contribute to the statistics we consider, so this choice is inconsequential. and Yji=Yij1Y_{ji}=Y_{ij}^{-1} to preserve the GG-Hermitian property. Here p=p(n)(0,1)p=p(n)\in(0,1) is some parameter governing the amount of information available in the observation YY. Our goal is to produce an estimator M^=M^(Y)\widehat{M}=\widehat{M}(Y) of MM. (Note that, for the same reason as in Remark 1.5, it is not sensible to try to recover the xix_{i} themselves.) Such problems in general are referred to as group synchronization, and this particular noise model is called the truth-or-Haar model, so named by [PWBM16].

To assess the quality of the estimator M^\widehat{M}, we will introduce a loss function :G×G\ell:G\times G\to\mathbb{R}. In principle in our results this can be nearly arbitrary, but to interpret our results naturally it should be viewed as a distance-like function of two elements of GG. We would then like to assess, given an estimator, the average loss

𝔼[1n2i,j=1n(Mij,M^(Y)ij)].\mathbb{E}\left[\frac{1}{n^{2}}\sum_{i,j=1}^{n}\ell(M_{ij},\widehat{M}(Y)_{ij})\right].

We expect that the average loss should also concentrate around its expectation under very mild assumptions, and we always observe strong concentration in our experiments, but we do not pursue establishing such results theoretically here. As a simple concrete example, for G=/KG=\mathbb{Z}/K we may take simply

(x,y):=𝟏{xy},\ell(x,y)\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\mathbf{1}\{x\neq y\},

in which case the average loss is

1n2i,j=1n(Mij,M^ij)=#{(i,j)[n]2:MijM^ij}n2,\frac{1}{n^{2}}\sum_{i,j=1}^{n}\ell(M_{ij},\widehat{M}_{ij})=\frac{\#\{(i,j)\in[n]^{2}:M_{ij}\neq\widehat{M}_{ij}\}}{n^{2}},

the fraction of incorrectly estimated entries of MM.

While quite general GG have been considered in the literature, we consider two particular examples that are well-suited to applying our theoretical results: the finite cyclic groups G=/KG=\mathbb{Z}/K for some K2K\geq 2, and the infinite circle group G=U(1)SO(2)G=U(1)\cong\mathrm{SO}(2), which we can identify with /2π\mathbb{R}/2\pi. Note that the cyclic groups are subgroups of U(1)U(1); all of these cases are sometimes referred to as angular synchronization problems.

The case K=2K=2 has an important additional interpretation: it is equivalent to a dense version of the stochastic block model, a much-studied model of community detection in random graphs. In this case, the xix_{i} are Boolean and may be identified with a label given to each vertex ii in a graph, and Mij{1,+1}M_{ij}\in\{-1,+1\} describes whether vertices ii and jj have the same label or different labels. In this case, our results describe the asymptotic rate of mislabeling for essentially arbitrary rounded spectral estimators of these pairwise community membership relations. Some of our numerical experiments will concern this case, but we do not focus on it otherwise since we will see that, because /2\mathbb{Z}/2 has only one non-trivial group character, multi-spectral algorithms (at least in our framework for them) are not applicable to this example.

Let us focus on G=U(1)G=U(1) for the purposes of exposition; our results below apply equally well to the finite cyclic groups as well. In this case, Singer in [Sin11] observed that a reasonable spectral algorithm is as follows. (See Section 1.4 below for further references.) Given YGn×nY\in G^{n\times n}, we may build Hhermn×nH\in\mathbb{C}^{n\times n}_{\mathrm{herm}} by taking Hij=χ(Yij)H_{ij}=\chi(Y_{ij}) entrywise for a non-trivial character (or one-dimensional representation) χ\chi of GG. A natural choice when G=U(1)G=U(1) identified with /2π\mathbb{R}/2\pi is to simply take

Hij=1nexp(𝒊Yij).H_{ij}=\frac{1}{\sqrt{n}}\exp(\bm{i}Y_{ij}).

This is close to a spiked matrix model, and thus the top eigenvector v^\widehat{v} of HH gives a good estimate of the vector vv of vi=exp(𝒊gi)v_{i}=\exp(\bm{i}g_{i}). We may then round the v^iv^j¯\widehat{v}_{i}\overline{\widehat{v}_{j}} to the unit circle of \mathbb{C} and take its argument as an estimate Mij=gigj1M_{ij}=g_{i}g_{j}^{-1}.

Subsequent literature has explored multi-frequency algorithms, which instead of applying a single χ\chi to GG apply several different characters. Yet it is less clear how to leverage these different frequencies in spectral algorithms; the authors of [PWBM18], who studied analogous approximate message passing algorithms, wrote:

“In sharp contrast to spectral methods, which offer no reasonable way to couple the frequencies together, AMP produces an estimate that is orders of magnitude more accurate than what is possible with a single frequency.”

An algorithm somewhat in the spirit of capturing multi-frequency information with spectral algorithms was later proposed in [GZ19], but here we explore and will obtain much more precise predictions concerning a direct implementation of this idea. We propose the following natural class of algorithm. First, from YGn×nY\in G^{n\times n}, we build a tuple of Hermitian matrices H(1),,H(L)H^{(1)},\ldots,H^{(L)} by applying several characters of GG to the entries of YY. To be concrete, for fixed integers k1,,kLk_{1},\ldots,k_{L} such that the characters below are nontrivial, we take:

χ(a)(x):={exp(2π𝒊kax/K)if G=/K,exp(𝒊kax)if G=U(1)},\chi^{(a)}(x)\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left\{\begin{array}[]{ll}\exp(2\pi\bm{i}k_{a}x/K)&\text{if }G=\mathbb{Z}/K,\\ \exp(\bm{i}k_{a}x)&\text{if }G=U(1)\end{array}\right\}, (1.14)

and set, for each a[L]a\in[L],

Hij(a):=1nχ(a)(Yij).H^{(a)}_{ij}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\frac{1}{\sqrt{n}}\chi^{(a)}(Y_{ij}).

Note that each H(a)H^{(a)} is Hermitian since we have chosen YY to be GG-Hermitian and χ(a)(x1)=χ(a)(x)¯\chi^{(a)}(x^{-1})=\overline{\chi^{(a)}(x)}. We will see that the covariance condition of Theorem 1.8 is satisfied provided we choose characters that are not only distinct but also not conjugate; this is sensible since applying conjugate characters to YY would yield conjugate matrices H(a)H^{(a)} containing essentially the same information. Thus we assume that, for aba\neq b, χ(a){χ(b),χ(b)¯}\chi^{(a)}\notin\{\chi^{(b)},\overline{\chi^{(b)}}\}. We will see in our proofs that this setup indeed makes (H(1),,H(L))(H^{(1)},\dots,H^{(L)}) a weakly Wigner tuple in the sense of Definition 1.7.

Next, let v^(a)=v1(H(a))\widehat{v}^{(a)}=v_{1}(H^{(a)}) for each a[L]a\in[L] be the top eigenvector of each of these matrices. Finally, let Round:LG\mathrm{Round}:\mathbb{C}^{L}\to G be a function that rounds tuples of complex numbers to GG. We think of Round\mathrm{Round} as an inverse of 𝝌=(χ(1),,χ(L))\bm{\chi}=(\chi^{(1)},\dots,\chi^{(L)}), extended to an “approximate inverse” on all of L\mathbb{C}^{L} rather than just 𝝌(G)L\bm{\chi}(G)\subset\mathbb{C}^{L}. Using this rounding function, we define the estimator

M^ij:=Round((nv^i(a)v^j(a)¯)a=1L).\widehat{M}_{ij}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\mathrm{Round}\left(\left(n\cdot\widehat{v}^{(a)}_{i}\overline{\widehat{v}^{(a)}_{j}}\right)_{a=1}^{L}\right).

One natural class that we will focus on in our experiments is

Roundϱ(z1,,zL):=argminxG(1La=1L|χ(a)(x)za|ϱ)1/ϱ,\mathrm{Round}_{\varrho}(z_{1},\dots,z_{L})\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\argmin_{x\in G}\left(\frac{1}{L}\sum_{a=1}^{L}\big|\chi^{(a)}(x)-z_{a}\big|^{\varrho}\right)^{1/\varrho},

the minimum of the generalized mean with parameter ϱ\varrho of the distance of zaz_{a} to a character value χ(a)(x)\chi^{(a)}(x) over the LL coordinates. We may take any ϱ[,]\varrho\in[-\infty,\infty], the generalized mean becoming a minimum and maximum at the respective extremes.

We note that, comparing again to Singer’s spectral algorithm, the case L=1L=1 and k1=1k_{1}=1 reduces, after identifying G=U(1)G=U(1) with the unit circle, to Roundϱ(z):=z/|z|\mathrm{Round}_{\varrho}(z)\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}z/|z| regardless of the choice of ϱ\varrho, and likewise for G=/KG=\mathbb{Z}/K to the function that rounds a complex number to the nearest KKth root of unity. In both cases, these are the natural rounding functions for producing an estimator M^\widehat{M} from a spectral estimate; our Roundϱ\mathrm{Round}_{\varrho} extends these to natural ways to round multi-frequency estimates.

The following uses our earlier results to describe, in considerable generality, the asymptotic average loss of rounded spectral estimators for spectral (L=1L=1) and multi-spectral or multi-frequency (L2L\geq 2) algorithms for group synchronization.

Theorem 1.11.

Suppose in the above setting that χ(a){χ(b),χ(b)¯}\chi^{(a)}\notin\{\chi^{(b)},\overline{\chi^{(b)}}\} for aba\neq b, that the map x𝛘(x)=(χ(1)(x),,χ(L)(x))x\mapsto\bm{\chi}(x)=(\chi^{(1)}(x),\dots,\chi^{(L)}(x)) is injective on GG, and that

p=p(n)=θnp=p(n)=\frac{\theta}{\sqrt{n}}

for some θ>1\theta>1. Let 𝔽a=\mathbb{F}_{a}=\mathbb{R} if χ(a)(G)\chi^{(a)}(G)\subseteq\mathbb{R} and 𝔽a=\mathbb{F}_{a}=\mathbb{C} otherwise, and view 𝔽1××𝔽L\mathbb{F}_{1}\times\cdots\times\mathbb{F}_{L} as a real Euclidean space. Let Round:LG\mathrm{Round}:\mathbb{C}^{L}\to G be a Borel function whose restriction to 𝔽1××𝔽L\mathbb{F}_{1}\times\cdots\times\mathbb{F}_{L} is continuous Lebesgue-almost everywhere. Let :G×G\ell:G\times G\to\mathbb{R} be arbitrary if G=/KG=\mathbb{Z}/K, or a smooth function if G=U(1)G=U(1). Let x,yHaar(G)x,y\sim\mathrm{Haar}(G) and ga,ha𝒩𝔽a(0,1)g_{a},h_{a}\sim\mathcal{N}_{\mathbb{F}_{a}}(0,1) for a[L]a\in[L] all be drawn independently. Then, we have

limn𝔼[1n2i,j=1n(Mij,M^ij)]\displaystyle\lim_{n\to\infty}\mathbb{E}\left[\frac{1}{n^{2}}\sum_{i,j=1}^{n}\ell(M_{ij},\widehat{M}_{ij})\right]
=𝔼[(xy1,Round(((ρ(θ)χ(a)(x)+τ(θ)ga)(ρ(θ)χ(a)(y)+τ(θ)ha)¯)a=1L))].\displaystyle\hskip 28.45274pt=\mathbb{E}\left[\ell\bigg(xy^{-1},\mathrm{Round}\left(\left((\rho(\theta)\chi^{(a)}(x)+\tau(\theta)g_{a})\overline{(\rho(\theta)\chi^{(a)}(y)+\tau(\theta)h_{a})}\right)_{a=1}^{L}\right)\bigg)\right]. (1.15)

As in the abstract case earlier, this gives a “nearly closed form” for the error rate of such estimators under various notions of error, provided we can estimate the integral of fixed dimension on the right-hand side. We note that our condition θ>1\theta>1 is natural here, since the results of [Kun25] give evidence that, at least for finite GG, this problem is computationally hard when θ<1\theta<1, requiring super-polynomial time to even distinguish an observation YGn×nY\in G^{n\times n} from a matrix of i.i.d. (up to the GG-Hermitian relations) Haar-distributed group elements. Similar results with the same θ=1\theta=1 threshold under additive Gaussian noise have also been obtained by [KBK26, Li26].

We give some illustrative numerical experiments here for the case G=U(1)G=U(1) rounding function Roundϱ\mathrm{Round}_{\varrho} with ϱ=2\varrho=2, and loss function (x,y)=1cos(xy)\ell(x,y)=1-\cos(x-y); in Section 5 we give additional supporting results for the claim of Gaussianity, for the quality of our predictions for different ϱ\varrho, and for discrete groups G=/KG=\mathbb{Z}/K. First, in Figure 2, for matrices of dimension n=500n=500 and numbers of frequencies L{1,2,5}L\in\{1,2,5\} we plot empirical average losses achieved under the truth-or-Haar noise model as well as the analogous Gaussian additive noise model, and the theoretical prediction of Theorem 1.11, observing excellent agreement in all cases (especially away from the transition point θ=1\theta=1). Second, in Figure 3, we plot on a single pair of axes only the theoretical predictions for each number of frequencies (the right-hand side of (1.15)). We observe that more frequencies lead to lower estimation error for a large signal strength θ\theta, confirming the superiority of multi-frequency algorithms over the direct spectral algorithm of [Sin11]. Further, as the number of frequencies LL increases, the value of θ\theta where the corresponding algorithm becomes superior to ones for smaller LL decreases; thus, in a sense the limit LL\to\infty appears to give the “best” of these algorithms (though their computational cost also increases with LL). It remains unclear to us how to explain why, when comparing two finite L1<L2L_{1}<L_{2}, there is a regime of θ\theta close to 1 for which the algorithm using only L1L_{1} frequencies achieves lower error. We leave this as an interesting question for future work.

We remark that, in both cases, the theoretical predictions given by evaluating the prediction in Theorem 1.11 above are approximated by Monte Carlo estimation of the expectation, and thus have an inherent error, but we take a sufficient number of trials that this is vanishingly small.

Figure 2: Agreement of empirical and predicted performance of synchronization algorithms. We plot the empirical average loss of the rounded multi-frequency spectral algorithm for group synchronization for G=U(1)G=U(1) (with other parameters as detailed at the end of Section 1.3) versus the theoretical prediction of Theorem 1.11 for various numbers of frequencies LL.
Figure 3: Predicted superiority of multi-frequency spectral algorithms. In the same setting as Figure 2, for group synchronization over G=U(1)G=U(1), we plot only the theoretical predictions of average entrywise loss per Theorem 1.11 for various numbers of frequencies LL.

1.4 Related work

Spiked matrix models

The spiked matrix model originates in statistics in the work of [Joh01], who proposed a version where the signal is applied to a covariance matrix of vector observations, though a special case of a model similar to ours was discussed much earlier in the foundational work [FK81]. Those predictions were proved in the seminal work of [BBAP05]. Subsequently, many variants of such results appeared; the first results to the effect of our Theorem 1.2 for models with additive noise appear to have been shown soon after by [Péc06, FP07]. Other more general versions are shown, for instance, in [CDMF09, BGN11], and an overview of this line of work from the mathematical point of view may be found in [Cap17], and from the statistical point of view in [PA14, JP18]. The general approach we take to analyzing spiked matrix models using the resolvent appears in [BGN11, BGGM11] and was used more systematically together with isotropic local laws by [KY13b, KY14, KY17], whose techniques we will follow closely. Another relevant work for studying eigenvectors by this method is the earlier one of [KY13a], though this concerns matrices like our WW with no low-rank perturbation.

Universality and non-universality

As we discussed briefly above, the general phenomenon of non-universality of fluctuations in spiked matrix models depending on the structure of the signal was gradually uncovered by works including [CDMF09, CDMF12, RS13, BGGM11, KY13b, KY14]. For eigenvectors, non-universality for localized signals was shown by [CDM21], with regard to a quantity similar to v^,v\langle\widehat{v},v\rangle. We note that this is a special projection of v^\widehat{v} to study because, for θ>1\theta>1, per Theorem 1.2 it is of magnitude Θ(1)\Theta(1); in contrast, for vv delocalized the entries v^i\widehat{v}_{i} are of magnitude o(1)o(1) and therefore considerably more delicate to understand. Non-asymptotic entrywise eigenvector bounds, including an application to /2\mathbb{Z}/2 synchronization, were obtained by [AFWZ20]. Distributions of similar quantities in asymmetric and rectangular matrix denoising models and in supercritical spiked covariance models are considered by [BDW21, BDWW22]. The results perhaps most similar to ours are those of [FFHL22, BCH26], but both consider, in our normalization, the case θ=θ(n)\theta=\theta(n)\to\infty, in which case v^\widehat{v} is very close to vv and its Gaussian fluctuations have vanishing magnitude.

Approximate message passing and state evolution

One important other branch of the literature where similar statements of Gaussian fluctuations appear is in the characterization of approximate message passing (AMP) algorithms. See, e.g., [FVRS22] for a general survey of this area. These algorithms perform a general kind of nonlinear power method, alternating multiplying a vector by a matrix like our HH and applying an entrywise nonlinearity. Various expressions of the Gaussian distribution of the result of such algorithms are known as state evolution descriptions. In principle, the power method itself for computing v^\widehat{v} is a special case of AMP where the nonlinearities are omitted, and thus various averages of nonlinear functions of v^\widehat{v} can be described as measurements of the state of a suitable AMP algorithm. However, there is an important caveat that much of the work on AMP concerns only asymptotic results as nn\to\infty for a finite number of iterations t=O(1)t=O(1) of such an algorithm, from a random initialization. Such an algorithm essentially would compute HtxH^{t}x for a random xx independent of HH, which does not faithfully approximate v^\widehat{v}, and thus such analysis cannot be used in our setting directly.

In the AMP literature there have been essentially two workarounds considered to this kind of issue (variants of which are also relevant to other, more sophisticated instances of AMP). First, one may perform non-asymptotic analysis of a number of iterations t=t(n)t=t(n) depending on nn, as pioneered by [RV18] and developed further by [LW23, LFW23, CR24]. However, such analysis requires strong assumptions on WW such as Gaussianity or orthogonal invariance. The recent work [Han25] is in a similar spirit to ours, but only can treat t(n)(logn)1/3t(n)\leq(\log n)^{1/3} iterations of an AMP algorithm, while t(n)=Θ(logn)t(n)=\Theta(\log n) are required to faithfully estimate v^\widehat{v} in spiked matrix models.

Second, a line of work of [MV21, MV22, MTV22], also discussed in Section 3.2 of [FVRS22], has sought to analyze AMP initialized not from a random vector xx independent of HH, but from precisely our v^\widehat{v}, a so-called spectral initialization from the top eigenvector. Of course, if this were possible, then we could study v^\widehat{v} very directly using such analysis. However, these works involve various tricks that rely deeply on the Gaussian structure of WW, and in their current form do not seem to imply any analysis of v^\widehat{v} comparable to ours for non-Gaussian WW, or indeed for any WW besides very symmetric Gaussian models like the GOE and GUE—the information about v^\widehat{v} that these approaches use stems from precisely the same invariance properties of the GOE and GUE that we take advantage of in our proof of Theorem 1.8.

Group synchronization

The group synchronization problem asks to recover group-valued labels from noisy pairwise relative measurements. Algorithmic approaches include Singer’s eigenvector and semidefinite programming methods for angular synchronization [Sin11], message passing over compact groups [PWBM18], and multi-frequency phase synchronization [GZ19]. Yang, Wee, and Fan [YWF25] characterize the information-theoretic limits of inference in several quadratic models over compact groups, including angular and phase synchronization. Other work proves lower bounds for low-degree polynomials and low-coordinate-degree algorithms in Gaussian multi-frequency and truth-or-Haar synchronization models [KBK26, Kun25, Li26], giving evidence for computational hardness below the threshold θ=1\theta=1. Our results complement these by giving precise formulas for the asymptotic performance of specific entrywise rounded spectral estimators when θ>1\theta>1.

1.5 Notation

We use the standard notation [n]={1,,n}[n]=\{1,\dots,n\}. We occasionally use the special function

φ(n):=(logn)loglogn.\varphi(n)\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}(\log n)^{\log\log n}.

We write 𝒊=1\bm{i}=\sqrt{-1} for the imaginary unit, not to be confused with the indices i[n]i\in[n]. For 𝔽{,}\mathbb{F}\in\{\mathbb{R},\mathbb{C}\}, we denote the associated unit sphere in dimension nn by 𝕊n1(𝔽)={x𝔽n:x=1}\mathbb{S}^{n-1}(\mathbb{F})=\{x\in\mathbb{F}^{n}:\|x\|=1\}. For a proposition PP about various variables defined in our arguments, we write 𝟏{P}\mathbf{1}\{P\} to equal 1 if PP is true and 0 if PP is false. Reusing this notation, in probabilistic arguments we also use 𝟏\mathbf{1}_{\mathcal{E}} for the indicator random variable of an event \mathcal{E}.

We write hermn×n\mathbb{C}^{n\times n}_{\mathrm{herm}} for the set of Hermitian matrices and symn×n\mathbb{R}^{n\times n}_{\mathrm{sym}} for the set of real symmetric matrices. I=InI=I_{n} denotes the n×nn\times n identity matrix; we omit the subscript if it is clear from context. For Xhermn×nX\in\mathbb{C}^{n\times n}_{\mathrm{herm}}, we write λ1(X)λ2(X)λn(X)\lambda_{1}(X)\geq\lambda_{2}(X)\geq\cdots\geq\lambda_{n}(X) for its ordered real eigenvalues and v1(X),,vn(X)v_{1}(X),\dots,v_{n}(X) for the associated eigenvectors, provided the corresponding eigenvalues are simple.

The asymptotic notations ,,O(),Θ(),o(),ω(),Ω()\lesssim,\gtrsim,O(\cdot),\Theta(\cdot),o(\cdot),\omega(\cdot),\Omega(\cdot) have their usual meanings, always with respect to the limit nn\to\infty. We sometimes write subscripts on these notations for parameters that the implicit constants depend on, but when proving a result whose statement gives this dependence explicitly, we omit those parameters for the sake of brevity. In general, all of the quantities we have called “parameters” in the statements, with subscripts of “vv” or “WW”, are viewed as constants for the purposes of our proofs.

2 Preliminaries

2.1 Probability

Let us give a name to the tail bound assumption (1.1) that we make on the magnitudes of the entries of our generalized Wigner matrices.

Definition 2.1.

We say that a random variable XX is ξ\xi-sub-Weibull if, for all t>0t>0,

[|X|t]ξ1exp(tξ).\mathbb{P}[|X|\geq t]\leq\xi^{-1}\exp(-t^{\xi}).
Proposition 2.2.

If XX is ξ\xi-sub-Weibull and p1p\geq 1, then there exists C(ξ,p)C(\xi,p) such that 𝔼|X|pC(ξ,p)\mathbb{E}|X|^{p}\leq C(\xi,p).

Proof.

We have

𝔼|X|p\displaystyle\mathbb{E}|X|^{p} =t=0[|X|pt]dt\displaystyle=\int_{t=0}^{\infty}\mathbb{P}[|X|^{p}\geq t]\,dt
=t=0[|X|t1/p]dt\displaystyle=\int_{t=0}^{\infty}\mathbb{P}[|X|\geq t^{1/p}]\,dt
ξ1t=0exp(tξ/p)𝑑t,\displaystyle\leq\xi^{-1}\int_{t=0}^{\infty}\exp(-t^{\xi/p})\,dt,

where the remaining integral is finite and depends only on ξ\xi and pp. ∎

The following notation for tail bounds up to a polynomial amount of “slack” will be useful throughout. See Appendix A of [AEK17] for further generalities about this definition.

Definition 2.3 (Polynomial stochastic domination).

For sequences of non-negative random variables X=XnX=X_{n} and Y=YnY=Y_{n}, we say XX is polynomially stochastically dominated by YY, written as XYX\prec Y or (Xn)(Yn)(X_{n})\prec(Y_{n}), if for every ε>0\varepsilon>0 and for all D>0D>0, there exists n0=n0(ε,D)n_{0}=n_{0}(\varepsilon,D) such that, for all nn0n\geq n_{0},

[Xn>nεYn]nD.\mathbb{P}[X_{n}>n^{\varepsilon}Y_{n}]\leq n^{-D}.

Also, if n\mathcal{E}_{n} is a sequence of events, we say that n\mathcal{E}_{n} occurs with polynomially high probability if, for every D>0D>0, there exists n0=n0(D)n_{0}=n_{0}(D) such that, for all nn0n\geq n_{0},

[nc]nD.\mathbb{P}[\mathcal{E}_{n}^{c}]\leq n^{-D}.

The following property of polynomial stochastic domination is elementary to verify by the union bound.

Proposition 2.4.

Suppose that Xn,1,,Xn,k,Yn,1,,Yn,kX_{n,1},\dots,X_{n,k},Y_{n,1},\dots,Y_{n,k} are non-negative random variables such that (Xn,i)(Yn,i)(X_{n,i})\prec(Y_{n,i}) for each fixed i[k]i\in[k]. Then,

i=1kXn,i\displaystyle\sum_{i=1}^{k}X_{n,i} i=1kYn,i,\displaystyle\prec\sum_{i=1}^{k}Y_{n,i},
i=1kXn,i\displaystyle\prod_{i=1}^{k}X_{n,i} i=1kYn,i.\displaystyle\prec\prod_{i=1}^{k}Y_{n,i}.

The following gives a convenient condition under which polynomial stochastic domination results can be converted to control of expectations.

Proposition 2.5.

Let (Xn)n1(X_{n})_{n\geq 1} be a sequence of non-negative random variables and (Yn)n1(Y_{n})_{n\geq 1} be a sequence of deterministic non-negative numbers. Suppose that the following hold:

  1. 1.

    XnYnX_{n}\prec Y_{n}.

  2. 2.

    supn1𝔼Xn2<\sup_{n\geq 1}\mathbb{E}X_{n}^{2}<\infty.

  3. 3.

    There exists C>0C>0 such that YnnCY_{n}\geq n^{-C} for all nn.

Then, for any ε>0\varepsilon>0,

𝔼XnεnεYn.\mathbb{E}X_{n}\lesssim_{\varepsilon}n^{\varepsilon}Y_{n}.
Proof.

Let K=supn1𝔼Xn2K=\sup_{n\geq 1}\mathbb{E}X_{n}^{2}. We have

𝔼Xn\displaystyle\mathbb{E}X_{n} =𝔼Xn𝟏XnnεYn+𝔼Xn𝟏Xn>nεYn\displaystyle=\mathbb{E}X_{n}\mathbf{1}_{X_{n}\leq n^{\varepsilon}Y_{n}}+\mathbb{E}X_{n}\mathbf{1}_{X_{n}>n^{\varepsilon}Y_{n}}
and using the bound from the indicator on the first term and the Cauchy-Schwarz inequality on the second,
nεYn+(𝔼Xn2)1/2([Xn>nεYn])1/2\displaystyle\leq n^{\varepsilon}Y_{n}+(\mathbb{E}X_{n}^{2})^{1/2}(\mathbb{P}[X_{n}>n^{\varepsilon}Y_{n}])^{1/2}
Now, choosing some D2CD\geq 2C, if nn0(ε,D)n\geq n_{0}(\varepsilon,D), then we have
nεYn+K1/2nD/2\displaystyle\leq n^{\varepsilon}Y_{n}+K^{1/2}n^{-D/2}
nεYn+K1/2Yn\displaystyle\leq n^{\varepsilon}Y_{n}+K^{1/2}Y_{n}
(1+K1/2)nεYn,\displaystyle\leq(1+K^{1/2})n^{\varepsilon}Y_{n},

completing the proof. ∎

2.2 Matrix resolvents and low-rank perturbations

We will work extensively with the resolvents of random matrices, whose basic properties we enumerate below.

Definition 2.6.

The resolvent of a matrix XX at the point zz\in\mathbb{C} is

Resz(X)=(zIX)1,\mathrm{Res}_{z}(X)=(zI-X)^{-1},

defined at any zz that is not an eigenvalue of XX. In particular, if Xhermn×nX\in\mathbb{C}^{n\times n}_{\mathrm{herm}}, then Resz(X)\mathrm{Res}_{z}(X) is defined away from a compact subset of \mathbb{R}.

Proposition 2.7 (Riesz projection formula).

Let Γ\Gamma be a counterclockwise simple contour in \mathbb{C} that does not pass through any eigenvalues of a matrix Xhermn×nX\in\mathbb{C}^{n\times n}_{\mathrm{herm}} and encircles a single eigenvalue λ\lambda of XX. Let PP be the orthogonal projection to the eigenspace of XX associated to λ\lambda. Then,

P=12πiΓResz(X)𝑑z.P=\frac{1}{2\pi i}\oint_{\Gamma}\mathrm{Res}_{z}(X)\,dz.
Proposition 2.8 (Resolvent identities).

Let X,Yhermn×nX,Y\in\mathbb{C}^{n\times n}_{\mathrm{herm}} and w,z(spec(X)spec(Y))w,z\in\mathbb{C}\setminus(\mathrm{spec}(X)\cup\mathrm{spec}(Y)). Then, the following hold:

  1. 1.

    (First resolvent identity) RX(z)RX(w)=(wz)RX(z)RX(w)R_{X}(z)-R_{X}(w)=(w-z)R_{X}(z)R_{X}(w).

  2. 2.

    (Second resolvent identity) RX(z)RY(z)=RX(z)(XY)RY(z)R_{X}(z)-R_{Y}(z)=R_{X}(z)(X-Y)R_{Y}(z).

  3. 3.

    (Resolvent expansion) RX(z)=s=0tRY(z)((XY)RY(z))s+RX(z)((XY)RY(z))t+1R_{X}(z)=\sum_{s=0}^{t}R_{Y}(z)((X-Y)R_{Y}(z))^{s}+R_{X}(z)((X-Y)R_{Y}(z))^{t+1} for any t0t\geq 0.

The two resolvent identities are standard, while the resolvent expansion follows from repeatedly applying the second resolvent identity.

One important use of resolvents is in characterizing the eigenvalues and eigenvectors of a low-rank perturbation of a matrix. We will use these results on random spiked matrix models, but they are true deterministically as well. We state the results for the general case

H=i=1kθivivi+W=VDiag(θ)V+W.H=\sum_{i=1}^{k}\theta_{i}v_{i}v_{i}^{*}+W=V\mathrm{Diag}(\theta)V^{*}+W.

To study the eigenvalues, we use the following device, used in derivations of the main phase transition phenomena of spiked matrix models by, e.g., [KY13b, KY14]; see also Section 5 of [BGN11]. This characterization of the eigenvalues of a low-rank perturbation is sometimes called the secular equation.

Proposition 2.9.

Suppose that θi0\theta_{i}\neq 0 for every i[k]i\in[k] and that λspec(W)\lambda\notin\mathrm{spec}(W). Then, λspec(H)\lambda\in\mathrm{spec}(H) if and only if

det(Diag(θ)1VRW(λ)V)=0.\det(\mathrm{Diag}(\theta)^{-1}-V^{*}R_{W}(\lambda)V)=0.

If k=1k=1, θ=θ1\theta=\theta_{1}, and v=v1v=v_{1}, then this reduces to

vRW(λ)v=1θ.v^{*}R_{W}(\lambda)v=\frac{1}{\theta}.
Proof.

Since λspec(W)\lambda\notin\mathrm{spec}(W), the matrix RW(λ)1=λIWR_{W}(\lambda)^{-1}=\lambda I-W is well-defined and non-singular. We therefore have

det(λIH)\displaystyle\det(\lambda I-H) =det(λIWVDiag(θ)V)\displaystyle=\det(\lambda I-W-V\mathrm{Diag}(\theta)V^{*})
=det(RW(λ)1)det(IRW(λ)VDiag(θ)V)\displaystyle=\det(R_{W}(\lambda)^{-1})\det(I-R_{W}(\lambda)V\mathrm{Diag}(\theta)V^{*})
=det(RW(λ)1)det(IDiag(θ)VRW(λ)V)\displaystyle=\det(R_{W}(\lambda)^{-1})\det(I-\mathrm{Diag}(\theta)V^{*}R_{W}(\lambda)V)
=det(RW(λ)1)det(Diag(θ))det(Diag(θ)1VRW(λ)V),\displaystyle=\det(R_{W}(\lambda)^{-1})\det(\mathrm{Diag}(\theta))\det(\mathrm{Diag}(\theta)^{-1}-V^{*}R_{W}(\lambda)V),

where the third line follows from Sylvester’s determinant identity. The first two determinant factors in the last line are non-zero, which gives the result. ∎

We also use the following calculation of the resolvent of a spiked matrix model as a low-rank perturbation of the resolvent of the underlying noise matrix.

Proposition 2.10.

Suppose that θi0\theta_{i}\neq 0 for every i[k]i\in[k] and that zspec(W)spec(H)z\notin\mathrm{spec}(W)\cup\mathrm{spec}(H). Then,

RH(z)=RW(z)+RW(z)V(Diag(θ)1VRW(z)V)1VRW(z).R_{H}(z)=R_{W}(z)+R_{W}(z)V(\mathrm{Diag}(\theta)^{-1}-V^{*}R_{W}(z)V)^{-1}V^{*}R_{W}(z).
Proof.

We calculate directly:

RW+VDiag(θ)V(z)\displaystyle R_{W+V\mathrm{Diag}(\theta)V^{*}}(z) =(zIWVDiag(θ)V)1\displaystyle=(zI-W-V\mathrm{Diag}(\theta)V^{*})^{-1}
=(IRW(z)VDiag(θ)V)1RW(z)\displaystyle=(I-R_{W}(z)V\mathrm{Diag}(\theta)V^{*})^{-1}R_{W}(z)
=(I+RW(z)V(Diag(θ)1VRW(z)V)1V)RW(z)\displaystyle=\left(I+R_{W}(z)V(\mathrm{Diag}(\theta)^{-1}-V^{*}R_{W}(z)V)^{-1}V^{*}\right)R_{W}(z)
=RW(z)+RW(z)V(Diag(θ)1VRW(z)V)1VRW(z),\displaystyle=R_{W}(z)+R_{W}(z)V(\mathrm{Diag}(\theta)^{-1}-V^{*}R_{W}(z)V)^{-1}V^{*}R_{W}(z),

where we have used the Woodbury matrix inverse identity in the middle. ∎

Using this, we may compute the eigenvectors of HH.

Proposition 2.11.

Suppose that λspec(H)spec(W)\lambda\in\mathrm{spec}(H)\setminus\mathrm{spec}(W) is an eigenvalue with some multiplicity \ell, and let PP be the orthogonal projection to the associated (\ell-dimensional) eigenspace. Then, the dimension of ker(Diag(θ)1VRW(λ)V)\ker(\mathrm{Diag}(\theta)^{-1}-V^{*}R_{W}(\lambda)V) is \ell. Letting Uk×U\in\mathbb{C}^{k\times\ell} have as its columns an orthonormal basis for this kernel, we have

P=RW(λ)VU(UVRW(λ)2VU)1UVRW(λ).P=R_{W}(\lambda)VU(U^{*}V^{*}R_{W}(\lambda)^{2}VU)^{-1}U^{*}V^{*}R_{W}(\lambda).

Equivalently, the eigenspace of HH associated to λ\lambda is the column space of RW(λ)VUR_{W}(\lambda)VU, spanned by RW(λ)VuR_{W}(\lambda)Vu over all uu such that (Diag(θ)1VRW(λ)V)u=0(\mathrm{Diag}(\theta)^{-1}-V^{*}R_{W}(\lambda)V)u=0.

If k=1k=1, θ=θ1\theta=\theta_{1}, v=v1v=v_{1}, and λ\lambda is a simple eigenvalue with eigenvector v^\widehat{v} (matching our earlier notation in the case of λ\lambda the largest eigenvalue of HH), then this reduces to

v^v^=1vRW(λ)2vRW(λ)vvRW(λ),\widehat{v}\widehat{v}^{*}=\frac{1}{v^{*}R_{W}(\lambda)^{2}v}R_{W}(\lambda)vv^{*}R_{W}(\lambda),

and v^\widehat{v} up to rescaling is RW(λ)vR_{W}(\lambda)v.

Proof.

Let Γ\Gamma be a closed contour encircling λ\lambda but no other eigenvalues of WW or HH. We have, by Propositions 2.7 and 2.10,

P\displaystyle P =12πiΓRH(z)𝑑z\displaystyle=\frac{1}{2\pi i}\oint_{\Gamma}R_{H}(z)\,dz
=12πiΓ(RW(z)+RW(z)V(Diag(θ)1VRW(z)V)1VRW(z))𝑑z\displaystyle=\frac{1}{2\pi i}\oint_{\Gamma}\big(R_{W}(z)+R_{W}(z)V(\mathrm{Diag}(\theta)^{-1}-V^{*}R_{W}(z)V)^{-1}V^{*}R_{W}(z)\big)\,dz
and, since Γ\Gamma avoids the eigenvalues of WW, RW(z)R_{W}(z) is analytic on an open neighborhood of the interior of Γ\Gamma, and thus by Cauchy’s integral theorem the first term integrates to zero and we have
=12πiΓRW(z)V(Diag(θ)1VRW(z)V)1VRW(z)𝑑z.\displaystyle=\frac{1}{2\pi i}\oint_{\Gamma}R_{W}(z)V(\mathrm{Diag}(\theta)^{-1}-V^{*}R_{W}(z)V)^{-1}V^{*}R_{W}(z)\,dz.
Now, by Proposition 2.9, the integrand has poles precisely zz equaling the eigenvalues of HH. Since the only one of these that Γ\Gamma encircles is λ\lambda, the above equals the (matrix-valued) residue at this pole. Again by our assumption, RW(z)R_{W}(z) is analytic in an open neighborhood of the interior of Γ\Gamma, and so we may take the (matrix-valued) residue of the inner term (Diag(θ)1VRW(z)V)1(\mathrm{Diag}(\theta)^{-1}-V^{*}R_{W}(z)V)^{-1} at z=λz=\lambda. Let Uk×U\in\mathbb{C}^{k\times\ell} have as its columns an orthonormal basis of ker(Diag(θ)1VRW(λ)V)\ker(\mathrm{Diag}(\theta)^{-1}-V^{*}R_{W}(\lambda)V). Since ddzRW(z)=RW(z)2\frac{d}{dz}R_{W}(z)=-R_{W}(z)^{2}, this is given by U(UVRW(λ)2VU)1UU(U^{*}V^{*}R_{W}(\lambda)^{2}VU)^{-1}U^{*}. Thus we may evaluate and we find
=RW(λ)VU(UVRW(λ)2VU)1UVRW(λ).\displaystyle=R_{W}(\lambda)VU(U^{*}V^{*}R_{W}(\lambda)^{2}VU)^{-1}U^{*}V^{*}R_{W}(\lambda).

Note that, since λ\lambda\in\mathbb{R}, RW(λ)R_{W}(\lambda) is Hermitian, and so this is just the formula for the orthogonal projection to the span of the vectors RW(λ)VUR_{W}(\lambda)VU. ∎

2.3 Random matrix theory

2.3.1 Semicircle limit theorems

Let us review some of the properties of generalized Wigner random matrices we have mentioned above.

Definition 2.12.

The semicircle distribution, denoted μsc\mu_{\mathrm{sc}}, is the probability measure with density

12π4x2 1{x[2,2]}dx\frac{1}{2\pi}\sqrt{4-x^{2}}\,\mathbf{1}\{x\in[-2,2]\}\,dx

with respect to Lebesgue measure.

The following general limit theorem is a direct consequence of the much more precise results of [EYY12] on generalized Wigner matrices.

Theorem 2.13.

Let W(n)hermn×nW^{(n)}\in\mathbb{C}^{n\times n}_{\mathrm{herm}} be a sequence of generalized Wigner matrices (Definition 1.1) with parameters (γW,ξW)(\gamma_{W},\xi_{W}) not depending on nn. Then, the following convergences hold in probability:

1ni=1nf(λi(W(n)))\displaystyle\frac{1}{n}\sum_{i=1}^{n}f(\lambda_{i}(W^{(n)})) fdμsc,\displaystyle\to\int fd\mu_{\mathrm{sc}},
λ1(W(n))\displaystyle\lambda_{1}(W^{(n)}) 2,\displaystyle\to 2,
λn(W(n))\displaystyle\lambda_{n}(W^{(n)}) 2,\displaystyle\to-2,
W(n)\displaystyle\|W^{(n)}\| 2.\displaystyle\to 2.

In the first claim, f:f:\mathbb{R}\to\mathbb{R} is any bounded continuous function or any polynomial.

Many of the calculations to follow will also involve the following object evaluated with the semicircle measure, a few of whose properties we establish now.

Definition 2.14.

For a probability measure μ\mu on \mathbb{R} supported on some KK\subseteq\mathbb{R}, its Cauchy transform is the function Gμ:KG_{\mu}:\mathbb{C}\setminus K\to\mathbb{C} defined by

Gμ(z)=1zx𝑑μ(x).G_{\mu}(z)=\int\frac{1}{z-x}d\mu(x).
Proposition 2.15.

The Cauchy transform Gμsc(z)G_{\mu_{\mathrm{sc}}}(z) of the semicircle distribution is defined on [2,2]\mathbb{C}\setminus[-2,2] and satisfies:

Gμsc(z)\displaystyle G_{\mu_{\mathrm{sc}}}(z) =12(zz24),\displaystyle=\frac{1}{2}(z-\sqrt{z^{2}-4}),
Gμsc(z)\displaystyle G_{\mu_{\mathrm{sc}}}^{\prime}(z) =12(1zz24).\displaystyle=\frac{1}{2}\left(1-\frac{z}{\sqrt{z^{2}-4}}\right).

In particular, on (2,)(2,\infty)\subset\mathbb{R}, GμscG_{\mu_{\mathrm{sc}}}^{\prime} increases monotonically from -\infty to 0.

2.3.2 Tail bounds

We follow works like [KY13b] in adopting the notation

φ(n):=(logn)loglogn\varphi(n)\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}(\log n)^{\log\log n}

for a nuisance factor that appears in the following results. For intuition, it is useful to keep in mind that φ(n)\varphi(n) grows slower than any polynomial of nn but faster than any polynomial of logn\log n: for any C>0C>0,

(logn)Cφ(n)n1/C.(\log n)^{C}\ll\varphi(n)\ll n^{1/C}.

The following concrete norm bounds will be useful.

Proposition 2.16 (Wigner matrix norm deviation bounds [EYY12]).

Let Whermn×nW\in\mathbb{C}^{n\times n}_{\mathrm{herm}} be a generalized Wigner matrix. Then,

[W2+φ(n)n2/3]Cexp(φ(n))\mathbb{P}[\|W\|\geq 2+\varphi(n)n^{-2/3}]\leq C\exp(-\varphi(n))

for a constant CC depending only on the parameters (γW,ξW)(\gamma_{W},\xi_{W}) of the generalized Wigner matrix.

Proposition 2.17 (Resolvent norm bounds).

Let Whermn×nW\in\mathbb{C}^{n\times n}_{\mathrm{herm}} be a generalized Wigner matrix, and let zz\in\mathbb{C} have Im(z)>nC\mathrm{Im}(z)>n^{-C} and Re(z)>2+c\mathrm{Re}(z)>2+c for some c,C>0c,C>0. Then,

𝔼RW(z)kK\mathbb{E}\|R_{W}(z)\|^{k}\leq K

for some constant KK depending only on c,C,kc,C,k, and the parameters (γW,ξW)(\gamma_{W},\xi_{W}) of the generalized Wigner matrix. The same also holds for WW replaced by the matrix QQ formed by setting entries (a,b)(a,b) and (b,a)(b,a) to zero in WW, for any 1abn1\leq a\leq b\leq n.

Proof.

We have

RW(z)\displaystyle\|R_{W}(z)\| =(zIW)1\displaystyle=\|(zI-W)^{-1}\|
1d(z,spec(W))\displaystyle\leq\frac{1}{d(z,\mathrm{spec}(W))}
1Im(z).\displaystyle\leq\frac{1}{\mathrm{Im}(z)}.

Fix some δ<c\delta<c. Define the event

={W2+δ}={spec(W)[2δ,2+δ]}.\mathcal{E}=\{\|W\|\leq 2+\delta\}=\{\mathrm{spec}(W)\subseteq[-2-\delta,2+\delta]\}.

Then, we have using Proposition 2.16 that

𝔼RW(z)k\displaystyle\mathbb{E}\|R_{W}(z)\|^{k} =𝔼𝟏RW(z)k+𝔼𝟏cRW(z)k\displaystyle=\mathbb{E}\mathbf{1}_{\mathcal{E}}\|R_{W}(z)\|^{k}+\mathbb{E}\mathbf{1}_{\mathcal{E}^{c}}\|R_{W}(z)\|^{k}
(1cδ)k+Im(z)k[c]\displaystyle\leq\left(\frac{1}{c-\delta}\right)^{k}+\mathrm{Im}(z)^{-k}\mathbb{P}[\mathcal{E}^{c}]
Oc,k(1)+nCkexp(φ(n)),\displaystyle\leq O_{c,k}(1)+n^{Ck}\exp(-\varphi(n)),

and the result for WW follows since φ(n)\varphi(n) grows faster than any power of logn\log n. The result for QQ follows by the same argument since WQ=|Wab|\|W-Q\|=|W_{ab}|, which satisfies a sub-Weibull tail bound. ∎

2.3.3 Local laws

We will use the following powerful result about deterministic quadratic forms with the resolvent of a generalized Wigner random matrix.

Theorem 2.18 (Isotropic local law outside the spectrum, Theorem 2.15 and Remark 2.6 of [BEK+14]).

Fix c,C>0c,C>0. Let Whermn×nW\in\mathbb{C}^{n\times n}_{\mathrm{herm}} be a generalized Wigner matrix, let x,y𝕊n1()x,y\in\mathbb{S}^{n-1}(\mathbb{C}) be deterministic, and define

Ω=Ω(c,C)={z:Re(z)>2+c,Im(z)0,|z|C}.\Omega=\Omega(c,C)=\{z\in\mathbb{C}:\mathrm{Re}(z)>2+c,\mathrm{Im}(z)\geq 0,|z|\leq C\}.

Then,

supzΩ|xRW(z)yGμsc(z)xy|n1/2,\sup_{z\in\Omega}\left|x^{*}R_{W}(z)y-G_{\mu_{\mathrm{sc}}}(z)\cdot x^{*}y\right|\prec n^{-1/2}, (2.1)

with the constants n0(ε,D)n_{0}(\varepsilon,D) in the polynomial stochastic domination depending only on c,Cc,C, and the parameters (γW,ξW)(\gamma_{W},\xi_{W}) of the generalized Wigner matrix.

Proof.

For Im(z)>0\mathrm{Im}(z)>0, the result follows from the cited theorem: the error away from the spectrum in the upper half-plane is of order n1/2n^{-1/2}, and Remark 2.6 of the reference makes this estimate uniform in zz. Next, by Proposition 2.16, with polynomially high probability the spectrum of WW is disjoint from Ω\Omega\cap\mathbb{R}. On this event, both RW(z)R_{W}(z) and Gμsc(z)G_{\mu_{\mathrm{sc}}}(z) are continuous as Im(z)0\mathrm{Im}(z)\downarrow 0, so the same estimate also holds on the real axis by the limiting argument of Remark 2.7 of [BEK+14]. ∎

Corollary 2.19.

Let Whermn×nW\in\mathbb{C}^{n\times n}_{\mathrm{herm}} be a generalized Wigner matrix. For some i,j[n]i,j\in[n], let QQ be formed by setting entries (i,j)(i,j) and (j,i)(j,i) in WW to zero. Then, the conclusion (2.1) of Theorem 2.18 holds with WW replaced by QQ.

Proof.

It suffices to show that

supzΩRW(z)RQ(z)n1/2.\sup_{z\in\Omega}\|R_{W}(z)-R_{Q}(z)\|\prec n^{-1/2}.

By the second resolvent identity, we have

RW(z)RQ(z)\displaystyle\|R_{W}(z)-R_{Q}(z)\| =RW(z)(WQ)RQ(z)\displaystyle=\|R_{W}(z)(W-Q)R_{Q}(z)\|
|Wij|RW(z)RQ(z)\displaystyle\leq|W_{ij}|\cdot\|R_{W}(z)\|\cdot\|R_{Q}(z)\|
|Wij|RW(z)(RW(z)+RW(z)RQ(z)),\displaystyle\leq|W_{ij}|\cdot\|R_{W}(z)\|\cdot(\|R_{W}(z)\|+\|R_{W}(z)-R_{Q}(z)\|),
and rearranging this gives
|Wij|RW(z)2(1|Wij|RW(z))\displaystyle\leq\frac{|W_{ij}|\cdot\|R_{W}(z)\|^{2}}{(1-|W_{ij}|\cdot\|R_{W}(z)\|)}

provided that |Wij|RW(z)<1|W_{ij}|\cdot\|R_{W}(z)\|<1. We have |Wij|n1/2|W_{ij}|\prec n^{-1/2} since by assumption it is sub-Weibull, while it follows from Proposition 2.16 that supzRW(z)1\sup_{z}\|R_{W}(z)\|\prec 1 over the range of zz in the statement, and the result follows by a suitable union bound. ∎

We will also use the following corollary, which essentially says that one may differentiate with respect to zz inside the expression appearing in the local law and still obtain the same quality of guarantee.

Corollary 2.20 (Derivative of isotropic local law).

In the setting of Theorem 2.18, we also have

supzΩ|xRW(z)2y+Gμsc(z)xy|n1/2,\sup_{z\in\Omega}\left|x^{*}R_{W}(z)^{2}y+G_{\mu_{\mathrm{sc}}}^{\prime}(z)x^{*}y\right|\prec n^{-1/2},

with the same dependence of constants.

Proof.

For deterministic unit vectors a,bna,b\in\mathbb{C}^{n}, define

fa,b(z):=aRW(z)bGμsc(z)ab.f_{a,b}(z)\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}a^{*}R_{W}(z)b-G_{\mu_{\mathrm{sc}}}(z)a^{*}b.

Let

Ω:={z:Re(z)>2+c/2,|z|C+c/2}.\Omega^{\prime}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\{z\in\mathbb{C}:\mathrm{Re}(z)>2+c/2,|z|\leq C+c/2\}.

By Proposition 2.16, with polynomially high probability spec(W)\mathrm{spec}(W) is disjoint from Ω\Omega^{\prime}, so fa,bf_{a,b} is analytic on an open neighborhood of this domain. Theorem 2.18, applied with parameters c/2c/2 and C+c/2C+c/2, controls fx,yf_{x,y} and fy,xf_{y,x} on the part of Ω\Omega^{\prime} in the closed upper half-plane. For zz in the lower half-plane, Hermitian symmetry gives

fx,y(z)=fy,x(z¯)¯.f_{x,y}(z)=\overline{f_{y,x}(\overline{z})}.

Consequently,

supzΩ|fx,y(z)|n1/2.\sup_{z\in\Omega^{\prime}}|f_{x,y}(z)|\prec n^{-1/2}.

For any wΩw\in\Omega, let CwC_{w} be the circle of radius c/4c/4 centered at ww. This circle is contained in Ω\Omega^{\prime}, including when ww lies on or near the real axis. Cauchy’s integral formula therefore gives

supwΩ|fx,y(w)|=supwΩ|12πiCwfx,y(z)(zw)2dz|n1/2.\sup_{w\in\Omega}|f_{x,y}^{\prime}(w)|=\sup_{w\in\Omega}\left|\frac{1}{2\pi i}\oint_{C_{w}}\frac{f_{x,y}(z)}{(z-w)^{2}}\,dz\right|\prec n^{-1/2}.

Since ddzRW(z)=RW(z)2\frac{d}{dz}R_{W}(z)=-R_{W}(z)^{2}, we have

fx,y(z)=xRW(z)2yGμsc(z)xy,f_{x,y}^{\prime}(z)=-x^{*}R_{W}(z)^{2}y-G_{\mu_{\mathrm{sc}}}^{\prime}(z)x^{*}y,

which proves the result. ∎

2.3.4 Spiked matrix model estimates

Combining the above results, we may prove estimates on the largest eigenvalue and associated eigenvector of spiked matrix models with generalized Wigner matrix noise that will be useful later. These arguments for deducing such bounds from isotropic local laws are standard; we outline them omitting some details for the sake of completeness.

Theorem 2.21.

Let WW be a generalized Wigner matrix, θ>1\theta>1, and v𝕊n1()v\in\mathbb{S}^{n-1}(\mathbb{C}) be deterministic. Write λ^\widehat{\lambda} for the largest eigenvalue of H=θvv+WH=\theta vv^{*}+W and v^\widehat{v} for the associated eigenvector. Then, we have

|λ^λ(θ)|\displaystyle|\widehat{\lambda}-\lambda(\theta)| n1/2,\displaystyle\prec n^{-1/2},
||v^,v|2ρ(θ)2|\displaystyle\big||\langle\widehat{v},v\rangle|^{2}-\rho(\theta)^{2}\big| n1/2,\displaystyle\prec n^{-1/2},

with the constants in the polynomial stochastic domination depending only on θ\theta and the parameters (γW,ξW)(\gamma_{W},\xi_{W}) of the generalized Wigner matrix. Further, for any ε>0\varepsilon>0, with polynomially high probability HH has at most one eigenvalue in the interval [2+ε,)[2+\varepsilon,\infty).

Proof.

Write λ=λ(θ)=θ+θ1\lambda=\lambda(\theta)=\theta+\theta^{-1} and choose δ>0\delta>0 such that

I:=[λδ,λ+δ](2,).I\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}[\lambda-\delta,\lambda+\delta]\subset(2,\infty).

By Proposition 2.16, with polynomially high probability λ1(W)<λδ\lambda_{1}(W)<\lambda-\delta. On this event, the function

mW(E):=vRW(E)vm_{W}(E)\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}v^{*}R_{W}(E)v

is defined and strictly decreasing on II, since mW(E)=vRW(E)2v<0m_{W}^{\prime}(E)=-v^{*}R_{W}(E)^{2}v<0. Theorem 2.18 gives

supEI|mW(E)Gμsc(E)|n1/2.\sup_{E\in I}|m_{W}(E)-G_{\mu_{\mathrm{sc}}}(E)|\prec n^{-1/2}.

Proposition 2.15 shows that GμscG_{\mu_{\mathrm{sc}}} is strictly decreasing on II and

Gμsc(λ)=1θ.G_{\mu_{\mathrm{sc}}}(\lambda)=\frac{1}{\theta}.

The values of GμscG_{\mu_{\mathrm{sc}}} at the two endpoints of II lie on opposite sides of 1/θ1/\theta, with fixed gaps. The uniform local law therefore gives the same inequalities for mWm_{W} with polynomially high probability, so continuity and strict monotonicity show that the equation mW(E)=1/θm_{W}(E)=1/\theta has a unique solution in II. By Proposition 2.9, this solution is an eigenvalue of HH. The interlacing inequality gives λ2(H)λ1(W)\lambda_{2}(H)\leq\lambda_{1}(W), so this solution is the simple largest eigenvalue λ^\widehat{\lambda} of HH. Since GμscG_{\mu_{\mathrm{sc}}}^{\prime} is bounded away from zero on II, the mean value theorem and the uniform local law give

|λ^λ|n1/2.|\widehat{\lambda}-\lambda|\prec n^{-1/2}.

Proposition 2.11 and the secular equation give

|v^,v|2=(vRW(λ^)v)2vRW(λ^)2v=θ2vRW(λ^)2v.|\langle\widehat{v},v\rangle|^{2}=\frac{(v^{*}R_{W}(\widehat{\lambda})v)^{2}}{v^{*}R_{W}(\widehat{\lambda})^{2}v}=\frac{\theta^{-2}}{v^{*}R_{W}(\widehat{\lambda})^{2}v}.

Corollary 2.20, evaluated uniformly at λ^I\widehat{\lambda}\in I, yields

|vRW(λ^)2v1θ21|n1/2,\left|v^{*}R_{W}(\widehat{\lambda})^{2}v-\frac{1}{\theta^{2}-1}\right|\prec n^{-1/2},

where we also used |λ^λ|n1/2|\widehat{\lambda}-\lambda|\prec n^{-1/2} and the smoothness of GμscG_{\mu_{\mathrm{sc}}}^{\prime} on II. Therefore,

||v^,v|2(1θ2)|n1/2,\big||\langle\widehat{v},v\rangle|^{2}-(1-\theta^{-2})\big|\prec n^{-1/2},

which is the claimed estimate because ρ(θ)2=1θ2\rho(\theta)^{2}=1-\theta^{-2}.

Finally, for any fixed ε>0\varepsilon>0, Proposition 2.16 gives λ1(W)<2+ε\lambda_{1}(W)<2+\varepsilon with polynomially high probability. The interlacing inequality λ2(H)λ1(W)\lambda_{2}(H)\leq\lambda_{1}(W) then shows that HH has at most one eigenvalue in [2+ε,)[2+\varepsilon,\infty). ∎

3 Universality of entrywise statistics: Proof of Theorem 1.4

Recall the setting: we have two tuples of generalized Wigner matrices,

𝐖\displaystyle\mathbf{W} =(W(1),,W(L)),\displaystyle=(W^{(1)},\ldots,W^{(L)}),
𝐗\displaystyle\mathbf{X} =(X(1),,X(L)),\displaystyle=(X^{(1)},\ldots,X^{(L)}),

where each W(a),X(a)hermn×nW^{(a)},X^{(a)}\in\mathbb{C}^{n\times n}_{\mathrm{herm}} for a[L]a\in[L]. For 1pqn1\leq p\leq q\leq n, we write

w(pq)=(ReWpq(1),ImWpq(1),,ReWpq(L),ImWpq(L))2L.w^{(pq)}=(\mathrm{Re}W^{(1)}_{pq},\mathrm{Im}W^{(1)}_{pq},\ldots,\mathrm{Re}W^{(L)}_{pq},\mathrm{Im}W^{(L)}_{pq})\in\mathbb{R}^{2L}.

Thus the whole tuple 𝐖\mathbf{W} is equivalently encoded by 𝒘=(w(pq))1pqn\bm{w}=\left(w^{(pq)}\right)_{1\leq p\leq q\leq n}.

For p<qp<q, the (p,q)(p,q)-th entry of aa-th matrix is recovered as

Wpq(a)=ReWpq(a)+𝒊ImWpq(a)=w2a1(pq)+𝒊w2a(pq)W^{(a)}_{pq}=\mathrm{Re}W^{(a)}_{pq}+\bm{i}\mathrm{Im}W^{(a)}_{pq}=w^{(pq)}_{2a-1}+\bm{i}w^{(pq)}_{2a}

The lower-triangular entries are determined by Hermitian symmetry, and on the diagonal we have

Wpp(a)\displaystyle W^{(a)}_{pp} =w2a1(pp),\displaystyle=w^{(pp)}_{2a-1},
w2a(pp)\displaystyle w^{(pp)}_{2a} =0\displaystyle=0

For each a[L]a\in[L], we are given a deterministic vector v(a)𝕊n1()v^{(a)}\in\mathbb{S}^{n-1}(\mathbb{C}). On v(a)v^{(a)} we have only assumed that v(a)=1\|v^{(a)}\|=1 and that, for each a[L]a\in[L],

v(a)Cvn1/2+εv,\|v^{(a)}\|_{\infty}\leq C_{v}n^{-1/2+\varepsilon_{v}},

for some εv(0,1/20)\varepsilon_{v}\in(0,1/20). For a[L]a\in[L], we define

H(W,a)\displaystyle H^{(W,a)} =H(W(a))=θv(a)v(a)+W(a),\displaystyle=H(W^{(a)})=\theta v^{(a)}v^{(a)*}+W^{(a)},
λ^(W,a)\displaystyle\widehat{\lambda}^{(W,a)} =λ^(W(a))=λ1(H(W(a))),\displaystyle=\widehat{\lambda}(W^{(a)})=\lambda_{1}(H(W^{(a)})),
v^(W,a)\displaystyle\widehat{v}^{(W,a)} =v^(W(a))=v1(H(W(a))),\displaystyle=\widehat{v}(W^{(a)})=v_{1}(H(W^{(a)})),
Fij(𝒘)\displaystyle F_{ij}(\bm{w}) =ϕ(nv^i(W,1)v^j(W,1)¯,,nv^i(W,L)v^j(W,L)¯)\displaystyle=\phi\left(n\cdot\widehat{v}^{(W,1)}_{i}\cdot\overline{\widehat{v}^{(W,1)}_{j}},\ldots,n\cdot\widehat{v}^{(W,L)}_{i}\cdot\overline{\widehat{v}^{(W,L)}_{j}}\right)

for some fixed i,j[n]i,j\in[n] and ϕ:L\phi:\mathbb{C}^{L}\to\mathbb{R} five times continuously differentiable and with bounded values and derivatives. We view ϕ\phi as a function on 2L\mathbb{R}^{2L}. The analogous statements hold for the 𝐗\mathbf{X}-tuple, with W,wW,w replaced by X,xX,x.

We will use the notation defined earlier

λ=λ(θ):=θ+θ1>2.\lambda=\lambda(\theta)\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\theta+\theta^{-1}>2.

Our goal is to show that, uniformly in i,ji,j,

𝔼Fij(𝒘)𝔼Fij(𝒙).\mathbb{E}F_{ij}(\bm{w})\approx\mathbb{E}F_{ij}(\bm{x}).

We proceed in two steps. First, we make some initial simplifications using our calculations related to the eigenspaces of HH from Proposition 2.11. In particular, we reduce Fij(𝒘)F_{ij}(\bm{w}) to a simpler resolvent statistic F~ij(𝒘)\widetilde{F}_{ij}(\bm{w}). Then, we complete the remaining proof by the Lindeberg method.

3.1 Initial simplifications

Let ε(0,εv)\varepsilon\in(0,\varepsilon_{v}) be a constant for the purposes of the proof to be chosen later, and fix 0<δ1<δ20<\delta_{1}<\delta_{2} such that λ(2+δ1,2+δ2)\lambda\in(2+\delta_{1},2+\delta_{2}). We view ε,δ1,δ2,L\varepsilon,\delta_{1},\delta_{2},L as well as the function ϕ\phi and the generalized Wigner matrix parameters (γW,ξW,γX,ξX)(\gamma_{W},\xi_{W},\gamma_{X},\xi_{X}), all as constants for the purposes of this proof, and the dependence of various asymptotic notations on these parameters is not mentioned from now on.

For each a[L]a\in[L], let (a)\mathcal{E}^{(a)} be the event that the following conditions hold for W(a)W^{(a)}:

  1. 1.

    There is a single eigenvalue of H(W,a)H^{(W,a)} in [2+δ1,2+δ2)[2+\delta_{1},2+\delta_{2}), of multiplicity 1, which is also the top eigenvalue λ^(W,a)=λ1(H(W(a)))\widehat{\lambda}^{(W,a)}=\lambda_{1}(H(W^{(a)})).

  2. 2.

    |λ^(W,a)λ|n1/2+ε|\widehat{\lambda}^{(W,a)}-\lambda|\leq n^{-1/2+\varepsilon}.

  3. 3.

    W(a)<2+δ1\|W^{(a)}\|<2+\delta_{1}, and in particular λ^(W,a)spec(W(a))\widehat{\lambda}^{(W,a)}\notin\mathrm{spec}(W^{(a)}).

  4. 4.

    The following bounds hold:

    |eiRW(a)(λ^(W,a))v(a)Gμsc(λ^(W,a))vi(a)|\displaystyle|e_{i}^{*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})v^{(a)}-G_{\mu_{\mathrm{sc}}}(\widehat{\lambda}^{(W,a)})v^{(a)}_{i}| n1/2+ε,\displaystyle\leq n^{-1/2+\varepsilon},
    |v(a)RW(a)(λ^(W,a))ejGμsc(λ^(W,a))vj(a)¯|\displaystyle|v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})e_{j}-G_{\mu_{\mathrm{sc}}}(\widehat{\lambda}^{(W,a)})\overline{v^{(a)}_{j}}| n1/2+ε,\displaystyle\leq n^{-1/2+\varepsilon},
    |v(a)RW(a)(λ^(W,a))2v(a)+Gμsc(λ^(W,a))|\displaystyle|v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})^{2}v^{(a)}+G_{\mu_{\mathrm{sc}}}^{\prime}(\widehat{\lambda}^{(W,a)})| n1/2+ε.\displaystyle\leq n^{-1/2+\varepsilon}.
Proposition 3.1.

:=a=1L(a)\mathcal{E}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\bigcap_{a=1}^{L}\mathcal{E}^{(a)} holds with polynomially high probability.

Proof.

The result follows by combining the isotropic local law (Theorem 2.18), its derivative (Corollary 2.20), and Theorem 2.21 on the eigenvalues of spiked matrix models. For each fixed a[L]a\in[L], this follows from the corresponding scalar estimates. Since LL is fixed, taking the intersection over a[L]a\in[L] only changes the implicit constants. ∎

Define

Λ1\displaystyle\Lambda_{1} :=Gμsc(2+δ2),\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}-G_{\mu_{\mathrm{sc}}}^{\prime}(2+\delta_{2}),
Λ2\displaystyle\Lambda_{2} :=Gμsc(2+δ1).\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}-G_{\mu_{\mathrm{sc}}}^{\prime}(2+\delta_{1}).

By Proposition 2.15, we have 0<Λ1<Λ20<\Lambda_{1}<\Lambda_{2} and Gμsc(λ)[Λ1,Λ2]-G_{\mu_{\mathrm{sc}}}^{\prime}(\lambda)\in[\Lambda_{1},\Lambda_{2}] since Gμsc(z)-G_{\mu_{\mathrm{sc}}}^{\prime}(z) is monotonically decreasing on the interval (2,)(2,\infty).

Since ϕ\phi is bounded, we have by Proposition 3.1 the bound:

|𝔼Fij(𝒘)𝔼𝟏Fij(𝒘)|ϕL[c]n1.|\mathbb{E}F_{ij}(\bm{w})-\mathbb{E}\mathbf{1}_{\mathcal{E}}F_{ij}(\bm{w})|\leq\|\phi\|_{L^{\infty}}\mathbb{P}[\mathcal{E}^{c}]\lesssim n^{-1}. (3.1)

Applying Proposition 2.11 to each coordinate a[L]a\in[L], we obtain:

𝔼𝟏Fij(𝒘)\displaystyle\mathbb{E}\mathbf{1}_{\mathcal{E}}F_{ij}(\bm{w}) =𝔼𝟏ϕ((nv^i(W,a)v^j(W,a)¯)a=1L)\displaystyle=\mathbb{E}\mathbf{1}_{\mathcal{E}}\phi\left(\left(n\cdot\widehat{v}^{(W,a)}_{i}\cdot\overline{\widehat{v}^{(W,a)}_{j}}\right)_{a=1}^{L}\right)
=𝔼𝟏ϕ((n(eiRW(a)(λ^(W,a))v(a))(v(a)RW(a)(λ^(W,a))ej)v(a)RW(a)(λ^(W,a))2v(a))a=1L).\displaystyle=\mathbb{E}\mathbf{1}_{\mathcal{E}}\phi\left(\left(n\cdot\frac{\left(e_{i}^{*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})v^{(a)}\right)\left(v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})e_{j}\right)}{v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})^{2}v^{(a)}}\right)_{a=1}^{L}\right).

We will now show that 𝔼Fij(𝒘)\mathbb{E}F_{ij}(\bm{w}) is close to what we get when we perform two operations on the above expression: (1) we replace the random eigenvalues λ^(W,a)\widehat{\lambda}^{(W,a)} by their deterministic typical location λ\lambda (which is the same for each a[L]a\in[L]), and (2) we apply the isotropic local law to the denominator. There is a small but important caveat that we must attend to:

Remark 3.2 (Resolvent for discrete models).

When the entries of W(a)W^{(a)} have discrete distributions, which our Definition 1.1 does not rule out, it is possible that λspec(W(a))\lambda\in\mathrm{spec}(W^{(a)}) with small but positive probability. On this event, RW(a)(λ)R_{W^{(a)}}(\lambda) is undefined, and in particular expectations involving RW(a)(λ)R_{W^{(a)}}(\lambda) are undefined.

To deal with this, consider the slight further perturbation

λ~=λ~(θ)=λ(θ)+𝒊n1/2.\widetilde{\lambda}=\widetilde{\lambda}(\theta)=\lambda(\theta)+\bm{i}n^{-1/2}.

For the above technical reason, we instead consider replacing λ^(W,a)\widehat{\lambda}^{(W,a)} by λ~\widetilde{\lambda}.

It is convenient to state this reduction in terms of the following auxiliary function:

ϕ~(t1,,tL):=ϕ(1Gμsc(λ~)t1,,1Gμsc(λ~)tL).\widetilde{\phi}(t_{1},\ldots,t_{L})\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\phi\left(\frac{-1}{G^{\prime}_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})}t_{1},\dots,\frac{-1}{G^{\prime}_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})}t_{L}\right).

Since 1/Gμsc(λ~)-1/G^{\prime}_{\mu_{\mathrm{sc}}}(\widetilde{\lambda}) is uniformly bounded for all sufficiently large nn, ϕ~\widetilde{\phi} satisfies the same boundedness and smoothness assumptions as ϕ\phi up to modifying the constants involved (in a way not depending on nn). We then define

F~ij(𝒘):=ϕ~((n(eiRW(a)(λ~)v(a))(v(a)RW(a)(λ~)ej))a=1L).\widetilde{F}_{ij}(\bm{w})\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\widetilde{\phi}\left(\left(n\cdot(e_{i}^{*}R_{W^{(a)}}(\widetilde{\lambda})v^{(a)})(v^{(a)*}R_{W^{(a)}}(\widetilde{\lambda})e_{j})\right)_{a=1}^{L}\right). (3.2)
Lemma 3.3.

For all i,j[n]i,j\in[n], the following bounds hold:

|𝔼Fij(𝒘)𝔼F~ij(𝒘)|\displaystyle|\mathbb{E}F_{ij}(\bm{w})-\mathbb{E}\widetilde{F}_{ij}(\bm{w})| n1/2+3εv,\displaystyle\lesssim n^{-1/2+3\varepsilon_{v}},
|𝔼Fij(𝒙)𝔼F~ij(𝒙)|\displaystyle|\mathbb{E}F_{ij}(\bm{x})-\mathbb{E}\widetilde{F}_{ij}(\bm{x})| n1/2+3εv,\displaystyle\lesssim n^{-1/2+3\varepsilon_{v}},
and therefore also
|𝔼Fij(𝒘)𝔼Fij(𝒙)|\displaystyle|\mathbb{E}F_{ij}(\bm{w})-\mathbb{E}F_{ij}(\bm{x})| |𝔼F~ij(𝒘)𝔼F~ij(𝒙)|+O(n1/2+3εv).\displaystyle\leq|\mathbb{E}\widetilde{F}_{ij}(\bm{w})-\mathbb{E}\widetilde{F}_{ij}(\bm{x})|+O(n^{-1/2+3\varepsilon_{v}}).
Proof.

We will show the first bound for 𝒘\bm{w}; the same argument will apply symmetrically for 𝒙\bm{x}, and the two bounds combined will give the third bound.

Throughout the proof, we further intersect \mathcal{E} with the polynomially high probability event on which Corollary 2.20 holds uniformly on the fixed domains used below, for the finitely many deterministic pairs

(x,y){(ei,v(a)),(v(a),ej):a[L]}.(x,y)\in\{(e_{i},v^{(a)}),(v^{(a)},e_{j}):a\in[L]\}.

Since LL is fixed, this new event still holds with polynomially high probability, and we keep the notation \mathcal{E} for it.

We first replace the denominator by its deterministic equivalent. By the conditions included in \mathcal{E}, we have that, on this event, for a[L]a\in[L],

|1v(a)RW(a)(λ^(W,a))2v(a)+1Gμsc(λ~)|\displaystyle\left|\frac{1}{v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})^{2}v^{(a)}}+\frac{1}{G^{\prime}_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})}\right|
|1v(a)RW(a)(λ^(W,a))2v(a)+1Gμsc(λ^(W,a))|+|1Gμsc(λ~)1Gμsc(λ^(W,a))|\displaystyle\hskip 56.9055pt\leq\left|\frac{1}{v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})^{2}v^{(a)}}+\frac{1}{G^{\prime}_{\mu_{\mathrm{sc}}}(\widehat{\lambda}^{(W,a)})}\right|+\left|\frac{1}{G^{\prime}_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})}-\frac{1}{G^{\prime}_{\mu_{\mathrm{sc}}}(\widehat{\lambda}^{(W,a)})}\right|
=|v(a)RW(a)(λ^(W,a))2v(a)+Gμsc(λ^(W,a))||v(a)RW(a)(λ^(W,a))2v(a)||Gμsc(λ^(W,a))|+O(n1/2+ε)\displaystyle\hskip 56.9055pt=\frac{|v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})^{2}v^{(a)}+G^{\prime}_{\mu_{\mathrm{sc}}}(\widehat{\lambda}^{(W,a)})|}{|v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})^{2}v^{(a)}|\cdot|G^{\prime}_{\mu_{\mathrm{sc}}}(\widehat{\lambda}^{(W,a)})|}+O(n^{-1/2+\varepsilon})
n1/2+ε(Λ1n1/2+ε)Λ1+O(n1/2+ε)\displaystyle\hskip 56.9055pt\leq\frac{n^{-1/2+\varepsilon}}{(\Lambda_{1}-n^{-1/2+\varepsilon})\Lambda_{1}}+O(n^{-1/2+\varepsilon})
=O(n1/2+ε),\displaystyle\hskip 56.9055pt=O(n^{-1/2+\varepsilon}),

where we also use that 1/Gμsc(z)1/G^{\prime}_{\mu_{\mathrm{sc}}}(z) is Lipschitz away from {±2}\{\pm 2\}, per Proposition 2.15.

Moreover, on \mathcal{E}, the isotropic local law bounds of Theorem 2.18 give

|eiRW(a)(λ^(W,a))v(a)||Gμsc(λ^(W,a))||vi(a)|+n1/2+εn1/2+εv.|e_{i}^{*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})v^{(a)}|\leq|G_{\mu_{\mathrm{sc}}}(\widehat{\lambda}^{(W,a)})||v_{i}^{(a)}|+n^{-1/2+\varepsilon}\lesssim n^{-1/2+\varepsilon_{v}}.

Similarly

|v(a)RW(a)(λ^(W,a))ej|n1/2+εv.|v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})e_{j}|\lesssim n^{-1/2+\varepsilon_{v}}.

Therefore, we also have that on \mathcal{E},

|n(eiRW(a)(λ^(W,a))v(a))(v(a)RW(a)(λ^(W,a))ej)v(a)RW(a)(λ^(W,a))2v(a)\displaystyle\Bigg|n\cdot\frac{\left(e_{i}^{*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})v^{(a)}\right)\left(v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})e_{j}\right)}{v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})^{2}v^{(a)}}
+n(eiRW(a)(λ^(W,a))v(a))(v(a)RW(a)(λ^(W,a))ej)Gμsc(λ~)|\displaystyle\hskip 28.45274pt+n\cdot\frac{\left(e_{i}^{*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})v^{(a)}\right)\left(v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})e_{j}\right)}{G_{\mu_{\mathrm{sc}}}^{\prime}(\widetilde{\lambda})}\Bigg|
nn1+2εvn1/2+ε=O(n1/2+ε+2εv).\displaystyle\hskip 56.9055pt\lesssim n\cdot n^{-1+2\varepsilon_{v}}\cdot n^{-1/2+\varepsilon}=O(n^{-1/2+\varepsilon+2\varepsilon_{v}}).

Since ϕ\phi is Lipschitz as a function on 2L\mathbb{R}^{2L} and LL is fixed, it follows that

𝔼𝟏Fij(𝒘)\displaystyle\mathbb{E}\mathbf{1}_{\mathcal{E}}F_{ij}(\bm{w}) =𝔼𝟏ϕ((n1Gμsc(λ~)(eiRW(a)(λ^(W,a))v(a))(v(a)RW(a)(λ^(W,a))ej))a=1L)\displaystyle=\mathbb{E}\mathbf{1}_{\mathcal{E}}\phi\left(\left(n\cdot\frac{-1}{G_{\mu_{\mathrm{sc}}}^{\prime}(\widetilde{\lambda})}(e_{i}^{*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})v^{(a)})(v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})e_{j})\right)_{a=1}^{L}\right)
+O(n1/2+ε+2εv).\displaystyle\hskip 56.9055pt+O(n^{-1/2+\varepsilon+2\varepsilon_{v}}).

We next replace λ^(W,a)\widehat{\lambda}^{(W,a)} by λ~\widetilde{\lambda} in the two numerator factors. Decompose

|eiRW(a)(λ~)v(a)\displaystyle|e_{i}^{*}R_{W^{(a)}}(\widetilde{\lambda})v^{(a)} eiRW(a)(λ^(W,a))v(a)||Gμsc(λ~)Gμsc(λ^(W,a))||vi(a)|\displaystyle-e_{i}^{*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})v^{(a)}|\leq|G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})-G_{\mu_{\mathrm{sc}}}(\widehat{\lambda}^{(W,a)})||v^{(a)}_{i}|
+|(eiRW(a)(λ~)v(a)Gμsc(λ~)vi(a))(eiRW(a)(λ^(W,a))v(a)Gμsc(λ^(W,a))vi(a))|.\displaystyle+\left|\left(e_{i}^{*}R_{W^{(a)}}(\widetilde{\lambda})v^{(a)}-G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})v^{(a)}_{i}\right)-\left(e_{i}^{*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})v^{(a)}-G_{\mu_{\mathrm{sc}}}(\widehat{\lambda}^{(W,a)})v^{(a)}_{i}\right)\right|.

Consider the neighborhood

Ω~:={z:|zλ|14(λ(2+δ1))}.\widetilde{\Omega}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\{z\in\mathbb{C}:|z-\lambda|\leq\frac{1}{4}(\lambda-(2+\delta_{1}))\}.

Choose c,C>0c,C>0 such that

Ω~{Imz>0}Ω(c,C)={z:Rez>2+c,Imz>0,|z|C}\widetilde{\Omega}\cap\{\mathrm{Im}z>0\}\subset\Omega(c,C)=\{z\in\mathbb{C}:\mathrm{Re}z>2+c,\mathrm{Im}z>0,|z|\leq C\}

On \mathcal{E}, for sufficiently large nn, the segment joining λ^(W,a)\widehat{\lambda}^{(W,a)} and λ~\widetilde{\lambda} except possibly the real endpoint λ^(W,a)\widehat{\lambda}^{(W,a)} is contained in Ω~{Imz>0}\widetilde{\Omega}\cap\{\mathrm{Im}z>0\} and is separated from spec(W(a))\mathrm{spec}(W^{(a)}) by a deterministic positive distance. The derivative local law, Corollary 2.20, applies on the part of the segment with positive imaginary part, and the same bound holds at the real endpoint by continuity.

By the fundamental theorem of calculus along the segment [λ^(W,a),λ~][\widehat{\lambda}^{(W,a)},\widetilde{\lambda}], and by Corollary 2.20,

|(eiRW(a)(λ~)v(a)Gμsc(λ~)vi(a))(eiRW(a)(λ^(W,a))v(a)Gμsc(λ^(W,a))vi(a))|\displaystyle\left|\left(e_{i}^{*}R_{W^{(a)}}(\widetilde{\lambda})v^{(a)}-G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})v^{(a)}_{i}\right)-\left(e_{i}^{*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})v^{(a)}-G_{\mu_{\mathrm{sc}}}(\widehat{\lambda}^{(W,a)})v^{(a)}_{i}\right)\right|
|λ~λ^(W,a)|supz[λ^(W,a),λ~]|eiRW(a)(z)2v(a)+Gμsc(z)vi(a)||λ~λ^(W,a)|n1/2+ε.\displaystyle\qquad\leq|\widetilde{\lambda}-\widehat{\lambda}^{(W,a)}|\sup_{z\in[\widehat{\lambda}^{(W,a)},\widetilde{\lambda}]}\left|e^{*}_{i}R_{W^{(a)}}(z)^{2}v^{(a)}+G_{\mu_{\mathrm{sc}}}^{\prime}(z)v^{(a)}_{i}\right|\lesssim|\widetilde{\lambda}-\widehat{\lambda}^{(W,a)}|\cdot n^{-1/2+\varepsilon}.

Using GμscG_{\mu_{\mathrm{sc}}} is Lipschitz and |λ~λ^(W,a)|n1/2+ε|\widetilde{\lambda}-\widehat{\lambda}^{(W,a)}|\lesssim n^{-1/2+\varepsilon}, we obtain

|eiRW(a)(λ~)v(a)eiRW(a)(λ^(W,a))v(a)|\displaystyle|e_{i}^{*}R_{W^{(a)}}(\widetilde{\lambda})v^{(a)}-e_{i}^{*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})v^{(a)}| |Gμsc(λ~)Gμsc(λ^(W,a))||vi(a)|+O(|λ~λ^(W,a)|n1/2+ε)\displaystyle\leq|G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})-G_{\mu_{\mathrm{sc}}}(\widehat{\lambda}^{(W,a)})||v^{(a)}_{i}|+O(|\widetilde{\lambda}-\widehat{\lambda}^{(W,a)}|n^{-1/2+\varepsilon})
n1/2+εn1/2+εv+n1/2+εn1/2+ε\displaystyle\lesssim n^{-1/2+\varepsilon}n^{-1/2+\varepsilon_{v}}+n^{-1/2+\varepsilon}n^{-1/2+\varepsilon}
n1+ε+εv\displaystyle\lesssim n^{-1+\varepsilon+\varepsilon_{v}}

Also applying this symmetrically to the v(a)RW(a)(λ^(W,a))ejv^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})e_{j} term, we find that,

n|(eiRW(a)(λ~)v(a))(v(a)RW(a)(λ~)ej)(eiRW(a)(λ^(W,a))v(a))(v(a)RW(a)(λ^(W,a))ej)|\displaystyle n\left|(e_{i}^{*}R_{W^{(a)}}(\widetilde{\lambda})v^{(a)})(v^{(a)*}R_{W^{(a)}}(\widetilde{\lambda})e_{j})-(e_{i}^{*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})v^{(a)})(v^{(a)*}R_{W^{(a)}}(\widehat{\lambda}^{(W,a)})e_{j})\right|
nn1+ε+εvn1/2+εv+nn1+ε+εvn1+ε+εv\displaystyle\qquad\lesssim n\cdot n^{-1+\varepsilon+\varepsilon_{v}}\cdot n^{-1/2+\varepsilon_{v}}+n\cdot n^{-1+\varepsilon+\varepsilon_{v}}\cdot n^{-1+\varepsilon+\varepsilon_{v}}
n1/2+ε+2εv.\displaystyle\qquad\lesssim n^{-1/2+\varepsilon+2\varepsilon_{v}}.

Since ϕ\phi is smooth and ϕ\phi^{\prime} is bounded, and combining with (3.1),

𝔼Fij(𝒘)\displaystyle\mathbb{E}F_{ij}(\bm{w}) =𝔼𝟏ϕ((n1Gμsc(λ~)(eiRW(a)(λ~)v(a))(v(a)RW(a)(λ~)ej))a=1L)+O(n1/2+ε+2εv)\displaystyle=\mathbb{E}\mathbf{1}_{\mathcal{E}}\phi\left(\left(n\cdot\frac{-1}{G^{\prime}_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})}\cdot(e_{i}^{*}R_{W^{(a)}}(\widetilde{\lambda})v^{(a)})(v^{(a)*}R_{W^{(a)}}(\widetilde{\lambda})e_{j})\right)_{a=1}^{L}\right)+O(n^{-1/2+\varepsilon+2\varepsilon_{v}})
and, undoing our first truncation step where we introduced the indicator of \mathcal{E}, we have
=𝔼ϕ((n1Gμsc(λ~)(eiRW(a)(λ~)v(a))(v(a)RW(a)(λ~)ej))a=1L)+O(n1/2+ε+2εv)\displaystyle=\mathbb{E}\phi\left(\left(n\cdot\frac{-1}{G^{\prime}_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})}\cdot(e_{i}^{*}R_{W^{(a)}}(\widetilde{\lambda})v^{(a)})(v^{(a)*}R_{W^{(a)}}(\widetilde{\lambda})e_{j})\right)_{a=1}^{L}\right)+O(n^{-1/2+\varepsilon+2\varepsilon_{v}})
=𝔼ϕ~((n(eiRW(a)(λ~)v(a))(v(a)RW(a)(λ~)ej))a=1L)+O(n1/2+ε+2εv),\displaystyle=\mathbb{E}\widetilde{\phi}\left(\left(n\cdot(e_{i}^{*}R_{W^{(a)}}(\widetilde{\lambda})v^{(a)})(v^{(a)*}R_{W^{(a)}}(\widetilde{\lambda})e_{j})\right)_{a=1}^{L}\right)+O(n^{-1/2+\varepsilon+2\varepsilon_{v}}), (3.3)

completing the proof, where we note that the stated error bound follows since we have taken ε<εv\varepsilon<\varepsilon_{v} by assumption. ∎

3.2 Lindeberg method

Now, we consider comparing 𝔼F~ij(𝒘)\mathbb{E}\widetilde{F}_{ij}(\bm{w}) and 𝔼F~ij(𝒙)\mathbb{E}\widetilde{F}_{ij}(\bm{x}) by the Lindeberg method. Our final result will be the following:

Lemma 3.4.

For F~ij\widetilde{F}_{ij} as defined above in (3.2), for any ε>0\varepsilon>0,

maxi,j[n]|𝔼F~ij(𝒙)𝔼F~ij(𝒘)|εn1/2+10εv+ε.\max_{i,j\in[n]}|\mathbb{E}\widetilde{F}_{ij}(\bm{x})-\mathbb{E}\widetilde{F}_{ij}(\bm{w})|\lesssim_{\varepsilon}n^{-1/2+10\varepsilon_{v}+\varepsilon}.

The proof of Theorem 1.4 is then completed by combining Lemma 3.3 with Lemma 3.4.

Following the prescription of the Lindeberg method (our use of the method is also quite similar to the specific one in [KY13a]), let us construct an interpolating path from tuple 𝐖\mathbf{W} to tuple 𝐗\mathbf{X}. Let N=n(n+1)2N=\frac{n(n+1)}{2} be the number of entries on and above the diagonal in a Hermitian matrix. Fix some enumeration (p1,q1),,(pN,qN)(p_{1},q_{1}),\dots,(p_{N},q_{N}) of all pairs (p,q)(p_{\ell},q_{\ell}) with 1pqn1\leq p\leq q\leq n. Throughout this argument, i,j[n]i,j\in[n] are fixed and refer to the entry statistic FijF_{ij}. By contrast, at the \ell-th Lindeberg step we write p,qp_{\ell},q_{\ell} for the matrix entry currently being swapped.

For 0N0\leq\ell\leq N, let

𝐖()=(W(,1),,W(,L))\mathbf{W}^{(\ell)}=\left(W^{(\ell,1)},\ldots,W^{(\ell,L)}\right)

be the tuple obtained from 𝐖\mathbf{W} by replacing the entries Wp1q1(a),,Wpq(a)W^{(a)}_{p_{1}q_{1}},\dots,W^{(a)}_{p_{\ell}q_{\ell}} with the corresponding entries of 𝐗\mathbf{X} for all a[L]a\in[L], and by their conjugates symmetrically below the diagonal. Thus, 𝐖(0)=𝐖\mathbf{W}^{(0)}=\mathbf{W}, 𝐖(N)=𝐗\mathbf{W}^{(N)}=\mathbf{X}. We emphasize that, while we work with LL different Hermitian matrices with a total of LNL\cdot N entries (possibly complex-valued), our Lindeberg procedure only involves NN steps: for each fixed upper-triangular coordinate (p,q)(p_{\ell},q_{\ell}), in one step the whole vector w(pq)2Lw^{(p_{\ell}q_{\ell})}\in\mathbb{R}^{2L} is replaced by x(pq)2Lx^{(p_{\ell}q_{\ell})}\in\mathbb{R}^{2L}. We swap the whole block at once because the coordinates across the LL matrices may be correlated within the same upper-triangular position, while the blocks are independent across different positions. This also ensures that, after the current block is removed, the tuple 𝐐()\mathbf{Q}^{(\ell)} introduced below is independent of both w(pq)w^{(p_{\ell}q_{\ell})} and x(pq)x^{(p_{\ell}q_{\ell})}.

We follow the usual prescription of the Lindeberg method. First, we bound by a telescoping sum,

|𝔼F~ij(𝒙)𝔼F~ij(𝒘)|\displaystyle|\mathbb{E}\widetilde{F}_{ij}(\bm{x})-\mathbb{E}\widetilde{F}_{ij}(\bm{w})| =|=1N[𝔼F~ij(𝐖())𝔼F~ij(𝐖(1))]|\displaystyle=\left|\sum_{\ell=1}^{N}\left[\mathbb{E}\widetilde{F}_{ij}(\mathbf{W}^{(\ell)})-\mathbb{E}\widetilde{F}_{ij}(\mathbf{W}^{(\ell-1)})\right]\right|
=1N|𝔼F~ij(𝐖())𝔼F~ij(𝐖(1))|.\displaystyle\leq\sum_{\ell=1}^{N}|\mathbb{E}\widetilde{F}_{ij}(\mathbf{W}^{(\ell)})-\mathbb{E}\widetilde{F}_{ij}(\mathbf{W}^{(\ell-1)})|. (3.4)

We now consider bounding the effect of each individual swap, i.e., the size of

|𝔼F~ij(𝐖())𝔼F~ij(𝐖(1))||\mathbb{E}\widetilde{F}_{ij}(\mathbf{W}^{(\ell)})-\mathbb{E}\widetilde{F}_{ij}(\mathbf{W}^{(\ell-1)})| (3.5)

for each fixed \ell. For the sake of brevity, let us write p=pp=p_{\ell} and q=qq=q_{\ell} while analyzing a single step. Define

𝐄pq(W):=(Epq(W,1),,Epq(W,L)),Epq(W,a):={Wpq(a)epeq+Wpq(a)¯eqepif pq,Wpp(a)epepif p=q.},\mathbf{E}^{(W)}_{pq}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left(E^{(W,1)}_{pq},\ldots,E^{(W,L)}_{pq}\right),\qquad\begin{aligned} E^{(W,a)}_{pq}&\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left\{\begin{array}[]{ll}W^{(a)}_{pq}e_{p}e_{q}^{*}+\overline{W^{(a)}_{pq}}e_{q}e_{p}^{*}&\text{if }p\neq q,\\ W^{(a)}_{pp}e_{p}e_{p}^{*}&\text{if }p=q.\end{array}\right\},\end{aligned}

and similarly 𝐄pq(X)\mathbf{E}^{(X)}_{pq}, Epq(X,a)E^{(X,a)}_{pq}. Then the original tuples decompose as

W(a)=1μνnEμν(W,a),X(a)=1μνnEμν(X,a),a[L].W^{(a)}=\sum_{1\leq\mu\leq\nu\leq n}E^{(W,a)}_{\mu\nu},\qquad X^{(a)}=\sum_{1\leq\mu\leq\nu\leq n}E^{(X,a)}_{\mu\nu},\qquad a\in[L].

Let

Q(,a)=W(,a)Epq(X,a)=W(1,a)Epq(W,a),a[L],Q^{(\ell,a)}=W^{(\ell,a)}-E^{(X,a)}_{pq}=W^{(\ell-1,a)}-E^{(W,a)}_{pq},\qquad a\in[L],

and write

𝐐():=(Q(,1),,Q(,L)).\mathbf{Q}^{(\ell)}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left(Q^{(\ell,1)},\ldots,Q^{(\ell,L)}\right).

For notational convenience, we suppress the superscript \ell and simply write

𝐐=(Q(1),,Q(L)).\mathbf{Q}=\left(Q^{(1)},\ldots,Q^{(L)}\right).

Thus,

𝐖()=𝐐+𝐄pq(X),𝐖(1)=𝐐+𝐄pq(W).\mathbf{W}^{(\ell)}=\mathbf{Q}+\mathbf{E}^{(X)}_{pq},\qquad\mathbf{W}^{(\ell-1)}=\mathbf{Q}+\mathbf{E}^{(W)}_{pq}.

In words, 𝐐\mathbf{Q} is the tuple after the first 1\ell-1 swaps have been performed but where the entries about to be swapped at step \ell (in positions (p=p,q=q)(p=p_{\ell},q=q_{\ell}) and (q,p)(q,p)) have been set to zero.

It is useful to pause to understand the distribution of 𝐐\mathbf{Q}. For each fixed a[L]a\in[L], the matrix Q(a)Q^{(a)} is obtained by taking the entries indexed by (pr,qr)(p_{r},q_{r}) with r<r<\ell from X(a)X^{(a)}, the entries indexed by (pr,qr)(p_{r},q_{r}) with r>r>\ell from W(a)W^{(a)}, and setting the current entries in positions (p=p,q=q)(p=p_{\ell},q=q_{\ell}) and (q,p)(q,p) to zero. Adding back either Epq(W,a)E^{(W,a)}_{pq} or Epq(X,a)E^{(X,a)}_{pq} gives respectively W(1,a)W^{(\ell-1,a)} or W(,a)W^{(\ell,a)}, both of which are generalized Wigner matrices with the same variance normalization, since the second moments of w(pq)w^{(pq)} and x(pq)x^{(pq)} match. Hence Corollary 2.19 applies to each Q(a)Q^{(a)}. Since LL is fixed, we may use these local law bounds simultaneously for all a[L]a\in[L].

Moreover, by independence over upper-triangular index pairs, 𝐐\mathbf{Q} is independent of the current entry vectors w(pq),x(pq)2Lw^{(pq)},x^{(pq)}\in\mathbb{R}^{2L}.

Rewriting our expression in (3.5) in terms of this 𝐐\mathbf{Q}, we have

|𝔼F~ij(𝐖())𝔼F~ij(𝐖(1))|=|𝔼F~ij(𝐐+𝐄pq(X))𝔼F~ij(𝐐+𝐄pq(W))|,\left|\mathbb{E}\widetilde{F}_{ij}(\mathbf{W}^{(\ell)})-\mathbb{E}\widetilde{F}_{ij}(\mathbf{W}^{(\ell-1)})\right|=\left|\mathbb{E}\widetilde{F}_{ij}\left(\mathbf{Q}+\mathbf{E}^{(X)}_{pq}\right)-\mathbb{E}\widetilde{F}_{ij}\left(\mathbf{Q}+\mathbf{E}^{(W)}_{pq}\right)\right|,

where the addition of tuples is understood coordinatewise.

We would like to understand the leading order effect of this difference, viewed as a small perturbation. We do this by first taking a perturbative expansion of each term separately. Recall that F~ij\widetilde{F}_{ij} is a function of 𝐖\mathbf{W} only through the resolvent value RW(a)(λ~)R_{W^{(a)}}(\widetilde{\lambda}) for a[L]a\in[L], where λ~=θ+θ1+𝒊n1/2\widetilde{\lambda}=\theta+\theta^{-1}+\bm{i}n^{-1/2} is a complex constant. Let us note in passing that this constant falls in the regions treated in the local laws (Theorem 2.18 and Corollary 2.19) for suitable choices of the parameters there. All resolvents are evaluated at this λ~\widetilde{\lambda} in the following calculations, so for W(a)W^{(a)} and all other matrices involved, let us abbreviate

RW(a):=RW(a)(λ~).R_{W^{(a)}}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}R_{W^{(a)}}(\widetilde{\lambda}).

We first compute the effect on the resolvent of the above perturbations of 𝐐\mathbf{Q}. By the resolvent expansion of Proposition 2.8, for each a[L]a\in[L], we have

RW(1,a)=RQ(a)+Epq(W,a)=s=04RQ(a)(Epq(W,a)RQ(a))s+RW(1,a)(Epq(W,a)RQ(a))5.R_{W^{(\ell-1,a)}}=R_{Q^{(a)}+E^{(W,a)}_{pq}}=\sum_{s=0}^{4}R_{Q^{(a)}}\left(E^{(W,a)}_{pq}R_{Q^{(a)}}\right)^{s}+R_{W^{(\ell-1,a)}}\left(E^{(W,a)}_{pq}R_{Q^{(a)}}\right)^{5}. (3.6)

We will see momentarily why taking the expansion to order 4 is the correct choice here.

To work with these expressions, let us set up some notation for what happens when we expand the Epq(W,a)E^{(W,a)}_{pq}.

Definition 3.5 (Conjugation words).

Write id:\mathrm{id}:\mathbb{C}\to\mathbb{C} for the identity map id(z)=z\mathrm{id}(z)=z, and conj:\mathrm{conj}:\mathbb{C}\to\mathbb{C} for the conjugation map conj(z)=z¯\mathrm{conj}(z)=\overline{z}. For a given 1pqn1\leq p\leq q\leq n, define

𝒮p,q,s={{id,conj}sif pq,{id}sif p=q}.\mathcal{S}_{p,q,s}=\left\{\begin{array}[]{ll}\{\mathrm{id},\mathrm{conj}\}^{s}&\text{if }p\neq q,\\ \{\mathrm{id}\}^{s}&\text{if }p=q\end{array}\right\}.

For S=(f1,,fs)𝒮p,q,sS=(f_{1},\dots,f_{s})\in\mathcal{S}_{p,q,s}, viewed as a string of functions of length ss, let

zS:=f1(z)fs(z).z^{S}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}f_{1}(z)\cdots f_{s}(z).

Also, associate to such SS indices αS,1,βS,1,,αS,s,βS,s{p,q}\alpha_{S,1},\beta_{S,1},\dots,\alpha_{S,s},\beta_{S,s}\in\{p,q\}, so that αS,i=p\alpha_{S,i}=p and βS,i=q\beta_{S,i}=q if fi=idf_{i}=\mathrm{id}, while αS,i=q\alpha_{S,i}=q and βS,i=p\beta_{S,i}=p if fi=conjf_{i}=\mathrm{conj}.

Then, for a[L]a\in[L] and s1s\geq 1, we may expand

(Epq(W,a)RQ(a))s=S𝒮p,q,s(Wpq(a))SeαS,1(eβS,1RQ(a)eαS,2)(eβS,s1RQ(a)eαS,s)eβS,sRQ(a),\left(E^{(W,a)}_{pq}R_{Q^{(a)}}\right)^{s}=\sum_{S\in\mathcal{S}_{p,q,s}}(W^{(a)}_{pq})^{S}e_{\alpha_{S,1}}\left(e_{\beta_{S,1}}^{*}R_{Q^{(a)}}e_{\alpha_{S,2}}\right)\cdots\left(e_{\beta_{S,{s-1}}}^{*}R_{Q^{(a)}}e_{\alpha_{S,s}}\right)e_{\beta_{S,s}}^{*}R_{Q^{(a)}},

where we note that the quantities in parentheses are just scalars giving certain entries of RQ(a)R_{Q^{(a)}}. Since such expressions will often come up, let us write epRYeq=(RY)pq=:RY(p,q)e_{p}^{*}R_{Y}e_{q}=(R_{Y})_{pq}\mathrel{{=}\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}}R_{Y}(p,q) for matrices Y{W(1,a),W(,a),Q(a)}Y\in\{W^{(\ell-1,a)},W^{(\ell,a)},Q^{(a)}\}. Plugging this into (3.6), we have

RW(1,a)\displaystyle R_{W^{(\ell-1,a)}} =s=04S𝒮p,q,s(Wpq(a))SRQ(a)(βS,1,αS,2)RQ(a)(βS,s1,αS,s)RQ(a)eαS,1eβS,sRQ(a)\displaystyle=\sum_{s=0}^{4}\sum_{S\in\mathcal{S}_{p,q,s}}(W^{(a)}_{pq})^{S}\cdot R_{Q^{(a)}}(\beta_{S,1},\alpha_{S,2})\cdots R_{Q^{(a)}}(\beta_{S,{s-1}},\alpha_{S,s})\cdot R_{Q^{(a)}}e_{\alpha_{S,1}}e_{\beta_{S,s}}^{*}R_{Q^{(a)}}
+S𝒮p,q,5(Wpq(a))SRQ(a)(βS,1,αS,2)RQ(a)(βS,t,αS,t+1)RW(1,a)eαS,1eβS,t+1RQ(a).\displaystyle\hskip 14.22636pt+\sum_{S\in\mathcal{S}_{p,q,5}}(W^{(a)}_{pq})^{S}\cdot R_{Q^{(a)}}(\beta_{S,1},\alpha_{S,2})\cdots R_{Q^{(a)}}(\beta_{S,t},\alpha_{S,t+1})\cdot R_{W^{(\ell-1,a)}}e_{\alpha_{S,1}}e_{\beta_{S,t+1}}^{*}R_{Q^{(a)}}.

The expressions we will finally be interested in are eiRYv(a)e_{i}^{*}R_{Y}v^{(a)} and v(a)RYejv^{(a)*}R_{Y}e_{j}. Let us extend the previous notation and write RY(i,v(a))R_{Y}(i,v^{(a)}) and RY(v(a),j)R_{Y}(v^{(a)},j) for these, respectively, and extending in the same way to other mixed quadratic forms of this kind, for all YY mentioned above.

Then, firstly we may rewrite

F~ij(𝐖(1))=ϕ~((nRW(1,a)(i,v(a))RW(1,a)(v(a),j))a=1L),\widetilde{F}_{ij}(\mathbf{W}^{(\ell-1)})=\widetilde{\phi}\left(\left(n\cdot R_{W^{(\ell-1,a)}}(i,v^{(a)})R_{W^{(\ell-1,a)}}(v^{(a)},j)\right)_{a=1}^{L}\right), (3.7)

and similarly for 1\ell-1 replaced by \ell, and for the above expansion of the inner terms, we have for instance

RW(1,a)(i,v(a))\displaystyle R_{W^{(\ell-1,a)}}(i,v^{(a)}) =M(,a)(i,v(a),𝐖)+Δ(,a)(i,v(a),𝐖),\displaystyle=M^{(\ell,a)}(i,v^{(a)},\mathbf{W})+\Delta^{(\ell,a)}(i,v^{(a)},\mathbf{W}),
M(,a)(i,v(a),𝐘)\displaystyle M^{(\ell,a)}(i,v^{(a)},\mathbf{Y}) =s=04S𝒮p,q,s(Ypq(a))SRQ(a)(i,αS,1)RQ(a)(βS,1,αS,2)\displaystyle=\sum_{s=0}^{4}\sum_{S\in\mathcal{S}_{p,q,s}}(Y^{(a)}_{pq})^{S}R_{Q^{(a)}}(i,\alpha_{S,1})R_{Q^{(a)}}(\beta_{S,1},\alpha_{S,2})\cdots
RQ(a)(βS,s1,αS,s)RQ(a)(βS,s,v(a))\displaystyle\hskip 113.81102ptR_{Q^{(a)}}(\beta_{S,{s-1}},\alpha_{S,s})R_{Q^{(a)}}(\beta_{S,s},v^{(a)})
=RQ(a)(i,v(a))+s=14S𝒮p,q,s(Ypq(a))SRQ(a)(i,αS,1)RQ(a)(βS,1,αS,2)+\displaystyle=R_{Q^{(a)}}(i,v^{(a)})+\sum_{s=1}^{4}\sum_{S\in\mathcal{S}_{p,q,s}}(Y^{(a)}_{pq})^{S}R_{Q^{(a)}}(i,\alpha_{S,1})R_{Q^{(a)}}(\beta_{S,1},\alpha_{S,2})\cdots
RQ(a)(βS,s1,αS,s)RQ(a)(βS,s,v(a)),\displaystyle\hskip 170.71652ptR_{Q^{(a)}}(\beta_{S,{s-1}},\alpha_{S,s})R_{Q^{(a)}}(\beta_{S,s},v^{(a)}),
Δ(,a)(i,v(a),𝐘)\displaystyle\Delta^{(\ell,a)}(i,v^{(a)},\mathbf{Y}) =S𝒮p,q,5(Ypq(a))SRQ(a)+Epq(Y,a)(i,αS,1)RQ(a)(βS,1,αS,2)\displaystyle=\sum_{S\in\mathcal{S}_{p,q,5}}(Y^{(a)}_{pq})^{S}R_{Q^{(a)}+E^{(Y,a)}_{pq}}(i,\alpha_{S,1})R_{Q^{(a)}}(\beta_{S,1},\alpha_{S,2})\cdots
RQ(a)(βS,4,αS,5)RQ(a)(βS,5,v(a)).\displaystyle\hskip 113.81102ptR_{Q^{(a)}}(\beta_{S,4},\alpha_{S,5})R_{Q^{(a)}}(\beta_{S,5},v^{(a)}).

We define these expressions taking p=pp=p_{\ell} and q=qq=q_{\ell} under our previous ordering, and allow for a general matrix input YY, so that we also have

RW(,a)(i,v(a))=M(,a)(i,v(a),𝐗)+Δ(,a)(i,v(a),𝐗),R_{W^{(\ell,a)}}(i,v^{(a)})=M^{(\ell,a)}(i,v^{(a)},\mathbf{X})+\Delta^{(\ell,a)}(i,v^{(a)},\mathbf{X}),

reusing the same definitions. Also, as a convention, we take the term s=0s=0 in this expression to contribute RQ(a)(i,v(a))R_{Q^{(a)}}(i,v^{(a)}), compatible with our previous calculation.

We will now show the following estimates, which imply that the various Δ\Delta terms in the input to ϕ~\widetilde{\phi} in (3.7) may be ignored.

Lemma 3.6.

In the above setting, for each [N]\ell\in[N], we have

|𝔼F~ij(𝐖(1))𝔼ϕ~((nM(,a)(i,v(a),𝐖)M(,a)(v(a),j,𝐖))a=1L)|n5/2+3εv,\displaystyle\left|\mathbb{E}\widetilde{F}_{ij}(\mathbf{W}^{(\ell-1)})-\mathbb{E}\widetilde{\phi}\left(\left(n\cdot M^{(\ell,a)}(i,v^{(a)},\mathbf{W})M^{(\ell,a)}(v^{(a)},j,\mathbf{W})\right)_{a=1}^{L}\right)\right|\lesssim n^{-5/2+3\varepsilon_{v}}, (3.8)
|𝔼F~ij(𝐖())𝔼ϕ~((nM(,a)(i,v(a),𝐗)M(,a)(v(a),j,𝐗))a=1L)|n5/2+3εv,\displaystyle\left|\mathbb{E}\widetilde{F}_{ij}(\mathbf{W}^{(\ell)})-\mathbb{E}\widetilde{\phi}\left(\left(n\cdot M^{(\ell,a)}(i,v^{(a)},\mathbf{X})M^{(\ell,a)}(v^{(a)},j,\mathbf{X})\right)_{a=1}^{L}\right)\right|\lesssim n^{-5/2+3\varepsilon_{v}}, (3.9)

and thus by (3.4)

|𝔼F~ij(𝒘)𝔼F~ij(𝒙)|\displaystyle|\mathbb{E}\widetilde{F}_{ij}(\bm{w})-\mathbb{E}\widetilde{F}_{ij}(\bm{x})|
=1N|𝔼ϕ~((nM(,a)(i,v(a),𝐖)M(,a)(v(a),j,𝐖))a=1L)\displaystyle\hskip 14.22636pt\lesssim\sum_{\ell=1}^{N}\bigg|\mathbb{E}\widetilde{\phi}\left(\left(n\cdot M^{(\ell,a)}(i,v^{(a)},\mathbf{W})M^{(\ell,a)}(v^{(a)},j,\mathbf{W})\right)_{a=1}^{L}\right)
𝔼ϕ~((nM(,a)(i,v(a),𝐗)M(,a)(v(a),j,𝐗))a=1L)|+O(n1/2+3εv).\displaystyle\hskip 96.73918pt-\mathbb{E}\widetilde{\phi}\left(\left(n\cdot M^{(\ell,a)}(i,v^{(a)},\mathbf{X})M^{(\ell,a)}(v^{(a)},j,\mathbf{X})\right)_{a=1}^{L}\right)\bigg|+O(n^{-1/2+3\varepsilon_{v}}). (3.10)

Having shown this, we will have again made useful progress. For the \ell-th term, the inputs into ϕ~\widetilde{\phi} are functions of the current entry vector w(pq)w^{(pq)}, x(pq)x^{(pq)}. More precisely, for each coordinate a[L]a\in[L], the quantity M(,a)(i,v(a),𝐖)M(,a)(v(a),j,𝐖)M^{(\ell,a)}(i,v^{(a)},\mathbf{W})M^{(\ell,a)}(v^{(a)},j,\mathbf{W}) is a polynomial in Wpq(a)W^{(a)}_{pq}, Wpq(a)¯\overline{W^{(a)}_{pq}}, and the entries of RQ(a)R_{Q^{(a)}}. Crucially, 𝐐\mathbf{Q} is independent of those quantities by definition.

To do this, it will be useful to establish some estimates on the sizes of resolvent quadratic forms and associated expectations.

Proposition 3.7 (Resolvent quadratic form bounds).

In the above setting, for any a[L]a\in[L], any distinct α,β[n]\alpha,\beta\in[n], and any Y{W(,a),W(1,a),Q(a)}Y\in\{W^{(\ell,a)},W^{(\ell-1,a)},Q^{(a)}\}, we have

|Y(α,α)|\displaystyle|Y(\alpha,\alpha)| n1/2,\displaystyle\prec n^{-1/2},
|Y(α,β)|\displaystyle|Y(\alpha,\beta)| n1/2,\displaystyle\prec n^{-1/2},
|RY(α,β)|\displaystyle|R_{Y}(\alpha,\beta)| n1/2,\displaystyle\prec n^{-1/2},
|RY(α,α)|\displaystyle|R_{Y}(\alpha,\alpha)| 1,\displaystyle\prec 1,
|RY(α,v(a))|\displaystyle|R_{Y}(\alpha,v^{(a)})| n1/2+εv,\displaystyle\prec n^{-1/2+\varepsilon_{v}},
|RY(v(a),v(a))|\displaystyle|R_{Y}(v^{(a)},v^{(a)})| 1.\displaystyle\prec 1.

Further, fix nonnegative integers s,ts,t, and consider a product of terms

Π=RY1(ζ11,ζ12)RYs(ζs1,ζs2)Ys+1(α11,α12)Ys+t(αt1,αt2).\Pi=R_{Y_{1}}(\zeta_{11},\zeta_{12})\cdots R_{Y_{s}}(\zeta_{s1},\zeta_{s2})\cdot Y_{s+1}(\alpha_{11},\alpha_{12})\cdots Y_{s+t}(\alpha_{t1},\alpha_{t2}).

For each m[s+t]m\in[s+t], suppose that there is an index am[L]a_{m}\in[L] such that

Ym{W(,am),W(1,am),Q(am)}.Y_{m}\in\{W^{(\ell,a_{m})},W^{(\ell-1,a_{m})},Q^{(a_{m})}\}.

For m[s]m\in[s], each argument ζm1\zeta_{m1}, ζm2\zeta_{m2} is either an index in [n][n], which represents the corresponding basis vector, or the formal symbol v(am)v^{(a_{m})}. Thus, for example

RYm(α,v(am))=eαRYmv(am),RYm(v(am),β)=v(am)RYmeβ.R_{Y_{m}}(\alpha,v^{(a_{m})})=e^{*}_{\alpha}R_{Y_{m}}v^{(a_{m})},\qquad R_{Y_{m}}(v^{(a_{m})},\beta)=v^{(a_{m})*}R_{Y_{m}}e_{\beta}.

For the entry factors, we have αm1,αm2[n]\alpha_{m1},\alpha_{m2}\in[n], m[t]m\in[t], so that Ys+m(αm1,αm2)Y_{s+m}(\alpha_{m1},\alpha_{m2}) denotes the (αm1,αm2)(\alpha_{m1},\alpha_{m2})-th entry of the matrix Ys+mY_{s+m}.

Let svs_{v} be the number of resolvent factors whose two arguments contain the corresponding formal symbol v(am)v^{(a_{m})} exactly once, and let soffdiags_{\mathrm{offdiag}} be the number of resolvent factors whose two arguments are distinct as formal symbols. Then, we have

|Π|nt/2soffdiag/2+svεv|\Pi|\prec n^{-t/2-s_{\mathrm{offdiag}}/2+s_{v}\varepsilon_{v}} (3.11)

and, for any ε>0\varepsilon>0,

𝔼|Π|εnt/2soffdiag/2+svεv+ε.\mathbb{E}|\Pi|\lesssim_{\varepsilon}n^{-t/2-s_{\mathrm{offdiag}}/2+s_{v}\varepsilon_{v}+\varepsilon}. (3.12)
Proof.

We first prove the individual stochastic domination bounds. The entry bounds follow immediately from the ξ\xi-sub-Weibull tail bound in (1.1), since every entry of any YY is either zero, an entry of W(a)W^{(a)}, or an entry of X(a)X^{(a)}.

Next, for fixed a[L]a\in[L], the matrices W(,a)W^{(\ell,a)} and W(1,a)W^{(\ell-1,a)} are generalized Wigner matrices with uniformly controlled parameters. Indeed, the second-moment matching of the vectors w(pq)w^{(pq)} and x(pq)x^{(pq)} implies in particular that W(a)W^{(a)} and X(a)X^{(a)} have the same variance profile. Moreover, Q(a)Q^{(a)} is obtained from such a matrix by setting one upper-triangular entry and its Hermitian conjugate to zero. Hence the local laws of Theorem 2.18 apply to W(,a)W^{(\ell,a)} and W(1,a)W^{(\ell-1,a)}, while Corollary 2.19 applies to Q(a)Q^{(a)}, since applying the triangle inequality to those results it suffices to observe that |Gμsc(λ~)||G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})| is just a constant, while |vα(a)|v(a)n1/2+εv|v_{\alpha}^{(a)}|\leq\|v^{(a)}\|_{\infty}\leq n^{-1/2+\varepsilon_{v}} by assumption. Since LL is fixed, all estimates are uniform in a[L]a\in[L].

The bound of (3.11) then follows by Proposition 2.4, since this is just a product of several polynomial stochastic dominations from the first set of results. It remains to pass to expectations. If s+t=0s+t=0, then Π=1\Pi=1, and there is nothing to show. Otherwise, (3.12) says that we may take expectations on either side of this polynomial stochastic domination, provided we insert a factor of nεn^{\varepsilon} on the right-hand side. By Proposition 2.5, to show this it suffices to check that 𝔼|Π|2\mathbb{E}|\Pi|^{2} is bounded independently of nn. But, we have

|Π|RY1RYs|Ys+1(α11,α12)||Ys+t(αt1,αt2)|,|\Pi|\leq\|R_{Y_{1}}\|\cdots\|R_{Y_{s}}\|\cdot|Y_{s+1}(\alpha_{11},\alpha_{12})|\cdots|Y_{s+t}(\alpha_{t1},\alpha_{t2})|,

and so by Hölder’s inequality,

𝔼|Π|2\displaystyle\mathbb{E}|\Pi|^{2} (𝔼RY12s+2t)1s+t(𝔼RYs2s+2t)1s+t\displaystyle\leq\left(\mathbb{E}\|R_{Y_{1}}\|^{2s+2t}\right)^{\frac{1}{s+t}}\cdots\left(\mathbb{E}\|R_{Y_{s}}\|^{2s+2t}\right)^{\frac{1}{s+t}}
(𝔼|Ys+1(α11,α12)|2s+2t)1s+t(𝔼|Ys+t(αt1,αt2)|2s+2t)1s+t\displaystyle\hskip 28.45274pt\left(\mathbb{E}|Y_{s+1}(\alpha_{11},\alpha_{12})|^{2s+2t}\right)^{\frac{1}{s+t}}\cdots\left(\mathbb{E}|Y_{s+t}(\alpha_{t1},\alpha_{t2})|^{2s+2t}\right)^{\frac{1}{s+t}}

and we then indeed find that 𝔼|Π|2=O(1)\mathbb{E}|\Pi|^{2}=O(1) by Proposition 2.17 and Proposition 2.2 which show that each factor here is O(1)O(1).

We note that a condition of that Proposition is that Im(z)>nΩ(1)\mathrm{Im}(z)>n^{-\Omega(1)}, so we are using that we are evaluating our resolvents at λ~\widetilde{\lambda} rather than λ\lambda\in\mathbb{R}, and this cannot be avoided for working with such expectations (without introducing other technical devices like restricting to particular events) for the reason discussed in Remark 3.2. ∎

Proof of Lemma 3.6.

It suffices to show (3.8), then (3.9) follows by a symmetric argument, and (3.10) follows from the two taken together and the triangle inequality since N=O(n2)N=O(n^{2}).

By Proposition 3.7, uniformly in a[L]a\in[L], we have that

M(,a)(i,v(a),𝐖)\displaystyle M^{(\ell,a)}(i,v^{(a)},\mathbf{W}) n1/2+εv,\displaystyle\prec n^{-1/2+\varepsilon_{v}},
with the dominant contribution coming from the s=0s=0 term, and
Δ(,a)(i,v(a),𝐖)\displaystyle\Delta^{(\ell,a)}(i,v^{(a)},\mathbf{W}) n3+εv,\displaystyle\prec n^{-3+\varepsilon_{v}},

where in the Δ(,a)\Delta^{(\ell,a)} bounds a factor of n5/2n^{-5/2} comes from 5 factors of entries of YY, and another factor of n1/2+εvn^{-1/2+\varepsilon_{v}} comes from one factor of the form RY(α,v)R_{Y}(\alpha,v) for some α[n]\alpha\in[n]. (We use here that the M(,a)M^{(\ell,a)} and Δ(,a)\Delta^{(\ell,a)} are finite sums of the form treated by Proposition 3.7, which may be combined over finite sums by Proposition 2.4.) The same bounds hold for (i,v(a))(i,v^{(a)}) replaced by (v(a),j)(v^{(a)},j) as well. Then, we may expand

F~ij(𝐖(1))\displaystyle\widetilde{F}_{ij}(\mathbf{W}^{(\ell-1)}) =ϕ~((nRW(1,a)(i,v(a))RW(1,a)(v(a),j))a=1L)\displaystyle=\widetilde{\phi}\left(\left(n\cdot R_{W^{(\ell-1,a)}}(i,v^{(a)})R_{W^{(\ell-1,a)}}(v^{(a)},j)\right)_{a=1}^{L}\right)
=ϕ~((nM(,a)(i,v(a),𝐖)M(,a)(v(a),j,𝐖)CLOSECLOSE\displaystyle=\widetilde{\phi}\bigg(\Big(n\cdot M^{(\ell,a)}(i,v^{(a)},\mathbf{W})M^{(\ell,a)}(v^{(a)},j,\mathbf{W})
+nM(,a)(i,v(a),𝐖)Δ(,a)(v(a),j,𝐖)\displaystyle\hskip 28.45274pt+n\cdot M^{(\ell,a)}(i,v^{(a)},\mathbf{W})\Delta^{(\ell,a)}(v^{(a)},j,\mathbf{W})
+nΔ(,a)(i,v(a),𝐖)M(,a)(v(a),j,𝐖)\displaystyle\hskip 28.45274pt+n\cdot\Delta^{(\ell,a)}(i,v^{(a)},\mathbf{W})M^{(\ell,a)}(v^{(a)},j,\mathbf{W})
+nΔ(,a)(i,v(a),𝐖)Δ(,a)(v(a),j,𝐖))a=1L)\displaystyle\hskip 28.45274pt+n\cdot\Delta^{(\ell,a)}(i,v^{(a)},\mathbf{W})\Delta^{(\ell,a)}(v^{(a)},j,\mathbf{W})\Big)_{a=1}^{L}\bigg)
=:ϕ~((nM(,a)(i,v(a),𝐖)M(,a)(v(a),j,𝐖)+Ξa)a=1L),\displaystyle\mathrel{{=}\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}}\widetilde{\phi}\left(\left(n\cdot M^{(\ell,a)}(i,v^{(a)},\mathbf{W})M^{(\ell,a)}(v^{(a)},j,\mathbf{W})+\Xi_{a}\right)_{a=1}^{L}\right),

where the above bounds imply that

|Ξa|n5/2+2εv.|\Xi_{a}|\prec n^{-5/2+2\varepsilon_{v}}.

Using that ϕ~\widetilde{\phi} is Lipschitz and LL is fixed, we then have

|F~ij(𝐖(1))ϕ~((nM(,a)(i,v(a),𝐖)M(,a)(v(a),j,𝐖))a=1L)|a=1L|Ξa|,\left|\widetilde{F}_{ij}(\mathbf{W}^{(\ell-1)})-\widetilde{\phi}\left(\left(n\cdot M^{(\ell,a)}(i,v^{(a)},\mathbf{W})M^{(\ell,a)}(v^{(a)},j,\mathbf{W})\right)_{a=1}^{L}\right)\right|\leq\sum_{a=1}^{L}|\Xi_{a}|,

and the result follows upon taking expectations and using the triangle inequality, since 𝔼|Ξa|\mathbb{E}|\Xi_{a}| may be bounded by another triangle inequality as a sum of terms each controlled by Proposition 3.7. Applying the Proposition costs another factor of nεn^{\varepsilon} for an arbitrarily small ε>0\varepsilon>0, and we take ε=εv\varepsilon=\varepsilon_{v} to obtain the result as stated. ∎

Remark 3.8 (Order of resolvent expansion).

We see in the above bounds why our choice of fourth order in the resolvent expansion of (3.6) was convenient: in order for N=O(n2)N=O(n^{2}) terms with the above error bound to not contribute in total to our bound on |𝔼F~ij(𝐰)𝔼F~ij(𝐱)||\mathbb{E}\widetilde{F}_{ij}(\bm{w})-\mathbb{E}\widetilde{F}_{ij}(\bm{x})|, we must have each term to be bounded by n2δn^{-2-\delta} for some δ>0\delta>0, which will no longer hold by the above argument if we take a third order (or shorter) expansion.33 3 It seems that with slightly more care in the combinatorics of how many terms of what “types” in terms of the intersection among the {p,q}\{p_{\ell},q_{\ell}\} and {i,j}\{i,j\} index pairs one could take an expansion to order 3, but we take the longer expansion that is more transparently correct for the sake of exposition.

We now continue to finish the proof of the main result of this section, modulo one technical result we will encounter at the end of our calculation that we defer to the following section.

To keep track of conjugated scalar resolvent forms arising from the multivariate Taylor expansion below, let σ{id,conj}\sigma\in\{\mathrm{id},\mathrm{conj}\} with the maps defined in Definition 3.5 and write

RYσ(ζ1,ζ2):=σ(RY(ζ1,ζ2)).R^{\sigma}_{Y}(\zeta_{1},\zeta_{2})\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\sigma\big(R_{Y}(\zeta_{1},\zeta_{2})\big).

Since |RYσ(ζ1,ζ2)|=|RY(ζ1,ζ2)||R^{\sigma}_{Y}(\zeta_{1},\zeta_{2})|=|R_{Y}(\zeta_{1},\zeta_{2})|, the same bounds in Proposition 3.7 hold if any scalar resolvent form RY(ζ1,ζ2)R_{Y}(\zeta_{1},\zeta_{2}) in the statement is replaced by either RYσ(ζ1,ζ2)R^{\sigma}_{Y}(\zeta_{1},\zeta_{2}).

Proof of Lemma 3.4.

Starting from the result of Lemma 3.6, we must control summands of the form

|\displaystyle\bigg| 𝔼ϕ~((nM(,a)(i,v(a),𝐖)M(,a)(v(a),j,𝐖))a=1L)\displaystyle\mathbb{E}\widetilde{\phi}\left(\left(n\cdot M^{(\ell,a)}(i,v^{(a)},\mathbf{W})M^{(\ell,a)}(v^{(a)},j,\mathbf{W})\right)_{a=1}^{L}\right)
𝔼ϕ~((nM(,a)(i,v(a),𝐗)M(,a)(v(a),j,𝐗))a=1L)|.\displaystyle\qquad-\mathbb{E}\widetilde{\phi}\left(\left(n\cdot M^{(\ell,a)}(i,v^{(a)},\mathbf{X})M^{(\ell,a)}(v^{(a)},j,\mathbf{X})\right)_{a=1}^{L}\right)\bigg|.

For a word 𝜿:=(κ1,,κu)[2L]u\bm{\kappa}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}(\kappa_{1},\ldots,\kappa_{u})\in[2L]^{u}, write

𝜿:=κ1κu,\partial_{\bm{\kappa}}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\partial_{\kappa_{1}}\ldots\partial_{\kappa_{u}},

and define

Φ𝜿,i,j(𝐐):=1u!𝜿ϕ~((nRQ(a)(i,v(a))RQ(a)(v(a),j))a=1L).\Phi_{\bm{\kappa},i,j}(\mathbf{Q})\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\frac{1}{u!}\partial_{\bm{\kappa}}\widetilde{\phi}\left(\left(n\cdot R_{Q^{(a)}}(i,v^{(a)})R_{Q^{(a)}}(v^{(a)},j)\right)_{a=1}^{L}\right).

We now take a multivariate Taylor expansion of ϕ~\widetilde{\phi} in each term in such a difference. We define, for each a[L]a\in[L],

Γa(𝐘):=M(,a)(i,v(a),𝐘)M(,a)(v(a),j,𝐘)RQ(a)(i,v(a))RQ(a)(v(a),j).\Gamma_{a}(\mathbf{Y})\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}M^{(\ell,a)}(i,v^{(a)},\mathbf{Y})M^{(\ell,a)}(v^{(a)},j,\mathbf{Y})-R_{Q^{(a)}}(i,v^{(a)})R_{Q^{(a)}}(v^{(a)},j).

We view ϕ~\widetilde{\phi} as a function on 2L\mathbb{R}^{2L} and define

𝚪(𝐘):=(ReΓ1(𝐘),ImΓ1(𝐘),,ReΓL(𝐘),ImΓL(𝐘)),\bm{\Gamma}(\mathbf{Y})\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}(\mathrm{Re}\Gamma_{1}(\mathbf{Y}),\mathrm{Im}\Gamma_{1}(\mathbf{Y}),\ldots,\mathrm{Re}\Gamma_{L}(\mathbf{Y}),\mathrm{Im}\Gamma_{L}(\mathbf{Y})),

so

𝚪2a1(𝐘)=ReΓa(𝐘)=id(Γa(𝐘))+conj(Γa(𝐘))2,\bm{\Gamma}_{2a-1}(\mathbf{Y})=\mathrm{Re}\Gamma_{a}(\mathbf{Y})=\frac{\mathrm{id}(\Gamma_{a}(\mathbf{Y}))+\mathrm{conj}(\Gamma_{a}(\mathbf{Y}))}{2},
𝚪2a(𝐘)=ImΓa(𝐘)=id(Γa(𝐘))conj(Γa(𝐘))2𝒊.\bm{\Gamma}_{2a}(\mathbf{Y})=\mathrm{Im}\Gamma_{a}(\mathbf{Y})=\frac{\mathrm{id}(\Gamma_{a}(\mathbf{Y}))-\mathrm{conj}(\Gamma_{a}(\mathbf{Y}))}{2\bm{i}}.

By the same application of Proposition 3.7 as above in the proof of Lemma 3.6, we have, uniformly in a[L]a\in[L],

|Γa(𝐘)|n3/2+2εv|\Gamma_{a}(\mathbf{Y})|\prec n^{-3/2+2\varepsilon_{v}}

for each 𝐘{𝐖,𝐗}\mathbf{Y}\in\{\mathbf{W},\mathbf{X}\}. Taking a Taylor expansion to fourth order around

(nRQ(a)(i,v(a))RQ(a)(v(a),j))a=1L\left(n\cdot R_{Q^{(a)}}(i,v^{(a)})R_{Q^{(a)}}(v^{(a)},j)\right)_{a=1}^{L}

in each expectation, we find

ϕ~((nM(,a)(i,v(a),𝐘)M(,a)(v(a),j,𝐘))a=1L)\displaystyle\hskip-28.45274pt\widetilde{\phi}\left(\left(n\cdot M^{(\ell,a)}(i,v^{(a)},\mathbf{Y})M^{(\ell,a)}(v^{(a)},j,\mathbf{Y})\right)_{a=1}^{L}\right)
=u=04𝜿[2L]uΦ𝜿,i,j(𝐐)nur=1u𝚪κr(𝐘)+Ξ\displaystyle=\sum_{u=0}^{4}\sum_{\bm{\kappa}\in[2L]^{u}}\Phi_{\bm{\kappa},i,j}(\mathbf{Q})\cdot n^{u}\cdot\prod_{r=1}^{u}\bm{\Gamma}_{\kappa_{r}}(\mathbf{Y})+\Xi^{\prime}
=u=04𝜿[2L]uΦ𝜿,i,j(𝐐)nur=1u(σr{id,conj}cσr,κrσr(Γκr2(𝐘)))+Ξ\displaystyle=\sum_{u=0}^{4}\sum_{\bm{\kappa}\in[2L]^{u}}\Phi_{\bm{\kappa},i,j}(\mathbf{Q})\cdot n^{u}\cdot\prod_{r=1}^{u}\left(\sum_{\sigma_{r}\in\{\mathrm{id},\mathrm{conj}\}}c_{\sigma_{r},\kappa_{r}}\cdot\sigma_{r}\big(\Gamma_{\lceil\frac{\kappa_{r}}{2}\rceil}(\mathbf{Y})\big)\right)+\Xi^{\prime}
=:u=04𝜿[2L]uΦ𝜿,i,j(𝐐)nu𝝈{id,conj}uc𝜿,𝝈r=1uσr(Γκr2(𝐘))+Ξ\displaystyle\mathrel{{=}\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}}\sum_{u=0}^{4}\sum_{\bm{\kappa}\in[2L]^{u}}\Phi_{\bm{\kappa},i,j}(\mathbf{Q})\cdot n^{u}\cdot\sum_{\bm{\sigma}\in\{\mathrm{id},\mathrm{conj}\}^{u}}c_{\bm{\kappa},\bm{\sigma}}\prod_{r=1}^{u}\sigma_{r}\big(\Gamma_{\lceil\frac{\kappa_{r}}{2}\rceil}(\mathbf{Y})\big)+\Xi^{\prime}

where |c𝜿,𝝈|=|r=1ucσr,κr|=2u1|c_{\bm{\kappa},\bm{\sigma}}|=|\prod_{r=1}^{u}c_{\sigma_{r},\kappa_{r}}|=2^{-u}\leq 1 and Ξ\Xi^{\prime} is a random error term (the prime mark distinguishing it from the one in the proof of Lemma 3.6 above) that is bounded by

|Ξ|n5(maxa[L]|Γa(𝐘)|)5n5/2+10εv.|\Xi^{\prime}|\lesssim n^{5}\cdot\left(\max_{a\in[L]}|\Gamma_{a}(\mathbf{Y})|\right)^{5}\prec n^{-5/2+10\varepsilon_{v}}.

Thus, again using the expectation bounds in Proposition 3.7, we have

|𝔼ϕ~((nM(,a)(i,v(a),𝐘)M(,a)(v(a),j,𝐘))a=1L)𝔼u=04𝜿[2L]uΦ𝜿,i,j(𝐐)nur=1u𝚪κr(𝐘)|\displaystyle\left|\mathbb{E}\widetilde{\phi}\left(\left(n\cdot M^{(\ell,a)}(i,v^{(a)},\mathbf{Y})M^{(\ell,a)}(v^{(a)},j,\mathbf{Y})\right)_{a=1}^{L}\right)-\mathbb{E}\sum_{u=0}^{4}\sum_{\bm{\kappa}\in[2L]^{u}}\Phi_{\bm{\kappa},i,j}(\mathbf{Q})\cdot n^{u}\cdot\prod_{r=1}^{u}\bm{\Gamma}_{\kappa_{r}}(\mathbf{Y})\right|
𝔼|Ξ|\displaystyle\hskip 28.45274pt\leq\mathbb{E}|\Xi^{\prime}|
εn5/2+10εv+ε\displaystyle\hskip 28.45274pt\lesssim_{\varepsilon}n^{-5/2+10\varepsilon_{v}+\varepsilon}

for ε>0\varepsilon>0 arbitrarily small.

Combining this with Lemma 3.6, we find

|𝔼F~ij(𝒘)𝔼F~ij(𝒙)|\displaystyle|\mathbb{E}\widetilde{F}_{ij}(\bm{w})-\mathbb{E}\widetilde{F}_{ij}(\bm{x})|
=1N|𝔼u=04𝜿[2L]uΦ𝜿,i,j(𝐐)nu(r=1u𝚪κr(𝐖)r=1u𝚪κr(𝐗))|+O(n1/2+10εv+ε).\displaystyle\hskip 14.22636pt\lesssim\sum_{\ell=1}^{N}\left|\mathbb{E}\sum_{u=0}^{4}\sum_{\bm{\kappa}\in[2L]^{u}}\Phi_{\bm{\kappa},i,j}(\mathbf{Q})\cdot n^{u}\cdot\left(\prod_{r=1}^{u}\bm{\Gamma}_{\kappa_{r}}(\mathbf{W})-\prod_{r=1}^{u}\bm{\Gamma}_{\kappa_{r}}(\mathbf{X})\right)\right|+O(n^{-1/2+10\varepsilon_{v}+\varepsilon}).

(In the same way as discussed in Remark 3.8, we see from this calculation that fourth order was the correct order of Taylor expansion for this argument to work.) Let us write

ε:=10εv+ε\varepsilon^{\prime}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}10\varepsilon_{v}+\varepsilon

for this remaining error exponent, which we will carry through to the end of the proof and which gives rise to the dependence on εv\varepsilon_{v} in Theorem 1.4.

The only properties of Φ𝜿,i,j(𝐐)\Phi_{\bm{\kappa},i,j}(\mathbf{Q}) we will need are that it is O(1)O(1) almost surely (by the boundedness of ϕ~\widetilde{\phi}) and its derivatives and that it is independent of w(pq)w^{(pq)} and x(pq)x^{(pq)} (since it is a function only of 𝐐\mathbf{Q}, where this entry is set to zero). Also, let us develop notation for the quantities appearing in powers of 𝚪(𝐘)\bm{\Gamma}(\mathbf{Y}). For a[L]a\in[L] and σ{id,conj}\sigma\in\{\mathrm{id},\mathrm{conj}\}, expanding σ(Γa(𝐘))\sigma(\Gamma_{a}(\mathbf{Y})) from the definition of Γa\Gamma_{a}, we have

σ(Γa(𝐘))\displaystyle\sigma(\Gamma_{a}(\mathbf{Y})) =(s1,s2){0,1,2,3,4}2(s1,s2)(0,0)S1𝒮p,q,s1S2𝒮p,q,s2σ((Ypq(a))S1(Ypq(a))S2)\displaystyle=\sum_{\begin{subarray}{c}(s_{1},s_{2})\in\{0,1,2,3,4\}^{2}\\ (s_{1},s_{2})\neq(0,0)\end{subarray}}\sum_{\begin{subarray}{c}S_{1}\in\mathcal{S}_{p,q,s_{1}}\\ S_{2}\in\mathcal{S}_{p,q,s_{2}}\end{subarray}}\sigma\left((Y^{(a)}_{pq})^{S_{1}}(Y^{(a)}_{pq})^{S_{2}}\right)
RQ(a)σ(i,αS1,1)RQ(a)σ(βS1,1,αS1,2)RQ(a)σ(βS1,s11,αS1,s1)RQ(a)σ(βS1,s1,v(a))\displaystyle\hskip 56.9055pt\cdot R^{\sigma}_{Q^{(a)}}(i,\alpha_{S_{1},1})R^{\sigma}_{Q^{(a)}}(\beta_{S_{1},1},\alpha_{S_{1},2})\cdots R^{\sigma}_{Q^{(a)}}(\beta_{S_{1},{s_{1}-1}},\alpha_{S_{1},s_{1}})R^{\sigma}_{Q^{(a)}}(\beta_{S_{1},s_{1}},v^{(a)})
RQ(a)σ(v(a),αS2,1)RQ(a)σ(βS2,1,αS2,2)RQ(a)σ(βS2,s21,αS2,s2)RQ(a)σ(βS2,s2,j)\displaystyle\hskip 56.9055pt\cdot R^{\sigma}_{Q^{(a)}}(v^{(a)},\alpha_{S_{2},1})R^{\sigma}_{Q^{(a)}}(\beta_{S_{2},1},\alpha_{S_{2},2})\cdots R^{\sigma}_{Q^{(a)}}(\beta_{S_{2},{s_{2}-1}},\alpha_{S_{2},s_{2}})R^{\sigma}_{Q^{(a)}}(\beta_{S_{2},s_{2}},j)
=:(s1,s2){0,1,2,3,4}2(s1,s2)(0,0)S1𝒮p,q,s1S2𝒮p,q,s2(Ypq(a))S1(Ypq(a))S2Ji,j,S1(a,σ)(Q(a))Ji,j,S2(a,σ)(Q(a)),\displaystyle\mathrel{{=}\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}}\sum_{\begin{subarray}{c}(s_{1},s_{2})\in\{0,1,2,3,4\}^{2}\\ (s_{1},s_{2})\neq(0,0)\end{subarray}}\sum_{\begin{subarray}{c}S_{1}\in\mathcal{S}_{p,q,s_{1}}\\ S_{2}\in\mathcal{S}_{p,q,s_{2}}\end{subarray}}(Y^{(a)}_{pq})^{S_{1}}(Y^{(a)}_{pq})^{S_{2}}J^{(a,\sigma)}_{i,j,S_{1}}(Q^{(a)})J^{(a,\sigma)}_{i,j,S_{2}}(Q^{(a)}),
for σ{id,conj}\sigma\in\{\mathrm{id},\mathrm{conj}\}. It will be useful to rearrange a little bit more, by defining the union 𝒮p,q:=𝒮p,q,0𝒮p,q,4\mathcal{S}_{p,q}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\mathcal{S}_{p,q,0}\sqcup\cdots\sqcup\mathcal{S}_{p,q,4}. For S𝒮p,qS\in\mathcal{S}_{p,q}, is a string of indeterminate length now, write s(S)s(S) for its length, so that S𝒮p,q,s(S)S\in\mathcal{S}_{p,q,s(S)}. Then, writing \varnothing for the empty string that is the only element of 𝒮p,q,0\mathcal{S}_{p,q,0}, we may write the above as a single sum,
=(S1,S2)𝒮p,q2{(,)}((Ypq(a))S1(Ypq(a))S2)σJi,j,S1(a,σ)(Q(a))Ji,j,S2(a,σ)(Q(a)).\displaystyle=\sum_{(S_{1},S_{2})\in\mathcal{S}_{p,q}^{2}\setminus\{(\varnothing,\varnothing)\}}\left((Y^{(a)}_{pq})^{S_{1}}(Y^{(a)}_{pq})^{S_{2}}\right)^{\sigma}J^{(a,\sigma)}_{i,j,S_{1}}(Q^{(a)})J^{(a,\sigma)}_{i,j,S_{2}}(Q^{(a)}).

We may then write the result of taking products of the real-coordinate components of 𝚪(𝐘)\bm{\Gamma}(\mathbf{Y}) in terms of matrices with entries taking values in 𝒮p,q\mathcal{S}_{p,q}. For each r[u]r\in[u], let ar[L]a_{r}\in[L] be the unique index such that κr{2ar1,2ar}\kappa_{r}\in\{2a_{r}-1,2a_{r}\}. Expanding the real and imaginary parts of Γar(𝐘)\Gamma_{a_{r}}(\mathbf{Y}) gives a finite linear combination of terms of the following form:

r=1u𝚪κr(𝐘)=𝝈{id,conj}uc𝜿,𝝈S𝒮p,qu×2no row of S equals (,)𝐘pq𝜿,S,𝝈J𝜿,i,j,S𝝈(𝐐),\prod_{r=1}^{u}\bm{\Gamma}_{\kappa_{r}}(\mathbf{Y})=\sum_{\bm{\sigma}\in\{\mathrm{id},\mathrm{conj}\}^{u}}c_{\bm{\kappa},\bm{\sigma}}\sum_{\begin{subarray}{c}S\in\mathcal{S}_{p,q}^{u\times 2}\\ \text{no row of }S\text{ equals }(\varnothing,\varnothing)\end{subarray}}\mathbf{Y}^{\bm{\kappa},S,\bm{\sigma}}_{pq}J^{\bm{\sigma}}_{\bm{\kappa},i,j,S}(\mathbf{Q}),

Lastly, for such a matrix SS, define

𝐘pq𝜿,S,𝝈\displaystyle\mathbf{Y}^{\bm{\kappa},S,\bm{\sigma}}_{pq} :=r=1uσr((Ypq(ar))Sr1(Ypq(ar))Sr2),\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\prod_{r=1}^{u}\sigma_{r}\left((Y^{(a_{r})}_{pq})^{S_{r1}}(Y^{(a_{r})}_{pq})^{S_{r2}}\right),
J𝜿,i,j,S𝝈(𝐐)\displaystyle J^{\bm{\sigma}}_{\bm{\kappa},i,j,S}(\mathbf{Q}) :=r=1uJi,j,Sr1(ar,σr)(Q(ar))Ji,j,Sr2(ar,σr)(Q(ar)),\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\prod_{r=1}^{u}J^{(a_{r},\sigma_{r})}_{i,j,S_{r1}}(Q^{(a_{r})})J^{(a_{r},\sigma_{r})}_{i,j,S_{r2}}(Q^{(a_{r})}),
|S|\displaystyle|S| :=r,fs(Srf),\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\sum_{r,f}s(S_{rf}),
|S|0\displaystyle|S|_{0} :=#{(r,f):Srf}.\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\#\{(r,f):S_{rf}\neq\varnothing\}.

Using this, we may rewrite our bound concisely as

|𝔼F~ij(𝒘)𝔼F~ij(𝒙)|\displaystyle|\mathbb{E}\widetilde{F}_{ij}(\bm{w})-\mathbb{E}\widetilde{F}_{ij}(\bm{x})|
=1N|𝔼u=04𝜿[2L]u𝝈{id,conj}u\displaystyle\lesssim\sum_{\ell=1}^{N}\bigg|\mathbb{E}\sum_{u=0}^{4}\sum_{\bm{\kappa}\in[2L]^{u}}\sum_{\bm{\sigma}\in\{\mathrm{id},\mathrm{conj}\}^{u}}
S𝒮p,qu×2no row of S equals (,)Φ𝜿,i,j(𝐐)J𝜿,i,j,S𝝈(𝐐)nu(𝐖pq𝜿,S,𝝈𝐗pq𝜿,S,𝝈)|+O(n1/2+ε)\displaystyle\hskip 28.45274pt\sum_{\begin{subarray}{c}S\in\mathcal{S}_{p,q}^{u\times 2}\\ \text{no row of }S\text{ equals }(\varnothing,\varnothing)\end{subarray}}\Phi_{\bm{\kappa},i,j}(\mathbf{Q})J^{\bm{\sigma}}_{\bm{\kappa},i,j,S}(\mathbf{Q})\cdot n^{u}\cdot\left(\mathbf{W}^{\bm{\kappa},S,\bm{\sigma}}_{pq}-\mathbf{X}^{\bm{\kappa},S,\bm{\sigma}}_{pq}\right)\bigg|+O(n^{-1/2+\varepsilon^{\prime}})
where we may use independence of 𝐐\mathbf{Q} and (Wpq(a),Xpq(a))(W^{(a)}_{pq},X^{(a)}_{pq}), giving when combined with a triangle inequality
=1Nu=04𝜿[2L]u𝝈{id,conj}u\displaystyle\leq\sum_{\ell=1}^{N}\sum_{u=0}^{4}\sum_{\bm{\kappa}\in[2L]^{u}}\sum_{\bm{\sigma}\in\{\mathrm{id},\mathrm{conj}\}^{u}}
S𝒮p,qu×2no row of S equals (,)|𝔼[Φ𝜿,i,j(𝐐)J𝜿,i,j,S𝝈(𝐐)]|nu|𝔼[𝐖pq𝜿,S,𝝈]𝔼[𝐗pq𝜿,S,𝝈]|+O(n1/2+ε)\displaystyle\hskip 28.45274pt\sum_{\begin{subarray}{c}S\in\mathcal{S}_{p,q}^{u\times 2}\\ \text{no row of }S\text{ equals }(\varnothing,\varnothing)\end{subarray}}\left|\mathbb{E}[\Phi_{\bm{\kappa},i,j}(\mathbf{Q})J^{\bm{\sigma}}_{\bm{\kappa},i,j,S}(\mathbf{Q})]\right|\cdot n^{u}\cdot\left|\mathbb{E}\big[\mathbf{W}^{\bm{\kappa},S,\bm{\sigma}}_{pq}\big]-\mathbb{E}\big[\mathbf{X}^{\bm{\kappa},S,\bm{\sigma}}_{pq}\big]\right|+O(n^{-1/2+\varepsilon^{\prime}})
Here, for each fixed 𝝈\bm{\sigma}, the expressions 𝐖pq𝜿,S,𝝈\mathbf{W}_{pq}^{\bm{\kappa},S,\bm{\sigma}} and 𝐗pq𝜿,S,𝝈\mathbf{X}_{pq}^{\bm{\kappa},S,\bm{\sigma}} are fixed complex linear combinations of monomials of total degree |S||S| in the real and imaginary parts of the block variables w(pq),x(pq)2Lw^{(pq)},x^{(pq)}\in\mathbb{R}^{2L}. Indeed, applying either id\mathrm{id} or conj\mathrm{conj} only decides whether we use an entry monomial or its complex conjugate, and this does not change the total degree. Therefore the difference 𝔼[𝐖pq𝜿,S,𝝈]𝔼[𝐗pq𝜿,S,𝝈]\mathbb{E}[\mathbf{W}_{pq}^{\bm{\kappa},S,\bm{\sigma}}]-\mathbb{E}[\mathbf{X}_{pq}^{\bm{\kappa},S,\bm{\sigma}}] vanishes whenever |S|2|S|\leq 2, by the assumed matching of joint moments up to order two. On the other hand, we have |J𝜿,i,j,S𝝈(𝐐)|nu+2uεvnu+8εv|J^{\bm{\sigma}}_{\bm{\kappa},i,j,S}(\mathbf{Q})|\prec n^{-u+2u\varepsilon_{v}}\prec n^{-u+8\varepsilon_{v}} for all S𝒮p,qu×2S\in\mathcal{S}^{u\times 2}_{p,q} by Proposition 3.7. Thus, also applying the expectation bounds from Proposition 3.7 and using the boundedness of Φ𝜿,i,j(𝐐)\Phi_{\bm{\kappa},i,j}(\mathbf{Q}), each term above has a simple a priori bound of O(nunu+8εv+εn|S|/2)=O(n|S|/2+9εv)O(n^{u}\cdot n^{-u+8\varepsilon_{v}+\varepsilon}\cdot n^{-|S|/2})=O(n^{-|S|/2+9\varepsilon_{v}}). (Here and below we let ε(0,εv)\varepsilon\in(0,\varepsilon_{v}) be a temporary parameter as needed.) In particular, if |S|5|S|\geq 5 then this is at most O(n5/2+9εv)O(n^{-5/2+9\varepsilon_{v}}), and since 9εv<ε9\varepsilon_{v}<\varepsilon^{\prime} this may be subsumed into the error term we already have, giving
=1Nu=04𝜿[2L]u𝝈{id,conj}u\displaystyle\leq\sum_{\ell=1}^{N}\sum_{u=0}^{4}\sum_{\bm{\kappa}\in[2L]^{u}}\sum_{\bm{\sigma}\in\{\mathrm{id},\mathrm{conj}\}^{u}}
S𝒮p,qu×2no row of S equals (,)|S|{3,4}nu|𝔼[Φ𝜿,i,j(𝐐)J𝜿,i,j,S𝝈(𝐐)]||𝔼[𝐖pq𝜿,S,𝝈]𝔼[𝐗pq𝜿,S,𝝈]|+O(n1/2+ε),\displaystyle\hskip 28.45274pt\sum_{\begin{subarray}{c}S\in\mathcal{S}_{p,q}^{u\times 2}\\ \text{no row of }S\text{ equals }(\varnothing,\varnothing)\\ |S|\in\{3,4\}\end{subarray}}n^{u}\cdot|\mathbb{E}[\Phi_{\bm{\kappa},i,j}(\mathbf{Q})J^{\bm{\sigma}}_{\bm{\kappa},i,j,S}(\mathbf{Q})]|\cdot\left|\mathbb{E}[\mathbf{W}^{\bm{\kappa},S,\bm{\sigma}}_{pq}]-\mathbb{E}[\mathbf{X}^{\bm{\kappa},S,\bm{\sigma}}_{pq}]\right|+O(n^{-1/2+\varepsilon^{\prime}}),
Now, for the remaining terms we will not have any particular control of the difference of expectations involving 𝐖\mathbf{W} and 𝐗\mathbf{X}, so a direct bound on these using Proposition 2.2 reduces this to
=1Nu=04𝜿[2L]u𝝈{id,conj}uS𝒮p,qu×2no row of S equals (,)|S|{3,4}nu|S|/2|𝔼[Φ𝜿,i,j(𝐐)J𝜿,i,j,S𝝈(𝐐)]|+O(n1/2+ε).\displaystyle\lesssim\sum_{\ell=1}^{N}\sum_{u=0}^{4}\sum_{\bm{\kappa}\in[2L]^{u}}\sum_{\bm{\sigma}\in\{\mathrm{id},\mathrm{conj}\}^{u}}\sum_{\begin{subarray}{c}S\in\mathcal{S}_{p,q}^{u\times 2}\\ \text{no row of }S\text{ equals }(\varnothing,\varnothing)\\ |S|\in\{3,4\}\end{subarray}}n^{u-|S|/2}\cdot\left|\mathbb{E}\left[\Phi_{\bm{\kappa},i,j}(\mathbf{Q})J^{\bm{\sigma}}_{\bm{\kappa},i,j,S}(\mathbf{Q})\right]\right|+O(n^{-1/2+\varepsilon^{\prime}}).

Now, consider grouping the terms in the outer sum over \ell according to, firstly, whether p=qp_{\ell}=q_{\ell} or not, and second according to the size of {p,q}{i,j}\{p_{\ell},q_{\ell}\}\cap\{i,j\}. The total number of terms where either p=qp_{\ell}=q_{\ell} or |{p,q}{i,j}|1|\{p_{\ell},q_{\ell}\}\cap\{i,j\}|\geq 1 is O(n)O(n). On the other hand, note first that, since |J𝜿,i,j,S𝝈(𝐐)|nu+8εv|J^{\bm{\sigma}}_{\bm{\kappa},i,j,S}(\mathbf{Q})|\prec n^{-u+8\varepsilon_{v}} as we derived earlier, by Proposition 3.7 (and the Cauchy-Schwarz inequality to extract the Φ𝜿,i,j(𝐐)\Phi_{\bm{\kappa},i,j}(\mathbf{Q}) factor) every term in the sum above is O(n|S|/2+8εv+ε)O(n^{-|S|/2+8\varepsilon_{v}+\varepsilon}). Further, whenever S𝒮p,qu×2S\in\mathcal{S}^{u\times 2}_{p,q} and no row of SS equals (,)(\varnothing,\varnothing), then in particular every row contributes at least 1 to |S||S|, so |S|3|S|\geq 3. Therefore, we have the simpler bound that every term in the sum above is O(n|S|/2+8εv+ε)O(n3/2+9εv)O(n^{-|S|/2+8\varepsilon_{v}+\varepsilon)}\leq O(n^{-3/2+9\varepsilon_{v}}). So, the sum of O(n)O(n) such terms is always O(n1/2+9εv)O(n^{-1/2+9\varepsilon_{v}}), and up to such error, which is again subsumed in our current error term, we may restrict our attention to the case pqp_{\ell}\neq q_{\ell} and {p,q}{i,j}=\{p_{\ell},q_{\ell}\}\cap\{i,j\}=\varnothing. Since there are O(n2)O(n^{2}) such terms, we have:

|𝔼F~ij(𝒘)𝔼F~ij(𝒙)|\displaystyle|\mathbb{E}\widetilde{F}_{ij}(\bm{w})-\mathbb{E}\widetilde{F}_{ij}(\bm{x})| maxp,q[n] distinct{p,q}{i,j}=0u4,𝜿[2L]u,𝝈{id,conj}uS𝒮p,qu×2no row of S equals (,)|S|{3,4}n2+u|S|/2|𝔼[Φ𝜿,i,j(𝐐)J𝜿,i,j,S𝝈(𝐐)]|\displaystyle\lesssim\max_{\begin{subarray}{c}p,q\in[n]\text{ distinct}\\ \{p,q\}\cap\{i,j\}=\varnothing\\ 0\leq u\leq 4,\ \bm{\kappa}\in[2L]^{u},\bm{\sigma}\in\{\mathrm{id},\mathrm{conj}\}^{u}\\ S\in\mathcal{S}^{u\times 2}_{p,q}\\ \text{no row of $S$ equals $(\varnothing,\varnothing)$}\\ |S|\in\{3,4\}\end{subarray}}n^{2+u-|S|/2}\left|\mathbb{E}\left[\Phi_{\bm{\kappa},i,j}(\mathbf{Q})J^{\bm{\sigma}}_{\bm{\kappa},i,j,S}(\mathbf{Q})\right]\right|
+O(n1/2+ε).\displaystyle\hskip 85.35826pt+O(n^{-1/2+\varepsilon^{\prime}}).

Now, we note that when {p,q}{i,j}=\{p,q\}\cap\{i,j\}=\varnothing, then we can improve our above bound to |J𝜿,i,j,S𝝈(𝐐)|nu|S|0/2+2|S|εv|J^{\bm{\sigma}}_{\bm{\kappa},i,j,S}(\mathbf{Q})|\prec n^{-u-|S|_{0}/2+2|S|\varepsilon_{v}}. Thus, every expression in the maximum above, for a given value of uu, is bounded by O(n2|S|/2|S|0/2+8εv+ε)O(n^{2-|S|/2-|S|_{0}/2+8\varepsilon_{v}+\varepsilon}). All SS in the maximum have |S|01|S|_{0}\geq 1 (since otherwise they would have a row identically zero), so if |S|=4|S|=4 then any such term in the maximum has value O(n1/2+9εv)O(n^{-1/2+9\varepsilon_{v}}), again smaller than our current error term, and so such terms can effectively be ignored.

So, we may restrict our attention to terms with |S|=3|S|=3. In this case, if |S|02|S|_{0}\geq 2, then again such a term has value O(n1/2+9εv)O(n^{-1/2+9\varepsilon_{v}}) since |S|/2+|S|0/25/2|S|/2+|S|_{0}/2\geq 5/2, and such terms can likewise be ignored. So, we may finally restrict our attention to the case |S|=3|S|=3 and |S|0=1|S|_{0}=1. In this case, we must have u=1u=1, and S=[T]S=[T\,\,\,\varnothing] or S=[T]S=[\varnothing\,\,\,T] for some T𝒮p,q,3T\in\mathcal{S}_{p,q,3}. Expanding the definitions back out, we find

|𝔼F~ij(𝒘)𝔼F~ij(𝒙)|\displaystyle|\mathbb{E}\widetilde{F}_{ij}(\bm{w})-\mathbb{E}\widetilde{F}_{ij}(\bm{x})|
maxi,j[n]a[L],κ{2a1,2a}p,q[n] distinct {p,q}{i,j}=T𝒮p,q,3σ{id,conj}n3/2|𝔼[Φκ,i,j(𝐐)RQ(a)σ(i,αT,1)RQ(a)σ(βT,1,αT,2)RQ(a)σ(βT,2,αT,3)RQ(a)σ(βT,3,v(a))\displaystyle\lesssim\max_{\begin{subarray}{c}i,j\in[n]\\ a\in[L],\kappa\in\{2a-1,2a\}\\ p,q\in[n]\text{ distinct }\\ \{p,q\}\cap\{i,j\}=\varnothing\\ T\in\mathcal{S}_{p,q,3}\\ \sigma\in\{\mathrm{id},\mathrm{conj}\}\end{subarray}}n^{3/2}\bigg|\mathbb{E}\bigg[\Phi_{\kappa,i,j}(\mathbf{Q})R^{\sigma}_{Q^{(a)}}(i,\alpha_{T,1})R^{\sigma}_{Q^{(a)}}(\beta_{T,1},\alpha_{T,2})R^{\sigma}_{Q^{(a)}}(\beta_{T,2},\alpha_{T,3})R^{\sigma}_{Q^{(a)}}(\beta_{T,3},v^{(a)})
RQ(a)σ(v(a),j)]|+O(n1/2+ε),\displaystyle\hskip 91.04872pt\cdot R^{\sigma}_{Q^{(a)}}(v^{(a)},j)\bigg]\bigg|+O(n^{-1/2+\varepsilon^{\prime}}),
where the case of S=[T]S=[\varnothing\,\,\,T] is treated analogously with the decoupling performed from the right. Now, recall that given T𝒮p,q,3T\in\mathcal{S}_{p,q,3}, we have αT,d,βT,d{p,q}\alpha_{T,d},\beta_{T,d}\in\{p,q\}, and αT,dβT,d\alpha_{T,d}\neq\beta_{T,d} for d=1,2,3d=1,2,3 (so one is pp and the other is qq). Thus, whenever βT,1αT,2\beta_{T,1}\neq\alpha_{T,2} or βT,2αT,3\beta_{T,2}\neq\alpha_{T,3}, by Proposition 3.7 the expression in the maximum is O(n3/2n2+2εv+ε)O(n1/2+2εv+ε)O(n^{3/2}\cdot n^{-2+2\varepsilon_{v}+\varepsilon})\leq O(n^{-1/2+2\varepsilon_{v}+\varepsilon}), and the corresponding term is absorbed into the current error term. So, we may further reduce to a specific combination of resolvent quadratic forms, removing the dependence on TT and arriving at an entirely concrete expression:
maxi,j[n]a[L],κ{2a1,2a}p,q[n] distinct {p,q}{i,j}=σ{id,conj}n3/2|𝔼[Φκ,i,j(𝐐)RQ(a)σ(i,p)RQ(a)σ(p,p)RQ(a)σ(q,q)RQ(a)σ(q,v(a))RQ(a)σ(v(a),j)]|\displaystyle\lesssim\max_{\begin{subarray}{c}i,j\in[n]\\ a\in[L],\kappa\in\{2a-1,2a\}\\ p,q\in[n]\text{ distinct }\\ \{p,q\}\cap\{i,j\}=\varnothing\\ \sigma\in\{\mathrm{id},\mathrm{conj}\}\end{subarray}}n^{3/2}\bigg|\mathbb{E}\bigg[\Phi_{\kappa,i,j}(\mathbf{Q})R^{\sigma}_{Q^{(a)}}(i,p)R^{\sigma}_{Q^{(a)}}(p,p)R^{\sigma}_{Q^{(a)}}(q,q)R^{\sigma}_{Q^{(a)}}(q,v^{(a)})R^{\sigma}_{Q^{(a)}}(v^{(a)},j)\bigg]\bigg|
+O(n1/2+ε).\displaystyle\hskip 85.35826pt+O(n^{-1/2+\varepsilon^{\prime}}).

Expanding out the definition of Φκ,i,j(𝐐)\Phi_{\kappa,i,j}(\mathbf{Q}), we see that, for κ{2a1,2a}\kappa\in\{2a-1,2a\}, this factor is one of the first real partial derivatives of ϕ~\widetilde{\phi} evaluated at (nRQ(c)(i,v(c))RQ(c)(v(c),j))c=1L(n\cdot R_{Q^{(c)}}(i,v^{(c)})R_{Q^{(c)}}(v^{(c)},j))_{c=1}^{L}. Thus, by the boundedness of the derivatives of ϕ~\widetilde{\phi}, it is a bounded Lipschitz function of this LL-tuple. The proof is then completed upon using the result of Lemma 3.9 below. ∎

The last technical ingredient we will need is the following. We remove some of the specific details of the form of the Φκ,i,j(𝐐)\Phi_{\kappa,i,j}(\mathbf{Q}) factor for the sake of brevity in the proof.

Lemma 3.9.

In the above setting, suppose i,j[n]i,j\in[n] and that p,q[n]{i,j}p,q\in[n]\setminus\{i,j\} are distinct. Fix a0[L]a_{0}\in[L] and σ{id,conj}\sigma\in\{\mathrm{id},\mathrm{conj}\}. Let f:Lf:\mathbb{C}^{L}\to\mathbb{C} be a bounded Lipschitz function. Then, for any ε>0\varepsilon>0,

|𝔼[f((nRQ(a)(i,v(a))RQ(a)(v(a),j))a=1L)\displaystyle\bigg|\mathbb{E}\bigg[f\left(\left(n\cdot R_{Q^{(a)}}(i,v^{(a)})R_{Q^{(a)}}(v^{(a)},j)\right)_{a=1}^{L}\right)
RQ(a0)σ(i,p)RQ(a0)σ(p,p)RQ(a0)σ(q,q)RQ(a0)σ(q,v(a0))RQ(a0)σ(v(a0),j)]|ε,fn2+4εv+ε.\displaystyle\hskip 56.9055pt\cdot R^{\sigma}_{Q^{(a_{0})}}(i,p)R^{\sigma}_{Q^{(a_{0})}}(p,p)R^{\sigma}_{Q^{(a_{0})}}(q,q)R^{\sigma}_{Q^{(a_{0})}}(q,v^{(a_{0})})R^{\sigma}_{Q^{(a_{0})}}(v^{(a_{0})},j)\bigg]\bigg|\lesssim_{\varepsilon,f}n^{-2+4\varepsilon_{v}+\varepsilon}.

We note that this will conclude the proof of Theorem 1.4: Lemma 3.9 completes the proof of Lemma 3.4, and combining this with Lemma 3.3 in turn completes the proof of Theorem 1.4, as described earlier.

3.3 Resolvent monomial expectation via decoupling: Proof of Lemma 3.9

It suffices to prove the case σ=id\sigma=\mathrm{id}. Fix i,j,p,qi,j,p,q and a0a_{0} as in Lemma 3.9. For each a[L]a\in[L], let Q(a,p)herm(n1)×(n1)Q^{(a,p)}\in\mathbb{C}^{(n-1)\times(n-1)}_{\mathrm{herm}} be obtained from Q(a)Q^{(a)} by removing the pp-th row and column. Also, denote by

(p):=σ{Qμν(a):a[L],p{μ,ν}}\mathcal{F}^{(p)}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\sigma\{Q^{(a)}_{\mu\nu}:a\in[L],p\notin\{\mu,\nu\}\}

the σ\sigma-algebra generated by all entries of the tuple 𝐐\mathbf{Q} whose matrix indices do not touch pp. We write 𝔼p[]:=𝔼[|(p)]\mathbb{E}_{p}[\ \cdot\ ]\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\mathbb{E}[\ \cdot\ |\mathcal{F}^{(p)}] for the operation of taking conditional expectation with respect to this σ\sigma-algebra. In simpler language, 𝔼p\mathbb{E}_{p} is the operation of averaging over the pp-th row and column of all matrices Q(a)Q^{(a)}, a[L]a\in[L], jointly.

Since in this section we will only work with resolvents of matrices related to 𝐐\mathbf{Q}, and the proof is the same for both signs, let us slightly change notation and define, for a[L]a\in[L],

Ra\displaystyle R_{a} :=RQ(a)(λ~),\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}R_{Q^{(a)}}(\widetilde{\lambda}),

and write Ra(p)R^{(p)}_{a} for RQ(a,p)σ(λ~)R^{\sigma}_{Q^{(a,p)}}(\widetilde{\lambda}) extended to have dimension n×nn\times n by adding a pp-th row and column equal to zero. Here λ~=θ+θ1+𝒊n1/2\widetilde{\lambda}=\theta+\theta^{-1}+\bm{i}n^{-1/2} is the same complex value at which we have been evaluating all resolvents in the previous section as well.

For the distinguished tuple coordinate a0a_{0} appearing in Lemma 3.9, we abbreviate

R:=Ra0,R(p):=R(p)a0,v:=v(a0)R\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}R_{a_{0}},\qquad R^{(p)}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}R^{(p)}_{a_{0}},\qquad v\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}v^{(a_{0})}

while the test function ff may still depend on the whole tuple (nRa(i,v(a))Ra(v(a),j))a=1L\left(n\cdot R_{a}(i,v^{(a)})R_{a}(v^{(a)},j)\right)_{a=1}^{L}. The identities below are stated for a general tuple coordinate a[L]a\in[L]; in the proof of Lemma 3.9 we will mostly use them with a=a0a=a_{0}, in which case they are written using the abbreviated notation R,R(p),vR,R^{(p)},v.

For the remaining analysis, we use the method of fluctuation averaging (see Section 7 of [BGK16]). This method is based on the observation that 𝔼pR(i,p)\mathbb{E}_{p}R(i,p) is much smaller than the typical value of R(i,p)R(i,p) with no averaging. The following are the main technical devices used to make this argument, identities relating resolvents to their minors.

Proposition 3.10.

For any a[L]a\in[L] and any μ,ν,k[n]\mu,\nu,k\in[n] with k{μ,ν}k\notin\{\mu,\nu\}, we have

Ra(μ,ν)=Ra(k)(μ,ν)+Ra(μ,k)Ra(k,ν)Ra(k,k).R_{a}(\mu,\nu)=R^{(k)}_{a}(\mu,\nu)+\frac{R_{a}(\mu,k)R_{a}(k,\nu)}{R_{a}(k,k)}.
Proof.

This is a special case of the formula for the inverse of the entries of a block matrix, where we view the (k,k)(k,k) entry as a 1×11\times 1 block. ∎

Proposition 3.11.

For any a[L]a\in[L] and any distinct μ,ν[n]\mu,\nu\in[n], we have

Ra(μ,ν)\displaystyle R_{a}(\mu,\nu) =Ra(μ,μ)kμQμk(a)Ra(μ)(k,ν),\displaystyle=R_{a}(\mu,\mu)\sum_{k\neq\mu}Q^{(a)}_{\mu k}R^{(\mu)}_{a}(k,\nu),
Ra(μ,ν)\displaystyle R_{a}(\mu,\nu) =Ra(ν,ν)kνRa(ν)(μ,k)Qkν(a).\displaystyle=R_{a}(\nu,\nu)\sum_{k\neq\nu}R^{(\nu)}_{a}(\mu,k)Q^{(a)}_{k\nu}.
Proof.

This is a special case of the same formula referenced above, but now where we view the (μ,μ)(\mu,\mu) or (ν,ν)(\nu,\nu) entry as the 1×11\times 1 block with respect to which we expand the inverse. ∎

The following is then the main estimate that we will use. Note that Proposition 3.7 implies only the a priori bound 𝔼|Ra(m,r)|2εn1+ε\mathbb{E}|R_{a}(m,r)|^{2}\lesssim_{\varepsilon}n^{-1+\varepsilon}, on which this result improves considerably.

Lemma 3.12.

For any a[L]a\in[L] and any distinct i,p[n]i,p\in[n] , for any ε>0\varepsilon>0,

𝔼|𝔼pRa(i,p)|2εn2+ε.\mathbb{E}|\mathbb{E}_{p}R_{a}(i,p)|^{2}\lesssim_{\varepsilon}n^{-2+\varepsilon}.
Proof.

Fix a[L]a\in[L] and distinct i,p[n]i,p\in[n]. Recall the abbreviation

R=Ra,R(p)=Ra(p),Q=Q(a).R=R_{a},\qquad R^{(p)}=R^{(p)}_{a},\qquad Q=Q^{(a)}.

By Proposition 3.11, for indices ipi\neq p,

R(i,p)=R(p,p)kpR(p)(i,k)Qkp.R(i,p)=R(p,p)\sum_{k\neq p}R^{(p)}(i,k)Q_{kp}.

Decompose R(p,p)=Gμsc(λ~)+(R(p,p)Gμsc(λ~))R(p,p)=G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})+\big(R(p,p)-G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})\big), where GμscG_{\mu_{\mathrm{sc}}} is the Cauchy transform of the semicircle law. Then

R(i,p)=Gμsc(λ~)kpR(p)(i,k)Qkp+(R(p,p)Gμsc(λ~))kpR(p)(i,k)Qkp.R(i,p)=G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})\cdot\sum_{k\neq p}R^{(p)}(i,k)Q_{kp}+\big(R(p,p)-G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})\big)\cdot\sum_{k\neq p}R^{(p)}(i,k)Q_{kp}.

Now we consider taking the conditional expectation 𝔼p\mathbb{E}_{p} (conditional on (p)\mathcal{F}^{(p)}). Since R(p)R^{(p)} is (p)\mathcal{F}^{(p)}-measurable (i.e., independent of row and column pp of Q=Q(a)Q=Q^{(a)}) and by the definition of generalized Wigner matrices we have 𝔼p[Qkp]=0\mathbb{E}_{p}[Q_{kp}]=0, we have

𝔼p[Gμsc(λ~)kpR(p)(i,k)Qkp]=Gμsc(λ~)kpR(p)(i,k)𝔼p[Qkp]=0.\mathbb{E}_{p}\Big[G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})\cdot\sum_{k\neq p}R^{(p)}(i,k)Q_{kp}\Big]=G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})\cdot\sum_{k\neq p}R^{(p)}(i,k)\mathbb{E}_{p}[Q_{kp}]=0.

So, we are left with just

𝔼p[R(i,p)]=𝔼p[(R(p,p)Gμsc(λ~))kpR(p)(i,k)Qkp].\mathbb{E}_{p}[R(i,p)]=\mathbb{E}_{p}\left[\big(R(p,p)-G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})\big)\cdot\sum_{k\neq p}R^{(p)}(i,k)Q_{kp}\right].

Applying the Cauchy-Schwarz inequality on this conditional expectation, we have

|𝔼p[R(i,p)]|2(𝔼p|R(p,p)Gμsc(λ~)|2)(𝔼p|kpR(p)(i,k)Qkp|2),\Big|\mathbb{E}_{p}[R(i,p)]\Big|^{2}\leq\Big(\mathbb{E}_{p}|R(p,p)-G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})|^{2}\Big)\left(\mathbb{E}_{p}\left|\sum_{k\neq p}R^{(p)}(i,k)Q_{kp}\right|^{2}\right), (3.13)

and taking the unconditional expectation,

𝔼|𝔼p[R(i,p)]|2𝔼[𝔼p|R(p,p)Gμsc(λ~)|2𝔼p|kpR(p)(i,k)Qkp|2].\mathbb{E}\Big|\mathbb{E}_{p}[R(i,p)]\Big|^{2}\leq\mathbb{E}\left[\mathbb{E}_{p}|R(p,p)-G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})|^{2}\ \cdot\mathbb{E}_{p}\left|\sum_{k\neq p}R^{(p)}(i,k)Q_{kp}\right|^{2}\right]. (3.14)

Expanding the second factor, we see that we may simplify it to

𝔼p|kpR(p)(i,k)Qkp|2\displaystyle\mathbb{E}_{p}\Big|\sum_{k\neq p}R^{(p)}(i,k)Q_{kp}\Big|^{2} =𝔼p[k,lpR(p)(i,k)R(p)(i,l)¯QkpQlp¯]\displaystyle=\mathbb{E}_{p}\bigg[\sum_{k,l\neq p}R^{(p)}(i,k)\overline{R^{(p)}(i,l)}\cdot Q_{kp}\overline{Q_{lp}}\bigg]
=k,lpR(p)(i,k)R(p)(i,l)¯𝔼p[QkpQlp¯]\displaystyle=\sum_{k,l\neq p}R^{(p)}(i,k)\overline{R^{(p)}(i,l)}\mathbb{E}_{p}\left[Q_{kp}\overline{Q_{lp}}\right]
=kp𝔼p|Qkp|2|R(p)(i,k)|2\displaystyle=\sum_{k\neq p}\mathbb{E}_{p}|Q_{kp}|^{2}\,|R^{(p)}(i,k)|^{2}

We have 𝔼p[|Qkp|2]=𝔼[|Qkp|2]Cn1\mathbb{E}_{p}\left[|Q_{kp}|^{2}\right]=\mathbb{E}\left[|Q_{kp}|^{2}\right]\leq Cn^{-1} for some C>0C>0 using our definition of generalized Wigner matrices and that QQ is formed from a matrix of this kind by setting some entries to zero. So, we have

𝔼p|kpR(p)(i,k)Qkp|2n1k|R(p)(i,k)|2,\mathbb{E}_{p}\Big|\sum_{k\neq p}R^{(p)}(i,k)Q_{kp}\Big|^{2}\lesssim n^{-1}\cdot\sum_{k}\big|R^{(p)}(i,k)\big|^{2},

which is an (p)\mathcal{F}^{(p)}-measurable random variable (not depending on row and column pp of QQ).

Substituting back, we may then rewrite as a full expectation,

𝔼|𝔼p[R(i,p)]|2\displaystyle\mathbb{E}\Big|\mathbb{E}_{p}[R(i,p)]\Big|^{2} n1𝔼[𝔼p[|R(p,p)Gμsc(λ~)|2]k|R(p)(i,k)|2]\displaystyle\lesssim n^{-1}\cdot\mathbb{E}\left[\mathbb{E}_{p}[|R(p,p)-G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})|^{2}]\cdot\sum_{k}\big|R^{(p)}(i,k)\big|^{2}\right]
=n1𝔼[|R(p,p)Gμsc(λ~)|2k|R(p)(i,k)|2]\displaystyle=n^{-1}\cdot\mathbb{E}\left[|R(p,p)-G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})|^{2}\cdot\sum_{k}\big|R^{(p)}(i,k)\big|^{2}\right]
using that k|R(p)(i,k)|2=(R(p))ei2R(p)2\sum_{k}|R^{(p)}(i,k)|^{2}=\|(R^{(p)})^{*}e_{i}\|^{2}\leq\|R^{(p)}\|^{2}, we have
n1𝔼[|R(p,p)Gμsc(λ~)|2R(p)2]\displaystyle\leq n^{-1}\cdot\mathbb{E}\left[|R(p,p)-G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})|^{2}\cdot\|R^{(p)}\|^{2}\right]
n1(𝔼|R(p,p)Gμsc(λ~)|4)1/2(𝔼R(p)4)1/2.\displaystyle\leq n^{-1}\cdot\left(\mathbb{E}|R(p,p)-G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})|^{4}\right)^{1/2}\cdot\left(\mathbb{E}\|R^{(p)}\|^{4}\right)^{1/2}.

By Proposition 2.17, the resolvent norm factor is O(1)O(1). By Corollary 2.19 (the isotropic local law for QQ), Proposition 2.5, and Proposition 2.17, the second factor is Oε(n1+ε)O_{\varepsilon}(n^{-1+\varepsilon}) for any ε>0\varepsilon>0. The result then follows from combining these bounds. ∎

Now we use this to give the proof of the main result of this section.

Proof of Lemma 3.9.

We note before continuing that the bounds on resolvent entries and quadratic forms with eαe_{\alpha} and v(a)v^{(a)} proved in Proposition 3.7 all apply equally well to Ra(p)R^{(p)}_{a}, for every a[L]a\in[L]. Indeed, Ra(p)R^{(p)}_{a} is zero outside of a principal (n1)×(n1)(n-1)\times(n-1) submatrix that equals the resolvent of Q(a,p)Q^{(a,p)}, where Q(a,p)Q^{(a,p)} is either a Wigner matrix of dimension n1n-1 or such a matrix with one or two entries set to zero, to which Proposition 3.7 applies directly.

We would like to move towards applying Lemma 3.12. To that end, define

𝒜\displaystyle\mathcal{A} :=f((nRa(i,v(a))Ra(v(a),j))a=1L)R(p,p)R(q,q)R(q,v)R(v,j),\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}f\bigg(\left(n\cdot R_{a}(i,v^{(a)})R_{a}(v^{(a)},j)\right)_{a=1}^{L}\bigg)\cdot R(p,p)R(q,q)R(q,v)R(v,j),
𝒜(p)\displaystyle\mathcal{A}^{(p)} :=f((nRa(p)(i,v(a))Ra(p)(v(a),j))a=1L)Gμsc(λ~)R(p)(q,q)R(p)(q,v)R(p)(v,j).\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}f\bigg(\left(n\cdot R_{a}^{(p)}(i,v^{(a)})R_{a}^{(p)}(v^{(a)},j)\right)_{a=1}^{L}\bigg)\cdot G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})R^{(p)}(q,q)R^{(p)}(q,v)R^{(p)}(v,j).

Recall that we are trying to prove a bound on 𝔼[𝒜R(i,p)]\mathbb{E}[\mathcal{A}\cdot R(i,p)]. Since 𝒜(p)\mathcal{A}^{(p)} is (p)\mathcal{F}^{(p)}-measurable, we will try to replace 𝒜\mathcal{A} by 𝒜(p)\mathcal{A}^{(p)}, noting that we expect these to be close by since we make only small perturbations by replacing various RaR_{a} by Ra(p)R_{a}^{(p)}, and another small perturbation in replacing R(p,p)R(p,p) by Gμsc(λ~)G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda}) according to the local law in Corollary 2.19. We have:

|𝔼[𝒜R(i,p)]|\displaystyle\big|\mathbb{E}[\mathcal{A}\cdot R(i,p)]\big| =|𝔼[(𝒜𝒜(p))R(i,p)]+𝔼[𝒜(p)R(i,p)]|\displaystyle=|\mathbb{E}[(\mathcal{A}-\mathcal{A}^{(p)})\cdot R(i,p)]+\mathbb{E}[\mathcal{A}^{(p)}\cdot R(i,p)]|
|𝔼[(𝒜𝒜(p))R(i,p)]|+|𝔼[𝒜(p)𝔼p[R(i,p)]]|,\displaystyle\leq\big|\mathbb{E}[(\mathcal{A}-\mathcal{A}^{(p)})\cdot R(i,p)]\big|+\big|\mathbb{E}[\mathcal{A}^{(p)}\cdot\mathbb{E}_{p}[R(i,p)]]\big|, (3.15)

where in the second term we have used that 𝒜(p)\mathcal{A}^{(p)} is (p)\mathcal{F}^{(p)}-measurable.

We now bound each term individually. Consider the second term of (3.15) first. Using Cauchy–Schwarz,

|𝔼[𝒜(p)𝔼p[R(i,p)]]|(𝔼|𝒜(p)|2)1/2(𝔼|𝔼p[R(i,p)]|2)1/2\Big|\mathbb{E}\big[\mathcal{A}^{(p)}\cdot\mathbb{E}_{p}[R(i,p)]\big]\Big|\leq\Big(\mathbb{E}|\mathcal{A}^{(p)}|^{2}\Big)^{1/2}\Big(\mathbb{E}\big|\mathbb{E}_{p}[R(i,p)]\big|^{2}\Big)^{1/2}

As shown in Lemma 3.12, we already have

𝔼|𝔼p[R(i,p)]|2n2+ε\mathbb{E}\big|\mathbb{E}_{p}[R(i,p)]\big|^{2}\lesssim n^{-2+\varepsilon}

for any ε>0\varepsilon>0. It remains to bound 𝔼|𝒜(p)|2\mathbb{E}|\mathcal{A}^{(p)}|^{2}. Since ff is bounded, we have by Proposition 3.7

|𝒜(p)|f|Gμsc(λ~)||R(p)(q,q)||R(p)(q,v)||R(p)(v,j)|n1+2εv,|\mathcal{A}^{(p)}|\leq\|f\|_{\infty}\cdot|G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})|\cdot|R^{(p)}(q,q)|\cdot|R^{(p)}(q,v)|\cdot|R^{(p)}(v,j)|\prec n^{-1+2\varepsilon_{v}},

and by the expectation bounds in Proposition 3.7 we have

𝔼|𝒜(p)|2n2+4εv+ε.\mathbb{E}|\mathcal{A}^{(p)}|^{2}\lesssim n^{-2+4\varepsilon_{v}+\varepsilon}.

for any ε>0\varepsilon>0. Thus, the entire second term of (3.15) is bounded by

|𝔼[𝒜(p)𝔼p[R(i,p)]]|n1+2εv+ε/2n1+ε/2=n2+2εv+ε.\big|\mathbb{E}[\mathcal{A}^{(p)}\cdot\mathbb{E}_{p}[R(i,p)]]\big|\lesssim n^{-1+2\varepsilon_{v}+\varepsilon/2}\cdot n^{-1+\varepsilon/2}=n^{-2+2\varepsilon_{v}+\varepsilon}. (3.16)

Now we consider the first term in (3.15). We first reorganize the expression 𝒜𝒜(p)\mathcal{A}-\mathcal{A}^{(p)}. Define

\displaystyle\mathcal{B} :=R(p,p)R(q,q)R(q,v)R(v,j),\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}R(p,p)R(q,q)R(q,v)R(v,j),
(p)\displaystyle\mathcal{B}^{(p)} :=Gμsc(λ~)R(p)(q,q)R(p)(q,v)R(p)(v,j).\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})R^{(p)}(q,q)R^{(p)}(q,v)R^{(p)}(v,j).

Then, we have

𝒜𝒜(p)=(f((nRa(i,v(a))Ra(v(a),j))a=1L)f((nRa(p)(i,v(a))Ra(p)(v(a),j))a=1L))\displaystyle\mathcal{A}-\mathcal{A}^{(p)}=\left(f\Big(\big(n\cdot R_{a}(i,v^{(a)})R_{a}(v^{(a)},j)\big)_{a=1}^{L}\Big)-f\Big(\big(n\cdot R_{a}^{(p)}(i,v^{(a)})R_{a}^{(p)}(v^{(a)},j)\big)_{a=1}^{L}\Big)\right)\cdot\mathcal{B} (3.17)
+f((nRa(p)(i,v(a))Ra(p)(v(a),j))a=1L)((p)).\displaystyle+f\left(\left(n\cdot R_{a}^{(p)}(i,v^{(a)})R_{a}^{(p)}(v^{(a)},j)\right)_{a=1}^{L}\right)\cdot(\mathcal{B}-\mathcal{B}^{(p)}).

We first bound the first term on the right-hand side of (3.17). For each a[L]a\in[L], Proposition 3.10 gives

Ra(i,v(a))Ra(p)(i,v(a))=Ra(i,p)Ra(p,v(a))Ra(p,p),R_{a}(i,v^{(a)})-R_{a}^{(p)}(i,v^{(a)})=\frac{R_{a}(i,p)R_{a}(p,v^{(a)})}{R_{a}(p,p)},

and

Ra(v(a),j)Ra(p)(v(a),j)=Ra(v(a),p)Ra(p,j)Ra(p,p).R_{a}(v^{(a)},j)-R_{a}^{(p)}(v^{(a)},j)=\frac{R_{a}(v^{(a)},p)R_{a}(p,j)}{R_{a}(p,p)}.

Therefore, by Proposition 3.7,

|Ra(i,v(a))Ra(p)(i,v(a))|n1+εv,|Ra(v(a),j)Ra(p)(v(a),j)|n1+εv.|R_{a}(i,v^{(a)})-R_{a}^{(p)}(i,v^{(a)})|\prec n^{-1+\varepsilon_{v}},\qquad|R_{a}(v^{(a)},j)-R_{a}^{(p)}(v^{(a)},j)|\prec n^{-1+\varepsilon_{v}}.

It follows that

|nRa(i,v(a))Ra(v(a),j)nRa(p)(i,v(a))Ra(p)(v(a),j)|n|Ra(i,v(a))Ra(p)(i,v(a))||Ra(v(a),j)|\displaystyle\left|n\cdot R_{a}(i,v^{(a)})R_{a}(v^{(a)},j)-n\cdot R_{a}^{(p)}(i,v^{(a)})R_{a}^{(p)}(v^{(a)},j)\right|\leq n|R_{a}(i,v^{(a)})-R_{a}^{(p)}(i,v^{(a)})||R_{a}(v^{(a)},j)|
+n|Ra(p)(i,v(a))||Ra(v(a),j)Ra(p)(v(a),j)|n1/2+2εv.\displaystyle\hskip 91.04872pt+n|R_{a}^{(p)}(i,v^{(a)})|\,|R_{a}(v^{(a)},j)-R_{a}^{(p)}(v^{(a)},j)|\prec n^{-1/2+2\varepsilon_{v}}.

Since ff is Lipschitz and LL is fixed

|f((nRa(i,v(a))Ra(v(a),j))a=1L)f((nRa(p)(i,v(a))Ra(p)(v(a),j))a=1L)|n1/2+2εv.\bigg|f\Big(\big(n\cdot R_{a}(i,v^{(a)})R_{a}(v^{(a)},j)\big)_{a=1}^{L}\Big)-f\Big(\big(n\cdot R_{a}^{(p)}(i,v^{(a)})R_{a}^{(p)}(v^{(a)},j)\big)_{a=1}^{L}\Big)\bigg|\prec n^{-1/2+2\varepsilon_{v}}.

Again by Proposition 3.7, we have

||n1+2εv,|\mathcal{B}|\prec n^{-1+2\varepsilon_{v}},

so by Proposition 2.4, the first term of (3.17) is polynomially stochastically dominated as

|(f((nRa(i,v(a))Ra(v(a),j))a=1L)f((nRa(p)(i,v(a))Ra(p)(v(a),j))a=1L))|n3/2+4εv.\bigg|\Big(f\Big(\big(n\cdot R_{a}(i,v^{(a)})R_{a}(v^{(a)},j)\big)_{a=1}^{L}\Big)-f\Big(\big(n\cdot R_{a}^{(p)}(i,v^{(a)})R_{a}^{(p)}(v^{(a)},j)\big)_{a=1}^{L}\Big)\Big)\cdot\mathcal{B}\bigg|\prec n^{-3/2+4\varepsilon_{v}}. (3.18)

Meanwhile, for the second term of (3.17), we may apply a telescoping expansion to (p)\mathcal{B}-\mathcal{B}^{(p)}, obtaining

(p)\displaystyle\mathcal{B}-\mathcal{B}^{(p)} =(R(p,p)Gμsc(λ~))R(q,q)R(q,v)R(v,j)\displaystyle=\big(R(p,p)-G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})\big)R(q,q)R(q,v)R(v,j)
+Gμsc(λ~)(R(q,q)R(p)(q,q))R(q,v)R(v,j)\displaystyle\hskip 28.45274pt+G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})\big(R(q,q)-R^{(p)}(q,q)\big)R(q,v)R(v,j)
+Gμsc(λ~)R(p)(q,q)(R(q,v)R(p)(q,v))R(v,j)\displaystyle\hskip 28.45274pt+G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})R^{(p)}(q,q)\big(R(q,v)-R^{(p)}(q,v)\big)R(v,j)
+Gμsc(λ~)R(p)(q,q)R(p)(q,v)(R(v,j)R(p)(v,j)).\displaystyle\hskip 28.45274pt+G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})R^{(p)}(q,q)R^{(p)}(q,v)\big(R(v,j)-R^{(p)}(v,j)\big).

We now control the individual differences that appear here. The first term above is

(R(p,p)Gμsc(λ~))R(q,q)R(q,v)R(v,j)n1/2+ε1n1/2+εvn1/2+εvn3/2+2εv+ε\big(R(p,p)-G_{\mu_{\mathrm{sc}}}(\widetilde{\lambda})\big)R(q,q)R(q,v)R(v,j)\prec n^{-1/2+\varepsilon}\cdot 1\cdot n^{-1/2+\varepsilon_{v}}\cdot n^{-1/2+\varepsilon_{v}}\prec n^{-3/2+2\varepsilon_{v}+\varepsilon}

For the remaining differences, Proposition 3.10 gives

R(q,q)R(p)(q,q)=R(q,p)R(p,q)R(p,p)n1,R(q,q)-R^{(p)}(q,q)=\frac{R(q,p)R(p,q)}{R(p,p)}\prec n^{-1},

and

R(q,v)R(p)(q,v)=R(q,p)R(p,v)R(p,p)n1+εv,R(q,v)-R^{(p)}(q,v)=\frac{R(q,p)R(p,v)}{R(p,p)}\prec n^{-1+\varepsilon_{v}},

with the analogous bound

R(v,j)R(p)(v,j)n1+εv.R(v,j)-R^{(p)}(v,j)\prec n^{-1+\varepsilon_{v}}.

Putting everything together, we find

|(p)|n3/2+2εv+ε,|\mathcal{B}-\mathcal{B}^{(p)}|\prec n^{-3/2+2\varepsilon_{v}+\varepsilon},

and therefore the whole second term in (3.17) is bounded as

|f((nRa(p)(i,v(a))Ra(p)(v(a),j))a=1L)((p))|n3/2+2εv+ε\bigg|f\Big(\big(n\cdot R_{a}^{(p)}(i,v^{(a)})R_{a}^{(p)}(v^{(a)},j)\big)_{a=1}^{L}\Big)\cdot(\mathcal{B}-\mathcal{B}^{(p)})\bigg|\prec n^{-3/2+2\varepsilon_{v}+\varepsilon}

as well, since ff is bounded. Combining this with (3.18) and plugging both bounds into (3.17), we have

|𝒜𝒜(p)|n3/2+4εv+ε.|\mathcal{A}-\mathcal{A}^{(p)}|\prec n^{-3/2+4\varepsilon_{v}+\varepsilon}.

Since R(i,p)n1/2R(i,p)\prec n^{-1/2}, we obtain

|(𝒜𝒜(p))R(i,p)|n2+4εv,|(\mathcal{A}-\mathcal{A}^{(p)})\cdot R(i,p)|\prec n^{-2+4\varepsilon_{v}},

and thus

|𝔼[(𝒜𝒜(p))R(i,p)]|ε,fn2+4εv+ε.|\mathbb{E}[(\mathcal{A}-\mathcal{A}^{(p)})\cdot R(i,p)]|\lesssim_{\varepsilon,f}n^{-2+4\varepsilon_{v}+\varepsilon}.

Finally, applying our bounds to either term of (3.15), we get

|𝔼[𝒜R(i,p)]|εn2+4εv+ε,|\mathbb{E}[\mathcal{A}\cdot R(i,p)]|\lesssim_{\varepsilon}n^{-2+4\varepsilon_{v}+\varepsilon},

completing the proof. ∎

4 Gaussian formula for entrywise averages: Proof of Theorem 1.8

Recall the setting of the Theorem: unlike the above proof of Theorem 1.4, the statement of this result depends on the fields to which the entries of the matrices in the tuple 𝐖\mathbf{W} belong. So, as in the Introduction, for each a[L]a\in[L], let us write

𝔽a={if W(a) is real symmetric,if W(a) is complex Hermitian}.\mathbb{F}_{a}=\left\{\begin{array}[]{ll}\mathbb{R}&\text{if }W^{(a)}\text{ is real symmetric},\\ \mathbb{C}&\text{if }W^{(a)}\text{ is complex Hermitian}\end{array}\right\}.

Recall that we write 𝒩𝔽a(0,1)\mathcal{N}_{\mathbb{F}_{a}}(0,1) for the standard Gaussian scalar distribution associated to the field 𝔽a\mathbb{F}_{a} (normalized so that 𝔼|Z|2=1\mathbb{E}|Z|^{2}=1 for Z𝒩𝔽a(0,1)Z\in\mathcal{N}_{\mathbb{F}_{a}}(0,1) in either case), and we write 𝒩𝔽a(0,In)\mathcal{N}_{\mathbb{F}_{a}}(0,I_{n}) for the law of a vector of nn i.i.d. such Gaussians.

In Theorem 1.8, we have ψ:L×L\psi:\mathbb{C}^{L}\times\mathbb{C}^{L}\to\mathbb{R} a 𝒞5\mathcal{C}^{5} function with bounded values and first five derivatives, and an associated function Ψ:(hermn×n)L×(hermn×n)L\Psi:(\mathbb{C}^{n\times n}_{\mathrm{herm}})^{L}\times(\mathbb{C}^{n\times n}_{\mathrm{herm}})^{L}\to\mathbb{R} defined by

Ψ(𝐀,𝐁):=1n2i,j=1nψ((Aij(a))a=1L,(Bij(a))a=1L).\Psi(\mathbf{A},\mathbf{B})\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\frac{1}{n^{2}}\sum_{i,j=1}^{n}\psi\left(\left(A^{(a)}_{ij}\right)_{a=1}^{L},\left(B^{(a)}_{ij}\right)_{a=1}^{L}\right).

Theorem 1.8 then bounds, for v^(W,a)=v1(θv(a)v(a)+W(a))\widehat{v}^{(W,a)}=v_{1}\left(\theta v^{(a)}v^{(a)*}+W^{(a)}\right) and g(a)𝒩𝔽a(0,In)g^{(a)}\sim\mathcal{N}_{\mathbb{F}_{a}}(0,I_{n}), with the Gaussian vectors g(a)g^{(a)} independent for a[L]a\in[L], the quantity

|𝔼Ψ((nv(a)v(a))a=1L,(nv^(W,a)v^(W,a))a=1L)\displaystyle\Bigg|\mathbb{E}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\widehat{v}^{(W,a)}\widehat{v}^{(W,a)*}\right)_{a=1}^{L}\right)
𝔼Ψ((nv(a)v(a))a=1L,((ρ(θ)nv(a)+τ(θ)g(a))(ρ(θ)nv(a)+τ(θ)g(a)))a=1L)|.\displaystyle\hskip 56.9055pt-\mathbb{E}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left((\rho(\theta)\sqrt{n}v^{(a)}+\tau(\theta)g^{(a)})(\rho(\theta)\sqrt{n}v^{(a)}+\tau(\theta)g^{(a)})^{*}\right)_{a=1}^{L}\right)\Bigg|.

Before proceeding, let us establish a few tools for working with such functions Ψ\Psi. First, it follows immediately from our assumptions that Ψ\Psi is O(n1)O(n^{-1})-Lipschitz:

Proposition 4.1.

For any 𝐀,𝐁,𝐀,𝐁(hermn×n)L\mathbf{A},\mathbf{B},\mathbf{A}^{\prime},\mathbf{B}^{\prime}\in(\mathbb{C}^{n\times n}_{\mathrm{herm}})^{L}, writing 𝐀=(A(a))a=1L\mathbf{A}=(A^{(a)})_{a=1}^{L}, 𝐁=(B(a))a=1L\mathbf{B}=(B^{(a)})_{a=1}^{L}, 𝐀=(A(a))a=1L\mathbf{A}^{\prime}=(A^{\prime(a)})_{a=1}^{L} and 𝐁=(B(a))a=1L\mathbf{B}^{\prime}=(B^{\prime(a)})_{a=1}^{L}, we have

|Ψ(𝐀,𝐁)Ψ(𝐀,𝐁)|Cψna=1L(A(a)A(a)F+B(a)B(a)F),|\Psi(\mathbf{A},\mathbf{B})-\Psi(\mathbf{A}^{\prime},\mathbf{B}^{\prime})|\leq\frac{C_{\psi}}{n}\sum_{a=1}^{L}\left(\|A^{(a)}-A^{\prime(a)}\|_{\mathrm{F}}+\|B^{(a)}-B^{\prime(a)}\|_{\mathrm{F}}\right),

for a constant CψC_{\psi} depending only on LL and the bounds on ψ\psi and its first derivatives.

Proof.

For each i,j[n]i,j\in[n], write 𝐀ij:=(Aij(a))a=1L\mathbf{A}_{ij}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}(A^{(a)}_{ij})_{a=1}^{L}, 𝐁ij:=(Bij(a))a=1L\mathbf{B}_{ij}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}(B^{(a)}_{ij})_{a=1}^{L}, and define 𝐀ij\mathbf{A}^{\prime}_{ij}, 𝐁ij\mathbf{B}^{\prime}_{ij} similarly. Then, the result follows since we have entrywise

|ψ(𝐀ij,𝐁ij)ψ(𝐀ij,𝐁ij)|\displaystyle\left|\psi(\mathbf{A}_{ij},\mathbf{B}_{ij})-\psi(\mathbf{A}^{\prime}_{ij},\mathbf{B}^{\prime}_{ij})\right| Cψa=1L(|Aij(a)Aij(a)|+|Bij(a)Bij(a)|)\displaystyle\leq C_{\psi}\sum_{a=1}^{L}\left(|A^{(a)}_{ij}-A^{\prime(a)}_{ij}|+|B^{(a)}_{ij}-B^{\prime(a)}_{ij}|\right)

for a suitable CψC_{\psi}. Thus, we have

|Ψ(𝐀,𝐁)Ψ(𝐀,𝐁)|\displaystyle|\Psi(\mathbf{A},\mathbf{B})-\Psi(\mathbf{A}^{\prime},\mathbf{B}^{\prime})| Cψn2i,j=1na=1L(|Aij(a)Aij(a)|+|Bij(a)Bij(a)|)\displaystyle\leq\frac{C_{\psi}}{n^{2}}\sum_{i,j=1}^{n}\sum_{a=1}^{L}\left(\big|A^{(a)}_{ij}-A^{\prime(a)}_{ij}\big|+\big|B^{(a)}_{ij}-B^{\prime(a)}_{ij}\big|\right)
Cψna=1L(A(a)A(a)F+B(a)B(a)F)\displaystyle\leq\frac{C_{\psi}}{n}\sum_{a=1}^{L}\left(\|A^{(a)}-A^{\prime(a)}\|_{\mathrm{F}}+\|B^{(a)}-B^{\prime(a)}\|_{\mathrm{F}}\right)

using the Cauchy-Schwarz inequality in the sum over (i,j)(i,j) for each fixed aa. ∎

Next, we give a standard bound on the change in eigenvector projections under small perturbations of a matrix; the result we give is a simple reformulation of the Davis-Kahan inequality; see [DK70, YWS15].

Proposition 4.2.

Let H,Δhermn×nH,\Delta\in\mathbb{C}^{n\times n}_{\mathrm{herm}}. Let λ1(H)\lambda_{1}(H) be a simple eigenvalue and suppose Δ<12(λ1(H)λ2(H))\|\Delta\|<\frac{1}{2}\big(\lambda_{1}(H)-\lambda_{2}(H)\big). Then,

v1(H)v1(H)v1(H+Δ)v1(H+Δ)F22Δλ1(H)λ2(H).\|v_{1}(H)v_{1}(H)^{*}-v_{1}(H+\Delta)v_{1}(H+\Delta)^{*}\|_{\mathrm{F}}\leq 2\sqrt{2}\cdot\frac{\|\Delta\|}{\lambda_{1}(H)-\lambda_{2}(H)}.

Combining the two results and rescaling per our setting gives the following useful corollary.

Corollary 4.3.

Let 𝐖=(W(a))a=1L\mathbf{W}=(W^{(a)})_{a=1}^{L} and 𝐖=(W(a))a=1L\mathbf{W}^{\prime}=(W^{\prime(a)})_{a=1}^{L} be two tuples of Hermitian matrices and for each a[L]a\in[L], consider H(a)=θv(a)v(a)+W(a)H^{(a)}=\theta v^{(a)}v^{(a)*}+W^{(a)} and H(a)=θv(a)v(a)+W(a)H^{\prime(a)}=\theta v^{(a)}v^{(a)*}+W^{\prime(a)}. Write v^(W,a)=v1(H(a))\widehat{v}^{(W,a)}=v_{1}(H^{(a)}) and v^(W,a)=v1(H(a))\widehat{v}^{(W^{\prime},a)}=v_{1}(H^{\prime(a)}) for the associated top eigenvectors. If, for every a[L]a\in[L], W(a)W(a)<12(λ1(H(a))λ2(H(a)))\|W^{(a)}-W^{\prime(a)}\|<\frac{1}{2}\big(\lambda_{1}(H^{(a)})-\lambda_{2}(H^{(a)})\big), then

|Ψ((nv(a)v(a))a=1L,(nv^(W,a)v^(W,a))a=1L)Ψ((nv(a)v(a))a=1L,(nv^(W,a)v^(W,a))a=1L)|\displaystyle\Bigg|\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\widehat{v}^{(W,a)}\widehat{v}^{(W,a)*}\right)_{a=1}^{L}\right)-\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\widehat{v}^{(W^{\prime},a)}\widehat{v}^{(W^{\prime},a)*}\right)_{a=1}^{L}\right)\Bigg|
22Cψa=1LW(a)W(a)λ1(H(a))λ2(H(a)),\displaystyle\hskip 56.9055pt\leq 2\sqrt{2}C_{\psi}\sum_{a=1}^{L}\frac{\|W^{(a)}-W^{\prime(a)}\|}{\lambda_{1}(H^{(a)})-\lambda_{2}(H^{(a)})},

where CψC_{\psi} is the constant from Proposition 4.1.

By Theorem 2.21, applied separately to each coordinate a[L]a\in[L], and since LL is fixed, we have that mina[L](λ1(H(a))λ2(H(a)))=Ω(1)\min_{a\in[L]}\left(\lambda_{1}(H^{(a)})-\lambda_{2}(H^{(a)})\right)=\Omega(1) with polynomially high probability in all settings below where the matrices W(a)W^{(a)} are generalized Wigner matrices with uniformly controlled parameters. Therefore, in effect, coordinatewise perturbations of the tuple 𝐖\mathbf{W} satisfying maxa[L]W(a)W(a)=o(1)\max_{a\in[L]}\|W^{(a)}-W^{\prime(a)}\|=o(1) have a small effect on the value of Ψ\Psi we are interested in. We will use this observation several times below.

Returning to the main proof, for each a[L]a\in[L], let us define

G𝔽Ea(n):={GOE(n)if 𝔽a=,GUE(n)if 𝔽a=}.\mathrm{G}\mathbb{F}\mathrm{E}_{a}(n)\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left\{\begin{array}[]{ll}\mathrm{GOE}(n)&\text{if }\mathbb{F}_{a}=\mathbb{R},\\ \mathrm{GUE}(n)&\text{if }\mathbb{F}_{a}=\mathbb{C}\end{array}\right\}.

We write 𝐗a=1LG𝔽Ea(n)\mathbf{X}\sim\prod_{a=1}^{L}\mathrm{G}\mathbb{F}\mathrm{E}_{a}(n) to mean that X(1),,X(L)X^{(1)},\ldots,X^{(L)} are independent and X(a)G𝔽Ea(n)X^{(a)}\sim\mathrm{G}\mathbb{F}\mathrm{E}_{a}(n) for each a[L]a\in[L].

We prove the Theorem in several steps. First, we show using the above argument that we may, at the cost of a small error, replace the weakly Wigner tuple 𝐖\mathbf{W} with a tuple 𝐖(reg)\mathbf{W}^{(\mathrm{reg})} such that each coordinate is a generalized Wigner matrix satisfying the exact variance normalization (1.3).

Theorem 4.4 (Special case of symmetric Sinkhorn scaling theorem, Theorem 5.4 of [Ide16]).

Let n2n\geq 2 and Ssymn×nS\in\mathbb{R}^{n\times n}_{\mathrm{sym}} be a non-negative matrix with Sij>0S_{ij}>0 for all i,j[n]i,j\in[n] distinct. Then, there exists a diagonal matrix X=Diag(x1,,xn)X=\mathrm{Diag}(x_{1},\ldots,x_{n}) with positive entries such that XSXXSX has row sums equal 11.

Lemma 4.5.

Let S:=(sij)i,j=1nsymn×nS\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}(s_{ij})_{i,j=1}^{n}\in\mathbb{R}^{n\times n}_{\mathrm{sym}} be a nonnegative matrix. Suppose that,

maxj=1ni[n]|sij1n|n1+nεW,\max_{i\in[n]}\sum_{j=1}^{n}\bigg|s_{ij}-\frac{1}{n}\bigg|\lesssim n^{-1}+n^{-\varepsilon_{W}},

and sij>0s_{ij}>0 for all iji\neq j. Then, for sufficiently large nn, there exist positive d1,,dnd_{1},\ldots,d_{n} such that

j=1nsijdidj=1,i[n]\sum_{j=1}^{n}\frac{s_{ij}}{d_{i}d_{j}}=1,\qquad i\in[n] (4.1)

with

maxi[n]|1di|n1+nεW.\max_{i\in[n]}|1-d_{i}|\lesssim n^{-1}+n^{-\varepsilon_{W}}. (4.2)
Proof.

By Theorem 4.4, there exists positive X=Diag(x1,,xn)X=\mathrm{Diag}(x_{1},\ldots,x_{n}) such that xij=1nsijxj=1x_{i}\sum_{j=1}^{n}s_{ij}x_{j}=1 for all i[n]i\in[n]. Setting di=xi1d_{i}=x_{i}^{-1} gives (4.1). Thus, for all i[n]i\in[n], di=j=1nsijdjd_{i}=\sum_{j=1}^{n}\frac{s_{ij}}{d_{j}}. To show (4.2), first define diM=maxi[n]did_{i_{M}}=\max_{i\in[n]}d_{i} and dim=mini[n]did_{i_{m}}=\min_{i\in[n]}d_{i}. Then |diMdim|=|j=1nsiMjsimjdj|n1+nεWdim.|d_{i_{M}}-d_{i_{m}}|=\left|\sum_{j=1}^{n}\frac{s_{i_{M}j}-s_{i_{m}j}}{d_{j}}\right|\lesssim\frac{n^{-1}+n^{-\varepsilon_{W}}}{d_{i_{m}}}. Also, 1O(n1+nεW)j=1nsij1+O(n1+nεW)1-O(n^{-1}+n^{-\varepsilon_{W}})\leq\sum_{j=1}^{n}s_{ij}\leq 1+O(n^{-1}+n^{-\varepsilon_{W}}). Hence, dim=j=1nsimjdj1O(n1+nεW)diMd_{i_{m}}=\sum_{j=1}^{n}\frac{s_{i_{m}j}}{d_{j}}\geq\frac{1-O(n^{-1}+n^{-\varepsilon_{W}})}{d_{i_{M}}} and diM=j=1nsiMjdj1+O(n1+nεW)dimd_{i_{M}}=\sum_{j=1}^{n}\frac{s_{i_{M}j}}{d_{j}}\leq\frac{1+O(n^{-1}+n^{-\varepsilon_{W}})}{d_{i_{m}}}, which implies diMdim=1+O(n1+nεW)d_{i_{M}}d_{i_{m}}=1+O(n^{-1}+n^{-\varepsilon_{W}}). Also, dim(diMdim)n1+nεWd_{i_{m}}(d_{i_{M}}-d_{i_{m}})\lesssim n^{-1}+n^{-\varepsilon_{W}}. Then, dim=(diMdimdim(diMdim))1/2=1+O(n1+nεW)d_{i_{m}}=(d_{i_{M}}d_{i_{m}}-d_{i_{m}}(d_{i_{M}}-d_{i_{m}}))^{1/2}=1+O(n^{-1}+n^{-\varepsilon_{W}}) and hence diM=1+O(n1+nεW)d_{i_{M}}=1+O(n^{-1}+n^{-\varepsilon_{W}}). The claim follows. ∎

Lemma 4.6.

Let L1L\geq 1 be fixed, θ>1\theta>1, and let 𝐖=(W(a))a=1L\mathbf{W}=(W^{(a)})_{a=1}^{L} be a weakly Wigner tuple with parameters (ξW,εW,CW)(\xi_{W},\varepsilon_{W},C_{W}). Let v(a)𝕊n1(𝔽a)v^{(a)}\in\mathbb{S}^{n-1}(\mathbb{F}_{a}) be deterministic for each a[L]a\in[L].

For each a[L]a\in[L], define

sij(a)\displaystyle s^{(a)}_{ij} :=𝔼|W(a)ij|2,\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\mathbb{E}|W^{(a)}_{ij}|^{2},
S(a)\displaystyle S^{(a)} :=(sij(a))i,j=1n,\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}(s^{(a)}_{ij})_{i,j=1}^{n},

Let D(a):=Diag(d1(a),,dn(a))D^{(a)}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\mathrm{Diag}(d^{(a)}_{1},\ldots,d^{(a)}_{n}) and write

W(reg,a):=(D(a))1/2W(a)(D(a))1/2,𝐖reg:=(W(reg,a))a=1LW^{(\mathrm{reg},a)}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}(D^{(a)})^{-1/2}W^{(a)}(D^{(a)})^{-1/2},\qquad\mathbf{W}^{\mathrm{reg}}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left(W^{(\mathrm{reg},a)}\right)_{a=1}^{L}

Then, for each a[L]a\in[L], W(reg,a)W^{(\mathrm{reg},a)} is a generalized Wigner matrix satisfying (1.3) with uniformly controlled parameters. Moreover, 𝐖(reg)\mathbf{W}^{(\mathrm{reg})} is again a weakly Wigner tuple with parameters depending only on (ξW,εW,CW)(\xi_{W},\varepsilon_{W},C_{W}) and LL. Also, we have, for any ε>0\varepsilon>0,

|𝔼Ψ((nv(a)v(a))a=1L,(nv^(W,a)v^(W,a))a=1L)\displaystyle\bigg|\mathbb{E}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\widehat{v}^{(W,a)}\widehat{v}^{(W,a)*}\right)_{a=1}^{L}\right)
𝔼Ψ((nv(a)v(a))a=1L,(nv^(W(reg),a)v^(W(reg),a))a=1L)|n1+ε+nεW+ε,\displaystyle\hskip 56.9055pt-\mathbb{E}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\widehat{v}^{(W^{(\mathrm{reg})},a)}\widehat{v}^{(W^{(\mathrm{reg})},a)*}\right)_{a=1}^{L}\right)\bigg|\lesssim n^{-1+\varepsilon}+n^{-\varepsilon_{W}+\varepsilon}, (4.3)

where

v^(W(reg),a):=v1(θv(a)v(a)+W(reg,a)).\widehat{v}^{(W^{(\mathrm{reg})},a)}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}v_{1}\left(\theta v^{(a)}v^{(a)*}+W^{(\mathrm{reg},a)}\right).
Proof.

For each a[L]a\in[L], by the weakly Wigner tuple assumption, we have sij(a)=𝔼|Wij(a)|2=1n+O(n1εW)>0s^{(a)}_{ij}=\mathbb{E}|W^{(a)}_{ij}|^{2}=\frac{1}{n}+O(n^{-1-\varepsilon_{W}})>0 for iji\neq j and sii(a)=O(n1)s^{(a)}_{ii}=O(n^{-1}) uniformly in aa. Then

maxj=1ni[n]|sij(a)1n|maxjii[n]|sij(a)1n|+maxi[n]|sii(a)1n|n1+nεW.\max_{i\in[n]}\sum_{j=1}^{n}\left|s^{(a)}_{ij}-\frac{1}{n}\right|\leq\max_{i\in[n]}\sum_{j\neq i}\left|s^{(a)}_{ij}-\frac{1}{n}\right|+\max_{i\in[n]}\left|s^{(a)}_{ii}-\frac{1}{n}\right|\lesssim n^{-1}+n^{-\varepsilon_{W}}.

Apply Lemma 4.5 to S(a)S^{(a)} for each a[L]a\in[L], we then have

j=1nsij(a)di(a)dj(a)=1for all i[n].\sum_{j=1}^{n}\frac{s^{(a)}_{ij}}{d^{(a)}_{i}d^{(a)}_{j}}=1\qquad\text{for all }i\in[n].

Choose D(a):=Diag(d1(a),,dn(a))D^{(a)}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\mathrm{Diag}(d^{(a)}_{1},\ldots,d^{(a)}_{n}) so that W(reg,a)W^{(\mathrm{reg},a)} satisfies (1.3).

By Lemma 4.5, we have maxa[L]maxi[n]|1di(a)|n1+nεW\max_{a\in[L]}\max_{i\in[n]}\big|1-d^{(a)}_{i}\big|\lesssim n^{-1}+n^{-\varepsilon_{W}}, and it then follows from elementary bounds that 𝐖(reg)\mathbf{W}^{(\mathrm{reg})} remains a weakly Wigner tuple as well.

For the final bound, note that we have by the above observation ID(a)+I(D(a))1/2n1+nεW\|I-D^{(a)}\|+\|I-(D^{(a)})^{-1/2}\|\lesssim n^{-1}+n^{-\varepsilon_{W}}, and Proposition 2.16 gives (D(a))1/2W(a)(D(a))1/21\|(D^{(a)})^{-1/2}W^{(a)}(D^{(a)})^{-1/2}\|\prec 1. So, we may bound

W(a)W(reg,a)\displaystyle\|W^{(a)}-W^{(\mathrm{reg},a)}\| =W(a)(D(a))1/2W(a)(D(a))1/2\displaystyle=\|W^{(a)}-(D^{(a)})^{-1/2}W^{(a)}(D^{(a)})^{-1/2}\|
W(a)(I(D(a))1/2)+(I(D(a))1/2)W(a)(D(a))1/2\displaystyle\leq\|W^{(a)}(I-(D^{(a)})^{-1/2})+(I-(D^{(a)})^{-1/2})W^{(a)}(D^{(a)})^{-1/2}\|
W(a)I(D(a))1/2(1+(D(a))1/2)\displaystyle\leq\|W^{(a)}\|\cdot\|I-(D^{(a)})^{-1/2}\|\cdot\big(1+\|(D^{(a)})^{-1/2}\|\big)
n1+nεW.\displaystyle\prec n^{-1}+n^{-\varepsilon_{W}}.

The result then follows from Corollary 4.3, with Theorem 2.21 used as mentioned after the Corollary. ∎

Next, we show that we may, at the cost of a small error, replace 𝐖(reg)\mathbf{W}^{(\mathrm{reg})} in the above quantity with a Gaussian tuple.

Lemma 4.7.

Let L1L\geq 1 be fixed, θ>1\theta>1, and let v(a)𝕊n1(𝔽a)v^{(a)}\in\mathbb{S}^{n-1}(\mathbb{F}_{a}) deterministic satisfying Assumption 1.3 uniformly for all a[L]a\in[L]. Let 𝐖(reg)=(W(reg,a))a=1L\mathbf{W}^{(\mathrm{reg})}=\left(W^{(\mathrm{reg},a)}\right)_{a=1}^{L} be a weakly Wigner tuple such that each coordinate W(reg,a)W^{(\mathrm{reg},a)} is a generalized Wigner matrix with uniformly controlled parameters and satisfies the exact variance normalization (1.3). Let 𝐗(0)=(X(0,a))a=1L\mathbf{X}^{(0)}=\left(X^{(0,a)}\right)_{a=1}^{L} be a centered Gaussian tuple with the same first two moments as 𝐖(reg)\mathbf{W}^{(\mathrm{reg})} such that for each 1pqn1\leq p\leq q\leq n, the real-coordinate block

x(0,pq)=(ReXpq(0,1),ImXpq(0,1),,ReXpq(0,L),ImXpq(0,L))x^{(0,pq)}=\left(\mathrm{Re}X^{(0,1)}_{pq},\mathrm{Im}X^{(0,1)}_{pq},\ldots,\mathrm{Re}X^{(0,L)}_{pq},\mathrm{Im}X^{(0,L)}_{pq}\right)

is centered Gaussian, the blocks are independent over 1pqn1\leq p\leq q\leq n, and

𝔼x(0,pq)x(0,pq)=𝔼w(reg,pq)w(reg,pq),\mathbb{E}x^{(0,pq)}x^{(0,pq)\top}=\mathbb{E}w^{(\mathrm{reg},pq)}w^{(\mathrm{reg},pq)\top},

where w(reg,pq)w^{(\mathrm{reg},pq)} is the corresponding real-coordinate block of 𝐖(reg)\mathbf{W}^{(\mathrm{reg})}. Write H(Y,a):=θv(a)v(a)+Y(a)H^{(Y,a)}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\theta v^{(a)}v^{(a)*}+Y^{(a)} for Y(a){W(reg,a),X(0,a)}Y^{(a)}\in\{W^{(\mathrm{reg},a)},X^{(0,a)}\} and v^(Y,a):=v1(H(Y,a))\widehat{v}^{(Y,a)}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}v_{1}(H^{(Y,a)}). Let ψ,Ψ\psi,\Psi be as above. Then, for any ε>0\varepsilon>0,

|𝔼Ψ((nv(a)v(a))a=1L,(nv^(W(reg),a)v^(W(reg),a))a=1L)\displaystyle\bigg|\mathbb{E}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\widehat{v}^{(W^{(\mathrm{reg})},a)}\widehat{v}^{(W^{(\mathrm{reg})},a)*}\right)_{a=1}^{L}\right)
𝔼Ψ((nv(a)v(a))a=1L,(nv^(X(0),a)v^(X(0),a))a=1L)|n1/2+10εv+ε.\displaystyle\hskip 56.9055pt-\mathbb{E}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\widehat{v}^{(X^{(0)},a)}\widehat{v}^{(X^{(0)},a)*}\right)_{a=1}^{L}\right)\bigg|\lesssim n^{-1/2+10\varepsilon_{v}+\varepsilon}.
Proof.

This is a simple consequence of Theorem 1.4: note that 𝐖(reg)\mathbf{W}^{(\mathrm{reg})} and 𝐗(0)\mathbf{X}^{(0)} satisfy the assumptions of that result, and we may bound

|𝔼Ψ((nv(a)v(a))a=1L,(nv^(W(reg),a)v^(W(reg),a))a=1L)\displaystyle\bigg|\mathbb{E}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\widehat{v}^{(W^{(\mathrm{reg})},a)}\widehat{v}^{(W^{(\mathrm{reg})},a)*}\right)_{a=1}^{L}\right)
𝔼Ψ((nv(a)v(a))a=1L,(nv^(X(0),a)v^(X(0),a))a=1L)|\displaystyle\hskip 99.58464pt-\mathbb{E}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\widehat{v}^{(X^{(0)},a)}\widehat{v}^{(X^{(0)},a)*}\right)_{a=1}^{L}\right)\bigg|
1n2i,j=1n|𝔼ψ((nvi(a)vj(a)¯)a=1L,(nv^i(W(reg),a)v^j(W(reg),a)¯)a=1L)\displaystyle\leq\frac{1}{n^{2}}\sum_{i,j=1}^{n}\Bigg|\mathbb{E}\psi\left(\left(nv^{(a)}_{i}\overline{v^{(a)}_{j}}\right)_{a=1}^{L},\left(n\widehat{v}^{(W^{(\mathrm{reg})},a)}_{i}\overline{\widehat{v}^{(W^{(\mathrm{reg})},a)}_{j}}\right)_{a=1}^{L}\right)
𝔼ψ((nvi(a)vj(a)¯)a=1L,(nv^i(X(0),a)v^j(X(0),a)¯)a=1L)|\displaystyle\hskip 99.58464pt-\mathbb{E}\psi\left(\left(nv^{(a)}_{i}\overline{v^{(a)}_{j}}\right)_{a=1}^{L},\left(n\widehat{v}^{(X^{(0)},a)}_{i}\overline{\widehat{v}^{(X^{(0)},a)}_{j}}\right)_{a=1}^{L}\right)\Bigg|
1n2i,j=1nn1/2+10εv+εn1/2+10εv+ε\displaystyle\lesssim\frac{1}{n^{2}}\sum_{i,j=1}^{n}n^{-1/2+10\varepsilon_{v}+\varepsilon}\lesssim n^{-1/2+10\varepsilon_{v}+\varepsilon}

as claimed. ∎

Note that above 𝐗(0)\mathbf{X}^{(0)} is itself a weakly Wigner tuple, and each X(0,a)X^{(0,a)} is a generalized Wigner matrix, since these conditions depend only on the first two moments and Gaussian entries have the required tail bounds. We next show that the top eigenvectors of any such 𝐗(0)\mathbf{X}^{(0)} are close to those of a Gaussian tuple 𝐗=(X(1),,X(L))a=1LG𝔽Ea(n)\mathbf{X}=(X^{(1)},\ldots,X^{(L)})\sim\bigotimes_{a=1}^{L}\mathrm{G}\mathbb{F}\mathrm{E}_{a}(n). The simple idea is that the weakly Wigner tuple assumption implies that we may couple 𝐗(0)\mathbf{X}^{(0)} and 𝐗\mathbf{X} so that maxa[L]X(0,a)X(a)n1/2+nεW/2\max_{a\in[L]}\|X^{(0,a)}-X^{(a)}\|\prec n^{-1/2}+n^{-\varepsilon_{W}/2}, whereby the top eigenvectors of outlier eigenvalues of rank-one perturbations of these matrices are close for all a[L]a\in[L] by standard eigenvector perturbation inequalities.

Lemma 4.8.

In the setting of Lemma 4.7, let 𝐗=(X(1),,X(L))a=1LG𝔽Ea(n)\mathbf{X}=(X^{(1)},\ldots,X^{(L)})\sim\bigotimes_{a=1}^{L}\mathrm{G}\mathbb{F}\mathrm{E}_{a}(n). Then, there exists a coupling between 𝐗(0)\mathbf{X}^{(0)} and 𝐗\mathbf{X} such that

a=1Lv^(X(0),a)v^(X(0),a)v^(X,a)v^(X,a)Fn1/2+nεW/2\sum_{a=1}^{L}\left\|\widehat{v}^{(X^{(0)},a)}\widehat{v}^{(X^{(0)},a)*}-\widehat{v}^{(X,a)}\widehat{v}^{(X,a)*}\right\|_{\mathrm{F}}\prec n^{-1/2}+n^{-\varepsilon_{W}/2} (4.4)
Proof.

From the condition of being weakly Wigner tuple, we may couple 𝐗(0)\mathbf{X}^{(0)} with a tuple 𝐗a=1LG𝔽Ea(n)\mathbf{X}\sim\bigotimes_{a=1}^{L}\mathrm{G}\mathbb{F}\mathrm{E}_{a}(n) so that, for each a[L]a\in[L],

X(0,a)X(a)=δXoffdiag(a)+Δoffdiag(a)+Δdiag(a),X^{(0,a)}-X^{(a)}=-\delta X^{(a)}_{\mathrm{offdiag}}+\Delta^{(a)}_{\mathrm{offdiag}}+\Delta^{(a)}_{\mathrm{diag}},

where δ=O(nεW)\delta=O(n^{-\varepsilon_{W}}), Xoffdiag(a)1\|X^{(a)}_{\mathrm{offdiag}}\|\prec 1, Δdiag(a)\Delta^{(a)}_{\mathrm{diag}} is diagonal with independent Gaussian diagonal entries having variance O(n1)O(n^{-1}), and Δoffdiag(a)\Delta^{(a)}_{\mathrm{offdiag}} is the Hermitian matrix formed from the aa-th coordinate of the off-diagonal residual blocks. To see that this decomposition exists, with δ=O(nεW)\delta=O(n^{-\varepsilon_{W}}), for each off-diagonal block, Σ0(pq)(1δ)2ΣG𝔽E(pq)=(Σ0(pq)ΣG𝔽E(pq))+(1(1δ)2)ΣG𝔽E(pq)0\Sigma^{(pq)}_{0}-(1-\delta)^{2}\Sigma^{(pq)}_{\mathrm{G}\mathbb{F}\mathrm{E}}=(\Sigma^{(pq)}_{0}-\Sigma^{(pq)}_{\mathrm{G}\mathbb{F}\mathrm{E}})+(1-(1-\delta)^{2})\Sigma^{(pq)}_{\mathrm{G}\mathbb{F}\mathrm{E}}\succeq 0, and this residual covariance is O(n1εW)O(n^{-1-\varepsilon_{W}}).

We next bound Δoffdiag(a)\Delta^{(a)}_{\mathrm{offdiag}}. From standard bounds such as those of Theorem 1.1 and Corollary 3.9 of [BvH16], it follows that Δ(a)offdiagnεW/2\|\Delta^{(a)}_{\mathrm{offdiag}}\|\prec n^{-\varepsilon_{W}/2}. Also, Δ(a)diagn1/2\|\Delta^{(a)}_{\mathrm{diag}}\|\prec n^{-1/2}. Thus, using a union bound over fixed LL, we have maxa[L]X(0,a)X(a)n1/2+nεW/2\max_{a\in[L]}\|X^{(0,a)}-X^{(a)}\|\prec n^{-1/2}+n^{-\varepsilon_{W}/2}.

By Theorem 2.21, again with a union bound over fixed LL, with polynomially high probability, the gap between the top two eigenvalues of θv(a)v(a)+X(a)\theta v^{(a)}v^{(a)*}+X^{(a)} is of size Ω(1)\Omega(1). The result then follows by Proposition 4.2. ∎

The third and more substantial part of the proof is a direct analysis of such expressions for G𝔽E\mathrm{G}\mathbb{F}\mathrm{E} tuples. For the sake of brevity below, for each a[L]a\in[L], let us define

y(a)=y(θ,v(a),g(a)):=ρ(θ)v(a)+τ(θ)ng(a),y^{(a)}=y(\theta,v^{(a)},g^{(a)})\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\rho(\theta)v^{(a)}+\frac{\tau(\theta)}{\sqrt{n}}g^{(a)},

where g(a)𝒩𝔽a(0,In)g^{(a)}\sim\mathcal{N}_{\mathbb{F}_{a}}(0,I_{n}) are independent over a[L]a\in[L] and 𝒚:=(y(a))a=1L\bm{y}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}(y^{(a)})_{a=1}^{L}.

Lemma 4.9.

In the setting of Lemma 4.8, let 𝐠=(g(1),,g(L))\bm{g}=(g^{(1)},\ldots,g^{(L)}) where g(a)𝒩𝔽a(0,In)g^{(a)}\sim\mathcal{N}_{\mathbb{F}_{a}}(0,I_{n}) are independent over a[L]a\in[L] and have i.i.d. entries. Then there exists a coupling between 𝐗\mathbf{X} and 𝐠\bm{g} such that

a=1Lv^(X,a)v^(X,a)y(a)y(a)Fn1/2.\sum_{a=1}^{L}\left\|\widehat{v}^{(X,a)}\widehat{v}^{(X,a)*}-y^{(a)}y^{(a)*}\right\|_{\mathrm{F}}\prec n^{-1/2}. (4.5)

This is our main technical tool, which we prove in the following section. For now, let us see how this together with Lemma 4.7 completes the proof of Theorem 1.8.

Corollary 4.10.

In the setting of Lemma 4.9, we have that, for any ε>0\varepsilon>0,

|𝔼Ψ((nv(a)v(a))a=1L,(nv^(X(0),a)v^(X(0),a))a=1L)𝔼Ψ((nv(a)v(a))a=1L,(ny(a)y(a))a=1L)|\displaystyle\Bigg|\mathbb{E}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\widehat{v}^{(X^{(0)},a)}\widehat{v}^{(X^{(0)},a)*}\right)_{a=1}^{L}\right)-\mathbb{E}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(ny^{(a)}y^{(a)*}\right)_{a=1}^{L}\right)\Bigg|
n1/2+ε+nεW/2+ε.\displaystyle\hskip 85.35826pt\lesssim n^{-1/2+\varepsilon}+n^{-\varepsilon_{W}/2+\varepsilon}.
Proof.

Combining Lemma 4.8 and Lemma 4.9, we may realize 𝐗(0)\mathbf{X}^{(0)}, 𝐗\mathbf{X} and 𝒈\bm{g} on the same probability space so that

a=1Lv^(X(0),a)v^(X(0),a)y(a)y(a)Fn1/2+nεW/2.\sum_{a=1}^{L}\left\|\widehat{v}^{(X^{(0)},a)}\widehat{v}^{(X^{(0)},a)*}-y^{(a)}y^{(a)*}\right\|_{\mathrm{F}}\prec n^{-1/2}+n^{-\varepsilon_{W}/2}.

Fix ε>0\varepsilon>0, and define the event

={a=1Lv^(X(0),a)v^(X(0),a)y(a)y(a)Fn1/2+ε+nεW/2+ε}.\mathcal{E}=\left\{\sum_{a=1}^{L}\left\|\widehat{v}^{(X^{(0)},a)}\widehat{v}^{(X^{(0)},a)*}-y^{(a)}y^{(a)*}\right\|_{\mathrm{F}}\leq n^{-1/2+\varepsilon}+n^{-\varepsilon_{W}/2+\varepsilon}\right\}.

Then, by polynomial stochastic domination, for every fixed D>0D>0 we have [c]nD\mathbb{P}[\mathcal{E}^{c}]\leq n^{-D} for all sufficiently large nn. Set

A\displaystyle A :=Ψ((nv(a)v(a))a=1L,(nv^(X(0),a)v^(X(0),a))a=1L),\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\widehat{v}^{(X^{(0)},a)}\widehat{v}^{(X^{(0)},a)*}\right)_{a=1}^{L}\right),
B\displaystyle B :=Ψ((nv(a)v(a))a=1L,(ny(a)y(a))a=1L).\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\Psi\left(\left(nv^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(ny^{(a)}y^{(a)*}\right)_{a=1}^{L}\right).

Since ψ\psi is bounded, Ψ\Psi is bounded. Then by Proposition 4.1, on \mathcal{E} we have

|AB|Cψa=1Lv^(X(0),a)v^(X(0),a)y(a)y(a)FCψ(n1/2+ε+nεW/2+ε),|A-B|\leq C_{\psi}\sum_{a=1}^{L}\left\|\widehat{v}^{(X^{(0)},a)}\widehat{v}^{(X^{(0)},a)*}-y^{(a)}y^{(a)*}\right\|_{\mathrm{F}}\leq C_{\psi}\left(n^{-1/2+\varepsilon}+n^{-\varepsilon_{W}/2+\varepsilon}\right),

where the factor nn in the inputs cancels by the 1/n1/n Lipschitz factor from Proposition 4.1. Using the boundedness of Ψ\Psi on c\mathcal{E}^{c},

|𝔼A𝔼B|𝔼𝟏|AB|+𝔼𝟏c|AB|n1/2+ε+nεW/2+ε+nD.|\mathbb{E}A-\mathbb{E}B|\leq\mathbb{E}\mathbf{1}_{\mathcal{E}}|A-B|+\mathbb{E}\mathbf{1}_{\mathcal{E}^{c}}|A-B|\lesssim n^{-1/2+\varepsilon}+n^{-\varepsilon_{W}/2+\varepsilon}+n^{-D}.

Taking DD large enough gives the claim. ∎

Theorem 1.8 then follows from combining Corollary 4.10 with Lemma 4.6 and Lemma 4.7. It remains to prove Lemma 4.9, which we do below.

4.1 Analysis of Gaussian noise models: Proof of Lemma 4.9

It will also be useful to define

𝒰𝔽(n)={𝒪(n), the n×n orthogonal groupif 𝔽=,𝒰(n), the n×n unitary groupif 𝔽=}.\mathcal{U}_{\mathbb{F}}(n)=\left\{\begin{array}[]{ll}\mathcal{O}(n),\text{ the $n\times n$ orthogonal group}&\text{if }\mathbb{F}=\mathbb{R},\\ \mathcal{U}(n),\text{ the $n\times n$ unitary group}&\text{if }\mathbb{F}=\mathbb{C}\end{array}\right\}.

The following is the main property of G𝔽E\mathrm{G}\mathbb{F}\mathrm{E} matrices that we will use.

Proposition 4.11.

For each 𝔽{,}\mathbb{F}\in\{\mathbb{R},\mathbb{C}\}, the following hold:

  • g𝒩𝔽(0,In)g\sim\mathcal{N}_{\mathbb{F}}(0,I_{n}) has a 𝒰𝔽(n)\mathcal{U}_{\mathbb{F}}(n) invariant law as a vector; that is, for all U𝒰𝔽(n)U\in\mathcal{U}_{\mathbb{F}}(n), Ug=(law)gUg\stackrel{{\scriptstyle\text{(law)}}}{{=}}g.

  • XG𝔽E(n)X\sim\mathrm{G}\mathbb{F}\mathrm{E}(n) has a 𝒰𝔽(n)\mathcal{U}_{\mathbb{F}}(n)-invariant law as Hermitian matrix; that is, for all U𝒰𝔽(n)U\in\mathcal{U}_{\mathbb{F}}(n), UXU=(law)XUXU^{*}\stackrel{{\scriptstyle\text{(law)}}}{{=}}X.

Proof of Lemma 4.9.

It suffices to prove the desired estimate for each fixed a[L]a\in[L], uniformly in aa, since LL is fixed. Fix a[L]a\in[L] and suppress the superscript aa throughout this proof, writing

𝔽=𝔽a,v=v(a),X=X(a),g=g(a).\mathbb{F}=\mathbb{F}_{a},\qquad v=v^{(a)},\qquad X=X^{(a)},\qquad g=g^{(a)}.

Let us write v^(X):=v1(θvv+X)\widehat{v}(X)\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}v_{1}(\theta vv^{*}+X). For U𝒰𝔽(n)U\in\mathcal{U}_{\mathbb{F}}(n), we have by Proposition 4.11 the equality of distributions

v^(X)=(law)v^(UXU)=v1(θvv+UXU)=Uv1(θ(Uv)(Uv)+X).\widehat{v}(X)\stackrel{{\scriptstyle\text{(law)}}}{{=}}\widehat{v}(UXU^{*})=v_{1}(\theta vv^{*}+UXU^{*})=Uv_{1}(\theta(U^{*}v)(U^{*}v)^{*}+X).

We may choose such a deterministic UU with Uv=e1U^{*}v=e_{1} (note here the delicate point that we are using that v𝔽nv\in\mathbb{F}^{n}, avoiding the case where 𝔽=\mathbb{F}=\mathbb{R} while vv has complex values), so we find that there is some such (deterministic) UU for which

v^(X)=(law)Uv1(θe1e1+X).\widehat{v}(X)\stackrel{{\scriptstyle\text{(law)}}}{{=}}Uv_{1}(\theta e_{1}e_{1}^{*}+X).

Define

y~:=ρ(θ)e1+τ(θ)ng.\widetilde{y}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\rho(\theta)e_{1}+\frac{\tau(\theta)}{\sqrt{n}}g.

By Proposition 4.11 again, we have

y~=(law)Uy.\widetilde{y}\stackrel{{\scriptstyle\text{(law)}}}{{=}}U^{*}y.

Let us couple the two sides of this equality so that, on the same probability space, we have y~=Uy\widetilde{y}=U^{*}y. Then, we have

v1(θe1e1+X)v1(θe1e1+X)y~y~F\displaystyle\left\|v_{1}(\theta e_{1}e_{1}^{*}+X)v_{1}(\theta e_{1}e_{1}^{*}+X)^{*}-\widetilde{y}\widetilde{y}^{*}\right\|_{\mathrm{F}} =Uv1(θe1e1+X)v1(θe1e1+X)UUy~y~UF\displaystyle=\left\|Uv_{1}(\theta e_{1}e_{1}^{*}+X)v_{1}(\theta e_{1}e_{1}^{*}+X)^{*}U^{*}-U\widetilde{y}\widetilde{y}^{*}U^{*}\right\|_{\mathrm{F}}
=(law)v^(X)v^(X)yyF\displaystyle\stackrel{{\scriptstyle\text{(law)}}}{{=}}\left\|\widehat{v}(X)\widehat{v}(X)^{*}-yy^{*}\right\|_{\mathrm{F}}

for a corresponding coupling of XX and gg. Thus, in effect we may take v=e1v=e_{1} without loss of generality, so let us make this assumption going forward, in which case we may also take y=y~y=\widetilde{y}.

We will only consider a single distribution of XX now, so let us take XG𝔽E(n)X\sim\mathrm{G}\mathbb{F}\mathrm{E}(n) and v^:=v^(X)=v1(θe1e1+X)\widehat{v}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\widehat{v}(X)=v_{1}(\theta e_{1}e_{1}^{*}+X). We remove the sign (for 𝔽=\mathbb{F}=\mathbb{R}) or phase (for 𝔽=\mathbb{F}=\mathbb{C}) ambiguity by choosing v^\widehat{v} to have its first coordinate real and non-negative. Write v^\widehat{v}^{\perp} for v^\widehat{v} with the first entry set to zero, so that we have

v^=|v^,e1|e1+v^.\widehat{v}=|\langle\widehat{v},e_{1}\rangle|e_{1}+\widehat{v}^{\perp}. (4.6)

Let 𝒰𝔽(1)(n)\mathcal{U}_{\mathbb{F}}^{(1)}(n) be the subgroup of matrices in 𝒰𝔽(n)\mathcal{U}_{\mathbb{F}}(n) that fix e1e_{1}, equivalently matrices of the form Diag(1,U0)\mathrm{Diag}(1,U_{0}) with U0𝒰𝔽(n1)U_{0}\in\mathcal{U}_{\mathbb{F}}(n-1). Then, θe1e1+X\theta e_{1}e_{1}^{*}+X is 𝒰𝔽(1)(n)\mathcal{U}_{\mathbb{F}}^{(1)}(n)-invariant, in the same sense as in the second part of Proposition 4.11, since the first summand is fixed by the action of these matrices and the distribution of the second summand is unchanged. Thus, its top eigenvector v^\widehat{v} is 𝒰𝔽(1)(n)\mathcal{U}_{\mathbb{F}}^{(1)}(n)-invariant, in the same sense as in the first part of Proposition 4.11. Since the two summands in (4.6) each belong to fixed subspaces of this group action, they are each individually invariant as well. In particular, v^/v^\widehat{v}^{\perp}/\|\widehat{v}^{\perp}\| has the law of an (n1)(n-1)-dimensional random vector drawn uniformly at random from 𝕊n2(𝔽)\mathbb{S}^{n-2}(\mathbb{F}), with a zero entry prepended to it.

Let us introduce g𝒩𝔽(0,In)g\sim\mathcal{N}_{\mathbb{F}}(0,I_{n}) and write gg^{\perp} for gg with the first entry set to zero. By the above observations, we have the equality of laws

v^v^=(law)gg.\frac{\widehat{v}^{\perp}}{\|\widehat{v}^{\perp}\|}\stackrel{{\scriptstyle\text{(law)}}}{{=}}\frac{g^{\perp}}{\|g^{\perp}\|}.

Let us couple XX and gg such that this equality holds. From this gg, we set

y=ρ(θ)e1+τ(θ)ng.y=\rho(\theta)e_{1}+\frac{\tau(\theta)}{\sqrt{n}}g.

Then, we have

v^v^yyF\displaystyle\hskip-3.55658pt\left\|\widehat{v}\widehat{v}^{*}-yy^{*}\right\|_{\mathrm{F}}
=(v^y)v^+y(v^y)F\displaystyle=\left\|(\widehat{v}-y)\widehat{v}^{*}+y(\widehat{v}-y)^{*}\right\|_{\mathrm{F}}
(1+y)v^y\displaystyle\leq(1+\|y\|)\|\widehat{v}-y\|
(1+ρ(θ)+τ(θ)gn)(||v^,e1|ρ(θ)|+τ(θ)ngv^gg)\displaystyle\leq\left(1+\rho(\theta)+\tau(\theta)\frac{\|g\|}{\sqrt{n}}\right)\left(\big||\langle\widehat{v},e_{1}\rangle|-\rho(\theta)\big|+\left\|\frac{\tau(\theta)}{\sqrt{n}}g-\frac{\|\widehat{v}^{\perp}\|}{\|g^{\perp}\|}g^{\perp}\right\|\right)
(1+ρ(θ)+τ(θ)gn)(||v^,e1|ρ(θ)|+τ(θ)|g,e1|n+|τ(θ)nv^g|g)\displaystyle\leq\left(1+\rho(\theta)+\tau(\theta)\frac{\|g\|}{\sqrt{n}}\right)\left(\big||\langle\widehat{v},e_{1}\rangle|-\rho(\theta)\big|+\tau(\theta)\frac{|\langle g,e_{1}\rangle|}{\sqrt{n}}+\left|\frac{\tau(\theta)}{\sqrt{n}}-\frac{\|\widehat{v}^{\perp}\|}{\|g^{\perp}\|}\right|\|g^{\perp}\|\right)
=(1+ρ(θ)+τ(θ)gn)(||v^,e1|ρ(θ)|+τ(θ)|g,e1|n+|τ(θ)gnv^|)\displaystyle=\left(1+\rho(\theta)+\tau(\theta)\frac{\|g\|}{\sqrt{n}}\right)\left(\big||\langle\widehat{v},e_{1}\rangle|-\rho(\theta)\big|+\tau(\theta)\frac{|\langle g,e_{1}\rangle|}{\sqrt{n}}+\left|\tau(\theta)\frac{\|g^{\perp}\|}{\sqrt{n}}-\|\widehat{v}^{\perp}\|\right|\right)
(1+ρ(θ)+τ(θ)gn)(||v^,e1|ρ(θ)|+τ(θ)|g,e1|n+|τ(θ)v^|gn+|gn1|).\displaystyle\leq\left(1+\rho(\theta)+\tau(\theta)\frac{\|g\|}{\sqrt{n}}\right)\left(\big||\langle\widehat{v},e_{1}\rangle|-\rho(\theta)\big|+\tau(\theta)\frac{|\langle g,e_{1}\rangle|}{\sqrt{n}}+\left|\tau(\theta)-\|\widehat{v}^{\perp}\|\right|\frac{\|g^{\perp}\|}{\sqrt{n}}+\left|\frac{\|g^{\perp}\|}{\sqrt{n}}-1\right|\right).

We recall here that τ(θ)=1ρ(θ)2\tau(\theta)=\sqrt{1-\rho(\theta)^{2}} and by definition v^=1|v^,e1|2\|\widehat{v}^{\perp}\|=\sqrt{1-|\langle\widehat{v},e_{1}\rangle|^{2}}, so quantities in the third term of the second factor are determined by those in the first term.

To complete the proof, it suffices to show that the above expression is small with high probability. To that end, fix ε>0\varepsilon>0 and define the events

1\displaystyle\mathcal{E}_{1} :={|gn1|n1/2+ε,|gn1|n1/2+ε},\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left\{\left|\frac{\|g\|}{\sqrt{n}}-1\right|\leq n^{-1/2+\varepsilon},\quad\left|\frac{\|g^{\perp}\|}{\sqrt{n}}-1\right|\leq n^{-1/2+\varepsilon}\right\},
2\displaystyle\mathcal{E}_{2} :={|g,e1|logn},\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left\{|\langle g,e_{1}\rangle|\leq\log n\right\},
3\displaystyle\mathcal{E}_{3} :={||v^,e1|ρ(θ)|n1/2+ε},\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left\{\big||\langle\widehat{v},e_{1}\rangle|-\rho(\theta)\big|\leq n^{-1/2+\varepsilon}\right\},
\displaystyle\mathcal{E} :=123.\displaystyle\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\mathcal{E}_{1}\cap\mathcal{E}_{2}\cap\mathcal{E}_{3}.

We see from elementary manipulations that, on the event \mathcal{E}, we have

v^v^yyFn1/2+ε.\|\widehat{v}\widehat{v}^{*}-yy^{*}\|_{\mathrm{F}}\lesssim n^{-1/2+\varepsilon}.

Fix some K>0K>0. We have by the tail bounds proved in [LM00] that

[1c]nK,\mathbb{P}[\mathcal{E}_{1}^{c}]\lesssim n^{-K},

and by well-known Gaussian tail bounds we also have

[2c]nK.\mathbb{P}[\mathcal{E}_{2}^{c}]\lesssim n^{-K}.

Finally, by Theorem 2.21, we also have ||v^,e1|2ρ(θ)2|n1/2\left||\langle\widehat{v},e_{1}\rangle|^{2}-\rho(\theta)^{2}\right|\prec n^{-1/2}, and thus

[3c]nK.\mathbb{P}[\mathcal{E}_{3}^{c}]\lesssim n^{-K}.

Thus [c]nK\mathbb{P}[\mathcal{E}^{c}]\lesssim n^{-K}. Hence, for this fixed aa,

v^(X,a)v^(X,a)y(a)y(a)Fn1/2.\left\|\widehat{v}^{(X,a)}\widehat{v}^{(X,a)*}-y^{(a)}y^{(a)*}\right\|_{\mathrm{F}}\prec n^{-1/2}.

After constructing the coupling for each aa, we take these couplings independently over a[L]a\in[L]. Since LL is fixed, the sum of the resulting LL errors is still n1/2\prec n^{-1/2}. ∎

4.2 Numerical evaluation of Gaussian fluctuations

As a simple test of the Gaussianity given as an informal approximation in (1.10) and made precise in Theorem 1.8, in Figure 4 we present Gaussian Q–Q plots of eigenvector fluctuations projected in a delocalized and in a localized direction. Both have a distribution very close to Gaussian, and this persists under several different noise models: we consider additive Gaussian and Rademacher noise (as in Figure 1) as well as the truth-or-Haar model of group synchronization as detailed in Section 1.3 for two cyclic groups.

Refer to caption
Figure 4: Q–Q plots of one-dimensional projections of eigenvector fluctuations. We present Q–Q plots justifying the approximate Gaussianity of v^ρ(θ)v\widehat{v}-\rho(\theta)v proposed in (1.10). In particular, we present such plots for the inner product of this vector with two test directions xx, one localized and one delocalized, and four different noise models. We find strong evidence of Gaussianity in all cases.

5 Applications

5.1 Angular synchronization: Proof of Theorem 1.11

We recall the construction of the random matrix tuple H(1),,H(L)H^{(1)},\ldots,H^{(L)} that this result concerns: we have an underlying family of group elements x1,,xni.i.d.Haar(G)x_{1},\dots,x_{n}\stackrel{{\scriptstyle\mathrm{i.i.d.}}}{{\sim}}\mathrm{Haar}(G), and define the group-valued matrix Mij=xixj1M_{ij}=x_{i}x_{j}^{-1}. From this, we define YY to be a noisy version of MM by (1.13), and H(a)H^{(a)} the entrywise image of YY under the character χ(a)\chi^{(a)} given in (1.14). Let us write

vi(a):=χ(a)(xi)n,a[L].v^{(a)}_{i}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\frac{\chi^{(a)}(x_{i})}{\sqrt{n}},\qquad a\in[L].

We have |v(a)i|=n1/2|v^{(a)}_{i}|=n^{-1/2}, so v(a)𝕊n1(𝔽a)v^{(a)}\in\mathbb{S}^{n-1}(\mathbb{F}_{a}), where 𝔽a=\mathbb{F}_{a}=\mathbb{R} if χ(a)(G)\chi^{(a)}(G)\subseteq\mathbb{R} and 𝔽a=\mathbb{F}_{a}=\mathbb{C} otherwise, v(a)v^{(a)} has v(a)=n1/2\|v^{(a)}\|_{\infty}=n^{-1/2}, thus satisfying Assumption 1.3 for any εv(0,1/20)\varepsilon_{v}\in(0,1/20).

Conditional on xx, then, the upper-triangular off-diagonal blocks of H(1),,H(L)H^{(1)},\ldots,H^{(L)} are distributed as, independently for 1i<jn1\leq i<j\leq n,

(Hij(a))a=1L{(nvi(a)vj(a)¯)a=1Lwith probability p,(χ(a)(y)/n)a=1L,yHaar(G)with probability 1p},\Big(H^{(a)}_{ij}\Big)_{a=1}^{L}\sim\left\{\begin{array}[]{ll}\Big(\sqrt{n}\cdot v^{(a)}_{i}\overline{v^{(a)}_{j}}\Big)_{a=1}^{L}&\text{with probability }p,\\ \Big(\chi^{(a)}(y)/\sqrt{n}\Big)_{a=1}^{L},y\sim\mathrm{Haar}(G)&\text{with probability }1-p\end{array}\right\}, (5.1)

where p=θ/np=\theta/\sqrt{n}. Since we construct YY to be GG-Hermitian, we also have H(a)=H(a)H^{(a)}=H^{(a)*}. Under our specific definition from the Introduction we will always have Hii(a)=1/n=χ(a)(e)/nH^{(a)}_{ii}=1/\sqrt{n}=\chi^{(a)}(e)/\sqrt{n} for ee the identity of GG, but this detail will be inconsequential as we have shown above.

We then have, conditional on the underlying randomness of the xix_{i}, for the off-diagonal entries,

𝔼[Hij(a)x]=pnvi(a)vj(a)¯=θvi(a)vj(a)¯,ij,a[L].\mathbb{E}\big[H^{(a)}_{ij}\mid x\big]=p\sqrt{n}\cdot v^{(a)}_{i}\overline{v^{(a)}_{j}}=\theta\cdot v^{(a)}_{i}\overline{v^{(a)}_{j}},\qquad i\neq j,\ a\in[L].

We then define

W(a):=H(a)θv(a)v(a),𝐖:=(W(a))a=1L,W^{(a)}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}H^{(a)}-\theta v^{(a)}v^{(a)*},\qquad\mathbf{W}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\Big(W^{(a)}\Big)_{a=1}^{L},

which will be centered conditional on xx, away from the inconsequential diagonal entries. Specifically, its entries conditional on xx are distributed as

(Wij(a))a=1L{((nθ)vi(a)vj(a)¯)a=1Lwith probability p,(χ(a)(y)/nθvi(a)vj(a)¯)a=1L,yHaar(G)with probability 1p}.\Big(W^{(a)}_{ij}\Big)_{a=1}^{L}\sim\left\{\begin{array}[]{ll}\Big((\sqrt{n}-\theta)\cdot v^{(a)}_{i}\overline{v^{(a)}_{j}}\Big)_{a=1}^{L}&\text{with probability }p,\\ \Big(\chi^{(a)}(y)/\sqrt{n}-\theta v^{(a)}_{i}\overline{v^{(a)}_{j}}\Big)_{a=1}^{L},y\sim\mathrm{Haar}(G)&\text{with probability }1-p\end{array}\right\}. (5.2)

This is a centered Hermitian tuple whose upper-triangular blocks are independent conditional on xx. We have |vi(a)vj(a)¯|=1/n\big|v^{(a)}_{i}\overline{v^{(a)}_{j}}\big|=1/n, and thus the covariance matrix of the vector of real coordinates

w(ij):=(ReWij(1),ImWij(1),,ReWij(L),ImWij(L))w^{(ij)}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left(\mathrm{Re}W^{(1)}_{ij},\mathrm{Im}W^{(1)}_{ij},\dots,\mathrm{Re}W^{(L)}_{ij},\mathrm{Im}W^{(L)}_{ij}\right)

is n1n^{-1} times that of

𝒳(y):=(Reχ(1)(y),Imχ(1)(y),,Reχ(L)(y),Imχ(L)(y)),yHaar(G),\mathcal{X}(y)\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\left(\mathrm{Re}\chi^{(1)}(y),\mathrm{Im}\chi^{(1)}(y),\dots,\mathrm{Re}\chi^{(L)}(y),\mathrm{Im}\chi^{(L)}(y)\right),\qquad y\sim\mathrm{Haar}(G),

up to an additive error O(n3/2)O(n^{-3/2}):

𝔼[w(ij)w(ij)x]=1n𝔼yHaar(G)𝒳(y)𝒳(y)+O(n3/2).\mathbb{E}\left[w^{(ij)}w^{(ij)\top}\mid x\right]=\frac{1}{n}\mathop{\mathbb{E}}_{y\sim\mathrm{Haar}(G)}\mathcal{X}(y)\mathcal{X}(y)^{\top}+O(n^{-3/2}).

Here and below, the O(n3/2)O(n^{-3/2}) term denotes a 2L×2L2L\times 2L matrix with entries O(n3/2)O(n^{-3/2}). Since LL is fixed, the same bound also holds in operator norm, up to changing the implicit constant.

We use the following group-theoretic observation, a slight variation on the classical orthogonality of characters, to verify the covariance assumption of the definition of weakly Wigner tuples. We note that this applies equally well to one-dimensional irreducible representations of any compact Lie group, but we focus on the cases that will appear in our applications for the sake of simplicity.

Proposition 5.1 (Real-coordinate character orthogonality).

Let G=/KG=\mathbb{Z}/K or G=U(1)G=U(1), let yHaar(G)y\sim\mathrm{Haar}(G) and let χ(1),,χ(L):GU(1)\chi^{(1)},\ldots,\chi^{(L)}:G\to U(1) be nontrivial characters such that χ(a){χ(b),χ(b)¯}\chi^{(a)}\notin\{\chi^{(b)},\overline{\chi^{(b)}}\} for aba\neq b. For a[L]a\in[L], let 𝔽a=\mathbb{F}_{a}=\mathbb{R} if χ(a)(G)\chi^{(a)}(G)\subseteq\mathbb{R} and 𝔽a=\mathbb{F}_{a}=\mathbb{C} otherwise. Then

𝔼𝒳(y)=0,𝔼[𝒳(y)𝒳(y)]=Diag(Σ𝔽1,,Σ𝔽L),\mathbb{E}\mathcal{X}(y)=0,\qquad\mathbb{E}\big[\mathcal{X}(y)\mathcal{X}(y)^{\top}\big]=\mathrm{Diag}(\Sigma_{\mathbb{F}_{1}},\ldots,\Sigma_{\mathbb{F}_{L}}),

where

Σ=(1000),Σ=(1/2001/2).\Sigma_{\mathbb{R}}=\begin{pmatrix}1&0\\ 0&0\end{pmatrix},\qquad\Sigma_{\mathbb{C}}=\begin{pmatrix}1/2&0\\ 0&1/2\end{pmatrix}.
Proof.

Using translation invariance of normalized Haar measure, we have

𝔼yHaar(G)χ(y)=𝟏{χ is trivial}.\mathop{\mathbb{E}}_{y\sim\mathrm{Haar}(G)}\chi(y)=\mathbf{1}\{\chi\text{ is trivial}\}.

By our assumption, each χ(a)\chi^{(a)} is nontrivial, so 𝔼𝒳(y)=0\mathbb{E}\mathcal{X}(y)=0. For a,b[L]a,b\in[L], set

Aab:=𝔼[χ(a)(y)χ(b)(y)¯],Bab:=𝔼[χ(a)(y)χ(b)(y)],A_{ab}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\mathbb{E}\big[\chi^{(a)}(y)\overline{\chi^{(b)}(y)}\big],\qquad B_{ab}\mathrel{\mathchoice{\vbox{\hbox{$\displaystyle:$}}}{\vbox{\hbox{$\textstyle:$}}}{\vbox{\hbox{$\scriptstyle:$}}}{\vbox{\hbox{$\scriptscriptstyle:$}}}{=}}\mathbb{E}\big[\chi^{(a)}(y)\chi^{(b)}(y)\big],

then

Aab=𝟏{χ(a)=χ(b)},Bab=𝟏{χ(a)=χ(b)¯}.A_{ab}=\mathbf{1}\{\chi^{(a)}=\chi^{(b)}\},\qquad B_{ab}=\mathbf{1}\{\chi^{(a)}=\overline{\chi^{(b)}}\}.

Then the (a,b)(a,b)-block of 𝔼[𝒳(y)𝒳(y)]\mathbb{E}\big[\mathcal{X}(y)\mathcal{X}(y)^{\top}\big] is

𝔼(Reχ(a)(y)Imχ(a)(y))(Reχ(b)(y)Imχ(b)(y))=12(Aab+Bab00AabBab).\mathbb{E}\begin{pmatrix}\mathrm{Re}\chi^{(a)}(y)\\ \mathrm{Im}\chi^{(a)}(y)\end{pmatrix}\begin{pmatrix}\mathrm{Re}\chi^{(b)}(y)&\mathrm{Im}\chi^{(b)}(y)\end{pmatrix}=\frac{1}{2}\begin{pmatrix}A_{ab}+B_{ab}&0\\ 0&A_{ab}-B_{ab}\end{pmatrix}. (5.3)

If aba\neq b, the assumptions give Aab=Bab=0A_{ab}=B_{ab}=0, so the corresponding off-diagonal block vanishes. If a=ba=b, then Aaa=1A_{aa}=1 while Baa=𝟏{χ(a)=χ(a)¯}=𝟏{χ(a)(G)}B_{aa}=\mathbf{1}\{\chi^{(a)}=\overline{\chi^{(a)}}\}=\mathbf{1}\{\chi^{(a)}(G)\subseteq\mathbb{R}\}. Thus, the (a,a)(a,a)-th block is Σ\Sigma_{\mathbb{R}} when 𝔽a=\mathbb{F}_{a}=\mathbb{R}, and Σ\Sigma_{\mathbb{C}} when 𝔽a=\mathbb{F}_{a}=\mathbb{C}. ∎

By Proposition 5.1, uniformly for 1i<jn1\leq i<j\leq n,

𝔼[w(ij)w(ij)x]=1nDiag(Σ𝔽1,,Σ𝔽L)+O(n3/2).\mathbb{E}\left[w^{(ij)}w^{(ij)\top}\mid x\right]=\frac{1}{n}\mathrm{Diag}(\Sigma_{\mathbb{F}_{1}},\ldots,\Sigma_{\mathbb{F}_{L}})+O(n^{-3/2}).

Replacing the diagonal of W(a)W^{(a)} by zero only subtracts the scalar matrix (n1/2θ/n)I(n^{-1/2}-\theta/n)I from H(a)H^{(a)}, and hence does not change its eigenvectors. After this diagonal modification, conditional on xx, 𝐖\mathbf{W} is a weakly Wigner tuple, with CWC_{W} an absolute constant and εW=1/2\varepsilon_{W}=1/2.

Now, let v^(a)\widehat{v}^{(a)} be the top eigenvector of this H(a)H^{(a)} for each a[L]a\in[L]. Choosing the parameters εv>0\varepsilon_{v}>0 in Assumption 1.3 and ε>0\varepsilon>0 in Theorem 1.8 sufficiently small, we then obtain that, provided that ψ\psi is 𝒞5\mathcal{C}^{5} with bounded values and first five derivatives, conditionally on xx (whereby v(a)v^{(a)} is not random), we have

|𝔼[Ψ((nv(a)v(a))a=1L,(nv^(a)v^(a))a=1L)|x]\displaystyle\Bigg|\mathbb{E}\left[\Psi\left(\left(n\cdot v^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\cdot\widehat{v}^{(a)}\widehat{v}^{(a)*}\right)_{a=1}^{L}\right)\middle|x\right]
𝔼g(a)𝒩𝔽a(0,In)a[L]Ψ((nv(a)v(a))a=1L,((ρ(θ)nv(a)+τ(θ)g(a))(ρ(θ)nv(a)+τ(θ)g(a)))a=1L)|\displaystyle-\mathop{\mathbb{E}}_{\begin{subarray}{c}g^{(a)}\sim\mathcal{N}_{\mathbb{F}_{a}}(0,I_{n})\\ a\in[L]\end{subarray}}\Psi\left(\left(n\cdot v^{(a)}v^{(a)*}\right)_{a=1}^{L},\left((\rho(\theta)\sqrt{n}\cdot v^{(a)}+\tau(\theta)g^{(a)})(\rho(\theta)\sqrt{n}\cdot v^{(a)}+\tau(\theta)g^{(a)})^{*}\right)_{a=1}^{L}\right)\Bigg|
n1/5.\displaystyle\hskip 28.45274pt\lesssim n^{-1/5}.

Next, we show that the expectation over xx (and thus over the randomness in v(a)v^{(a)}) of the second expression converges to a single-letter formula. Indeed, expanding the definition of Ψ\Psi, we have

𝔼x1,,xni.i.d.Haar(G)𝔼g(a)𝒩𝔽a(0,In)a[L]Ψ((nv(a)v(a))a=1L,\displaystyle\mathop{\mathbb{E}}_{x_{1},\dots,x_{n}\stackrel{{\scriptstyle\mathrm{i.i.d.}}}{{\sim}}\mathrm{Haar}(G)}\mathop{\mathbb{E}}_{\begin{subarray}{c}g^{(a)}\sim\mathcal{N}_{\mathbb{F}_{a}}(0,I_{n})\\ a\in[L]\end{subarray}}\Psi\Bigg(\left(n\cdot v^{(a)}v^{(a)*}\right)_{a=1}^{L},
OPEN((ρ(θ)nv(a)+τ(θ)g(a))(ρ(θ)nv(a)+τ(θ)g(a)))a=1L)\displaystyle\hskip 85.35826pt\left((\rho(\theta)\sqrt{n}\cdot v^{(a)}+\tau(\theta)g^{(a)})(\rho(\theta)\sqrt{n}\cdot v^{(a)}+\tau(\theta)g^{(a)})^{*}\right)_{a=1}^{L}\Bigg)
=𝔼x1,,xni.i.d.Haar(G)𝔼g(a)𝒩𝔽a(0,In)a[L]1n2i,j=1nψ((χ(a)(xi)χ(a)(xj)¯)a=1LCLOSE,\displaystyle=\mathop{\mathbb{E}}_{x_{1},\dots,x_{n}\stackrel{{\scriptstyle\mathrm{i.i.d.}}}{{\sim}}\mathrm{Haar}(G)}\mathop{\mathbb{E}}_{\begin{subarray}{c}g^{(a)}\sim\mathcal{N}_{\mathbb{F}_{a}}(0,I_{n})\\ a\in[L]\end{subarray}}\frac{1}{n^{2}}\sum_{i,j=1}^{n}\psi\Bigg(\left(\chi^{(a)}(x_{i})\overline{\chi^{(a)}(x_{j})}\right)_{a=1}^{L},
OPEN((ρ(θ)χ(a)(xi)+τ(θ)gi(a))(ρ(θ)χ(a)(xj)+τ(θ)gj(a))¯)a=1L)\displaystyle\hskip 85.35826pt\left((\rho(\theta)\chi^{(a)}(x_{i})+\tau(\theta)g_{i}^{(a)})\overline{(\rho(\theta)\chi^{(a)}(x_{j})+\tau(\theta)g_{j}^{(a)})}\right)_{a=1}^{L}\Bigg)
=1n2ij𝔼xi,xjHaar(G)gi(a),gj(a)𝒩𝔽a(0,1),a[L]ψ((χ(a)(xi)χ(a)(xj)¯)a=1LCLOSE,\displaystyle=\frac{1}{n^{2}}\sum_{i\neq j}\mathop{\mathbb{E}}_{\begin{subarray}{c}x_{i},x_{j}\sim\mathrm{Haar}(G)\\ g_{i}^{(a)},g_{j}^{(a)}\sim\mathcal{N}_{\mathbb{F}_{a}}(0,1),\ a\in[L]\end{subarray}}\psi\Bigg(\left(\chi^{(a)}(x_{i})\overline{\chi^{(a)}(x_{j})}\right)_{a=1}^{L},
OPEN((ρ(θ)χ(a)(xi)+τ(θ)gi(a))(ρ(θ)χ(a)(xj)+τ(θ)gj(a))¯)a=1L)+O(n1)\displaystyle\hskip 85.35826pt\left((\rho(\theta)\chi^{(a)}(x_{i})+\tau(\theta)g_{i}^{(a)})\overline{(\rho(\theta)\chi^{(a)}(x_{j})+\tau(\theta)g_{j}^{(a)})}\right)_{a=1}^{L}\Bigg)+O(n^{-1})
=𝔼x,yHaar(G)ga,ha𝒩𝔽a(0,1),a[L]ψ((χ(a)(x)χ(a)(y)¯)a=1LCLOSE,\displaystyle=\mathop{\mathbb{E}}_{\begin{subarray}{c}x,y\sim\mathrm{Haar}(G)\\ g_{a},h_{a}\sim\mathcal{N}_{\mathbb{F}_{a}}(0,1),\ a\in[L]\end{subarray}}\psi\Bigg(\left(\chi^{(a)}(x)\overline{\chi^{(a)}(y)}\right)_{a=1}^{L},
OPEN((ρ(θ)χ(a)(x)+τ(θ)ga)(ρ(θ)χ(a)(y)+τ(θ)ha)¯)a=1L)+O(n1),\displaystyle\hskip 85.35826pt\left((\rho(\theta)\chi^{(a)}(x)+\tau(\theta)g_{a})\overline{(\rho(\theta)\chi^{(a)}(y)+\tau(\theta)h_{a})}\right)_{a=1}^{L}\Bigg)+O(n^{-1}),

since the diagonal terms contribute negligibly and all other expectations are equal.

To derive the actual statement of Theorem 1.11, we choose ψ\psi such that

ψ((χ(a)(x)χ(a)(y)¯)a=1L,𝒛)=(xy1,Round(𝒛)),\psi\left(\left(\chi^{(a)}(x)\overline{\chi^{(a)}(y)}\right)_{a=1}^{L},\bm{z}\right)=\ell(xy^{-1},\mathrm{Round}(\bm{z})),

for x,yGx,y\in G and 𝒛L\bm{z}\in\mathbb{C}^{L}. Since (χ(a)(x)χ(a)(y)¯)a=1L=𝝌(xy1)\left(\chi^{(a)}(x)\overline{\chi^{(a)}(y)}\right)_{a=1}^{L}=\bm{\chi}(xy^{-1}) and 𝝌\bm{\chi} is injective, this is always possible. With this choice,

Ψ((nv(a)v(a))a=1L,(nv^(a)v^(a))a=1L)=1n2i,j=1n(Mij,M^ij),\Psi\left(\left(n\cdot v^{(a)}v^{(a)*}\right)_{a=1}^{L},\left(n\cdot\widehat{v}^{(a)}\widehat{v}^{(a)*}\right)_{a=1}^{L}\right)=\frac{1}{n^{2}}\sum_{i,j=1}^{n}\ell(M_{ij},\widehat{M}_{ij}),

and the single-letter expression above becomes

𝔼x,yHaar(G)ga,ha𝒩𝔽a(0,1),a[L](xy1,Round(((ρ(θ)χ(a)(x)+τ(θ)ga)(ρ(θ)χ(a)(y)+τ(θ)ha)¯)a=1L)).\mathop{\mathbb{E}}_{\begin{subarray}{c}x,y\sim\mathrm{Haar}(G)\\ g_{a},h_{a}\sim\mathcal{N}_{\mathbb{F}_{a}}(0,1),\ a\in[L]\end{subarray}}\ell\bigg(xy^{-1},\mathrm{Round}\left(\left((\rho(\theta)\chi^{(a)}(x)+\tau(\theta)g_{a})\overline{(\rho(\theta)\chi^{(a)}(y)+\tau(\theta)h_{a})}\right)_{a=1}^{L}\right)\bigg).

The resulting ψ\psi may not be smooth. However, conditional on xx and yy, the random vector

((ρ(θ)χ(a)(x)+τ(θ)ga)(ρ(θ)χ(a)(y)+τ(θ)ha)¯)a=1L\left((\rho(\theta)\chi^{(a)}(x)+\tau(\theta)g_{a})\overline{(\rho(\theta)\chi^{(a)}(y)+\tau(\theta)h_{a})}\right)_{a=1}^{L}

has a density with respect to Lebesgue measure on 𝔽1××𝔽L\mathbb{F}_{1}\times\cdots\times\mathbb{F}_{L}. It is therefore a continuity point of Round\mathrm{Round} almost surely. Theorem 1.8 gives weak convergence against bounded smooth test functions, and the standard extension of weak convergence to bounded functions that are continuous almost surely under the limiting measure gives the result.

5.2 Discussion and numerical experiments

Let us give some additional discussion and present numerical experiments verifying these results. We first validate that the prediction for the average entrywise error of Theorem 1.11 is sound with some extra results extending those presented for the circle group G=U(1)G=U(1) in Figure 2 above. Then we study the implications of this prediction for the choice of the parameters LL and ϱ\varrho in a multi-frequency spectral algorithm.

First, in Figure 5, we present experiments of the same kind as in Figure 2 comparing the prediction of Theorem 1.11 with empirical distributions of average entrywise loss, now considering several groups GG, several numbers of frequencies LL, and several parameters ϱ\varrho of the rounding function. For G=U(1)G=U(1) we use the loss function (x,y)=1cos(xy)\ell(x,y)=1-\cos(x-y) as before, while for G=/KG=\mathbb{Z}/K (for K{2,5}K\in\{2,5\} in the experiments) we use the indicator loss function (x,y)=𝟏{xy}\ell(x,y)=\mathbf{1}\{x\neq y\}. We note that, for the indicator loss function, there is a “baseline” fraction of errors K1K\frac{K-1}{K} achieved by the trivial estimator of MM that guesses each entry uniformly at random, which is indeed what our M^\widehat{M} achieves for θ<1\theta<1 and which we plot with a dotted line labelled “Random guess” in the figures below. Agreement with the theoretical prediction is strong in all cases.

Figure 5: Empirical versus predicted synchronization performance. We plot the empirical average loss of the rounded multi-frequency spectral algorithm for group synchronization for several choices of group GG, number of frequencies LL, and rounding function parameter ϱ\varrho.
Figure 6: Predicted superiority of multiple frequencies for cyclic groups. We plot the theoretical predictions of average entrywise loss per Theorem 1.11 for group synchronization in two cyclic groups for various numbers of frequencies LL.

Next, taking as a given the accuracy of the prediction of Theorem 1.11 tested in these experiments, we consider how to best choose the values of the parameters LL and ϱ\varrho.

5.2.1 Choice of LL parameter: Superiority of multi-frequency algorithms

We showed the effect of different choices of LL on the formula of Theorem 1.11 in the case G=U(1)G=U(1) in Figure 3. In Figure 6, we show analogous results for two cyclic groups /K\mathbb{Z}/K (note that for multi-frequency algorithms to be defined, we must take K3K\geq 3, so we do not include the case K=2K=2 here unlike in the previous experiments). In both cases, at all values of θ\theta, algorithms using more frequencies and ϱ=2\varrho=2 (see below for discussion of this choice) achieve lower loss.

We note that the phenomenon observed earlier for the case of G=U(1)G=U(1) with the cosine loss function (x,y)=1cos(xy)\ell(x,y)=1-\cos(x-y) does not appear here: for cyclic groups with the indicator loss function, our calculations indicate that using more frequencies gives a uniformly superior algorithm for all values of signal strength θ\theta. On the other hand, we do observe (in experiments omitted here) the same non-monotonicity using the cosine loss function for cyclic groups, suggesting that the phenomenon is associated to the choice of loss function \ell rather than to the underlying group GG.

5.2.2 Choice of ϱ\varrho parameter: Superiority of Euclidean rounding

For the parameter ϱ\varrho, in Figure 7 we fix a cyclic group G=/KG=\mathbb{Z}/K, fix various numbers of frequencies LL, and consider the average loss predicted by Theorem 1.11 over a range of values of ϱ[,]\varrho\in[-\infty,\infty]. Perhaps surprisingly, we find that in all cases, among the values of ϱ\varrho we include, ϱ=2\varrho=2 achieves the lowest average loss for all values of θ\theta, lower than choices of ϱ\varrho both larger (more “max-like” loss functions) and smaller (more “min-like” loss functions) than ϱ=2\varrho=2.

Remark 5.2.

Another special property of the choice ϱ=2\varrho=2 is that, if we consider including all frequencies associated to characters χ1,,χK\chi_{1},\dots,\chi_{K}, in the objective expression i=1K|χi(x)zi|ϱ\sum_{i=1}^{K}|\chi_{i}(x)-z_{i}|^{\varrho} for some xGx\in G, when ϱ=2\varrho=2 this is the same as the squared distance between the inverse discrete Fourier transform of the vector zz and the indicator vector associated to xGx\in G, by the GG-valued version of Parseval’s theorem. Still, it is unclear what this isometry might have to do with the superiority of the associated rounding scheme for a spectral algorithm.

Figure 7: Predicted performance of different rounding functions. We plot the theoretical predictions of average entrywise loss per Theorem 1.11 for group synchronization in G=/17G=\mathbb{Z}/17 for various numbers of frequencies LL and, in each plot, for various choices of the parameter ϱ\varrho controlling the choice of rounding function.

5.2.3 Sketch of extension to non-abelian groups

We do not pursue it here, but we remark that it is in principle straightforward (though technically tedious) to extend these results to multi-frequency synchronization over non-abelian groups GG. Let us sketch how such an extension might look.

In this case, following for instance the ideas of [PWBM18], we would like to apply possibly higher-dimensional irreducible representations entrywise to the matrix YY to form a block matrix HH; if the underlying representation is of dimension dd, then HH will be dn×dndn\times dn. By the Peter-Weyl theorem on orthogonality of matrix coefficients, such matrices will obey a covariance property analogous to that in Definition 1.7, and thus can be compared to spiked matrices with independent Gaussian noise via an analog of Theorem 1.8. We expect that a version of Theorem 1.11 should therefore apply, where now if our χ1,,χL\chi_{1},\dots,\chi_{L} were replaced with representations of dimensions d1,,dLd_{1},\dots,d_{L}, then our rounding function would map from d1×d1××dL×dL\mathbb{C}^{d_{1}\times d_{1}}\times\cdots\times\mathbb{C}^{d_{L}\times d_{L}} to GG, and accordingly the appropriate version of Theorem 1.11 would involve expectations of this dimension. One subtle point is that the higher-dimensional matrix-valued replacements for the gag_{a} and hah_{a} in the statement (scalars in our setting) must be matched to the type (real, complex, or quaternionic) of the representation χa\chi_{a}. The main contours of our proof technique should also still apply; perhaps the main technical subtlety is that we would need to invoke local laws dealing with block matrices, such as those of [EHR25].

Acknowledgments

We thank Afonso Bandeira, Tatiana Brailovskaya, and Ke Wang for helpful suggestions during the course of this project.

References

  • [AEK17] Oskari H Ajanki, László Erdős, and Torben Krüger. Universality for general Wigner-type matrices. Probability Theory and Related Fields, 169(3):667–727, 2017.
  • [AFWZ20] Emmanuel Abbe, Jianqing Fan, Kaizheng Wang, and Yiqiao Zhong. Entrywise eigenvector analysis of random matrices with low expected rank. The Annals of Statistics, 48(3):1452–1474, 2020.
  • [AGZ09] Greg W. Anderson, Alice Guionnet, and Ofer Zeitouni. An Introduction to Random Matrices, volume 118 of Cambridge Studies in Advanced Mathematics. Cambridge University Press, 2009.
  • [BBAP05] Jinho Baik, Gérard Ben Arous, and Sandrine Péché. Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices. The Annals of Probability, 33(5):1643–1697, 2005.
  • [BCH26] Bishakh Bhattacharya, Arijit Chakrabarty, and Rajat Subhra Hazra. Outlier eigenvalues and eigenvectors of generalized Wigner matrices with finite-rank perturbations, 2026.
  • [BDW21] Zhigang Bao, Xiucai Ding, and Ke Wang. Singular vector and singular subspace distribution for the matrix denoising model. The Annals of Statistics, 49(1):370–392, 2021.
  • [BDWW22] Zhigang Bao, Xiucai Ding, Jingming Wang, and Ke Wang. Statistical inference for principal components of spiked covariance matrices. The Annals of Statistics, 50(2):1144–1169, 2022.
  • [BEK+14] Alex Bloemendal, László Erdős, Antti Knowles, Horng-Tzer Yau, and Jun Yin. Isotropic local laws for sample covariance and generalized Wigner matrices. Electronic Journal of Probability, 19(33):1–53, 2014.
  • [BGGM11] Florent Benaych-Georges, Alice Guionnet, and Mylène Maïda. Fluctuations of the extreme eigenvalues of finite rank deformations of random matrices. Electronic Journal of Probability, 16(60):1621–1662, 2011.
  • [BGK16] Florent Benaych-Georges and Antti Knowles. Lectures on the local semicircle law for Wigner matrices, 2016.
  • [BGN11] Florent Benaych-Georges and Raj Rao Nadakuditi. The eigenvalues and eigenvectors of finite, low rank perturbations of large random matrices. Advances in Mathematics, 227(1):494–521, 2011.
  • [BvH16] Afonso S. Bandeira and Ramon van Handel. Sharp nonasymptotic bounds on the norm of random matrices with independent entries. The Annals of Probability, 44(4):2479–2506, 2016.
  • [Cap17] Mireille Capitaine. Deformed ensembles, polynomials in random matrices and free probability theory. Habilitation à diriger des recherches, Université Paul Sabatier – Toulouse 3, 2017. HAL Id: tel-01978065, version 1.
  • [CDM21] Mireille Capitaine and Catherine Donati-Martin. Non universality of fluctuations of outlier eigenvectors for block diagonal deformations of Wigner matrices. ALEA. Latin American Journal of Probability and Mathematical Statistics, 18:129–165, 2021.
  • [CDMF09] Mireille Capitaine, Catherine Donati-Martin, and Delphine Féral. The largest eigenvalues of finite rank deformation of large Wigner matrices: Convergence and nonuniversality of the fluctuations. The Annals of Probability, 37(1):1–47, January 2009.
  • [CDMF12] Mireille Capitaine, Catherine Donati-Martin, and Delphine Féral. Central limit theorems for eigenvalues of deformations of Wigner matrices. Annales de l’Institut Henri Poincaré, Probabilités et Statistiques, 48(1):107–133, 2012.
  • [CH13] Romain Couillet and Walid Hachem. Fluctuations of spiked random matrix models and failure diagnosis in sensor networks. IEEE Transactions on Information Theory, 59(1):509–525, 2013.
  • [CR24] Collin Cademartori and Cynthia Rush. A non-asymptotic analysis of generalized vector approximate message passing algorithms with rotationally invariant designs. IEEE Transactions on Information Theory, 70(8):5811–5856, August 2024.
  • [CSC12] Mihai Cucuringu, Amit Singer, and David Cowburn. Eigenvector synchronization, graph rigidity and the molecule problem. Information and Inference: A Journal of the IMA, 1(1):21–67, December 2012.
  • [CT22] Mihai Cucuringu and Hemant Tyagi. An extension of the angular synchronization problem to the heterogeneous setting. Foundations of Data Science, 4(1):71–122, 2022.
  • [DK70] Chandler Davis and W. M. Kahan. The rotation of eigenvectors by a perturbation. III. SIAM Journal on Numerical Analysis, 7(1):1–46, March 1970.
  • [EHR25] László Erdős, Sven Joscha Henheik, and Volodymyr Riabov. Cusp universality for correlated random matrices. Communications in Mathematical Physics, 406(10):253, 2025.
  • [EYY12] László Erdős, Horng-Tzer Yau, and Jun Yin. Rigidity of eigenvalues of generalized Wigner matrices. Advances in Mathematics, 229(3):1435–1515, 2012.
  • [FFHL22] Jianqing Fan, Yingying Fan, Xiao Han, and Jinchi Lv. Asymptotic theory of eigenvectors for random matrices with diverging spikes. Journal of the American Statistical Association, 117(538):996–1009, 2022.
  • [FK81] Zoltán Füredi and János Komlós. The eigenvalues of random symmetric matrices. Combinatorica, 1(3):233–241, September 1981.
  • [FP07] Delphine Féral and Sandrine Péché. The largest eigenvalue of rank one deformation of large Wigner matrices. Communications in Mathematical Physics, 272(1):185–228, March 2007.
  • [FVRS22] Oliver Y. Feng, Ramji Venkataramanan, Cynthia Rush, and Richard J. Samworth. A unifying tutorial on approximate message passing. Foundations and Trends in Machine Learning, 15(4):335–536, 2022.
  • [GZ19] Tingran Gao and Zhizhen Zhao. Multi-frequency phase synchronization. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 2132–2141. PMLR, 09–15 Jun 2019.
  • [Han25] Qiyang Han. Entrywise dynamics and universality of general first order methods. The Annals of Statistics, 53(4):1783–1807, 2025.
  • [Ide16] Martin Idel. A review of matrix scaling and Sinkhorn’s normal form for matrices and positive maps, 2016.
  • [Joh01] Iain M. Johnstone. On the distribution of the largest eigenvalue in principal components analysis. The Annals of Statistics, 29(2):295–327, 2001.
  • [JP18] Iain M. Johnstone and Debashis Paul. PCA in high dimensions: An orientation. Proceedings of the IEEE, 106(8):1277–1292, 2018.
  • [KBK26] Anastasia Kireeva, Afonso S. Bandeira, and Dmitriy Kunisky. Computational lower bounds for multi-frequency group synchronization. Applied and Computational Harmonic Analysis, 84:101880, 2026.
  • [Kun25] Dmitriy Kunisky. Low coordinate degree algorithms II: Categorical signals and generalized stochastic block models. In Nika Haghtalab and Ankur Moitra, editors, Proceedings of the Thirty-Eighth Conference on Learning Theory, volume 291 of Proceedings of Machine Learning Research, pages 3486–3526. PMLR, 2025.
  • [KY13a] Antti Knowles and Jun Yin. Eigenvector distribution of Wigner matrices. Probability Theory and Related Fields, 155(3–4):543–582, 2013.
  • [KY13b] Antti Knowles and Jun Yin. The isotropic semicircle law and deformation of Wigner matrices. Communications on Pure and Applied Mathematics, 66(11):1663–1750, 2013.
  • [KY14] Antti Knowles and Jun Yin. The outliers of a deformed Wigner matrix. The Annals of Probability, 42(5):1980–2031, September 2014.
  • [KY17] Antti Knowles and Jun Yin. Anisotropic local laws for random matrices. Probability Theory and Related Fields, 169(1–2):257–352, 2017.
  • [LCC24] Hugo Lebeau, Florent Chatelain, and Romain Couillet. Asymptotic gaussian fluctuations of eigenvectors in spectral clustering. IEEE Signal Processing Letters, 31:1920–1924, 2024.
  • [LFW23] Gen Li, Wei Fan, and Yuting Wei. Approximate message passing from random initialization with applications to 2\mathbb{Z}_{2} synchronization. Proceedings of the National Academy of Sciences, 120(31):e2302930120, 2023.
  • [Li26] Zhangsong Li. Improved computational lower bound of estimation for multi-frequency group synchronization. arXiv preprint arXiv:2601.20522, 2026.
  • [LM00] B. Laurent and P. Massart. Adaptive estimation of a quadratic functional by model selection. The Annals of Statistics, 28(5):1302–1338, 2000.
  • [LW23] Gen Li and Yuting Wei. A non-asymptotic framework for approximate message passing in spiked models, 2023.
  • [MTV22] Marco Mondelli, Christos Thrampoulidis, and Ramji Venkataramanan. Optimal combination of linear and spectral estimators for generalized linear models. Foundations of Computational Mathematics, 22(5):1513–1566, 2022.
  • [MV21] Andrea Montanari and Ramji Venkataramanan. Estimation of low-rank matrices via approximate message passing. The Annals of Statistics, 49(1):321–345, 2021.
  • [MV22] Marco Mondelli and Ramji Venkataramanan. Approximate message passing with spectral initialization for generalized linear models. Journal of Statistical Mechanics: Theory and Experiment, 2022(11):114003, 2022.
  • [MY22] Jake Marcinek and Horng-Tzer Yau. High dimensional normality of noisy eigenvectors. Communications in Mathematical Physics, 395(3):1007–1096, 2022.
  • [PA14] Debashis Paul and Alexander Aue. Random matrix theory in statistics: A review. Journal of Statistical Planning and Inference, 150:1–29, 2014.
  • [Péc06] Sandrine Péché. The largest eigenvalue of small rank perturbations of Hermitian random matrices. Probability Theory and Related Fields, 134(1):127–173, January 2006.
  • [PWBM16] Amelia Perry, Alexander S. Wein, Afonso S. Bandeira, and Ankur Moitra. Optimality and sub-optimality of PCA for spiked random matrices and synchronization, 2016.
  • [PWBM18] Amelia Perry, Alexander S. Wein, Afonso S. Bandeira, and Ankur Moitra. Message-passing algorithms for synchronization problems over compact groups. Communications on Pure and Applied Mathematics, 71(11):2275–2322, April 2018.
  • [RG20] Elad Romanov and Matan Gavish. The noise-sensitivity phase transition in spectral group synchronization over compact groups. Applied and Computational Harmonic Analysis, 49(3):935–970, November 2020.
  • [RS13] David Renfrew and Alexander Soshnikov. On finite rank deformations of Wigner matrices II: Delocalized perturbations. Random Matrices: Theory and Applications, 2(1):1250015, 2013.
  • [RV18] Cynthia Rush and Ramji Venkataramanan. Finite sample analysis of approximate message passing algorithms. IEEE Transactions on Information Theory, 64(11):7264–7286, November 2018.
  • [Sin11] Amit Singer. Angular synchronization by eigenvectors and semidefinite programming. Applied and Computational Harmonic Analysis, 30(1):20–36, 2011.
  • [YWF25] Kaylee Y. Yang, Timothy L. H. Wee, and Zhou Fan. Asymptotic mutual information in quadratic estimation problems over compact groups. Information and Inference: A Journal of the IMA, 14(3):iaaf024, 2025.
  • [YWS15] Yi Yu, Tengyao Wang, and Richard J. Samworth. A useful variant of the Davis–Kahan theorem for statisticians. Biometrika, 102(2):315–323, 2015.