arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2603.15792v1 [quant-ph] 16 Mar 2026

Almost-iid information theory

Giulia Mazzola Affiliation: 1Institute for Theoretical Physics, ETH Zurich
2IBM Research Europe – Zurich
   David Sutter Affiliation: 1Institute for Theoretical Physics, ETH Zurich
2IBM Research Europe – Zurich
Affiliation: 1Institute for Theoretical Physics, ETH Zurich
2IBM Research Europe – Zurich
   Renato Renner Affiliation: 1Institute for Theoretical Physics, ETH Zurich
2IBM Research Europe – Zurich
Abstract

Information-theoretic techniques are based on the assumption that resources are well characterized by independent and identically distributed (iid) states. This assumption cannot be justified operationally, since, for example, correlations between subsequent systems emitted by a source cannot be detected by any practical tomographic protocol. Operationally motivated symmetry assumptions still imply, via de Finetti theorems, that the resources are described by almost-iid states. This raises the question: Are almost-iid resources as effective as perfect iid resources for information-processing tasks? Here we address this question and prove that the conditional entropy of almost-iid states asymptotically coincides with that of iid states. As an application, this implies that squashed entanglement is robust for almost-iid states, asymptotically matching its value on iid states.

1 Introduction

A common assumption in physics is that the same experiment can be repeated many times independently. More precisely, one often assumes that the outcomes X1,,Xn+kX_{1},\ldots,X_{n+k} of running the experiment n+kn+k times are described by independent and identically distributed (iid) random variables. In technical terms, this means that their joint distribution PX1n+kP_{X_{1}^{n+k}} is factorized, i.e., PX1n+k=(PX)n+kP_{X_{1}^{n+k}}=(P_{X})^{n+k}. The Italian mathematician Bruno de Finetti cautioned that the “most common and misleading error” in probability theory is treating the iid assumption as fundamental [19]. He provided a mathematically convincing argument to operationally justify the iid assumption, relating it to the permutation invariance of the outcomes [18, 30]. For many realistic systems, permutation invariance follows from natural assumptions such as the indistinguishability of subsystems (see [34] for a more detailed discussion).11 1 Alternatively, permutation invariance can also be enforced by randomly permuting the subsystems.

Suppose that we have a source that generates random variables X1n+kX_{1}^{n+k}, where we assume that the underlying joint distribution PX1n+kP_{X_{1}^{n+k}} is permutation-invariant and that each random variable takes values in the set 𝒳\mathcal{X}. When selecting nn of these n+kn+k random variables, i.e. ignoring kk random variables, de Finetti’s theorem [20] shows that there exists a probability measure μ\mu on the set of distributions on 𝒳\mathcal{X} such that

PX1n(QX)nμ(𝑑Q)12dnn+k,\displaystyle\left\lVert P_{X_{1}^{n}}-\int(Q_{X})^{n}\mu(\mathrm{d}Q)\right\rVert_{1}\leq\frac{2dn}{n+k}\,, (1)

where 1\left\lVert\cdot\right\rVert_{1} denotes the 1\ell^{1}-norm and d=|𝒳|d=|\mathcal{X}|. This result justifies that when ignoring kk random variables the remaining ones are approximately a convex combination of iid random variables.

De Finetti theorems have been generalized to the quantum case. In [12] a quantum de Finetti theorem has been obtained as an implication of the classical result in the case where n=n=\infty. This yields a result that is applicable under the assumption of having an infinite number of samples. Later in [25, 13] quantum de Finetti theorems have been derived that work for a finite number of samples — analogous to Equation 1. For any state |Φ(n+k)|\Phi^{(n+k)}\rangle on n+k\mathcal{H}^{\otimes n+k} that is symmetric (i.e. invariant under permutations of the subsystems), there exists a probability measure μ\mu on the unit sphere ():={|θ:θ=1}\mathcal{B}(\mathcal{H}):=\{|\theta\rangle\in\mathcal{H}:\left\lVert\theta\right\rVert=1\} such that

trk[|Φ(n+k)Φ(n+k)|]|θθ|nμ(𝑑θ)12d2nn+k=:εd(n,k),\displaystyle\left\lVert\mathrm{tr}_{k}[|\Phi^{(n+k)}\rangle\!\langle\Phi^{(n+k)}|]-\int|\theta\rangle\!\langle\theta|^{\otimes n}\mu(\mathrm{d}\theta)\right\rVert_{1}\leq\frac{2d^{2}n}{n+k}=:\varepsilon_{d}(n,k)\,, (2)

where 1\left\lVert\cdot\right\rVert_{1} is the trace norm, trk[]\mathrm{tr}_{k}[\cdot] denotes the partial trace of kk subsystems, and d=dim()d=\dim(\mathcal{H}). It is worth mentioning that the results stated in Equations 1 and 2 are essentially tight [20, 13].

The main obstacle with these de Finetti results (classical and quantum) is that they require kk to be large. Specifically, if we want the error term εd(n,k)\varepsilon_{d}(n,k) to vanish in the limit nn\to\infty, we need kk to be superlinear in nn, i.e. k=ω(n)k=\omega(n).22 2 Note that, by definition, f(n)=ω(g(n))g(n)=o(f(n))f(n)=\omega(g(n))\iff g(n)=o(f(n)). Hence, k=ω(n)n=o(k)k=\omega(n)\iff n=o(k), or in other words, limnnk=0\lim_{n\to\infty}\frac{n}{k}=0. Therefore, we need to ignore a large fraction of our data. This may be justified in certain scenarios where the number of data is by default huge33 3 For example, random samples of a coin toss, which in principle can be repeated arbitrarily many times., however, it can be prohibitive in other settings.44 4 For example, in the setting of quantum key distribution. A second, but often less drastic, disadvantage is that, by choosing k=poly(n)k=\mathrm{poly}(n), the error term εd(n,k)\varepsilon_{d}(n,k) is vanishing at a slow rate of order 1/poly(n)1/\mathrm{poly}(n).

The insight of [33, 34] was that we can overcome these limitations when relaxing the iid assumption. More precisely, the exponential de Finetti theorem [33, 34] shows that there exist a probability measure ν\nu on the unit sphere ()\mathcal{B}(\mathcal{H}) and a family {|Ψr,θ(n)}θ\{|\Psi_{r,\theta}^{(n)}\rangle\}_{\theta} of almost-iid states in |θ|\theta\rangle with a defect of size rr such that

trk[|Φ(n+k)Φ(n+k)|]|Ψr,θ(n)Ψr,θ(n)|ν(𝑑θ)13kdexp(k(r+1)n+k)=:εd(n,k).\displaystyle\left\lVert\mathrm{tr}_{k}[|\Phi^{(n+k)}\rangle\!\langle\Phi^{(n+k)}|]-\int|\Psi_{r,\theta}^{(n)}\rangle\!\langle\Psi_{r,\theta}^{(n)}|\nu(\mathrm{d}\theta)\right\rVert_{1}\leq 3k^{d}\exp\Big(-\frac{k(r+1)}{n+k}\Big)=:\varepsilon^{\prime}_{d}(n,k)\,. (3)

The precise statement is given in Theorem 3.1. Almost-iid states in |θ|\theta\rangle are superpositions of states that are equal to |θnr|\theta\rangle^{\otimes n-r} on nrn-r subsystems and arbitrary on the remaining rr subsystems. Note that in each element of the superposition the positions of the defects may be different. Usually, the number of defects is sublinear in nn, i.e. r=o(n)r=o(n). The precise definition of almost-iid states is given in Section 2. Equation 3 has a drastically better parameter scaling compared to Equation 2. It allows us to choose kk sublinear in nn, i.e. k=o(n)k=o(n), and still have an error term that vanishes in the limit nn\to\infty — even at an exponential rate. To see this, note that, for example, by choosing k=n34k=n^{\frac{3}{4}} and r=nr=\sqrt{n}, we obtain εd(n,k)=3n3d4exp(n5/4+n3/4n+n3/4)=O(exp(n14))\varepsilon^{\prime}_{d}(n,k)=3\,n^{\frac{3d}{4}}\exp(-\frac{n^{5/4}+n^{3/4}}{n+n^{3/4}})=O(\exp(-n^{\frac{1}{4}})).

While the exponential de Finetti theorem justifies the importance of almost-iid states, it is natural to ask if these states are also relevant in a purely classical scenario. One may define almost-iid distributions as a convex combination of an iid distribution on nrn-r subsystems and an arbitrary distribution on the remaining rr subsystems, where the position of the defects can be different in each element of the convex combination. For such distributions, we show that no classical equivalent to Equation 3 exists. In other words, it is not possible to approximate a classical permutation-invariant distribution by a convex combination of almost-iid distributions for kk sublinear in nn. This is made precise in Section 3.

The lack of a classical exponential de Finetti theorem shows that almost-iid states are considerably more powerful than classical almost-iid distributions. This is due to the allowed superpositions which can contain long-range entanglement. For that reason, almost-iid states are more complicated to analyze. In particular, it is unclear if certain (robust) applications could behave differently for almost-iid and iid states. However, it is expected that they should not behave differently for practical purposes. Quantum tomography, which is used to infer the characteristics of a system, cannot distinguish between almost-iid and iid states for practically feasible scenarios. We refer to Proposition 2.7 for a mathematically rigorous statement.

Given the operational relevance of almost-iid states (ensured by the exponential de Finetti theorem), it is crucial to understand which applications behave asymptotically equally under almost-iid and iid states. In this work, we introduce the definition of mixed almost-iid states and prove that for conditional entropies, there is no asymptotic difference between almost-iid and iid states. For nn\in\mathbb{N} and ρA1nB1n\rho_{A_{1}^{n}B_{1}^{n}} a mixed almost-iid state in σAB\sigma_{AB} with a defect of size r=o(n)r=o(n), we show that

1nH(A1n|B1n)ρ=H(A|B)σ+o(n)n,\displaystyle\frac{1}{n}H(A_{1}^{n}|B_{1}^{n})_{\rho}=H(A|B)_{\sigma}+\frac{o(n)}{n}\,, (4)

where H(A|B)σ:=H(AB)σH(B)σH(A|B)_{\sigma}:=H(AB)_{\sigma}-H(B)_{\sigma} is the conditional entropy and H(B)σ:=tr[σBlogσB]H(B)_{\sigma}:=-\mathrm{tr}[\sigma_{B}\log\sigma_{B}] for σB=trA[σAB]\sigma_{B}=\mathrm{tr}_{A}[\sigma_{AB}] is the von Neumann entropy. The exact result is given in Theorem 4.1. To prove our results, we use information-theoretic tools that have been developed to go beyond the traditionally assumed iid structure [33, 17, 36].

We conclude in Section 5 with a discussion of whether popular entanglement measures asymptotically coincide for almost-iid and iid states. We prove that this holds for the squashed entanglement (see Corollary 5.1); however, it remains an open question for other entanglement measures such as the distillable entanglement, the entanglement cost, and the relative entropy of entanglement.

2 Almost-iid states

In this section, we formally define almost-iid states and discuss some of their properties. Historically [33], almost-iid states were introduced as pure and symmetric states with additional structure, as explained in the following. This family of states appears, for example, in exponential de Finetti theorems such as those presented in Equation 3. However, we will relax the definition from [33] to capture more general families of states, including mixed ones, that have a similar almost-iid structure.

Let 𝒮n\mathcal{S}_{n} be the set of permutations on {1,,n}\{1,...,n\} and let S()\mathrm{S}(\mathcal{H}) denote the set of density matrices on a Hilbert space \mathcal{H} whose dimension is denoted by dd. The symmetric subspace is given by Symn():=span{|ϕn:|ϕ}\mathrm{Sym}^{n}(\mathcal{H}):=\mathrm{span}\{|\phi\rangle^{\otimes n}:|\phi\rangle\in\mathcal{H}\}. For a fixed |θ|\theta\rangle, let 𝒱(n,|θm):={π(|θm|Ω(nm)):π𝒮n,|Ω(nm)nm}\mathcal{V}(\mathcal{H}^{\otimes n},|\theta\rangle^{\otimes m}):=\{\pi(|\theta\rangle^{\otimes m}\otimes|\Omega^{(n-m)}\rangle):\pi\in\mathcal{S}_{n},|\Omega^{(n-m)}\rangle\in\mathcal{H}^{\otimes n-m}\} and

Symn(,|θm):=Symn()span𝒱(n,|θm).\displaystyle\mathrm{Sym}^{n}(\mathcal{H},|\theta\rangle^{\otimes m}):=\mathrm{Sym}^{n}(\mathcal{H})\cap\mathrm{span}\,\mathcal{V}(\mathcal{H}^{\otimes n},|\theta\rangle^{\otimes m})\,. (5)

Let n,rn,r\in\mathbb{N} such that rnr\leq n and |θ|\theta\rangle\in\mathcal{H}. In [33], an (nr)\binom{n}{r}-almost-iid state in |θ|\theta\rangle was defined as a pure state |Ψ(n)Symn(,|θnr)|\Psi^{(n)}\rangle\in\mathrm{Sym}^{n}(\mathcal{H},|\theta\rangle^{\otimes n-r}).

However, for pure states outside the symmetric subspace or mixed states, this definition needs to be relaxed in order to capture states that intuitively have an almost-iid structure such as those mentioned in Example 2.2 below. We therefore next introduce a novel definition for mixed almost-iid states.

Definition 2.1 (Almost-iid states).

Let A\mathcal{H}_{A} be a Hilbert space, σAS(A)\sigma_{A}\in\mathrm{S}(\mathcal{H}_{A}), and n,rn,r\in\mathbb{N} such that rnr\leq n. Then, ρA1nS(An)\rho_{A_{1}^{n}}\in\mathrm{S}(\mathcal{H}_{A}^{\otimes n}) is called a (nr)\binom{n}{r}-almost-iid state in σA\sigma_{A} if there exists a purification |θAE|\theta\rangle_{AE} of σA\sigma_{A} and an extension ρA1nE1n\rho_{A_{1}^{n}E_{1}^{n}} of ρA1n\rho_{A_{1}^{n}} such that

  1. (i)

    ρA1nE1n\rho_{A_{1}^{n}E_{1}^{n}} is permutation-invariant with respect to (Ai,Ei)(Aj,Ej)(A_{i},E_{i})\leftrightarrow(A_{j},E_{j});

  2. (ii)

    supp(ρA1nE1n)span𝒱(AEn,|θAEnr)\mathrm{supp}(\rho_{A_{1}^{n}E_{1}^{n}})\subseteq\mathrm{span}\,\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\theta\rangle_{AE}^{\otimes n-r}).

We then write ρA1nSn(A,σAnr)\rho_{A_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}).

Here, supp(X)\mathrm{supp}(X) denotes the support of a linear operator XX and is defined as supp(X)=ker(X)\mathrm{supp}(X)=\ker(X)^{\perp}. We next discuss a few pedagogical examples of mixed almost-iid states.

Example 2.2.

A few simple instances of almost-iid states include:

  1. (a)

    Tensor power states ρA1n=σAn\rho_{A_{1}^{n}}=\sigma_{A}^{\otimes n}. These are almost-iid states with defect size r=0r=0, i.e. ρA1nSn(A,σAn)\rho_{A_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n}).

  2. (b)

    Convex combinations of tensor power states with a small number of defects, i.e., states of the form ρA1n=1n!π𝒮nπ(σAnrωA1r)π\rho_{A_{1}^{n}}=\frac{1}{n!}\sum_{\pi\in\mathcal{S}_{n}}\pi(\sigma_{A}^{\otimes n-r}\otimes\omega_{A_{1}^{r}})\pi^{\dagger}, where ωA1r\omega_{A_{1}^{r}} denotes an arbitrary density operator of size rr representing the defects. We have ρA1nSn(A,σAnr)\rho_{A_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}). To see this, let |θAE|\theta\rangle_{AE} and |ωA1rE1r|\omega\rangle_{A_{1}^{r}E_{1}^{r}} be purifications of θA\theta_{A} and ωA1r\omega_{A_{1}^{r}}, respectively. Then the extension

    ρA1nE1n=1n!π𝒮nπ(|θθ|AEnr|ωω|A1rE1r)π\displaystyle\rho_{A_{1}^{n}E_{1}^{n}}=\frac{1}{n!}\sum_{\pi\in\mathcal{S}_{n}}\pi\big(|\theta\rangle\!\langle\theta|^{\otimes n-r}_{AE}\otimes|\omega\rangle\!\langle\omega|_{A_{1}^{r}E_{1}^{r}}\big)\pi^{\dagger} (6)

    of ρA1n\rho_{A_{1}^{n}} clearly satisfies property (i) since it is permutation-invariant. It also satisfies the property (ii) as it can be written as ρA1nE1n=i,j𝒯βi,j|ΨiΨj|\rho_{A_{1}^{n}E_{1}^{n}}=\sum_{i,j\in\mathcal{T}}\beta_{i,j}|\Psi_{i}\rangle\langle\Psi_{j}| for βi,j=1n!δi,j\beta_{i,j}=\frac{1}{n!}\delta_{i,j} and |Ψi𝒱(AEn,|θAEnr)|\Psi_{i}\rangle\in\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\theta\rangle_{AE}^{\otimes n-r}) for all i𝒯i\in\mathcal{T}.

  3. (c)

    The antisymmetric Bell state ρA12=|ΨΨ|A12\rho_{A_{1}^{2}}=|\Psi^{-}\rangle\!\langle\Psi^{-}|_{A_{1}^{2}} with |ΨA12:=12(|0|1A12|1|0A12)|\Psi^{-}\rangle_{A_{1}^{2}}:=\frac{1}{\sqrt{2}}(|0\rangle|1\rangle_{A_{1}^{2}}-|1\rangle|0\rangle_{A_{1}^{2}}) is a (21)\binom{2}{1}-almost-iid state in |0A|0\rangle_{A}.

We want to emphasize that the class of almost-iid states contains many more families than the ones presented in Example 2.2. In particular, it allows for superpositions of product states with a small number of defects. These superpositions are crucial for an exponential de Finetti theorem as in Equation 3 to hold. On the other hand, classical almost-iid distributions have a simpler structure as they do not have superpositions and hence are just convex combinations of almost-product distributions. This is discussed in Section 3.

Remark 2.3.

A more restrictive definition of mixed almost-iid states was introduced in [9] by calling ρA1nS(An)\rho_{A_{1}^{n}}\in\mathrm{S}(\mathcal{H}_{A}^{\otimes n}) an (nr)\binom{n}{r}-almost-iid state in σA\sigma_{A} if there exists a purification |Ψ(n)A1nE1n|\Psi^{(n)}\rangle_{A_{1}^{n}E_{1}^{n}} such that |Ψ(n)Symn(AE,|θAEnr)|\Psi^{(n)}\rangle\in\mathrm{Sym}^{n}(\mathcal{H}_{A}\otimes\mathcal{H}_{E},|\theta\rangle_{AE}^{\otimes n-r}) for some purification |θAE|\theta\rangle_{AE} of σA\sigma_{A}. Such states clearly satisfy the assumptions of Definition 2.1. However, Definition 2.1 is strictly broader as it includes states that do not meet the definition given in [9]. One such example is given by Example 2.2 (c) as explained in Appendix A.

We call ρA1n\rho_{A_{1}^{n}} a (nr)\binom{n}{r}-generalized-almost-iid-state in σA\sigma_{A} if it satisfies (ii) but not necessarily (i). We then write ρA1nS¯n(A,σAnr)\rho_{A_{1}^{n}}\in\bar{\mathrm{S}}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}). Trivially, we have Sn(A,σAnr)S¯n(A,σAnr)\mathrm{S}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r})\subseteq\bar{\mathrm{S}}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}). With the following trace-preserving completely positive map

PERM:X1n!π𝒮nπXπ,\displaystyle\mathrm{PERM}:X\mapsto\frac{1}{n!}\sum_{\pi\in\mathcal{S}_{n}}\pi X\pi^{\dagger}\,, (7)

which symmetrizes the input, it is possible to convert a generalized-almost-iid-state into an almost-iid state in the sense that

ρA1nS¯n(A,σAnr)PERM(ρA1n)Sn(A,σAnr).\displaystyle\rho_{A_{1}^{n}}\in\bar{\mathrm{S}}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r})\implies\mathrm{PERM}(\rho_{A_{1}^{n}})\in\mathrm{S}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r})\,. (8)
Remark 2.4 (Properties of almost-iid states).

Let ρA1nSn(A,σAnr)\rho_{A_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}) be a mixed almost-iid state according to Definition 2.1. Then it satisfies the following properties:

  1. (a)

    For any\emph{any} purification |θAE|\theta\rangle_{AE} of σA\sigma_{A} there exists an extension ρA1nE1n\rho_{A_{1}^{n}E_{1}^{n}} of ρA1n\rho_{A_{1}^{n}} that satisfies (i) and (ii). See Lemma B.1.

  2. (b)

    ρA1n\rho_{A_{1}^{n}}is permutation-invariant.

  3. (c)

    There exists an orthonormal basis {|Ψt}t𝒯\{|\Psi_{t}\rangle\}_{t\in\mathcal{T}} of span𝒱(AEn,|θnr)\mathrm{span}\,\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\theta\rangle^{\otimes n-r}) with vectors |Ψt𝒱(AEn,|θnr)|\Psi_{t}\rangle\in\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\theta\rangle^{\otimes n-r}) and with

    |𝒯|(nr)dAEr2nh(r/n)dAEr,\displaystyle|\mathcal{T}|\leq\binom{n}{r}\,d_{AE}^{\,r}\leq 2^{nh(r/n)}\,d_{AE}^{\,r}, (9)

    where h(x):=xlogx(1x)log(1x)h(x):=-x\log x-(1-x)\log(1-x) for x[0,1]x\in[0,1] is the binary entropy function. With respect to this basis, condition (ii) can be rewritten as

    ρA1nE1n=i,j𝒯βi,j|ΨiΨj|,\displaystyle\rho_{A_{1}^{n}E_{1}^{n}}=\sum_{i,j\in\mathcal{T}}\beta_{i,j}|\Psi_{i}\rangle\langle\Psi_{j}|\,, (10)

    for coefficients βi,j\beta_{i,j}\in\mathbb{C} with βi,i[0,1]\beta_{i,i}\in[0,1] and i𝒯βi,i=1\sum_{i\in\mathcal{T}}\beta_{i,i}=1. See Lemma B.2.

  4. (d)

    ρA1n+mSn+m(A,σAn+mr)\rho_{A_{1}^{n+m}}\in\mathrm{S}^{n+m}(\mathcal{H}_{A},\sigma_{A}^{\otimes n+m-r}) implies ρA1nSn(A,σAnr)\rho_{A_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}) for any m,nm,n\in\mathbb{N}. See Lemma B.3.

  5. (e)

    ρA1nB1nSn(AB,σABnr)\rho_{A_{1}^{n}B_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{AB},\sigma_{AB}^{\otimes n-r}) implies ρA1nSn(A,σAnr)\rho_{A_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}). This follows directly from the definition of mixed almost-iid states.

  6. (f)

    For any ρA1nS¯n(A,σAnr)\rho_{A_{1}^{n}}\in\bar{\mathrm{S}}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}) we have for any ss\in\mathbb{N} that (ρA1n)sS¯ns(A,σAnsrs)(\rho_{A_{1}^{n}})^{\otimes s}\in\bar{\mathrm{S}}^{ns}(\mathcal{H}_{A},\sigma_{A}^{\otimes ns-rs}). See Lemma B.4.

  7. (g)

    The set of almost-iid states is convex. This follows directly from the definition of mixed almost-iid states.

In the analysis of almost-iid states, the representation in Equation 10 is particularly useful because it exploits the product structure of the space 𝒱(CEn,|θCEnr)\mathcal{V}(\mathcal{H}_{CE}^{\otimes n},|\theta\rangle_{CE}^{\otimes n-r}). However, it is sometimes preferable to go to a purified description, i.e., to work with a purification of Equation 10. The following lemma establishes that the underlying structure of 𝒱(CEn,|θCEnr)\mathcal{V}(\mathcal{H}_{CE}^{\otimes n},|\theta\rangle_{CE}^{\otimes n-r}) persists even in this purified formulation.

Lemma 2.5 (Purification in a preferred basis).

Let |ΘAHAH|\Theta\rangle_{AH}\in\mathcal{H}_{AH} for some composite Hilbert space AH\mathcal{H}_{AH}, and let 𝒱AA\mathcal{V}_{A}\subseteq\mathcal{H}_{A} be any subset of vectors such that span𝒱A\,\mathrm{span}\,\mathcal{V}_{A} admits an orthonormal basis {|ΨjA}j𝒯\big\{|\Psi_{j}\rangle_{A}\big\}_{j\in\mathcal{T}} with vectors |ΨjA𝒱Aj𝒯|\Psi_{j}\rangle_{A}\in\mathcal{V}_{A}\ \ \forall\,j\in\mathcal{T}. Then, the following are equivalent:

  1. (i)

    supp(trH[|ΘΘ|AH])span𝒱A\mathrm{supp}\big(\mathrm{tr}_{H}\big[|\Theta\rangle\!\langle\Theta|_{AH}\big]\big)\,\subseteq\,\mathrm{span}\,\mathcal{V}_{A}

  2. (ii)

    |ΘAH=j𝒯αj|ΨjA|h~jH|\Theta\rangle_{AH}=\sum\limits_{j\in\mathcal{T}}\alpha_{j}|\Psi_{j}\rangle_{A}|\tilde{h}_{j}\rangle_{H} for coefficients αj\alpha_{j}\in\mathbb{C} and normalized vectors |h~jHH|\tilde{h}_{j}\rangle_{H}\in\mathcal{H}_{H} for all j𝒯j\in\mathcal{T}.

In addition, if |ΘAH|\Theta\rangle_{AH} is normalized and satisfies (ii)(ii), we have 1=Θ|ΘAH=j𝒯|αj|21=\langle\Theta|\Theta\rangle_{AH}=\sum_{j\in\mathcal{T}}|\alpha_{j}|^{2}.

Proof.

(i)(ii)\eqref{it_1}\implies\eqref{it_2}: By the Schmidt decomposition, there exist orthonormal sets of vectors {|χkA}kA\big\{|\chi_{k}\rangle_{A}\big\}_{k}\subset\mathcal{H}_{A} and {|ϕkH}kH\big\{|\phi_{k}\rangle_{H}\big\}_{k}\subset\mathcal{H}_{H}, and scalars γk0\gamma_{k}\geq 0 such that |ΘAH=kγk|χkA|ϕkH|\Theta\rangle_{AH}\,=\,\sum_{k}\gamma_{k}|\chi_{k}\rangle_{A}|\phi_{k}\rangle_{H}. Taking the partial trace, we obtain

trH[|ΘΘ|AH]=mk,lγkγl|χkχl|Aϕm|ϕkHϕl|ϕmH=kγk2|χkχk|A.\displaystyle\mathrm{tr}_{H}\big[|\Theta\rangle\!\langle\Theta|_{AH}\big]\,=\,\sum_{m}\sum_{k,l}\gamma_{k}\gamma_{l}|\chi_{k}\rangle\langle\chi_{l}|_{A}\langle\phi_{m}|\phi_{k}\rangle_{H}\langle\phi_{l}|\phi_{m}\rangle_{H}\,=\,\sum_{k}\gamma_{k}^{2}|\chi_{k}\rangle\!\langle\chi_{k}|_{A}\,. (11)

By assumption (i), it follows that |χkspan𝒱A|\chi_{k}\rangle\in\mathrm{span}\,\mathcal{V}_{A} for all kk such that γk0\gamma_{k}\neq 0. (If, by contradiction, there were |χlspan𝒱A|\chi_{l}\rangle\notin\mathrm{span}\,\mathcal{V}_{A} for some ll with γl0\gamma_{l}\neq 0, then we would have trH[|ΘΘ|AH]|χl=γl2|χl0supp()span𝒱A\mathrm{tr}_{H}\big[|\Theta\rangle\!\langle\Theta|_{AH}\big]|\chi_{l}\rangle=\gamma_{l}^{2}|\chi_{l}\rangle\neq 0\implies\mathrm{supp}(\dots)\not\subseteq\mathrm{span}\,\mathcal{V}_{A}. \lightning) Therefore, we may write |χkA=j𝒯cj(k)|ΨjA|\chi_{k}\rangle_{A}=\sum_{j\in\mathcal{T}}c^{(k)}_{j}|\Psi_{j}\rangle_{A} for some coefficients cj(k)jc^{(k)}_{j}\in\mathbb{C}\ \forall\,j depending on kk. With that, we find

|ΘAH\displaystyle|\Theta\rangle_{AH} =kγk|χkA|ϕkH=kγkj𝒯cj(k)|ΨjA|ϕkH=j𝒯|ΨjA(kcj(k)γk|ϕkH)\displaystyle=\sum_{k}\gamma_{k}|\chi_{k}\rangle_{A}|\phi_{k}\rangle_{H}=\sum_{k}\gamma_{k}\sum_{j\in\mathcal{T}}c^{(k)}_{j}|\Psi_{j}\rangle_{A}|\phi_{k}\rangle_{H}=\sum_{j\in\mathcal{T}}|\Psi_{j}\rangle_{A}\Big(\sum_{k}c^{(k)}_{j}\gamma_{k}|\phi_{k}\rangle_{H}\Big) (12)
=j𝒯|ΨjA|hjH=j𝒯αj|ΨjA|h~jH,\displaystyle=\sum_{j\in\mathcal{T}}|\Psi_{j}\rangle_{A}|{h}_{j}\rangle_{H}=\sum_{j\in\mathcal{T}}\alpha_{j}|\Psi_{j}\rangle_{A}|\tilde{h}_{j}\rangle_{H}\,, (13)

where we have introduced coefficients αj0\alpha_{j}\geq 0 to normalize the vector |hjH:=kcj(k)γk|ϕkH|{h}_{j}\rangle_{H}:=\sum_{k}c^{(k)}_{j}\gamma_{k}|\phi_{k}\rangle_{H} (with the convention that αj=0\alpha_{j}=0 and |h~jH|\tilde{h}_{j}\rangle_{H} are normalized but arbitrary if |hjH=0|{h}_{j}\rangle_{H}=0).

(ii)(i)\eqref{it_2}\implies\eqref{it_1}: For any orthonormal basis {|ϕmH}k\big\{|\phi_{m}\rangle_{H}\big\}_{k} of H\mathcal{H}_{H}, we simply evaluate

trH[|ΘΘ|AH]\displaystyle\mathrm{tr}_{H}\big[|\Theta\rangle\!\langle\Theta|_{AH}\big]\, =i,j𝒯αiαjmϕm|h~iHh~j|ϕmH|ΨiΨj|A,\displaystyle=\,\sum_{i,j\in\mathcal{T}}\alpha_{i}\alpha^{*}_{j}\sum_{m}\langle\phi_{m}|\tilde{h}_{i}\rangle_{H}\langle\tilde{h}_{j}|\phi_{m}\rangle_{H}\,|\Psi_{i}\rangle\langle\Psi_{j}|_{A}\,, (14)

which implies supp(trH[|ΘΘ|AH])span𝒱A\mathrm{supp}\big(\mathrm{tr}_{H}\big[|\Theta\rangle\!\langle\Theta|_{AH}\big]\big)\,\subseteq\,\mathrm{span}\,\mathcal{V}_{A} since |ΨjA𝒱Aj𝒯|\Psi_{j}\rangle_{A}\in\mathcal{V}_{A}\ \ \forall\,j\in\mathcal{T}. In addition, if |ΘAH|\Theta\rangle_{AH} is normalized, we find 1=Θ|ΘAH=j𝒯|αj|21=\langle\Theta|\Theta\rangle_{AH}=\sum_{j\in\mathcal{T}}|\alpha_{j}|^{2} since the vectors {|ΨjA}j𝒯\big\{|\Psi_{j}\rangle_{A}\big\}_{j\in\mathcal{T}} are orthonormal and the vectors |h~jH|\tilde{h}_{j}\rangle_{H} are normalized for each j𝒯j\in\mathcal{T}. ∎

An important property of almost-iid states is that when tracing out many subsystems, we obtain a state that is close to an iid state.

Proposition 2.6.

Let n,r,sn,r,s\in\mathbb{N} such that r,snr,s\leq n, σAS(A)\sigma_{A}\in\mathrm{S}(\mathcal{H}_{A}), and ρA1nSn(A,σAnr)\rho_{A_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}). Then

trns[ρA1n]σAs14rsn.\displaystyle\left\lVert\mathrm{tr}_{n-s}[\rho_{A_{1}^{n}}]-\sigma_{A}^{\otimes s}\right\rVert_{1}\leq 4\sqrt{\frac{rs}{n}}\,. (15)
Proof.

Let n~:=nssn\tilde{n}:=\lfloor\frac{n}{s}\rfloor s\leq n and define ρA1n~:=trnn~[ρA1n]\rho_{A_{1}^{\tilde{n}}}:=\mathrm{tr}_{n-\tilde{n}}[\rho_{A_{1}^{n}}]. Let |θAE|\theta\rangle_{AE} be a purification of σA\sigma_{A}. By assumption ρA1nSn(A,σAnr)\rho_{A_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}) which according to Remark 2.4 implies ρA1n~Sn~(A,σAn~r)\rho_{A_{1}^{\tilde{n}}}\in\mathrm{S}^{\tilde{n}}(\mathcal{H}_{A},\sigma_{A}^{\otimes\tilde{n}-r}). Hence, due to the definition of this vector space, there exists a family {|Ψt}t𝒯\{|\Psi_{t}\rangle\}_{t\in\mathcal{T}} of orthonormal vectors from 𝒱(AEn~,|θAEn~r)\mathcal{V}(\mathcal{H}_{AE}^{\otimes\tilde{n}},|\theta\rangle_{AE}^{\otimes\tilde{n}-r}) such that

ρA1n~E1n~=t,t𝒯βt,t|ΨtΨt|,\displaystyle\rho_{A_{1}^{\tilde{n}}E_{1}^{\tilde{n}}}=\sum_{t,t^{\prime}\in\mathcal{T}}\beta_{t,t^{\prime}}|\Psi_{t}\rangle\langle\Psi_{t^{\prime}}|\,, (16)

with t𝒯βt,t=1\sum_{t\in\mathcal{T}}\beta_{t,t}=1, where ρA1n~E1n~\rho_{A_{1}^{\tilde{n}}E_{1}^{\tilde{n}}} is an extension of ρA1n~\rho_{A_{1}^{\tilde{n}}}. Now, we may interpret |Ψt|\Psi_{t}\rangle as a state on the ns\lfloor\tfrac{n}{s}\rfloor blocks. Since |Ψt𝒱(AEn~,|θAEn~r)|\Psi_{t}\rangle\in\mathcal{V}(\mathcal{H}_{AE}^{\otimes\tilde{n}},|\theta\rangle_{AE}^{\otimes\tilde{n}-r}), we know that |Ψt|\Psi_{t}\rangle is at least in nsr\lfloor\tfrac{n}{s}\rfloor-r blocks of the form |θs|\theta\rangle^{\otimes s} and in the remaining rr blocks arbitrary. For any t𝒯t\in\mathcal{T}, let t\mathcal{F}_{t} denote the set of block indices i{1,,ns}i\in\{1,\ldots,\lfloor\frac{n}{s}\rfloor\} on which |Ψt|\Psi_{t}\rangle is not of the form |θs|\theta\rangle^{\otimes s} and note that

|t|r.\displaystyle|\mathcal{F}_{t}|\leq r\,. (17)

Then, for any block specified by ii, the total weight of the vectors |Ψt|\Psi_{t}\rangle that deviate from |θs|\theta\rangle^{\otimes s} for that block is given by

wi:=t𝒯βt,tδit.\displaystyle w_{i}:=\sum_{t\in\mathcal{T}}\beta_{t,t}\delta_{i\in\mathcal{F}_{t}}\,. (18)

Summing over all blocks yields

i=1nswi=Equation 18i=1nst𝒯βt,tδit=t𝒯βt,ti=1nsδit=t𝒯βt,t|t|Equation 17t𝒯βt,tr=r,\displaystyle\sum_{i=1}^{\lfloor\frac{n}{s}\rfloor}w_{i}\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_wi}}}}{=}\sum_{i=1}^{\lfloor\frac{n}{s}\rfloor}\sum_{t\in\mathcal{T}}\beta_{t,t}\delta_{i\in\mathcal{F}_{t}}=\sum_{t\in\mathcal{T}}\beta_{t,t}\sum_{i=1}^{\lfloor\frac{n}{s}\rfloor}\delta_{i\in\mathcal{F}_{t}}=\sum_{t\in\mathcal{T}}\beta_{t,t}|\mathcal{F}_{t}|\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_size_error}}}}{\leq}\sum_{t\in\mathcal{T}}\beta_{t,t}r=r\,, (19)

where the final step uses t𝒯βt,t=1\sum_{t\in\mathcal{T}}\beta_{t,t}=1. Equation 19 implies that

wjrnsfor somej{1,,ns}.\displaystyle w_{j}\leq\frac{r}{\lfloor\frac{n}{s}\rfloor}\quad\textnormal{for some}\quad j\in\left\{1,\ldots,\left\lfloor\frac{n}{s}\right\rfloor\right\}\,. (20)

For a fixed j{1,,ns}j\in\left\{1,\ldots,\left\lfloor\frac{n}{s}\right\rfloor\right\}, let Π\Pi be the projector onto the subspace spanned by {|Ψt:jt}\{|\Psi_{t}\rangle:j\not\in\mathcal{F}_{t}\}, i.e.

Π=t𝒯 s.t. jt|ΨtΨt|andΠ=t𝒯 s.t. jt|ΨtΨt|.\displaystyle\Pi=\sum_{t\in\mathcal{T}\textnormal{ s.t. }j\not\in\mathcal{F}_{t}}|\Psi_{t}\rangle\!\langle\Psi_{t}|\qquad\textnormal{and}\qquad\Pi^{\perp}=\sum_{t\in\mathcal{T}\textnormal{ s.t. }j\in\mathcal{F}_{t}}|\Psi_{t}\rangle\!\langle\Psi_{t}|\,. (21)

Note that ΠΠ=0\Pi\,\Pi^{\perp}=0 and Π+Π=idspan𝒱(AEn~,|θAEn~r)\Pi+\Pi^{\perp}=\mathrm{id}_{\mathrm{span}\,\mathcal{V}(\mathcal{H}_{AE}^{\otimes\tilde{n}},|\theta\rangle_{AE}^{\otimes\tilde{n}-r})}. The operator ρA1n~E1n~:=ΠρA1n~E1n~Π\rho^{\prime}_{A_{1}^{\tilde{n}}E_{1}^{\tilde{n}}}:=\Pi\rho_{A_{1}^{\tilde{n}}E_{1}^{\tilde{n}}}\Pi^{\dagger} is subnormalized, i.e.,

tr[ρ]=tr[ΠρΠ]1,\displaystyle\mathrm{tr}[\rho^{\prime}]=\mathrm{tr}[\Pi\rho\Pi^{\dagger}]\leq 1\,, (22)

which can be seen, for example, via Hölder’s inequality [35, Proposition 2.5]. The fidelity between two density operators is defined as F(ρ,σ):=ρσ12F(\rho,\sigma):=\left\lVert\sqrt{\rho}\sqrt{\sigma}\right\rVert^{2}_{1}. Hence, for any ω0\omega\geq 0 we have

F(ω,ρ)=ωρ12=tr[ρ]ωρtr[ρ]12Equation 22ωρtr[ρ]12=F(ω,ρtr[ρ]).\displaystyle F(\omega,\rho^{\prime})=\left\lVert\sqrt{\omega}\sqrt{\rho^{\prime}}\right\rVert^{2}_{1}=\mathrm{tr}[\rho^{\prime}]\left\lVert\sqrt{\omega}\sqrt{\frac{\rho^{\prime}}{\mathrm{tr}[\rho^{\prime}]}}\right\rVert^{2}_{1}\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_subnormalized_ds}}}}{\leq}\left\lVert\sqrt{\omega}\sqrt{\frac{\rho^{\prime}}{\mathrm{tr}[\rho^{\prime}]}}\right\rVert^{2}_{1}=F\Big(\omega,\frac{\rho^{\prime}}{\mathrm{tr}[\rho^{\prime}]}\Big)\,. (23)

Using the permutation invariance of ρA1n~E1n~\rho_{A_{1}^{\tilde{n}}E_{1}^{\tilde{n}}} we obtain

F(ρA1sE1s,|θθ|s)Equation 23F(ρA(j1)s+1jsE(j1)s+1js,ρA(j1)s+1jsE(j1)s+1js)DPIF(ρA1n~E1n~,ρA1n~E1n~),\displaystyle F(\rho_{A_{1}^{s}E_{1}^{s}},|\theta\rangle\!\langle\theta|^{\otimes s})\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_monotonicity_operator}}}}{\geq}F(\rho_{A_{(j-1)s+1}^{js}E_{(j-1)s+1}^{js}},\rho^{\prime}_{A_{(j-1)s+1}^{js}E_{(j-1)s+1}^{js}})\overset{\textnormal{DPI}}{\geq}F(\rho_{A_{1}^{\tilde{n}}E_{1}^{\tilde{n}}},\rho^{\prime}_{A_{1}^{\tilde{n}}E_{1}^{\tilde{n}}})\,, (24)

where the final step uses the data-processing inequality for the fidelity [21, Lemma B.4]. By definition of the fidelity we have

F(ρA1n~E1n~,ρA1n~E1n~)=ρρ1=tr[ρρρ]=tr[ρΠρΠρ]=tr[Πρ].\displaystyle\sqrt{F(\rho_{A_{1}^{\tilde{n}}E_{1}^{\tilde{n}}},\rho^{\prime}_{A_{1}^{\tilde{n}}E_{1}^{\tilde{n}}})}=\left\lVert\sqrt{\rho}\sqrt{\rho^{\prime}}\right\rVert_{1}=\mathrm{tr}\Big[\sqrt{\sqrt{\rho}\rho^{\prime}\sqrt{\rho}}\,\Big]=\mathrm{tr}\Big[\sqrt{\sqrt{\rho}\Pi\rho\Pi^{\dagger}\sqrt{\rho}}\,\Big]=\mathrm{tr}[\Pi\rho]\,. (25)

By definition of Π\Pi we have

tr[Πρ]=1tr[Πρ]=Equation 211t𝒯,jtΨt|ρ|Ψt=Equation 161t𝒯,jtβt,t=Equation 181wj.\displaystyle\mathrm{tr}[\Pi\rho]=1-\mathrm{tr}[\Pi^{\perp}\rho]\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_def_PI}}}}{=}1-\sum_{t\in\mathcal{T},j\in\mathcal{F}_{t}}\langle\Psi_{t}|\rho|\Psi_{t}\rangle\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_ONB_dec}}}}{=}1-\sum_{t\in\mathcal{T},j\in\mathcal{F}_{t}}\beta_{t,t}\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_wi}}}}{=}1-w_{j}\,. (26)

The Fuchs-van der Graaf inequality [22] then gives

ρA1sE1s|θθ|s1\displaystyle\left\lVert\rho_{A_{1}^{s}E_{1}^{s}}-|\theta\rangle\!\langle\theta|^{\otimes s}\right\rVert_{1} 21F(ρA1sE1s,|θθ|s)\displaystyle\leq 2\sqrt{1-F(\rho_{A_{1}^{s}E_{1}^{s}},|\theta\rangle\!\langle\theta|^{\otimes s})} (27)
21(1wj)2\displaystyle{\leq}2\sqrt{1-(1-w_{j})^{2}} (28)
22rns.\displaystyle{\leq}2\sqrt{2}\sqrt{\frac{r}{\lfloor\frac{n}{s}\rfloor}}\,. (29)

Recalling that the trace distance is contractive under the partial trace [43, Theorem 8.16] implies

ρA1sσAs1ρA1sE1s|θθ|s1Equation 2922rns.\displaystyle\left\lVert\rho_{A_{1}^{s}}-\sigma_{A}^{\otimes s}\right\rVert_{1}\leq\left\lVert\rho_{A_{1}^{s}E_{1}^{s}}-|\theta\rangle\!\langle\theta|^{\otimes s}\right\rVert_{1}\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_almostDONE-dist}}}}{\leq}2\sqrt{2}\sqrt{\frac{r}{\lfloor\frac{n}{s}\rfloor}}\,. (30)

To conclude the proof of the assertion, note that for q=nsq=\lfloor\frac{n}{s}\rfloor we have

sq=snsns(q+1).\displaystyle sq=s\left\lfloor\frac{n}{s}\right\rfloor\leq n\leq s(q+1)\,. (31)

Hence

r/qrs/n=nqsEquation 31s(q+1)qs=1+1q2,\displaystyle\frac{\sqrt{r/q}}{\sqrt{rs/n}}=\sqrt{\frac{n}{qs}}\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_trivial_step}}}}{\leq}\sqrt{\frac{s(q+1)}{qs}}=\sqrt{1+\frac{1}{q}}\leq\sqrt{2}\,, (32)

where the final step uses q1q\geq 1. Putting everything together yields

ρA1sσAs1Equation 3022rnsEquation 324rsn.\displaystyle\left\lVert\rho_{A_{1}^{s}}-\sigma_{A}^{\otimes s}\right\rVert_{1}\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_almostDONE2-dist}}}}{\leq}2\sqrt{2}\sqrt{\frac{r}{\lfloor\frac{n}{s}\rfloor}}\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_trivial_step2}}}}{\leq}4\sqrt{\frac{rs}{n}}\,. (33)

Another property concerns the statistics of almost-iid and iid states when performing an iid measurement (as is used in a tomography procedure, for example) [33, Theorem 4.5.2]. Let σS()\sigma\in\mathrm{S}(\mathcal{H}) and assume that nn independent measurements with respect to a POVM ={Mx}x𝒳\mathcal{M}=\{M_{x}\}_{x\in\mathcal{X}} are performed on nn iid copies of σ\sigma, giving the outcomes 𝐱=(x1,,xn){\bf{x}}=(x_{1},\dots,x_{n}) with xi𝒳ix_{i}\in\mathcal{X}\ \forall\,i. The outcomes 𝐱{\bf{x}} can be characterized by a frequency distribution (or type) λ𝐱\lambda_{\bf{x}}, which is defined as the following probability distribution on 𝒳\mathcal{X},

λ𝐱(y):=1n|{i:xi=y}|,\displaystyle\lambda_{\bf{x}}(y):=\frac{1}{n}\big|\{i:\,x_{i}=y\}\big|\,, (34)

for all y𝒳y\in\mathcal{X}. Intuitively, λ𝐱(y)\lambda_{\bf{x}}(y) corresponds to the relative number of occurrences of yy in the sequence of outcomes 𝐱=(x1,,xn){\bf{x}}=(x_{1},\dots,x_{n}). For large values of nn, it follows by the law of large numbers that the frequency distribution λ𝐱\lambda_{\bf{x}} is close to the probability distribution PXP_{X} defined by PX(y):=tr(Myσ)P_{X}(y):=\mathrm{tr}(M_{y}\sigma). Proposition 2.7 shows that a similar statement holds when performing nn independent measurements on a (nr)\binom{n}{r}-almost-iid state in σ\sigma for small enough values of rr. This proves that almost-iid and iid states cannot be distinguished by any practically feasible measurement. This is discussed in more detail in [29].

Proposition 2.7 (Statistics of almost-iid states).

Let n,rn,r\in\mathbb{N} such that r12nr\leq\tfrac{1}{2}n, σS()\sigma\in\mathrm{S}(\mathcal{H}), and ρ(n)Sn(,σnr)\rho^{(n)}\in\mathrm{S}^{n}(\mathcal{H},\sigma^{\otimes n-r}). Let ={Mx}x𝒳\mathcal{M}=\{M_{x}\}_{x\in\mathcal{X}} be a POVM on \mathcal{H}, and PX(x)=tr[Mxσ]P_{X}(x)=\mathrm{tr}[M_{x}\sigma] for all x𝒳x\in\mathcal{X}. Then, for any ε>0\varepsilon>0, we have

𝐱[λ𝐱PX1>f(ε,r,n)]ε,\displaystyle\underset{{\bf{x}}}{\mathbb{P}}\big[\left\lVert\lambda_{\bf{x}}-P_{X}\right\rVert_{1}>f(\varepsilon,r,n)\big]\leq\varepsilon\,, (35)

where the probability is taken over the outcomes 𝐱=(x1,,xn){\bf{x}}=(x_{1},\dots,x_{n}) of the product measurement n\mathcal{M}^{\otimes n} applied to ρ(n)\rho^{(n)}, and

f(ε,r,n)=2log(1ε)n+|𝒳|nlog(n2+1)+h(rn)+2rnlog(d)+2rn,\displaystyle f(\varepsilon,r,n)=2\sqrt{\frac{\log{\big(\frac{1}{\varepsilon}\big)}}{n}+\frac{|\mathcal{X}|}{n}\log\Big(\frac{n}{2}+1\Big)+h\Big(\frac{r}{n}\Big)+\frac{2r}{n}\log(d)}+\frac{2r}{n}\,, (36)

where h()h(\cdot) is the binary entropy and d=dimd=\dim{\mathcal{H}}. Furthermore, if r=o(n)r=o(n), then limnf(ε,r,n)=0\lim_{n\to\infty}f(\varepsilon,r,n)=0.

The proof is given in Appendix C.

3 Quantum vs. classical exponential de Finetti theorem

In this section, we formally state the quantum exponential de Finetti theorem [34] which justifies the definition of almost-iid states. We then show that for classical almost-iid distributions, defined as convex combinations of product distributions with a small number of defects, no classical exponential de Finetti theorem can hold.

Theorem 3.1 (Exponential de Finetti [34, Theorem 1]).

Let n,k,rn,k,r\in\mathbb{N} and let \mathcal{H} be a dd-dimensional Hilbert space. For any |Φ(n+k)Symn+k()|\Phi^{(n+k)}\rangle\in\mathrm{Sym}^{n+k}(\mathcal{H}) there exists a probability measure ν\nu on the unit sphere ()\mathcal{B}(\mathcal{H}) and a family {|Ψr,θ(n)}θ\{|\Psi^{(n)}_{r,\theta}\rangle\}_{\theta} of states such that |Ψr,θ(n)Symn(,|θnr)|\Psi^{(n)}_{r,\theta}\rangle\in\mathrm{Sym}^{n}(\mathcal{H},|\theta\rangle^{\otimes n-r}) and

trk[|Φ(n+k)Φ(n+k)|]|Ψr,θ(n)Ψr,θ(n)|ν(𝑑θ)13kdek(r+1)n+k.\displaystyle\left\lVert\mathrm{tr}_{k}[|\Phi^{(n+k)}\rangle\!\langle\Phi^{(n+k)}|]-\int|\Psi^{(n)}_{r,\theta}\rangle\!\langle\Psi^{(n)}_{r,\theta}|\nu(\mathrm{d}\theta)\right\rVert_{1}\leq 3k^{d}\mathrm{e}^{-\frac{k(r+1)}{n+k}}\,. (37)

We next define almost-iid distributions which is the classical counterpart to almost-iid states. For a finite alphabet 𝒳\mathcal{X}, let Sym(𝒳n)\mathrm{Sym}(\mathcal{X}^{n}) denote the set of permutation-invariant distributions on 𝒳n\mathcal{X}^{n}. Furthermore, for a distribution qq on 𝒳\mathcal{X} let

𝒱(𝒳n,qm):={π(qm×RX1nm):π𝒮n,RX1nm distribution on 𝒳nm}\displaystyle\mathcal{V}(\mathcal{X}^{n},q^{m}):=\{\pi(q^{m}\times R_{X_{1}^{n-m}}):\pi\in\mathcal{S}_{n},\,R_{X_{1}^{n-m}}\textnormal{ distribution on }\mathcal{X}^{n-m}\} (38)

and

Sym(𝒳n,qm):=Sym(𝒳n)conv𝒱(𝒳n,qm).\displaystyle\mathrm{Sym}(\mathcal{X}^{n},q^{m}):=\mathrm{Sym}(\mathcal{X}^{n})\cap\mathrm{conv}\,\mathcal{V}(\mathcal{X}^{n},q^{m})\,. (39)

Note that the set Sym(𝒳n,qm)\mathrm{Sym}(\mathcal{X}^{n},q^{m}) is considerably simpler compared to Symn(,|θm)\mathrm{Sym}^{n}(\mathcal{H},|\theta\rangle^{\otimes m}) since the former set consists of convex combinations of product distributions, whereas the latter set contains superpositions of product states.

The following proposition states that no classical version of the quantum de Finetti theorem can exist where almost-iid states are replaced with almost-iid distributions.

Proposition 3.1 (No classical exponential de Finetti result).

Let nn\in\mathbb{N}, k=o(n)k=o(n), r=o(n)r=o(n), 𝒳\mathcal{X} be a finite alphabet of size dd. There exists PX1n+kSym(𝒳n+k)P_{X_{1}^{n+k}}\in\mathrm{Sym}(\mathcal{X}^{n+k}) such that it is not possible to approximate PX1nP_{X_{1}^{n}} with a probabilistic mixture of distributions {QX1n(q)}q\{Q_{X_{1}^{n}}^{(q)}\}_{q} with QX1n(q)Sym(𝒳n,qnr)Q_{X_{1}^{n}}^{(q)}\in\mathrm{Sym}(\mathcal{X}^{n},q^{n-r}) in the 1\ell^{1}-norm up to εd(n)\varepsilon_{d}(n) such that limnεd(n)=0\lim_{n\to\infty}\varepsilon_{d}(n)=0.

The proof of Proposition 3.1 is given in Appendix D. As mentioned in Equation 1, if we are willing to choose kk large, more precisely k=ω(n)k=\omega(n), then a classical de Finetti theorem holds. However, Proposition 3.1 states that this is no longer the case for k=o(n)k=o(n), even if the error term would decrease at a non-exponential rate. Further note that Proposition 3.1 does not prohibit the existence of a classical exponential de Finetti theorem if the definition of almost-iid distributions, Equation 39, is relaxed. For example, one could define the set 𝒱(𝒳n,qm)\mathcal{V}(\mathcal{X}^{n},q^{m}) in Equation 38 by only requiring the marginals of the distributions to yield an iid state instead of imposing a product structure between the iid part and the defects. However, such a structure would no longer be covered by the quantum version in Definition 2.1 and therefore, may potentially be difficult to work with.

One may still wonder what Theorem 3.1 produces if it is applied to a classical symmetric distribution. Can the resulting family of almost-iid states that approximates the classical distribution become classical as well? To analyze this, let PX1n+kSym(𝒳n+k)P_{X_{1}^{n+k}}\in\mathrm{Sym}(\mathcal{X}^{n+k}) be a symmetric classical distribution. We can apply Theorem 3.1 to PX1n+kP_{X_{1}^{n+k}} by embedding it into a permutation-invariant density matrix,

ρA1n+k=x1n+kPX1n+k(x1n+k)|x1x1||xn+kxn+k|.\displaystyle\rho_{A_{1}^{n+k}}=\sum_{x^{n+k}_{1}}P_{X_{1}^{n+k}}(x_{1}^{n+k})|x_{1}\rangle\!\langle x_{1}|\cdots|x_{n+k}\rangle\!\langle x_{n+k}|\,. (40)

Let |Φ(n+k)A1n+kE1n+k|\Phi^{(n+k)}\rangle_{A_{1}^{n+k}E_{1}^{n+k}} be a symmetric purification of ρA1n+k\rho_{A_{1}^{n+k}}. Consider the measurement channel on the AA-system

:YAEx|xx|AYAE|xx|A.\displaystyle\mathcal{M}:Y_{AE}\mapsto\sum_{x}|x\rangle\!\langle x|_{A}Y_{AE}|x\rangle\!\langle x|_{A}\,. (41)

Because ρA1n\rho_{A_{1}^{n}} is classical, we have

ρA1n=n(ρA1n)=n(trE1n[ρA1nE1n])\displaystyle\rho_{A_{1}^{n}}=\mathcal{M}^{\otimes n}(\rho_{A_{1}^{n}})=\mathcal{M}^{\otimes n}\big(\mathrm{tr}_{E_{1}^{n}}[\rho_{A_{1}^{n}E_{1}^{n}}]\big) =trE1n[n(ρA1nE1n)]\displaystyle=\mathrm{tr}_{E_{1}^{n}}\big[\mathcal{M}^{\otimes n}(\rho_{A_{1}^{n}E_{1}^{n}})\big] (42)
=trE1n[n(trk[|Φ(n+k)Φ(n+k)|])],\displaystyle=\mathrm{tr}_{E_{1}^{n}}\big[\mathcal{M}^{\otimes n}(\mathrm{tr}_{k}[|\Phi^{(n+k)}\rangle\!\langle\Phi^{(n+k)}|])\big]\,, (43)

where ρA1nE1n:=trk[|Φ(n+k)Φ(n+k)|]\rho_{A_{1}^{n}E_{1}^{n}}:=\mathrm{tr}_{k}[|\Phi^{(n+k)}\rangle\!\langle\Phi^{(n+k)}|], and the penultimate step uses that \mathcal{M} only acts on the AA-system and hence commutes with the partial trace over the EE-system. Theorem 3.1 shows that there exist a probability measure ν\nu on the unit sphere ()\mathcal{B}(\mathcal{H}) and, for each |θ()|\theta\rangle\in\mathcal{B}(\mathcal{H}), a family {|Ψr,θ(n)}θ\{|\Psi^{(n)}_{r,\theta}\rangle\}_{\theta} of states such that |Ψr,θ(n)Symn(,|θnr)|\Psi^{(n)}_{r,\theta}\rangle\in\mathrm{Sym}^{n}(\mathcal{H},|\theta\rangle^{\otimes n-r}), and such that for σθ(n):=trE1n[n(|Ψr,θ(n)Ψr,θ(n)|)]\sigma^{(n)}_{\theta}:=\mathrm{tr}_{E_{1}^{n}}[\mathcal{M}^{\otimes n}(|\Psi^{(n)}_{r,\theta}\rangle\!\langle\Psi^{(n)}_{r,\theta}|)] and ν¯\bar{\nu} denoting the induced measure obtained by taking the partial trace, we have

ρA1nσθ(n)ν¯(𝑑θ)1\displaystyle\left\lVert\rho_{A_{1}^{n}}\!-\!\int\!\sigma^{(n)}_{\theta}\!\bar{\nu}(\mathrm{d}\theta)\right\rVert_{1}\!\! =trE1n[n(trk[|Φ(n+k)Φ(n+k)|])]trE1n[n(|Ψr,θ(n)Ψr,θ(n)|ν(𝑑θ))]1\displaystyle{=}\!\!\left\lVert\mathrm{tr}_{E_{1}^{n}}\big[\mathcal{M}^{\otimes n}(\mathrm{tr}_{k}[|\Phi^{(n+k)}\rangle\!\langle\Phi^{(n+k)}|])\big]\!-\!\mathrm{tr}_{E_{1}^{n}}\Big[\mathcal{M}^{\otimes n}\Big(\!\int|\Psi^{(n)}_{r,\theta}\rangle\!\langle\Psi^{(n)}_{r,\theta}|\nu(\mathrm{d}\theta)\Big)\Big]\right\rVert_{1}
trk[|Φ(n+k)Φ(n+k)|])|Ψr,θ(n)Ψr,θ(n)|ν(dθ)1\displaystyle\leq\left\lVert\mathrm{tr}_{k}[|\Phi^{(n+k)}\rangle\!\langle\Phi^{(n+k)}|])-\int|\Psi^{(n)}_{r,\theta}\rangle\!\langle\Psi^{(n)}_{r,\theta}|\nu(\mathrm{d}\theta)\right\rVert_{1} (44)
3kdek(r+1)n+k,\displaystyle{\leq}3k^{d}\mathrm{e}^{-\frac{k(r+1)}{n+k}}\,, (45)

where the second step uses the fact that the trace distance is contractive under trace-preserving completely positive maps [43, Theorem 8.16]. Proposition 3.1 now implies that the density operators σθ(n)\sigma^{(n)}_{\theta} cannot be given almost-iid distributions in the sense of Equation 39. Intuitively, this happens since the states |θ|\theta\rangle appearing in Equation 44 are not necessarily classical and therefore, the almost-iid structure is not preserved after the measurement.

4 Conditional entropy of almost-iid states

In this section, we prove that the conditional entropy of almost-iid states asymptotically coincides with the conditional entropy of iid states. In the proof, we develop technical tools that may be of independent interest.

For ρABS(AB)\rho_{AB}\in\mathrm{S}(\mathcal{H}_{AB}) the conditional entropy is defined as H(A|B)ρ=H(AB)ρH(B)ρH(A|B)_{\rho}=H(AB)_{\rho}-H(B)_{\rho}, where H(B)ρ:=tr[ρBlogρB]H(B)_{\rho}:=-\mathrm{tr}[\rho_{B}\log\rho_{B}] is the von Neumann entropy. It is straightforward to see that the conditional entropy is additive for iid states, i.e.,

1nH(A1n|B1n)ρn=H(A|B)ρ.\displaystyle\frac{1}{n}H(A_{1}^{n}|B_{1}^{n})_{\rho^{\otimes n}}=H(A|B)_{\rho}\,. (46)

We next show that this property is preserved for almost-iid states in the limit nn\to\infty.

Theorem 4.1.

Let σABS(AB)\sigma_{AB}\in\mathrm{S}(\mathcal{H}_{AB}) and ρA1nB1nSn(AB,σABnr)\rho_{A_{1}^{n}B_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{AB},\sigma_{AB}^{\otimes n-r}) for r=o(n)r=o(n). Then

1nH(A1n|B1n)ρ=H(A|B)σ+o(n)n.\displaystyle\frac{1}{n}H(A_{1}^{n}|B_{1}^{n})_{\rho}=H(A|B)_{\sigma}+\frac{o(n)}{n}\,. (47)

The main difficulty in proving the assertion of Theorem 4.1 is the fact that almost-iid states are more general than just convex mixtures of product states where each element in the convex sum has a certain number of defects (see Example 2.2). A look at Definition 2.1 reveals that almost-iid states may contain superpositions which store long range correlations and entanglement. Dealing with these superpositions is the main technical challenge in the proof. We do this utilizing tools from one-shot information theory such as Rényi and smooth entropies. We also want to emphasize that the superpositions in the definition of almost-iid states are crucial for making the exponential de Finetti theorem (Theorem 3.1) possible. As shown in Proposition 3.1, no exponential de Finetti theorem can exist without superpositions.

Before presenting the proof of Theorem 4.1, which is given in Section 4.4, we need to define some entropic quantities. We note that an alternative proof using different techniques that may therefore be of independent interest is given in Appendix E.

4.1 Entropic quantities

For α[1/2,1)(1,)\alpha\in[1/2,1)\cup(1,\infty) the sandwiched Rényi divergence [31, 41] is given by

Dα(ρσ):=1α1logtr[(σ1α2αρσ1α2α)α].\displaystyle D_{\alpha}(\rho\|\sigma):=\frac{1}{\alpha-1}\log\mathrm{tr}\big[(\sigma^{\frac{1-\alpha}{2\alpha}}\rho\,\sigma^{\frac{1-\alpha}{2\alpha}})^{\alpha}\big]\,. (48)

For α=12\alpha=\frac{1}{2} we have D12(ρσ)=logF(ρ,σ)D_{\frac{1}{2}}(\rho\|\sigma)=-\log F(\rho,\sigma). In the limits α1\alpha\to 1 and α\alpha\to\infty the sandwiched Rényi divergence converges to the relative entropy D(ρσ)D(\rho\|\sigma) and the max-relative entropy [33, 16]

Dmax(ρσ):=inf{λ:ρ2λσ},\displaystyle D_{\max}(\rho\|\sigma):=\inf\{\lambda\in\mathbb{R}:\rho\leq 2^{\lambda}\sigma\}\,, (49)

respectively. For ρABS(AB)\rho_{AB}\in\mathrm{S}(\mathcal{H}_{AB}), the Rényi divergence can be used to define a conditional Rényi entropy [36]

Hα(A|B)ρ:=minσBS(B)Dα(ρABidAσB),\displaystyle H_{\alpha}(A|B)_{\rho}:=-\min_{\sigma_{B}\in\mathrm{S}(\mathcal{H}_{B})}D_{\alpha}(\rho_{AB}\|\mathrm{id}_{A}\otimes\sigma_{B})\,, (50)

which converges to the conditional entropy H(A|B)ρH(A|B)_{\rho} for α1\alpha\to 1 and to the conditional min-entropy Hmin(A|B)ρH_{\min}(A|B)_{\rho} for α\alpha\to\infty. For α=1/2\alpha=1/2 we obtain the conditional max-entropy Hmax(A|B)ρH_{\max}(A|B)_{\rho}. The trace distance between two states ρ,σS()\rho,\sigma\in\mathrm{S}(\mathcal{H}) is given by Δ(ρ,σ):=12ρσ1\Delta(\rho,\sigma):=\frac{1}{2}\left\lVert\rho-\sigma\right\rVert_{1} and the purified distance [36] is defined as P(ρ,σ):=1F(ρ,σ)P(\rho,\sigma):=\sqrt{1-F(\rho,\sigma)}. The Fuchs-van der Graaf inequality [22] implies P(ρ,σ)Δ(ρ,σ)P(\rho,\sigma)\geq\Delta(\rho,\sigma). For ρS()\rho\in\mathrm{S}(\mathcal{H}) and ε(0,1)\varepsilon\in(0,1) define the ε\varepsilon-ball around ρ\rho by ε(ρ):={ρS():P(ρ,ρ)ε}\mathcal{B}_{\varepsilon}(\rho):=\{\rho^{\prime}\in\mathrm{S}(\mathcal{H}):P(\rho,\rho^{\prime})\leq\varepsilon\}. We then define a smooth variant of the min- and max-entropy by

Hminε(A|B)ρ:=maxρε(ρ)Hmin(A|B)ρandHmaxε(A|B)ρ:=minρε(ρ)Hmax(A|B)ρ.\displaystyle H^{\varepsilon}_{\min}(A|B)_{\rho}:=\max_{\rho^{\prime}\in\mathcal{B}_{\varepsilon}(\rho)}H_{\min}(A|B)_{\rho^{\prime}}\qquad\textnormal{and}\qquad H^{\varepsilon}_{\max}(A|B)_{\rho}:=\min_{\rho^{\prime}\in\mathcal{B}_{\varepsilon}(\rho)}H_{\max}(A|B)_{\rho^{\prime}}\,. (51)

4.2 Asymptotic equipartition property for almost-iid states

Another statement which can be proven using similar techniques and may be of independent interest is a strong asymptotic equipartition property (AEP) for almost-iid states. To understand this, let us recall the AEP for iid states [37, 36]. This fundamental result ensures that for any density operator ρABS(AB)\rho_{AB}\in\mathrm{S}(\mathcal{H}_{AB}) and any ε(0,1)\varepsilon\in(0,1) we have

1nHminε(A1n|B1n)ρn=H(A|B)ρ+o(n)nand1nHmaxε(A1n|B1n)ρn=H(A|B)ρ+o(n)n.\displaystyle\frac{1}{n}H^{\varepsilon}_{\min}(A_{1}^{n}|B_{1}^{n})_{\rho^{{\otimes n}}}=H(A|B)_{\rho}+\frac{o(n)}{n}\qquad\textnormal{and}\qquad\frac{1}{n}H^{\varepsilon}_{\max}(A_{1}^{n}|B_{1}^{n})_{\rho^{{\otimes n}}}=H(A|B)_{\rho}+\frac{o(n)}{n}\,. (52)

Note that Equation 52 is called a strong AEP as the error term o(n)n\frac{o(n)}{n} vanishes in the limit nn\to\infty for any fixed ε(0,1)\varepsilon\in(0,1). Furthermore, it is understood how fast the error term vanishes for finite values of nn [38]. We show that Equation 52 remains valid when replacing iid states with almost-iid states.

Proposition 4.2 (Strong AEP for almost-iid states).

Let ε(0,1)\varepsilon\in(0,1), σABS(AB)\sigma_{AB}\in\mathrm{S}(\mathcal{H}_{AB}), and ρA1nB1nSn(AB,σABnr)\rho_{A_{1}^{n}B_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{AB},\sigma_{AB}^{\otimes n-r}) for r=o(n)r=o(n). Then

1nHminε(A1n|B1n)ρ=H(A|B)σ+o(n)nand1nHmaxε(A1n|B1n)ρ=H(A|B)σ+o(n)n.\displaystyle\frac{1}{n}H^{\varepsilon}_{\min}(A_{1}^{n}|B_{1}^{n})_{\rho}=H(A|B)_{\sigma}+\frac{o(n)}{n}\qquad\textnormal{and}\qquad\frac{1}{n}H^{\varepsilon}_{\max}(A_{1}^{n}|B_{1}^{n})_{\rho}=H(A|B)_{\sigma}+\frac{o(n)}{n}\,. (53)

In [33, Theorem 4.4.1] it was shown that the conditional smooth min-entropy for almost-iid states, which are classical on one subsystem, asymptotically coincides with the conditional entropy of iid states. This was crucial to prove security of quantum key distribution via a de Finetti argument. Using the duality of conditional entropy, the result can be lifted to the smooth max-entropy of almost-iid states [44]. In addition, for the case of pure almost-iid states, a similar result has been proven based on [33] in [44, Lemma 11]. Here, the presented proof of Proposition 4.2 is more general as it applies for mixed almost-iid states and is also more modular allowing one to distill other results such as Theorem 4.1. Beyond these results, to the best of our knowledge, little is known about how entropic functions behave for almost-iid states.

4.3 Proof of Proposition 4.2

Let ρA1nB1nSn(AB,σABnr)\rho_{A_{1}^{n}B_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{AB},\sigma_{AB}^{\otimes n-r}). By definition, there exist a purification |θABE|\theta\rangle_{ABE} of σAB\sigma_{AB} and an extension ρA1nB1nE1n\rho_{A_{1}^{n}B_{1}^{n}E_{1}^{n}} of ρA1nB1n\rho_{A_{1}^{n}B_{1}^{n}} that can be written as

ρA1nB1nE1n=t,t𝒯βt,t|ΨtΨt|A1nB1nE1n,\displaystyle\rho_{A_{1}^{n}B_{1}^{n}E_{1}^{n}}=\sum_{t,t^{\prime}\in\mathcal{T}}\beta_{t,t^{\prime}}|\Psi_{t}\rangle\langle\Psi_{t^{\prime}}|_{A_{1}^{n}B_{1}^{n}E_{1}^{n}}\,, (54)

for a family {|ΨtA1nB1nE1n}t𝒯\{|\Psi_{t}\rangle_{A_{1}^{n}B_{1}^{n}E_{1}^{n}}\}_{t\in\mathcal{T}} of orthonormal vectors from 𝒱(ABEn,|θABEnr)\mathcal{V}(\mathcal{H}_{ABE}^{\otimes n},|\theta\rangle_{ABE}^{\otimes n-r}) with βt,t\beta_{t,t^{\prime}}\in\mathbb{C} satisfying t𝒯βt,t=1\sum_{t\in\mathcal{T}}\beta_{t,t}=1 and

log|𝒯|nh(rn)+rlogdABE.\displaystyle\log|\mathcal{T}|\leq n\,h\Big(\frac{r}{n}\Big)+r\log d_{ABE}\,. (55)

Let

ρ~A1nB1nE1nT:=t𝒯βt,t|ΨtΨt|A1nB1nE1n=:ρ~(t)|tt|T.\displaystyle\tilde{\rho}_{A_{1}^{n}B_{1}^{n}E_{1}^{n}T}:=\sum_{t\in\mathcal{T}}\beta_{t,t}\underbrace{|\Psi_{t}\rangle\!\langle\Psi_{t}|_{A_{1}^{n}B_{1}^{n}E_{1}^{n}}}_{=:\tilde{\rho}^{(t)}}\otimes|t\rangle\!\langle t|_{T}\,. (56)
Lemma 4.3.

For the setting defined above, we have

ρA1nB1nE1n|𝒯|ρ~A1nB1nE1n,and henceρA1nB1n|𝒯|ρ~A1nB1n.\displaystyle\rho_{A_{1}^{n}B_{1}^{n}E_{1}^{n}}\leq|\mathcal{T}|\tilde{\rho}_{A_{1}^{n}B_{1}^{n}E_{1}^{n}}\,,\quad\textnormal{and hence}\quad\rho_{A_{1}^{n}B_{1}^{n}}\leq|\mathcal{T}|\tilde{\rho}_{A_{1}^{n}B_{1}^{n}}\,. (57)
Proof.

The proof idea is similar to [33, Proof of Lemma 3.1.13] but is based on pinching maps and therefore works for a more general setup. Consider the pinching map

𝒫:Zt𝒯|ΨtΨt|Z|ΨtΨt|.\displaystyle\mathcal{P}:Z\mapsto\sum_{t\in\mathcal{T}}|\Psi_{t}\rangle\!\langle\Psi_{t}|Z|\Psi_{t}\rangle\!\langle\Psi_{t}|\,. (58)

Using the fact that {|Ψt}t𝒯\{|\Psi_{t}\rangle\}_{t\in\mathcal{T}} is orthonormal, we have

𝒫(ρA1nB1nE1n)=Equation 54t,k,𝒯βk,|ΨtΨt||ΨkΨ||ΨtΨt|=t𝒯βt,t|ΨtΨt|=Equation 56ρ~A1nB1nE1n.\displaystyle\mathcal{P}(\rho_{A_{1}^{n}B_{1}^{n}E_{1}^{n}})\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_decomp_psi}}}}{=}\sum_{t,k,\ell\in\mathcal{T}}\beta_{k,\ell}|\Psi_{t}\rangle\!\langle\Psi_{t}||\Psi_{k}\rangle\langle\Psi_{\ell}||\Psi_{t}\rangle\!\langle\Psi_{t}|=\sum_{t\in\mathcal{T}}\beta_{t,t}|\Psi_{t}\rangle\!\langle\Psi_{t}|\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_tilde_rho}}}}{=}\tilde{\rho}_{A_{1}^{n}B_{1}^{n}E_{1}^{n}}\,. (59)

Hence, we find

ρ~A1nB1nE1n=Equation 59𝒫(ρA1nB1nE1n)pinching inequality1|𝒯|ρA1nB1nE1n.\displaystyle\tilde{\rho}_{A_{1}^{n}B_{1}^{n}E_{1}^{n}}\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_pinching_ds1}}}}{=}\mathcal{P}(\rho_{A_{1}^{n}B_{1}^{n}E_{1}^{n}})\overset{\textnormal{pinching inequality}}{\geq}\frac{1}{|\mathcal{T}|}\rho_{A_{1}^{n}B_{1}^{n}E_{1}^{n}}\,. (60)

The interested reader can find more information on pinching maps, including a proof of the pinching inequality in [35, Lemma 3.5]. Since the partial trace is a completely positive map Equation 60 implies

ρ~A1nB1n1|𝒯|ρA1nB1n.\displaystyle\tilde{\rho}_{A_{1}^{n}B_{1}^{n}}\geq\frac{1}{|\mathcal{T}|}\rho_{A_{1}^{n}B_{1}^{n}}\,. (61)

Lemma 4.4.

Let n,rn,r\in\mathbb{N} such that rnr\leq n, σABS(AB)\sigma_{AB}\in\mathrm{S}(\mathcal{H}_{AB}), ρA1nB1nSn(AB,σABnr)\rho_{A_{1}^{n}B_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{AB},\sigma_{AB}^{\otimes n-r}), dAB=dimABd_{AB}=\dim\mathcal{H}_{AB}, and dA=dimAd_{A}=\dim\mathcal{H}_{A}. Then

1nHα(A1n|B1n)ρHα(A|B)σ2rnlogdAαα1(h(rn)+2rnlogdAB)α>1\displaystyle\frac{1}{n}H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\rho}\geq H_{\alpha}(A|B)_{\sigma}-\frac{2r}{n}\log d_{A}-\frac{\alpha}{\alpha-1}\left(h\Big(\frac{r}{n}\Big)+\frac{2r}{n}\log d_{AB}\right)\quad\forall\alpha>1\, (62)

and

1nHα(A1n|B1n)ρHα(A|B)σ+2rnlogdA+11α(h(rn)+2rnlogdAB)α[12,1).\displaystyle\frac{1}{n}H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\rho}\leq H_{\alpha}(A|B)_{\sigma}+\frac{2r}{n}\log d_{A}+\frac{1}{1-\alpha}\left(h\Big(\frac{r}{n}\Big)+\frac{2r}{n}\log d_{AB}\right)\quad\forall\alpha\in[\tfrac{1}{2},1)\,. (63)
Proof.

We start by proving Equation 62. For α>1\alpha>1 and ρ\rho, ρ~\tilde{\rho} defined above, we have

α1αlog|𝒯|+Hα(A1n|B1n)ρ~\displaystyle\frac{\alpha}{1-\alpha}\log|\mathcal{T}|+H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\tilde{\rho}} =maxσS(Bn)11αlogtr[(σ1α2αρ~A1nB1n|𝒯|σ1α2α)α]\displaystyle{=}\max_{\sigma\in\mathrm{S}(\mathcal{H}_{B}^{\otimes n})}\frac{1}{1-\alpha}\log\mathrm{tr}\big[(\sigma^{\frac{1-\alpha}{2\alpha}}\tilde{\rho}_{A_{1}^{n}B_{1}^{n}}|\mathcal{T}|\sigma^{\frac{1-\alpha}{2\alpha}})^{\alpha}\big] (64)
maxσS(Bn)11αlogtr[(σ1α2αρA1nB1nσ1α2α)α]\displaystyle{\leq}\max_{\sigma\in\mathrm{S}(\mathcal{H}_{B}^{\otimes n})}\frac{1}{1-\alpha}\log\mathrm{tr}\big[(\sigma^{\frac{1-\alpha}{2\alpha}}\rho_{A_{1}^{n}B_{1}^{n}}\sigma^{\frac{1-\alpha}{2\alpha}})^{\alpha}\big] (65)
=Hα(A1n|B1n)ρ,\displaystyle{=}H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\rho}\,, (66)

where the inequality step used that the function Xtr[Xα]X\mapsto\mathrm{tr}[X^{\alpha}] is monotone [11, Theorem 2.10]. Furthermore, we have

Hα(A1n|B1n)ρ~\displaystyle H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\tilde{\rho}} Hα(A1n|B1nT)ρ~\displaystyle{\geq}H_{\alpha}(A_{1}^{n}|B_{1}^{n}T)_{\tilde{\rho}} (67)
=α1αlog(t𝒯βt,texp(1ααHα(A1n|B1n)ρ~(t)))\displaystyle{=}\frac{\alpha}{1-\alpha}\log\left(\sum_{t\in\mathcal{T}}\beta_{t,t}\exp\Big(\frac{1-\alpha}{\alpha}H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\tilde{\rho}^{(t)}}\Big)\right) (68)
mint𝒯Hα(A1n|B1n)ρ~(t),\displaystyle\geq\min_{t\in\mathcal{T}}H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\tilde{\rho}^{(t)}}\,, (69)

where the final step uses that the logarithm is a quasi-linear function and that βt,t[0,1]\beta_{t,t}\in[0,1] with t𝒯βt,t=1\sum_{t\in\mathcal{T}}\beta_{t,t}=1. Recalling that ρ~(t)=|ΨtΨt|\tilde{\rho}^{(t)}=|\Psi_{t}\rangle\!\langle\Psi_{t}| for |Ψt𝒱(ABEn,|θABEnr)|\Psi_{t}\rangle\in\mathcal{V}(\mathcal{H}_{ABE}^{\otimes n},|\theta\rangle_{ABE}^{\otimes n-r}) and using the additivity of the Rényi entropies under tensor products allows us to write for any t𝒯t\in\mathcal{T}

Hα(A1n|B1n)ρ~(t)=(nr)Hα(A|B)σ+Hα(A1r|B1r)Ω[36, Lem. 5.11](nr)Hα(A|B)σrlogdA.\displaystyle H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\tilde{\rho}^{(t)}}=(n-r)H_{\alpha}(A|B)_{\sigma}+H_{\alpha}(A_{1}^{r}|B_{1}^{r})_{\Omega}\overset{\textnormal{\cite[cite]{[\@@bibref{}{marco_book}{}{}, Lem.~5.11]}}}{\geq}(n-r)H_{\alpha}(A|B)_{\sigma}-r\log d_{A}\,. (70)

Putting everything together yields

1nHα(A1n|B1n)ρ\displaystyle\frac{1}{n}H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\rho} nrnHα(A|B)σrnlogdA1nαα1log|𝒯|\displaystyle\geq\frac{n-r}{n}H_{\alpha}(A|B)_{\sigma}-\frac{r}{n}\log d_{A}-\frac{1}{n}\frac{\alpha}{\alpha-1}\log|\mathcal{T}| (71)
Hα(A|B)σ2rnlogdAαα1(h(rn)+2rnlogdAB),\displaystyle{\geq}H_{\alpha}(A|B)_{\sigma}-\frac{2r}{n}\log d_{A}-\frac{\alpha}{\alpha-1}\left(h\Big(\frac{r}{n}\Big)+\frac{2r}{n}\log d_{AB}\right)\,, (72)

where in the final step, we used dABE=dAB2d_{ABE}=d_{AB}^{2}. This proves Equation 62.

The statement from Equation 63 follows similarly. For α[1/2,1)\alpha\in[1/2,1) we have

α1αlog|𝒯|+Hα(A1n|B1n)ρ~\displaystyle\frac{\alpha}{1-\alpha}\log|\mathcal{T}|+H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\tilde{\rho}} =maxσS(Bn)11αlogtr[(σ1α2αρ~A1nB1n|𝒯|σ1α2α)α]\displaystyle{=}\max_{\sigma\in\mathrm{S}(\mathcal{H}_{B}^{\otimes n})}\frac{1}{1-\alpha}\log\mathrm{tr}\big[(\sigma^{\frac{1-\alpha}{2\alpha}}\tilde{\rho}_{A_{1}^{n}B_{1}^{n}}|\mathcal{T}|\sigma^{\frac{1-\alpha}{2\alpha}})^{\alpha}\big] (73)
maxσS(Bn)11αlogtr[(σ1α2αρA1nB1nσ1α2α)α]\displaystyle{\geq}\max_{\sigma\in\mathrm{S}(\mathcal{H}_{B}^{\otimes n})}\frac{1}{1-\alpha}\log\mathrm{tr}\big[(\sigma^{\frac{1-\alpha}{2\alpha}}\rho_{A_{1}^{n}B_{1}^{n}}\sigma^{\frac{1-\alpha}{2\alpha}})^{\alpha}\big] (74)
=Hα(A1n|B1n)ρ,\displaystyle{=}H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\rho}\,, (75)

where the inequality step used that the function Xtr[Xα]X\mapsto\mathrm{tr}[X^{\alpha}] is monotone [11, Theorem 2.10]. In addition, we have

Hα(A1n|B1n)ρ~\displaystyle H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\tilde{\rho}} Hα(A1n|B1nT)ρ~+log|𝒯|\displaystyle{\leq}H_{\alpha}(A_{1}^{n}|B_{1}^{n}T)_{\tilde{\rho}}+\log|\mathcal{T}| (76)
=α1αlog(t𝒯βt,texp(1ααHα(A1n|B1n)ρ~(t)))+log|𝒯|\displaystyle{=}\frac{\alpha}{1-\alpha}\log\left(\sum_{t\in\mathcal{T}}\beta_{t,t}\exp\Big(\frac{1-\alpha}{\alpha}H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\tilde{\rho}^{(t)}}\Big)\right)+\log|\mathcal{T}| (77)
maxt𝒯Hα(A1n|B1n)ρ~(t)+log|𝒯|,\displaystyle\leq\max_{t\in\mathcal{T}}H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\tilde{\rho}^{(t)}}+\log|\mathcal{T}|\,, (78)

where the final step uses that the logarithm is a quasi-linear function and that βt,t[0,1]\beta_{t,t}\in[0,1] with t𝒯βt,t=1\sum_{t\in\mathcal{T}}\beta_{t,t}=1. Since ρ~(t)=|ΨtΨt|\tilde{\rho}^{(t)}=|\Psi_{t}\rangle\!\langle\Psi_{t}| for |Ψt𝒱(ABEn,|θABEnr)|\Psi_{t}\rangle\in\mathcal{V}(\mathcal{H}_{ABE}^{\otimes n},|\theta\rangle_{ABE}^{\otimes n-r}), the additivity of the Rényi entropies under tensor products implies for any t𝒯t\in\mathcal{T}

Hα(A1n|B1n)ρ~(t)=(nr)Hα(A|B)σ+Hα(A1r|B1r)Ω[36, Lem. 5.11](nr)Hα(A|B)σ+rlogdA.\displaystyle H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\tilde{\rho}^{(t)}}=(n-r)H_{\alpha}(A|B)_{\sigma}+H_{\alpha}(A_{1}^{r}|B_{1}^{r})_{\Omega}\overset{\textnormal{\cite[cite]{[\@@bibref{}{marco_book}{}{}, Lem.~5.11]}}}{\leq}(n-r)H_{\alpha}(A|B)_{\sigma}+r\log d_{A}\,. (79)

Combining Equations 75, 78 and 79 yields

1nHα(A1n|B1n)ρ\displaystyle\frac{1}{n}H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\rho} nrnHα(A|B)σ+rnlogdA+11α1nlog|𝒯|\displaystyle\leq\frac{n-r}{n}H_{\alpha}(A|B)_{\sigma}+\frac{r}{n}\log d_{A}+\frac{1}{1-\alpha}\frac{1}{n}\log|\mathcal{T}| (80)
Hα(A|B)σ+2rnlogdA+11α(h(rn)+2rnlogdAB),\displaystyle{\leq}H_{\alpha}(A|B)_{\sigma}+\frac{2r}{n}\log d_{A}+\frac{1}{1-\alpha}\left(h\Big(\frac{r}{n}\Big)+\frac{2r}{n}\log d_{AB}\right)\,, (81)

where in the final step we used that dE=dABd_{E}=d_{AB}. ∎

We are now equipped with all the tools we need to prove the assertion of Proposition 4.2. This will be done in four steps, by proving two inequalities (direct and converse part) for both the smooth min- and max-entropy.

  1. (i)

    Direct part for smooth min-entropy: For α>1\alpha>1 consider the error term

    δε(n,r,α):=2rnlogdA+αα1(h(rn)+2rnlogdAB)+1n1α1log1ε2+1nlog11ε2.\displaystyle\delta_{\varepsilon}(n,r,\alpha):=\frac{2r}{n}\log d_{A}+\frac{\alpha}{\alpha-1}\left(h\Big(\frac{r}{n}\Big)+\frac{2r}{n}\log d_{AB}\right)+\frac{1}{n}\frac{1}{\alpha-1}\log\frac{1}{\varepsilon^{2}}+\frac{1}{n}\log\frac{1}{1-\varepsilon^{2}}\,. (82)

    We can write

    1nHminε(A1n|B1n)ρ\displaystyle\frac{1}{n}H_{\min}^{\varepsilon}(A_{1}^{n}|B_{1}^{n})_{\rho} 1nHα(A1n|B1n)ρ1n1α1log1ε21nlog11ε2\displaystyle{\geq}\frac{1}{n}H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\rho}-\frac{1}{n}\frac{1}{\alpha-1}\log\frac{1}{\varepsilon^{2}}-\frac{1}{n}\log\frac{1}{1-\varepsilon^{2}} (83)
    Hα(A|B)σδε(n,r,α)\displaystyle{\geq}H_{\alpha}(A|B)_{\sigma}-\delta_{\varepsilon}(n,r,\alpha) (84)
    H(A|B)σδε(n,r,α)4(α1)(logη)2,\displaystyle{\geq}H(A|B)_{\sigma}-\delta_{\varepsilon}(n,r,\alpha)-4(\alpha-1)(\log\eta)^{2}\,, (85)

    for a constant η=2Hmin(A|B)σ+2Hmax(A|B)σ+1\eta=\sqrt{2^{-H_{\min}(A|B)_{\sigma}}}+\sqrt{2^{H_{\max}(A|B)_{\sigma}}}+1. For a choice α=1+1/log(r/n)\alpha=1+1/\log(r/n) and recalling that r=o(n)r=o(n) we see that

    δε(n,r,α)=o(n)n,\displaystyle\delta_{\varepsilon}(n,r,\alpha)=\frac{o(n)}{n}\,, (86)

    where we used that limx0(log(x)+1)h(x)=0\lim_{x\to 0}(\log(x)+1)h(x)=0 and limx0(log(x)+1)x=0\lim_{x\to 0}(\log(x)+1)x=0 . Hence, we obtain

    1nHminε(A1n|B1n)ρH(A|B)σo(n)n.\displaystyle\frac{1}{n}H_{\min}^{\varepsilon}(A_{1}^{n}|B_{1}^{n})_{\rho}\geq H(A|B)_{\sigma}-\frac{o(n)}{n}\,. (87)
  2. (ii)

    Direct part for smooth max-entropy: For α[1/2,1)\alpha\in[1/2,1) consider the error term

    δε(n,r,α):=2rnlogdA+11α(h(rn)+2rnlogdAB)+1nα1αlog1ε.\displaystyle\delta^{\prime}_{\varepsilon}(n,r,\alpha):=\frac{2r}{n}\log d_{A}+\frac{1}{1-\alpha}\left(h\Big(\frac{r}{n}\Big)+\frac{2r}{n}\log d_{AB}\right)+\frac{1}{n}\frac{\alpha}{1-\alpha}\log\frac{1}{\varepsilon}\,. (88)

    We can write

    1nHmaxε(A1n|B1n)ρ\displaystyle\frac{1}{n}H_{\max}^{\varepsilon}(A_{1}^{n}|B_{1}^{n})_{\rho} 1nHα(A1n|B1n)ρ+1nα1αlog1ε\displaystyle{\leq}\frac{1}{n}H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\rho}+\frac{1}{n}\frac{\alpha}{1-\alpha}\log\frac{1}{\varepsilon} (89)
    Hα(A|B)σ+δε(n,r,α)\displaystyle{\leq}H_{\alpha}(A|B)_{\sigma}+\delta^{\prime}_{\varepsilon}(n,r,\alpha) (90)
    H(A|B)σ+δε(n,r,α)+4(1α)(logη)2.\displaystyle{\leq}H(A|B)_{\sigma}+\delta^{\prime}_{\varepsilon}(n,r,\alpha)+4(1-\alpha)(\log\eta)^{2}\,. (91)

    Similarly as above, choosing α=11/log(r/n)\alpha=1-1/\log(r/n) yields δε(n,r,α)=o(n)n\delta^{\prime}_{\varepsilon}(n,r,\alpha)=\frac{o(n)}{n} and hence

    1nHmaxε(A1n|B1n)ρH(A|B)σ+o(n)n.\displaystyle\frac{1}{n}H_{\max}^{\varepsilon}(A_{1}^{n}|B_{1}^{n})_{\rho}\leq H(A|B)_{\sigma}+\frac{o(n)}{n}\,. (92)
  3. (iii)

    Converse part for smooth min-entropy: For a fixed ε(0,1)\varepsilon\in(0,1) consider an arbitrary ε(0,1ε)\varepsilon^{\prime}\in(0,1-\varepsilon). Then

    1nHminε(A1n|B1n)ρ[36, Eq. 6.107]1nHmaxε(A1n|B1n)ρ+1nlog1(ε+ε)2Equation 92H(A|B)σ+o(n)n.\displaystyle\frac{1}{n}H_{\min}^{\varepsilon}(A_{1}^{n}|B_{1}^{n})_{\rho}\!\overset{\textnormal{\cite[cite]{[\@@bibref{}{marco_book}{}{}, Eq.~6.107]}}}{\leq}\!\frac{1}{n}H_{\max}^{\varepsilon^{\prime}}(A_{1}^{n}|B_{1}^{n})_{\rho}\!+\!\frac{1}{n}\log\frac{1}{1\!-\!(\varepsilon\!+\!\varepsilon^{\prime})^{2}}\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_direct_Hmax}}}}{\leq}H(A|B)_{\sigma}\!+\!\frac{o(n)}{n}\,. (93)
  4. (iv)

    Converse part for smooth max-entropy: For a fixed ε(0,1)\varepsilon\in(0,1) consider an arbitrary ε(0,1ε)\varepsilon^{\prime}\in(0,1-\varepsilon). Then

    1nHmaxε(A1n|B1n)ρ[36, Eq. 6.107]1nHminε(A1n|B1n)ρ1nlog1(ε+ε)2Equation 87H(A|B)σo(n)n.\displaystyle\frac{1}{n}H_{\max}^{\varepsilon^{\prime}}(A_{1}^{n}|B_{1}^{n})_{\rho}\!\overset{\textnormal{\cite[cite]{[\@@bibref{}{marco_book}{}{}, Eq.~6.107]}}}{\geq}\!\frac{1}{n}H_{\min}^{\varepsilon}(A_{1}^{n}|B_{1}^{n})_{\rho}\!-\!\frac{1}{n}\log\frac{1}{1\!-\!(\varepsilon\!+\!\varepsilon^{\prime})^{2}}\!\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_direct_Hmin}}}}{\geq}\!H(A|B)_{\sigma}\!-\!\frac{o(n)}{n}\,. (94)

Combining Equations 87, 92, 93 and 94 completes the proof.∎

4.4 Proof of Theorem 4.1

The monotonicity of the Rényi divergence in α\alpha [31] implies that for α>1\alpha>1 we have

1nH(A1n|B1n)ρ\displaystyle\frac{1}{n}H(A_{1}^{n}|B_{1}^{n})_{\rho} 1nHα(A1n|B1n)ρ\displaystyle\geq\frac{1}{n}H_{\alpha}(A_{1}^{n}|B_{1}^{n})_{\rho} (95)
Hα(A|B)σ2rnlogdAαα1(h(rn)+2rnlogdAB)\displaystyle{\geq}H_{\alpha}(A|B)_{\sigma}-\frac{2r}{n}\log d_{A}-\frac{\alpha}{\alpha-1}\left(h\Big(\frac{r}{n}\Big)+\frac{2r}{n}\log d_{AB}\right) (96)
H(A|B)σ2rnlogdAαα1(h(rn)+2rnlogdAB)4(α1)(logη)2,\displaystyle{\geq}H(A|B)_{\sigma}\!-\!\frac{2r}{n}\log d_{A}\!-\!\frac{\alpha}{\alpha-1}\left(h\Big(\frac{r}{n}\Big)\!+\!\frac{2r}{n}\log d_{AB}\right)\!-\!4(\alpha\!-\!1)(\log\eta)^{2}\,, (97)

for a constant η=2Hmin(A|B)σ+2Hmax(A|B)σ+1\eta=\sqrt{2^{-H_{\min}(A|B)_{\sigma}}}+\sqrt{2^{H_{\max}(A|B)_{\sigma}}}+1. Choosing α=1+1/log(r/n)\alpha=1+1/\log(r/n) and recalling that r=o(n)r=o(n) yields55 5 Note that limx0(log(x)+1)h(x)=0\lim_{x\to 0}(\log(x)+1)h(x)=0 and limx0(log(x)+1)x=0\lim_{x\to 0}(\log(x)+1)x=0.

1nH(A1n|B1n)ρH(A|B)σo(n)n.\displaystyle\frac{1}{n}H(A_{1}^{n}|B_{1}^{n})_{\rho}\geq H(A|B)_{\sigma}-\frac{o(n)}{n}\,. (98)

To see the other direction, note that for εn=2rn\varepsilon_{n}=2\sqrt{\frac{r}{n}} with r=o(n)r=o(n)

1nH(A1n|B1n)ρ\displaystyle\frac{1}{n}H(A_{1}^{n}|B_{1}^{n})_{\rho} =1ni=1nH(Ai|A1i1B1n)ρ\displaystyle{=}\frac{1}{n}\sum_{i=1}^{n}H(A_{i}|A_{1}^{i-1}B_{1}^{n})_{\rho} (99)
1ni=1nH(Ai|Bi)ρ\displaystyle{\leq}\frac{1}{n}\sum_{i=1}^{n}H(A_{i}|B_{i})_{\rho} (100)
=H(A1|B1)ρ\displaystyle{=}H(A_{1}|B_{1})_{\rho} (101)
H(A|B)σ+2εnlogdA+(1+εn)h(εn1+εn)\displaystyle{\leq}H(A|B)_{\sigma}+2\varepsilon_{n}\log d_{A}+(1+\varepsilon_{n})h\Big(\frac{\varepsilon_{n}}{1+\varepsilon_{n}}\Big) (102)
=H(A|B)σ+o(n)n,\displaystyle=H(A|B)_{\sigma}+\frac{o(n)}{n}\,, (103)

where the penultimate step uses the continuity of entropy [42]. Combining Equations 98 and 103 completes the proof.

5 Robustness of information measures for almost-iid states

In this work, we justified the importance of almost-iid states. This prompts the question if almost-iid states are as effective as perfect iid states for information-processing tasks. To answer this, it is crucial to understand if certain functionals (that characterize specific information-processing tasks) behave equally or differently for almost-iid and perfect iid states.

In Section 4 we have seen that the conditional entropy is robust for almost-iid states in the sense that it asymptotically coincides with the entropy of iid states. This implies that also the mutual information is robust for almost-iid states. To make this precise, recall that for a bipartite density matrix σAB\sigma_{AB} the mutual information is defined as

I(A:B)σ:=H(A)σH(A|B)σ.\displaystyle I(A:B)_{\sigma}:=H(A)_{\sigma}-H(A|B)_{\sigma}\,. (104)

The mutual information is a popular correlation measure in the sense that it satisfies (i) I(A:B)σ0I(A:B)_{\sigma}\geq 0, (ii) I(A:B)σ=0I(A:B)_{\sigma}=0 iff σAB=σAσB\sigma_{AB}=\sigma_{A}\otimes\sigma_{B} and (iii) I(A:BC)σI(A:B)I(A:BC)_{\sigma}\geq I(A:B).66 6 Properties (i) and (ii) follow by noting that I(A:B)ρ=D(ρABρAρB)I(A:B)_{\rho}=D(\rho_{AB}\|\rho_{A}\otimes\rho_{B}). The third property follows from strong subadditivity together with the chain rule as I(A:BC)σ=I(A:B)σ+I(A:C|B)σI(A:B)σI(A:BC)_{\sigma}=I(A:B)_{\sigma}+I(A:C|B)_{\sigma}\geq I(A:B)_{\sigma}. Let σABS(AB)\sigma_{AB}\in\mathrm{S}(\mathcal{H}_{AB}) and ρA1nB1nSn(AB,σABnr)\rho_{A_{1}^{n}B_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{AB},\sigma_{AB}^{\otimes n-r}) for r=o(n)r=o(n). Then

1nI(A1n:B1n)ρ\displaystyle\frac{1}{n}I(A_{1}^{n}:B_{1}^{n})_{\rho} =1nH(A1n)ρ1nH(A1n|B1n)ρ\displaystyle{=}\frac{1}{n}H(A_{1}^{n})_{\rho}-\frac{1}{n}H(A_{1}^{n}|B_{1}^{n})_{\rho} (105)
=H(A)σH(A|B)σ+o(n)n\displaystyle{=}H(A)_{\sigma}-H(A|B)_{\sigma}+\frac{o(n)}{n} (106)
=I(A:B)σ+o(n)n.\displaystyle{=}I(A:B)_{\sigma}+\frac{o(n)}{n}\,. (107)

It is natural to ask if popular entanglement measures are also robust for almost-iid states. In abstract terms, let E()E(\cdot) be an arbitrary entanglement measure. Let σABS(AB)\sigma_{AB}\in\mathrm{S}(\mathcal{H}_{AB}) and ρA1nB1nSn(AB,σABnr)\rho_{A_{1}^{n}B_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{AB},\sigma_{AB}^{\otimes n-r}) for r=o(n)r=o(n). Is it true that

1nE(A1n:B1n)ρ=?1nE(A1n:B1n)σn+o(n)n\displaystyle\frac{1}{n}E(A_{1}^{n}:B_{1}^{n})_{\rho}\overset{?}{=}\frac{1}{n}E(A_{1}^{n}:B_{1}^{n})_{\sigma^{\otimes n}}+\frac{o(n)}{n} (108)

In the following, we discuss the robustness of (a) squashed entanglement, (b) entanglement distillation, (c) entanglement cost, and (d) relative entropy of entanglement.

5.1 Robustness of squashed entanglement

Above we have seen that the mutual information is robust under almost-iid states. The same argument can be extended to see that the conditional mutual information also coincides for almost-iid and iid states. The squashed entanglement [14] is an entanglement measure that is based on the conditional mutual information. Given a biparitite density matrix ρABS(AB)\rho_{AB}\in\mathrm{S}(\mathcal{H}_{AB}), the squashed entanglement is defined as

Esq(A:B)ρ:=12infρABES(ABE){I(A:B|E)ρ:trE[ρABE]=ρAB},\displaystyle E_{sq}(A:B)_{\rho}:=\frac{1}{2}\inf_{\rho_{ABE}\in\mathrm{S}(\mathcal{H}_{ABE})}\big\{I(A:B|E)_{\rho}:\mathrm{tr}_{E}[\rho_{ABE}]=\rho_{AB}\big\}\,, (109)

where there is no bound on the dimension of EE. It features many desirable properties such as being additive on tensor products and superadditive in general. We next show that the squashed entanglement for almost-iid and iid states coincide.

Corollary 5.1.

Let σABS(AB)\sigma_{AB}\in\mathrm{S}(\mathcal{H}_{AB}) and ρA1nB1nSn(AB,σABnr)\rho_{A_{1}^{n}B_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{AB},\sigma_{AB}^{\otimes n-r}) for r=o(n)r=o(n). Then

1nEsq(A1n:B1n)ρ=Esq(A:B)σ+o(n)n.\displaystyle\frac{1}{n}E_{sq}(A_{1}^{n}:B_{1}^{n})_{\rho}=E_{sq}(A:B)_{\sigma}+\frac{o(n)}{n}\,. (110)
Proof.

Let dA:=dim(A)d_{A}:=\dim(\mathcal{H}_{A}) and dB:=dim(B)d_{B}:=\dim(\mathcal{H}_{B}). We can employ the permutation invariance of ρ\rho and the superadditivity of the squashed entanglement [14, Proposition 4] to write for εn:=4rn\varepsilon_{n}:=4\sqrt{\frac{r}{n}}

1nEsq(A1n:B1n)ρ\displaystyle\frac{1}{n}E_{sq}(A_{1}^{n}:B_{1}^{n})_{\rho} 1ni=1nEsq(Ai:Bi)ρ\displaystyle{\geq}\frac{1}{n}\sum_{i=1}^{n}E_{sq}(A_{i}:B_{i})_{\rho} (111)
=Esq(A:B)ρ\displaystyle{=}E_{sq}(A:B)_{\rho} (112)
Esq(A:B)σ12εnlog(dAdB)6h(εn)\displaystyle{\geq}E_{sq}(A:B)_{\sigma}-12\varepsilon_{n}\log(d_{A}d_{B})-6h(\varepsilon_{n}) (113)
=Esq(A:B)σ+o(n)n,\displaystyle=E_{sq}(A:B)_{\sigma}+\frac{o(n)}{n}\,, (114)

where the continuity of squashed entanglement follows from the continuity of the conditional entropy [1] as explained in [14, Section IV].

It thus remains to prove the other direction. For any ξ>0\xi>0 there exists an extension σABE\sigma_{ABE} or σAB\sigma_{AB} with dE:=dim(E)<d_{E}:=\dim(\mathcal{H}_{E})<\infty such that

|Esq(A:B)σ12I(A:B|E)σ|ξ.\displaystyle\Big|E_{sq}(A:B)_{\sigma}-\frac{1}{2}I(A:B|E)_{\sigma}\Big|\leq\xi\,. (115)

To see this, note that by the definition of the squashed entanglement there exists an extension σABE\sigma^{\prime}_{ABE} of σAB\sigma_{AB} (with possibly unbounded EE-system) such that

|Esq(A:B)σ12I(A:B|E)σ|ξ2.\displaystyle\Big|E_{sq}(A:B)_{\sigma}-\frac{1}{2}I(A:B|E)_{\sigma^{\prime}}\Big|\leq\frac{\xi}{2}\,\,. (116)

Choose a finite-dimensional projector ΠE\Pi_{E} on the EE-system such that tr[ΠEσABE]=1ε\mathrm{tr}[\Pi_{E}\sigma^{\prime}_{ABE}]=1-\varepsilon for some ε>0\varepsilon>0. Let

σABE:=ΠEσABEΠE+(1tr[ΠEσABE])|ee|ABE,\displaystyle\sigma_{ABE}:=\Pi_{E}\sigma^{\prime}_{ABE}\Pi_{E}^{\dagger}+\big(1-\mathrm{tr}[\Pi_{E}\sigma^{\prime}_{ABE}]\big)|e\rangle\!\langle e|_{ABE}\,, (117)

where |e|e\rangle is a state orthogonal to the support of ΠE\Pi_{E}. By the continuity of the conditional entropy [1] we can choose ε>0\varepsilon>0 such that

|I(A:B|E)σI(A:B|E)σ|ξ,\displaystyle|I(A:B|E)_{\sigma^{\prime}}-I(A:B|E)_{\sigma}|\leq\xi\,, (118)

where we used that the continuity of the conditional entropy does not depend on the dimension of the conditioning system. The triangle inequality implies

|Esq(A:B)σ12I(A:B|E)σ|\displaystyle\Big|E_{sq}(A:B)_{\sigma}-\frac{1}{2}I(A:B|E)_{\sigma}\Big| |Esq(A:B)σ12I(A:B|E)σ|+|12I(A:B|E)σ12I(A:B|E)σ|\displaystyle\leq\Big|E_{sq}(A:B)_{\sigma}-\frac{1}{2}I(A:B|E)_{\sigma^{\prime}}\Big|+\Big|\frac{1}{2}I(A:B|E)_{\sigma^{\prime}}-\frac{1}{2}I(A:B|E)_{\sigma}\Big|
ξ,\displaystyle{\leq}\xi\,, (119)

which thus justifies Equation 115.

Due to Lemma B.5, there exists an extension ρA1nB1nE1n\rho_{A_{1}^{n}B_{1}^{n}E_{1}^{n}} of ρA1nB1n\rho_{A_{1}^{n}B_{1}^{n}} which is an (nr)\binom{n}{r}-almost-iid state in σABE\sigma_{ABE}. Hence,

1nEsq(A1n:B1n)ρ\displaystyle\frac{1}{n}E_{sq}(A_{1}^{n}:B_{1}^{n})_{\rho} 12nI(A1n:B1n|E1n)ρ\displaystyle\leq\frac{1}{2n}I(A_{1}^{n}:B_{1}^{n}|E_{1}^{n})_{\rho} (120)
=12nH(A1n|E1n)ρ12nH(A1n|B1nE1n)ρ\displaystyle=\frac{1}{2n}H(A_{1}^{n}|E_{1}^{n})_{\rho}-\frac{1}{2n}H(A_{1}^{n}|B_{1}^{n}E_{1}^{n})_{\rho} (121)
=12H(A|E)σ12H(A|BE)σ+o(n)n\displaystyle{=}\frac{1}{2}H(A|E)_{\sigma}-\frac{1}{2}H(A|BE)_{\sigma}+\frac{o(n)}{n} (122)
=12I(A:B|E)σ+o(n)n\displaystyle=\frac{1}{2}I(A:B|E)_{\sigma}+\frac{o(n)}{n} (123)
Esq(A:B)σ+ξ+o(n)n.\displaystyle{\leq}E_{sq}(A:B)_{\sigma}+\xi+\frac{o(n)}{n}\,. (124)

Since this holds for any ξ>0\xi>0 we can consider ξ0\xi\to 0, which concludes the proof. ∎

5.2 Robustness of entanglement distillation and entanglement cost

Let |ΦAB|\Phi\rangle_{AB} denote an entangled Bell state. Given a bipartite density matrix ρABS(AB)\rho_{AB}\in\mathrm{S}(\mathcal{H}_{AB}), recall the definitions of entanglement distillation [3, 4, 5]

ED(A:B)ρ:=limε0limnsup{mn:inf𝒫nLOCC12𝒫n(ρn)|ΦΦ|m1ε}\displaystyle E_{D}(A:B)_{\rho}:=\lim_{\varepsilon\to 0}\lim_{n\to\infty}\sup\Big\{\frac{m}{n}:\inf_{\mathcal{P}_{n}\in\mathrm{LOCC}}\frac{1}{2}\left\lVert\mathcal{P}_{n}(\rho^{\otimes n})-|\Phi\rangle\!\langle\Phi|^{\otimes m}\right\rVert_{1}\leq\varepsilon\Big\} (125)

and entanglement cost [24]

EC(A:B)ρ:=limε0limninf{mn:inf𝒫nLOCC12𝒫n(|ΦΦ|m)ρn1ε}.\displaystyle E_{C}(A:B)_{\rho}:=\lim_{\varepsilon\to 0}\lim_{n\to\infty}\inf\Big\{\frac{m}{n}:\inf_{\mathcal{P}_{n}\in\mathrm{LOCC}}\frac{1}{2}\left\lVert\mathcal{P}_{n}(|\Phi\rangle\!\langle\Phi|^{\otimes m})-\rho^{\otimes n}\right\rVert_{1}\leq\varepsilon\Big\}\,. (126)
Question 5.2.

Let σABS(AB)\sigma_{AB}\in\mathrm{S}(\mathcal{H}_{AB}) and ρA1nB1nSn(AB,σABnr)\rho_{A_{1}^{n}B_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{AB},\sigma_{AB}^{\otimes n-r}) for r=o(n)r=o(n). Is it true that

1nED(A1n:B1n)ρ=?ED(A:B)σ+o(n)n\displaystyle\frac{1}{n}E_{D}(A_{1}^{n}:B_{1}^{n})_{\rho}\overset{?}{=}E_{D}(A:B)_{\sigma}+\frac{o(n)}{n} (127)

As discussed in [29], there are strong indications that one direction of Equation 127 holds, namely that

1nED(A1n:B1n)ρED(A:B)σ+o(n)n.\displaystyle\frac{1}{n}E_{D}(A_{1}^{n}:B_{1}^{n})_{\rho}\geq E_{D}(A:B)_{\sigma}+\frac{o(n)}{n}\,. (128)

Whether the other direction holds also remains an open question.

The equivalent question for the entanglement cost asks:

Question 5.3.

Let σABS(AB)\sigma_{AB}\in\mathrm{S}(\mathcal{H}_{AB}) and ρA1nB1nSn(AB,σABnr)\rho_{A_{1}^{n}B_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{AB},\sigma_{AB}^{\otimes n-r}) for r=o(n)r=o(n). Is it true that

1nEC(A1n:B1n)ρ=?EC(A:B)σ+o(n)n\displaystyle\frac{1}{n}E_{C}(A_{1}^{n}:B_{1}^{n})_{\rho}\overset{?}{=}E_{C}(A:B)_{\sigma}+\frac{o(n)}{n} (129)

However, none of the two directions of Equation 129 are known to hold.

At this point, we emphasize that, unlike squashed entanglement, already the definitions of both entanglement cost and entanglement distillation rely on a tensor power (iid) structure. This may suggest that, rather than asking about the robustness of Equations 125 and 126, one should incorporate robustness directly into the definition itself. One may therefore wonder how these notions would change if an almost-iid structure were built into the definition from the outset. In [29], the authors investigate this question by introducing new asymptotic state transformation rates that avoid the standard iid assumption. We refer the interested reader to that paper for further details.

5.3 Robustness of relative entropy of entanglement

A popular measure to quantify the amount of entanglement is the relative entropy of entanglement [39] defined as

ER(A:B)ρ:=minσABSEP(A:B)D(ρABσAB),\displaystyle E_{R}(A:B)_{\rho}:=\min_{\sigma_{AB}\in\mathrm{SEP}(A:B)}D(\rho_{AB}\|\sigma_{AB})\,, (130)

where SEP(A:B):=conv{|ϕϕ|A|φφ|B:|ϕAA,|φBB,ϕ|ϕ=φ|φ=1}\mathrm{SEP}(A:B):=\mathrm{conv}\{|\phi\rangle\!\langle\phi|_{A}\otimes|\varphi\rangle\!\langle\varphi|_{B}:|\phi\rangle_{A}\in\mathcal{H}_{A},|\varphi\rangle_{B}\in\mathcal{H}_{B},\langle\phi|\phi\rangle=\langle\varphi|\varphi\rangle=1\}. It is known [40] that the relative entropy of entanglement is not additive under the tensor product, which justifies the definition of a regularized version ER(A:B)ρ:=limk1kER(A1k:B1k)ρkE^{\infty}_{R}(A:B)_{\rho}:=\lim_{k\to\infty}\frac{1}{k}E_{R}(A_{1}^{k}:B_{1}^{k})_{\rho^{\otimes k}}. The limit in the regularization exists due to Fekete’s subadditivity lemma.

Question 5.4.

Let σABS(AB)\sigma_{AB}\in\mathrm{S}(\mathcal{H}_{AB}) and ρA1nB1nSn(AB,σABnr)\rho_{A_{1}^{n}B_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{AB},\sigma_{AB}^{\otimes n-r}) for r=o(n)r=o(n). Is it true that

1nER(A1n:B1n)ρ=?1nER(A1n:B1n)σn+o(n)n\displaystyle\frac{1}{n}E_{R}(A_{1}^{n}:B_{1}^{n})_{\rho}\overset{?}{=}\frac{1}{n}E_{R}(A_{1}^{n}:B_{1}^{n})_{\sigma^{\otimes n}}+\frac{o(n)}{n} (131)

If Equation 131 were true, this would save the original proof of the generalized quantum Stein’s lemma [23, 26] by Brandão and Plenio [9] (see also [6, 7]). One direction of Equation 131 follows from the results developed in this paper. To see this, choose sns_{n} a monotonically increasing sequence of integers such that snr=o(n)s_{n}r=o(n) and limnsn=\lim_{n\to\infty}s_{n}=\infty. Let kn=n/snk_{n}=\lceil n/s_{n}\rceil. Note that nsnknn+snn\leq s_{n}k_{n}\leq n+s_{n}. Hence, using the monotonicity of the entanglement of formation under partial trace,

1nER(A1n:B1n)ρ\displaystyle\frac{1}{n}E_{R}(A_{1}^{n}:B_{1}^{n})_{\rho} 1nER(A1snkn:B1snkn)ρ\displaystyle\leq\frac{1}{n}E_{R}\big(A_{1}^{s_{n}k_{n}}:B_{1}^{s_{n}k_{n}}\big)_{\rho} (132)
1snknER(A1snkn:B1snkn)ρ+snnlogdAdB,\displaystyle\leq\frac{1}{s_{n}k_{n}}E_{R}\big(A_{1}^{s_{n}k_{n}}:B_{1}^{s_{n}k_{n}}\big)_{\rho}+\frac{s_{n}}{n}\log d_{A}d_{B}\,, (133)

where the final step uses ER(Am:Bm)ρmlog(dAdB)E_{R}(A^{m}:B^{m})_{\rho}\leq m\log(d_{A}d_{B}) and sn=o(n)s_{n}=o(n). In the following steps, we omit the subscripts nn for better readability. Let ωsargminτA1sB1sSEPD(ρA1sB1sτA1sB1s)\omega_{s}\in\arg\min_{\tau_{A_{1}^{s}B_{1}^{s}}\in\mathrm{SEP}}D(\rho_{A_{1}^{s}B_{1}^{s}}\|\tau_{A_{1}^{s}B_{1}^{s}}). Then for ρs=trns[ρn]\rho_{s}=\mathrm{tr}_{n-s}[\rho_{n}] we have

1nER(A1n:B1n)ρ\displaystyle\frac{1}{n}E_{R}(A_{1}^{n}:B_{1}^{n})_{\rho} 1skER(A1sk:B1sk)ρ+o(n)n\displaystyle{\leq}\frac{1}{sk}E_{R}\big(A_{1}^{sk}:B_{1}^{sk}\big)_{\rho}+\frac{o(n)}{n} (134)
=1skminτA1skB1skSEPD(ρA1skB1skτA1skB1sk)+o(n)n\displaystyle=\frac{1}{sk}\min_{\tau_{A_{1}^{sk}B_{1}^{sk}}\in\mathrm{SEP}}D(\rho_{A_{1}^{sk}B_{1}^{sk}}\|\tau_{A_{1}^{sk}B_{1}^{sk}})+\frac{o(n)}{n} (135)
1skD(ρsk(ωs)k)+o(n)n\displaystyle\leq\frac{1}{sk}D\big(\rho_{sk}\|(\omega_{s})^{\otimes k}\big)+\frac{o(n)}{n} (136)
=1skH(ρsk)1str[ρslogωs]+o(n)n\displaystyle=-\frac{1}{sk}H(\rho_{sk})-\frac{1}{s}\mathrm{tr}[\rho_{s}\log\omega_{s}]+\frac{o(n)}{n} (137)
=1sH(σs)1str[ρslogωs]+o(n)n\displaystyle{=}-\frac{1}{s}H(\sigma^{\otimes s})-\frac{1}{s}\mathrm{tr}[\rho_{s}\log\omega_{s}]+\frac{o(n)}{n} (138)
1sH(ρs)1str[ρslogωs]+o(n)n\displaystyle{\leq}\frac{1}{s}H(\rho_{s})-\frac{1}{s}\mathrm{tr}[\rho_{s}\log\omega_{s}]+\frac{o(n)}{n} (139)
=1sER(A1s:B1s)ρ+o(n)n\displaystyle=\frac{1}{s}E_{R}(A_{1}^{s}:B_{1}^{s})_{\rho}+\frac{o(n)}{n} (140)
1sER(A1s:B1s)σs+o(n)n\displaystyle{\leq}\frac{1}{s}E_{R}(A_{1}^{s}:B_{1}^{s})_{\sigma^{\otimes s}}+\frac{o(n)}{n} (141)
=1nER(A1n:B1n)σn+o(n)n,\displaystyle=\frac{1}{n}E_{R}(A_{1}^{n}:B_{1}^{n})_{\sigma^{\otimes n}}+\frac{o(n)}{n}\,, (142)

where the continuity of the von Neumann entropy and the relative entropy of entanglement can be found in [2, 32, 42]. In the last step, we use that limnsn=\lim_{n\to\infty}s_{n}=\infty and that the limit in the RHS of Equation 131 exists due to Fekete’s subadditivity lemma.

The other direction of Equation 131 appears more complicated and remains an open question.

Acknowledgements

We thank Fernando Brandão for his talk on almost-iid states at the SwissMAP Research Station (SRS) conference in Les Diablerets 2024, which motivated us to write this paper. We further thank Frédéric Dupuis and Ludovico Lami for insightful discussions on this topic at the same conference. GM and RR acknowledge support from the NCCR SwissMAP, the ETH Zurich Quantum Center, the SNSF project No. 20QU-1 225171, and the CHIST-ERA project MoDIC.

Appendix

Appendix A Justification of Remark 2.3

To justify the assertion of Remark 2.3, we need to show that for all purifications |ψA12E12|\psi\rangle_{A_{1}^{2}E_{1}^{2}} of ρA12\rho_{A_{1}^{2}} and for all purifications |θAE|\theta\rangle_{AE} of |0A|0\rangle_{A} it follows that |ψA12E12Sym2(AE)span𝒱(AE2,|θAE)|\psi\rangle_{A_{1}^{2}E_{1}^{2}}\not\in\mathrm{Sym}^{2}(\mathcal{H}_{AE})\cap\mathrm{span}\,\mathcal{V}(\mathcal{H}_{AE}^{\otimes 2},|\theta\rangle_{AE})

To see this note that, because |0A|0\rangle_{A} is pure, |θAE=|0A|ϑE|\theta\rangle_{AE}=|0\rangle_{A}\otimes|\vartheta\rangle_{E}. Similarly, because ρA12=|ΨΨ|\rho_{A_{1}^{2}}=|\Psi^{-}\rangle\!\langle\Psi^{-}| is pure, any purification on EE must be of the form |ΨA1A2|φE1E2=:|ψA12E12|\Psi^{-}\rangle_{A_{1}A_{2}}\otimes|\varphi\rangle_{E_{1}E_{2}}=:|\psi\rangle_{A_{1}^{2}E_{1}^{2}}. To ensure that ρA12\rho_{A_{1}^{2}} is an almost-iid state according to the definition from [9], we need |ψA12E12Sym2(AE)span𝒱(AE2,|θAE)|\psi\rangle_{A_{1}^{2}E_{1}^{2}}\in\mathrm{Sym}^{2}(\mathcal{H}_{AE})\cap\mathrm{span}\,\mathcal{V}(\mathcal{H}_{AE}^{\otimes 2},|\theta\rangle_{AE}). Thus, the swap operation π\pi must satisfy

π|ψA12E12=(π|ΨA1A2)(π|φE1E2)=|ΨA1A2(π|φE1E2)=!|ΨA1A2|φE1E2.\displaystyle\pi|\psi\rangle_{A_{1}^{2}E_{1}^{2}}=(\pi|\Psi^{-}\rangle_{A_{1}A_{2}})\otimes(\pi|\varphi\rangle_{E_{1}E_{2}})=-|\Psi^{-}\rangle_{A_{1}A_{2}}\otimes(\pi|\varphi\rangle_{E_{1}E_{2}})\overset{!}{=}|\Psi^{-}\rangle_{A_{1}A_{2}}\otimes|\varphi\rangle_{E_{1}E_{2}}\,. (143)

This yields π|φE1E2=|φE1E2\pi|\varphi\rangle_{E_{1}E_{2}}=-|\varphi\rangle_{E_{1}E_{2}}, i.e., |φ|\varphi\rangle is anti-symmetric. Expressing this in an orthonormal basis {|iE}i\{|i\rangle_{E}\}_{i} of EE with |0E=|ϑE|0\rangle_{E}=|\vartheta\rangle_{E} gives |φE1E2=i,j=0dE1αi,j|iE1|jE2|\varphi\rangle_{E_{1}E_{2}}=\sum_{i,j=0}^{d_{E}-1}\alpha_{i,j}|i\rangle_{E_{1}}|j\rangle_{E_{2}} with αi,j=αj,i\alpha_{i,j}=-\alpha_{j,i} for all i,j{0,,dE1}i,j\in\{0,\ldots,d_{E}-1\}. Hence,

|ψA12E12\displaystyle|\psi\rangle_{A_{1}^{2}E_{1}^{2}} =12i<jαi,j(|0A1|iE1|1A2|jE2|0A1|jE1|1A2|iE2\displaystyle=\frac{1}{\sqrt{2}}\sum_{i<j}\alpha_{i,j}\big(|0\rangle_{A_{1}}|i\rangle_{E_{1}}|1\rangle_{A_{2}}|j\rangle_{E_{2}}-|0\rangle_{A_{1}}|j\rangle_{E_{1}}|1\rangle_{A_{2}}|i\rangle_{E_{2}}
|1A1|iE1|0A2|jE2+|1A1|jE1|0A2|iE2),\displaystyle\hskip 71.13188pt-|1\rangle_{A_{1}}|i\rangle_{E_{1}}|0\rangle_{A_{2}}|j\rangle_{E_{2}}+|1\rangle_{A_{1}}|j\rangle_{E_{1}}|0\rangle_{A_{2}}|i\rangle_{E_{2}}\big)\,, (144)

which cannot be inside span𝒱(AE2,|θAE)\mathrm{span}\,\mathcal{V}(\mathcal{H}_{AE}^{\otimes 2},|\theta\rangle_{AE}). Thus, the state is not an almost-iid state according to the definition from [9].

Appendix B Properties of almost-iid states

Lemma B.1.

Let ρA1nSn(A,σAnr)\rho_{A_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}) with purification |θAE|\theta\rangle_{AE} of σA\sigma_{A} according to Definition 2.1. Then, any other purification |θ~AR|\tilde{\theta}\rangle_{AR} of σA\sigma_{A} would also satisfy Definition 2.1.

Proof.

Without loss of generality, assume dim(E)=dim(R)\dim(E)=\dim(R). This can be done since the rank of the reduced state of any purification on the purifying system is bounded by the rank of σA\sigma_{A}. Since all purifications are then equal up to unitaries on the purifying system, we have |θ~AE=(idAUE)|θAE|\tilde{\theta}\rangle_{AE}=(\mathrm{id}_{A}\otimes U_{E})|\theta\rangle_{AE} for some unitary UEU_{E} on EE. For ρA1nE1n\rho_{A_{1}^{n}E_{1}^{n}} being the extension of ρA1n\rho_{A_{1}^{n}} according to Definition 2.1 we define

ρ~A1nE1n:=(idA1nUEn)ρA1nE1n(idA1n(UE)n).\displaystyle\tilde{\rho}_{A_{1}^{n}E_{1}^{n}}:=(\mathrm{id}_{A_{1}^{n}}\otimes U^{\otimes n}_{E})\rho_{A_{1}^{n}E_{1}^{n}}(\mathrm{id}_{A_{1}^{n}}\otimes(U_{E}^{\dagger})^{\otimes n})\,. (145)

Clearly ρ~A1nE1n\tilde{\rho}_{A_{1}^{n}E_{1}^{n}} is an extension of ρA1n\rho_{A_{1}^{n}}, i.e., ρ~A1nE1nS(AEn)\tilde{\rho}_{A_{1}^{n}E_{1}^{n}}\in\mathrm{S}(\mathcal{H}_{AE}^{\otimes n}) and trE1n[ρ~A1nE1n]=ρA1n\mathrm{tr}_{E_{1}^{n}}[\tilde{\rho}_{A_{1}^{n}E_{1}^{n}}]=\rho_{A_{1}^{n}}. Thus, it remains to show that this extension satisfies the two properties in Definition 2.1.

First, we note that ρ~A1nE1n\tilde{\rho}_{A_{1}^{n}E_{1}^{n}} is permutation-invariant. To see this, let π\pi be a permutation that swaps (Ai,Ei)(Aj,Ej)(A_{i},E_{i})\leftrightarrow(A_{j},E_{j}). Then

πρ~A1nE1nπ\displaystyle\pi\tilde{\rho}_{A_{1}^{n}E_{1}^{n}}\pi^{\dagger} =(idA1nUEn)πρA1nE1nπ(idA1n(UE)n)\displaystyle=(\mathrm{id}_{A_{1}^{n}}\otimes U^{\otimes n}_{E})\pi\rho_{A_{1}^{n}E_{1}^{n}}\pi^{\dagger}(\mathrm{id}_{A_{1}^{n}}\otimes(U_{E}^{\dagger})^{\otimes n}) (146)
=(idA1nUEn)ρA1nE1n(idA1n(UE)n)\displaystyle=(\mathrm{id}_{A_{1}^{n}}\otimes U^{\otimes n}_{E})\rho_{A_{1}^{n}E_{1}^{n}}(\mathrm{id}_{A_{1}^{n}}\otimes(U_{E}^{\dagger})^{\otimes n}) (147)
=ρ~A1nE1n.\displaystyle=\tilde{\rho}_{A_{1}^{n}E_{1}^{n}}\,. (148)

Second, we have

ρ~A1nE1n\displaystyle\tilde{\rho}_{A_{1}^{n}E_{1}^{n}} =(idA1nUEn)ρA1nE1n(idA1n(UE)n)\displaystyle=(\mathrm{id}_{A_{1}^{n}}\otimes U^{\otimes n}_{E})\rho_{A_{1}^{n}E_{1}^{n}}(\mathrm{id}_{A_{1}^{n}}\otimes(U_{E}^{\dagger})^{\otimes n}) (149)
=k,𝒯βk,(idA1nUEn)|Ψk|Ψ~kΨ|(idA1n(UE)n)Ψ~|.\displaystyle=\sum_{k,\ell\in\mathcal{T}}\beta_{k,\ell}\underbrace{(\mathrm{id}_{A_{1}^{n}}\otimes U^{\otimes n}_{E})|\Psi_{k}\rangle}_{|\tilde{\Psi}_{k}\rangle}\underbrace{\langle\Psi_{\ell}|(\mathrm{id}_{A_{1}^{n}}\otimes(U_{E}^{\dagger})^{\otimes n})}_{\langle\tilde{\Psi}_{\ell}|}\,. (150)

Note that |Ψ~k,|Ψ~𝒱(AEn,|θ~AE)|\tilde{\Psi}_{k}\rangle,|\tilde{\Psi}_{\ell}\rangle\in\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\tilde{\theta}\rangle_{AE}). ∎

Lemma B.2.

There exists an orthonormal basis {|Ψt}t𝒯\{|\Psi_{t}\rangle\}_{t\in\mathcal{T}} of span𝒱(AEn,|θnr)\mathrm{span}\,\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\theta\rangle^{\otimes n-r}) with vectors |Ψt𝒱(AEn,|θnr)|\Psi_{t}\rangle\in\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\theta\rangle^{\otimes n-r}) for all t𝒯t\in\mathcal{T}, and with

|𝒯|(nr)dAEr2nh(r/n)dAEr,\displaystyle|\mathcal{T}|\leq\binom{n}{r}\,d_{AE}^{\,r}\leq 2^{nh(r/n)}\,d_{AE}^{\,r}, (151)

where dAE=dim(AE)d_{AE}=\dim(\mathcal{H}_{AE}), and h(x):=xlogx(1x)log(1x)h(x):=-x\log x-(1-x)\log(1-x) for x[0,1]x\in[0,1] is the binary entropy function.

Proof.

Let {|x}x𝒳\{|x\rangle\}_{x\in\mathcal{X}} denote an orthonormal basis of AE\mathcal{H}_{AE} such that |θ=|x¯|\theta\rangle=|\bar{x}\rangle for some x𝒳x\in\mathcal{X}. Any vector |Ψspan𝒱(AEn,|θnr)|\Psi\rangle\in\mathrm{span}\,\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\theta\rangle^{\otimes n-r}) can be expanded as

|Ψ=jαj|Ψ~j,\displaystyle|\Psi\rangle=\sum_{j}\alpha_{j}|\tilde{\Psi}_{j}\rangle\,, (152)

for coefficients αj\alpha_{j}\in\mathbb{C} and vectors |Ψ~j𝒱(AEn,|θnr)|\tilde{\Psi}_{j}\rangle\in\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\theta\rangle^{\otimes n-r}). By definition, |Ψ~j=πj(|θnr|Ωj(r))|\tilde{\Psi}_{j}\rangle=\pi_{j}(|\theta\rangle^{\otimes n-r}\otimes|\Omega_{j}^{(r)}\rangle) for some permutation πj𝒮n\pi_{j}\in\mathcal{S}_{n} and vector |Ωj(r)AEr|\Omega_{j}^{(r)}\rangle\in\mathcal{H}_{AE}^{\otimes r}. Since {|x}x𝒳\{|x\rangle\}_{x\in\mathcal{X}} is a basis of AE\mathcal{H}_{AE}, we can write

|Ωj(r)=𝐱β𝐱(j)|𝐱\displaystyle|\Omega_{j}^{(r)}\rangle=\sum_{\bf{x}}\beta^{(j)}_{\bf{x}}|\bf{x}\rangle (153)

for all jj, where 𝐱=(x1,,xr){\bf{x}}=(x_{1},\dots,x_{r}) is an rr-tuple of elements from 𝒳\mathcal{X}, |𝐱=|x1|xr|{\bf{x}}\rangle=|x_{1}\rangle\otimes\cdots\otimes|x_{r}\rangle, and β𝐱(j)\beta^{(j)}_{\bf{x}}\in\mathbb{C}. Then

|Ψ=j,𝐱αjβ𝐱(j)πj(|θnr|𝐱)=:tγt|Ψt\displaystyle|\Psi\rangle=\sum_{j,\bf{x}}\alpha_{j}\beta^{(j)}_{\bf{x}}\,\pi_{j}(|\theta\rangle^{\otimes n-r}\otimes|{\bf{x}}\rangle)=:\sum_{t}\gamma_{t}|\Psi_{t}\rangle (154)

with t=(j,𝐱)t=(j,{\bf x}), γt=αjβ𝐱(j)\gamma_{t}=\alpha_{j}\beta^{(j)}_{\bf{x}}, and |Ψt=πj(|θnr|𝐱)|\Psi_{t}\rangle=\pi_{j}(|\theta\rangle^{\otimes n-r}\otimes|{\bf{x}}\rangle). Intuitively, jj is labeling the positions of the defects, and 𝐱{\bf x} labels the state at these defected positions. Note that the vectors |Ψt|\Psi_{t}\rangle are normalized for all tt, and, for different values of tt, they are either pairwise orthogonal, or they are equal (e.g. if |𝐱=|θr|{\bf{x}}\rangle=|\theta\rangle^{\otimes r}). Thus, restricting to a maximal set of pairwise orthogonal vectors {|Ψt}t𝒯\{|\Psi_{t}\rangle\}_{t\in\mathcal{T}}Equation 154 shows that this set forms an orthonormal basis of span𝒱(AEn,|θnr)\mathrm{span}\,\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\theta\rangle^{\otimes n-r}).

To determine the size |𝒯||\mathcal{T}|77 7 More precisely, one finds |𝒯|=k=0r(nk)(dAE1)k|\mathcal{T}|=\sum_{k=0}^{r}\binom{n}{k}(d_{AE}-1)^{k}. Since (nk)(nk)(nkrk)=(nr)(rk)\binom{n}{k}\leq\binom{n}{k}\binom{n-k}{r-k}=\binom{n}{r}\binom{r}{k} for all krk\leq r, it follows that |𝒯|(nr)k=0r(rk)(dAE1)k=(nr)dAEr|\mathcal{T}|\leq\binom{n}{r}\sum_{k=0}^{r}\binom{r}{k}(d_{AE}-1)^{k}=\binom{n}{r}d_{AE}^{r}., note that for any vector in 𝒱(AEn,|θnr)\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\theta\rangle^{\otimes n-r}) there are (nr)\binom{n}{r} possible combinations for the positions of the defects, and at each such position, the dimension of the Hilbert space is given by dAEd_{AE}. Hence, we get the upper bound stated in Equation 151, where the first bound would correspond to the case where the defects are different from |θ|\theta\rangle. The second inequality then follows from the relation (nr)2nh(r/n)\binom{n}{r}\leq 2^{nh(r/n)}, e.g. see [15, Example 11.1.3]. ∎

Lemma B.3.

ρA1n+mSn+m(A,σAn+mr)\rho_{A_{1}^{n+m}}\in\mathrm{S}^{n+m}(\mathcal{H}_{A},\sigma_{A}^{\otimes n+m-r}) implies ρA1nSn(A,σAnr)\rho_{A_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}) for any m,nm,n\in\mathbb{N}.

Proof.

Since ρA1n+mSn+m(A,σAn+mr)\rho_{A_{1}^{n+m}}\in\mathrm{S}^{n+m}(\mathcal{H}_{A},\sigma_{A}^{\otimes n+m-r}), there exists an extension ρA1n+mE1n+m\rho_{A_{1}^{n+m}E_{1}^{n+m}} that is permutation invariant. We first show that ρA1nE1n\rho_{A_{1}^{n}E_{1}^{n}} is then also permutation-invariant. To see this, let π𝒮n\pi\in\mathcal{S}_{n} denote an arbitrary permutation. Then

πρA1nE1nπ\displaystyle\pi\rho_{A_{1}^{n}E_{1}^{n}}\pi^{\dagger} =πtrAn+1n+mEn+1n+m[ρA1n+mE1n+m]π\displaystyle=\pi\mathrm{tr}_{A_{n+1}^{n+m}E_{n+1}^{n+m}}[\rho_{A_{1}^{n+m}E_{1}^{n+m}}]\pi^{\dagger} (155)
=trAn+1n+mEn+1n+m[(πidm)ρA1n+mE1n+m(πidm)]\displaystyle=\mathrm{tr}_{A_{n+1}^{n+m}E_{n+1}^{n+m}}[(\pi\otimes\mathrm{id}_{m})\rho_{A_{1}^{n+m}E_{1}^{n+m}}(\pi^{\dagger}\otimes\mathrm{id}_{m})] (156)
=trAn+1n+mEn+1n+m[ρA1n+mE1n+m]\displaystyle=\mathrm{tr}_{A_{n+1}^{n+m}E_{n+1}^{n+m}}[\rho_{A_{1}^{n+m}E_{1}^{n+m}}] (157)
=ρA1nE1n.\displaystyle=\rho_{A_{1}^{n}E_{1}^{n}}\,. (158)

Thus, it remains to show that ρA1nE1n\rho_{A_{1}^{n}E_{1}^{n}} satisfies Property (ii) in Definition 2.1. By definition ρA1n+mE1n+m\rho_{A_{1}^{n+m}E_{1}^{n+m}} is such that supp(ρA1n+mE1n+m)span𝒱(AEn+m,|θAEn+mr)\mathrm{supp}(\rho_{A_{1}^{n+m}E_{1}^{n+m}})\subseteq\mathrm{span}\,\mathcal{V}(\mathcal{H}^{\otimes n+m}_{AE},|\theta\rangle_{AE}^{\otimes n+m-r}) for a purification |θAE|\theta\rangle_{AE} of σA\sigma_{A}. We need to show that supp(ρA1nE1n)span𝒱(AEn,|θAEnr)\mathrm{supp}(\rho_{A_{1}^{n}E_{1}^{n}})\subseteq\mathrm{span}\,\mathcal{V}(\mathcal{H}^{\otimes n}_{AE},|\theta\rangle_{AE}^{\otimes n-r}). To see this, let {|ψt}t𝒯\{|\psi_{t}\rangle\}_{t\in\mathcal{T}} be the “standard” tensor product basis of 𝒱(AEn+m,|θAEn+mr)\mathcal{V}(\mathcal{H}^{\otimes n+m}_{AE},|\theta\rangle_{AE}^{\otimes n+m-r}), i.e., for a basis {|u}u\{|u\rangle\}_{u} of AE\mathcal{H}_{AE} we have |ψt=πt(|u1A1E1|urArEr|θAEn+mr)|\psi_{t}\rangle=\pi_{t}(|u_{1}\rangle_{A_{1}E_{1}}\otimes\ldots\otimes|u_{r}\rangle_{A_{r}E_{r}}\otimes|\theta\rangle_{AE}^{\otimes n+m-r}) for some permutation πt\pi_{t}. We can observe that

ρA1nE1n=trAn+1n+mEn+1n+m[ρA1n+mE1n+m]=k,𝒯βk,trAn+1n+mEn+1n+m[|ψkψ|]\displaystyle\rho_{A_{1}^{n}E_{1}^{n}}=\mathrm{tr}_{A_{n+1}^{n+m}E_{n+1}^{n+m}}[\rho_{A_{1}^{n+m}E_{1}^{n+m}}]=\sum_{k,\ell\in\mathcal{T}}\beta_{k,\ell}\mathrm{tr}_{A_{n+1}^{n+m}E_{n+1}^{n+m}}[|\psi_{k}\rangle\langle\psi_{\ell}|] (159)

and

trAn+1n+mEn+1n+m[|ψkψ|]=s,t𝒯¯cs,t(k,)|ψ¯sψ¯t|,\displaystyle\mathrm{tr}_{A_{n+1}^{n+m}E_{n+1}^{n+m}}[|\psi_{k}\rangle\langle\psi_{\ell}|]=\sum_{s,t\in\bar{\mathcal{T}}}c^{(k,\ell)}_{s,t}|\bar{\psi}_{s}\rangle\langle\bar{\psi}_{t}|\,, (160)

where {|ψ¯t}t𝒯¯\{|\bar{\psi}_{t}\rangle\}_{t\in\bar{\mathcal{T}}} denotes the standard basis of 𝒱(AEn,|θAEnr)\mathcal{V}(\mathcal{H}^{\otimes n}_{AE},|\theta\rangle_{AE}^{\otimes n-r}). Note that these basis elements in Equation 160 are independent of kk and \ell. Hence, we find

ρA1nE1n\displaystyle\rho_{A_{1}^{n}E_{1}^{n}} =k,𝒯βk,trAn+1n+mEn+1n+m[|ψkψ|]\displaystyle{=}\sum_{k,\ell\in\mathcal{T}}\beta_{k,\ell}\mathrm{tr}_{A_{n+1}^{n+m}E_{n+1}^{n+m}}[|\psi_{k}\rangle\langle\psi_{\ell}|] (161)
=k,𝒯,s,t𝒯¯βk,cs,t(k,)|ψ¯sψ¯t|\displaystyle{=}\sum_{k,\ell\in\mathcal{T},s,t\in\bar{\mathcal{T}}}\beta_{k,\ell}c^{(k,\ell)}_{s,t}|\bar{\psi}_{s}\rangle\langle\bar{\psi}_{t}| (162)
=s,t𝒯¯γs,t|ψ¯sψ¯t|,\displaystyle=\sum_{s,t\in\bar{\mathcal{T}}}\gamma_{s,t}|\bar{\psi}_{s}\rangle\langle\bar{\psi}_{t}|\,, (163)

where the final step uses γs,t:=k,𝒯βk,cs,t(k,)\gamma_{s,t}:=\sum_{k,\ell\in\mathcal{T}}\beta_{k,\ell}c^{(k,\ell)}_{s,t}. This shows that supp(ρA1nE1n)span𝒱(AEn,|θAEnr)\mathrm{supp}(\rho_{A_{1}^{n}E_{1}^{n}})\subseteq\mathrm{span}\,\mathcal{V}(\mathcal{H}^{\otimes n}_{AE},|\theta\rangle_{AE}^{\otimes n-r}) and hence completes the proof. ∎

Lemma B.4.

Let ρA1nS¯n(A,σAnr)\rho_{A_{1}^{n}}\in\bar{\mathrm{S}}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}). Then for any ss\in\mathbb{N} we have (ρA1n)sS¯ns(A,σAnsrs)(\rho_{A_{1}^{n}})^{\otimes s}\in\bar{\mathrm{S}}^{ns}(\mathcal{H}_{A},\sigma_{A}^{\otimes ns-rs}).

Proof.

Let |θAE|\theta\rangle_{AE} denote a purification of σA\sigma_{A}. Since ρA1nS¯n(A,σAnr)\rho_{A_{1}^{n}}\in\bar{\mathrm{S}}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}) there exists an extension ρA1nE1n\rho_{A_{1}^{n}E_{1}^{n}} that can be written as

ρA1nE1n=i,jβi,j|ΨiΨj|,\displaystyle\rho_{A_{1}^{n}E_{1}^{n}}=\sum_{i,j}\beta_{i,j}|\Psi_{i}\rangle\langle\Psi_{j}|\,, (164)

for |Ψt𝒱(AEn,|θnr)|\Psi_{t}\rangle\in\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\theta\rangle^{\otimes n-r}). Furthermore, by Equation 164 we have

(ρA1nE1n)s=i1,j1,,is,jsβi1,j1βis,js|Ψi1Ψj1||ΨisΨjs|,\displaystyle(\rho_{A_{1}^{n}E_{1}^{n}})^{\otimes s}=\sum_{i_{1},j_{1},\ldots,i_{s},j_{s}}\beta_{i_{1},j_{1}}\ldots\beta_{i_{s},j_{s}}|\Psi_{i_{1}}\rangle\langle\Psi_{j_{1}}|\otimes\ldots\otimes|\Psi_{i_{s}}\rangle\langle\Psi_{j_{s}}|\,, (165)

where |Ψi1|Ψis𝒱(AEns,|θnsrs)|\Psi_{i_{1}}\rangle\otimes\ldots\otimes|\Psi_{i_{s}}\rangle\in\mathcal{V}(\mathcal{H}_{AE}^{\otimes ns},|\theta\rangle^{\otimes ns-rs}). This shows that supp(ρA1nE1ns)𝒱(AEns,|θnsrs)\mathrm{supp}(\rho_{A_{1}^{n}E_{1}^{n}}^{\otimes s})\subseteq\mathcal{V}(\mathcal{H}_{AE}^{\otimes ns},|\theta\rangle^{\otimes ns-rs}), which completes the proof. ∎

Lemma B.5.

Let nn\in\mathbb{N}, rnr\leq n, σABES(ABE)\sigma_{ABE}\in\mathrm{S}(A\otimes B\otimes E), and ρA1nB1nSn(AB,σABnr)\rho_{A_{1}^{n}B_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{AB},\sigma_{AB}^{\otimes n-r}). Then, there exists an extension ρA1nB1nE1n\rho_{A_{1}^{n}B_{1}^{n}E_{1}^{n}} of ρA1nB1n\rho_{A_{1}^{n}B_{1}^{n}} that is (nr)\binom{n}{r}-almost-iid in σABE\sigma_{ABE}, i.e. ρA1nB1nE1nSn(ABE,σABEnr)\rho_{A_{1}^{n}B_{1}^{n}E_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{ABE},\sigma_{ABE}^{\otimes n-r}).

Proof.

This is essentially a consequence of Lemma B.1. Formally, there exist a purification |θABG|\theta\rangle_{ABG} of σAB\sigma_{AB} and an extension ρA1nB1nG1n\rho_{A_{1}^{n}B_{1}^{n}G_{1}^{n}} of ρA1nB1n\rho_{A_{1}^{n}B_{1}^{n}} such that supp(ρA1nB1nG1n)span𝒱(ABGn,|θABGnr)\mathrm{supp}(\rho_{A_{1}^{n}B_{1}^{n}G_{1}^{n}})\subseteq\mathrm{span}\,\mathcal{V}(\mathcal{H}_{ABG}^{\otimes n},|\theta\rangle_{ABG}^{\otimes n-r}), by definition. Now fix a purification |θ~ABEF|\tilde{\theta}\rangle_{ABEF} of σABE\sigma_{ABE} such that dim(EF)dim(G)\dim(EF)\geq\dim(G). Since |θ~ABEF|\tilde{\theta}\rangle_{ABEF} is also a purification of σAB\sigma_{AB}, there exists an isometry V:GEFV:\,\mathcal{H}_{G}\to\mathcal{H}_{EF} such that V|θABG=|θ~ABEFV|\theta\rangle_{ABG}=|\tilde{\theta}\rangle_{ABEF}. Then we define

ρ~A1nB1nE1nF1n:=(idA1nB1nVn)ρA1nB1nG1n(idA1nB1n(V)n),\displaystyle\tilde{\rho}_{A_{1}^{n}B_{1}^{n}E_{1}^{n}F_{1}^{n}}:=(\mathrm{id}_{A_{1}^{n}B_{1}^{n}}\otimes V^{\otimes n})\rho_{A_{1}^{n}B_{1}^{n}G_{1}^{n}}(\mathrm{id}_{A_{1}^{n}B_{1}^{n}}\otimes(V^{\dagger})^{\otimes n})\,, (166)

and

ρA1nB1nE1n:=trF1n[ρ~A1nB1nE1nF1n].\displaystyle\rho_{A_{1}^{n}B_{1}^{n}E_{1}^{n}}:=\mathrm{tr}_{F_{1}^{n}}[\tilde{\rho}_{A_{1}^{n}B_{1}^{n}E_{1}^{n}F_{1}^{n}}]\,. (167)

Clearly, ρ~A1nB1nE1nF1n\tilde{\rho}_{A_{1}^{n}B_{1}^{n}E_{1}^{n}F_{1}^{n}} is an extension of ρA1nB1n\rho_{A_{1}^{n}B_{1}^{n}}. Furthermore, this extension satisfies the two properties in Definition 2.1, namely, it is permutation invariant and supp(ρ~A1nB1nE1nF1n)span𝒱(ABEFn,|θ~ABEFnr)\mathrm{supp}(\tilde{\rho}_{A_{1}^{n}B_{1}^{n}E_{1}^{n}F_{1}^{n}})\subseteq\mathrm{span}\,\mathcal{V}(\mathcal{H}_{ABEF}^{\otimes n},|\tilde{\theta}\rangle_{ABEF}^{\otimes n-r}), which can be shown by following the same arguments as in the proof of Lemma B.1. On the other hand, ρ~A1nB1nE1nF1n\tilde{\rho}_{A_{1}^{n}B_{1}^{n}E_{1}^{n}F_{1}^{n}} is also an extension of ρA1nB1nE1n\rho_{A_{1}^{n}B_{1}^{n}E_{1}^{n}}, which means that ρA1nB1nE1n\rho_{A_{1}^{n}B_{1}^{n}E_{1}^{n}} is (nr)\binom{n}{r}-almost-iid in σABE\sigma_{ABE}. Since ρA1nB1nE1n\rho_{A_{1}^{n}B_{1}^{n}E_{1}^{n}} is an extension of ρA1nB1n\rho_{A_{1}^{n}B_{1}^{n}}, the claim follows. ∎

Appendix C Proof of Proposition 2.7

We generalize the statement for pure almost-iid states from [33, Theorem 4.5.2] to the more general setting of mixed almost-iid states. The proof works similarly to the one from [33, Theorem 4.5.2] with a few modifications.

Let |θAE|\theta\rangle_{AE} be a purification of ρ\rho, where EE denotes the purifying system of dimension d=dimd=\dim\mathcal{H}. Since ρ(n)\rho^{(n)} is an almost-iid state, there exists an extension ρA1nE1n\rho_{A_{1}^{n}E_{1}^{n}} which can be written as

ρA1nE1n=i,j𝒯βi,j|ΨiΨj|,\displaystyle\rho_{A_{1}^{n}E_{1}^{n}}=\sum_{i,j\in\mathcal{T}}\beta_{i,j}|\Psi_{i}\rangle\langle\Psi_{j}|\,, (168)

where {|Ψt}t𝒯\{|\Psi_{t}\rangle\}_{t\in\mathcal{T}} is an orthonormal basis of span𝒱(AEn,|θnr)\mathrm{span}\,\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\theta\rangle^{\otimes n-r}) with vectors |Ψt𝒱(AEn,|θnr)|\Psi_{t}\rangle\in\mathcal{V}(\mathcal{H}_{AE}^{\otimes n},|\theta\rangle^{\otimes n-r}). For any fixed t𝒯t\in\mathcal{T} we can assume without loss of generality that |Ψt=|θnr|Ωr|\Psi_{t}\rangle=|\theta\rangle^{\otimes n-r}\otimes|\Omega_{r}\rangle, where |Ωr|\Omega_{r}\rangle represents the defects. Let 𝐱=(x1,xn)\mathbf{x}=(x_{1},\ldots x_{n}) be the outcomes of the measurement n\mathcal{M}^{\otimes n} applied to |ΨtΨt||\Psi_{t}\rangle\!\langle\Psi_{t}|. Furthermore, let 𝐱=(x1,,xnr)\mathbf{x^{\prime}}=(x_{1},\ldots,x_{n-r}) and 𝐱′′=(xnr+1,,xn)\mathbf{x^{\prime\prime}}=(x_{n-r+1},\ldots,x_{n}). Clearly 𝐱\mathbf{x^{\prime}} is distributed according to the product distribution (PX)nr(P_{X})^{n-r}. Hence, for any δ>0\delta>0 we have

[λ𝐱PX1>2(ln2)(δ+|𝒳|log(nr+1)nr)][33, Cor. B.3.3]2(nr)δ.\displaystyle\mathbb{P}\Big[\left\lVert\lambda_{\mathbf{x^{\prime}}}-P_{X}\right\rVert_{1}>\sqrt{2(\ln 2)\Big(\delta+\frac{|\mathcal{X}|\log(n-r+1)}{n-r}\Big)}\Big]\overset{\textnormal{\cite[cite]{[\@@bibref{}{renner_phd}{}{}, Cor.~B.3.3]}}}{\leq}2^{-(n-r)\delta}\,. (169)

Using rn2r\leq\frac{n}{2} this can be simplified to

[λ𝐱PX1>2(ln2)(δ+2|𝒳|log(n2+1)n)]2nδ2.\displaystyle\mathbb{P}\Big[\left\lVert\lambda_{\mathbf{x^{\prime}}}-P_{X}\right\rVert_{1}>\sqrt{2(\ln 2)\Big(\delta+\frac{2|\mathcal{X}|\log(\frac{n}{2}+1)}{n}\Big)}\Big]\leq 2^{-\frac{n\delta}{2}}\,. (170)

Using λ𝐱=nrnλ𝐱+rnλ𝐱′′\lambda_{\mathbf{x}}=\frac{n-r}{n}\lambda_{\mathbf{x^{\prime}}}+\frac{r}{n}\lambda_{\mathbf{x^{\prime\prime}}} yields

λ𝐱PX1trianglenrnλ𝐱PX1+rnλ𝐱′′PX1λ𝐱PX1+2rn.\displaystyle\left\lVert\lambda_{\mathbf{x}}-P_{X}\right\rVert_{1}\overset{\textnormal{triangle}}{\leq}\frac{n-r}{n}\left\lVert\lambda_{\mathbf{x^{\prime}}}-P_{X}\right\rVert_{1}+\frac{r}{n}\left\lVert\lambda_{\mathbf{x^{\prime\prime}}}-P_{X}\right\rVert_{1}\leq\left\lVert\lambda_{\mathbf{x^{\prime}}}-P_{X}\right\rVert_{1}+\frac{2r}{n}\,. (171)

Hence,

[λ𝐱PX1>2(ln2)(δ+2|𝒳|log(n2+1)n)+2rn]\displaystyle\mathbb{P}\Big[\left\lVert\lambda_{\mathbf{x}}-P_{X}\right\rVert_{1}>\sqrt{2(\ln 2)\Big(\delta+\frac{2|\mathcal{X}|\log(\frac{n}{2}+1)}{n}\Big)}+\frac{2r}{n}\Big]
Equation 171[λ𝐱PX1>2(ln2)(δ+2|𝒳|log(n2+1)n)]\displaystyle\hskip 99.58464pt\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_red_star_stat}}}}{\leq}\mathbb{P}\Big[\left\lVert\lambda_{\mathbf{x^{\prime}}}-P_{X}\right\rVert_{1}>\sqrt{2(\ln 2)\Big(\delta+\frac{2|\mathcal{X}|\log(\frac{n}{2}+1)}{n}\Big)}\Big] (172)
Equation 1702nδ2.\displaystyle\hskip 99.58464pt\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_green_star_stat}}}}{\leq}2^{-\frac{n\delta}{2}}\,. (173)

This can be rewritten as

𝐱|Ψt[𝐱𝒲δ]2nδ2,\displaystyle\underset{\mathbf{x}\leftarrow|\Psi_{t}\rangle}{\mathbb{P}}[\mathbf{x}\in\mathcal{W}_{\delta}]\leq 2^{-\frac{n\delta}{2}}\,, (174)

for

𝒲δ={𝐱𝒳n:λ𝐱PX1>2(ln2)(δ+2|𝒳|log(n2+1)n)+2rn}.\displaystyle\mathcal{W}_{\delta}=\Big\{\mathbf{x}\in\mathcal{X}^{n}:\left\lVert\lambda_{\mathbf{x}}-P_{X}\right\rVert_{1}>\sqrt{2(\ln 2)\Big(\delta+\frac{2|\mathcal{X}|\log(\frac{n}{2}+1)}{n}\Big)}+\frac{2r}{n}\Big\}\,. (175)

The notation 𝐱|Ψt\mathbf{x}\leftarrow|\Psi_{t}\rangle indicates that 𝐱\mathbf{x} is distributed according to the outcomes of the measurement applied to |Ψt|\Psi_{t}\rangle.

For M𝐱=Mx1MxnM_{\mathbf{x}}=M_{x_{1}}\otimes\ldots\otimes M_{x_{n}} we find

𝐱ρA1nE1n[𝐱𝒲δ]\displaystyle\underset{\mathbf{x}\leftarrow\rho_{A_{1}^{n}E_{1}^{n}}}{\mathbb{P}}[\mathbf{x}\in\mathcal{W}_{\delta}] =𝐱𝒲δtr[ρA1nE1nM𝐱]\displaystyle=\sum_{\mathbf{x}\in\mathcal{W}_{\delta}}\mathrm{tr}[\rho_{A_{1}^{n}E_{1}^{n}}M_{\mathbf{x}}] (176)
𝐱𝒲δ|𝒯|t𝒯βt,tΨt|M𝐱|Ψt\displaystyle{\leq}\sum_{\mathbf{x}\in\mathcal{W}_{\delta}}|\mathcal{T}|\sum_{t\in\mathcal{T}}\beta_{t,t}\langle\Psi_{t}|M_{\mathbf{x}}|\Psi_{t}\rangle (177)
=|𝒯|t𝒯βt,t𝐱|Ψt[𝐱𝒲δ]\displaystyle=|\mathcal{T}|\sum_{t\in\mathcal{T}}\beta_{t,t}\underset{\mathbf{x}\leftarrow|\Psi_{t}\rangle}{\mathbb{P}}[\mathbf{x}\in\mathcal{W}_{\delta}] (178)
|𝒯|2nδ2\displaystyle{\leq}|\mathcal{T}|2^{-\frac{n\delta}{2}} (179)
2n(δ2h(rn))d2r.\displaystyle{\leq}2^{-n(\frac{\delta}{2}-h(\frac{r}{n}))}d^{2r}\,. (180)

Choosing δ=2log(1ε)n+2h(rn)+4rnlog(d)\delta=\frac{2\log(\frac{1}{\varepsilon})}{n}+2h(\frac{r}{n})+\frac{4r}{n}\log(d) yields

[λ𝐱PX1>4(ln2)(log(1ε)n+h(rn)+2rnlog(d)+|𝒳|log(n2+1)n)]Equation 180ε,\displaystyle\mathbb{P}\Big[\left\lVert\lambda_{\mathbf{x}}-P_{X}\right\rVert_{1}>\sqrt{4(\ln 2)\left(\frac{\log(\frac{1}{\varepsilon})}{n}+h\Big(\frac{r}{n}\Big)+\frac{2r}{n}\log(d)+\frac{|\mathcal{X}|\log(\frac{n}{2}+1)}{n}\right)}\Big]\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_square_stats}}}}{\leq}\varepsilon\,, (181)

which completes the proof. ∎

Appendix D Proof of Proposition 3.1

Consider a (n+k)(n+k)-bit string with mm ones and n+kmn+k-m zeros. If we throw away kk bits, then the probability of having jj ones in the remaining nn-bit string is

pj=(nj)(kmj)(n+km).\displaystyle p_{j}=\frac{\binom{n}{j}\binom{k}{m-j}}{\binom{n+k}{m}}\,. (182)

Note that j=0mpj=1\sum_{j=0}^{m}p_{j}=1 follows from a known identity due to Vandermonde which states that for all r,s,tr,s,t\in\mathbb{N} we have

(r+st)==0t(r)(st).\displaystyle\binom{r+s}{t}=\sum_{\ell=0}^{t}\binom{r}{\ell}\binom{s}{t-\ell}\,. (183)

To see Equation 182, let v{0,1,,k}v\in\{0,1,\ldots,k\} denote the number of ones that have been thrown away. Hence,

pj=1c(nj)(kv)=1c(nj)(kmj).\displaystyle p_{j}=\frac{1}{c}\binom{n}{j}\binom{k}{v}=\frac{1}{c}\binom{n}{j}\binom{k}{m-j}\,. (184)

for some normalization constant

c=j=0m(nj)(kmj)=Equation 183(n+km).\displaystyle c=\sum_{j=0}^{m}\binom{n}{j}\binom{k}{m-j}\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_Vandermonde}}}}{=}\binom{n+k}{m}\,. (185)
Fact D.1.

For the setting above, we have

Varp[X]=kmn(n+km)(n+k1)(n+k)2.\displaystyle\mathrm{Var}_{p}[X]=\frac{kmn(n+k-m)}{(n+k-1)(n+k)^{2}}\,. (186)
Proof.

Recall that

j(nj)=n(n1j1).\displaystyle j\binom{n}{j}=n\binom{n-1}{j-1}\,. (187)

With this we can write

𝔼p[X]\displaystyle\mathbb{E}_{p}[X] =j=0mjpj\displaystyle=\sum_{j=0}^{m}jp_{j} (188)
=j=0mj(nj)(kmj)(n+km)\displaystyle{=}\sum_{j=0}^{m}j\frac{\binom{n}{j}\binom{k}{m-j}}{\binom{n+k}{m}} (189)
=n(n+km)j=1m(n1j1)(kmj)\displaystyle{=}\frac{n}{\binom{n+k}{m}}\sum_{j=1}^{m}\binom{n-1}{j-1}\binom{k}{m-j} (190)
=n(n+km)j=0m1(n1j)(km1j)\displaystyle=\frac{n}{\binom{n+k}{m}}\sum_{j=0}^{m-1}\binom{n-1}{j}\binom{k}{m-1-j} (191)
=n(n+km)(n+k1m1)\displaystyle{=}\frac{n}{\binom{n+k}{m}}\binom{n+k-1}{m-1} (192)
=nmn+k.\displaystyle{=}\frac{nm}{n+k}\,. (193)

Similarly, we find

𝔼p[X2]\displaystyle\mathbb{E}_{p}[X^{2}] =j=0mj2pj\displaystyle=\sum_{j=0}^{m}j^{2}p_{j} (194)
=1(n+km)j=0mj2(nj)(kmj)\displaystyle{=}\frac{1}{{\binom{n+k}{m}}}\sum_{j=0}^{m}j^{2}\binom{n}{j}\binom{k}{m-j} (195)
=n(n+km)j=0m1(j+1)(n1j)(km1j)\displaystyle=\frac{n}{{\binom{n+k}{m}}}\sum_{j=0}^{m-1}(j+1)\binom{n-1}{j}\binom{k}{m-1-j} (196)
=n(n+km)(j=0m1j(n1j)(km1j)+j=0m1(n1j)(km1j))\displaystyle=\frac{n}{{\binom{n+k}{m}}}\left(\sum_{j=0}^{m-1}j\binom{n-1}{j}\binom{k}{m-1-j}+\sum_{j=0}^{m-1}\binom{n-1}{j}\binom{k}{m-1-j}\right) (197)
=n(n+km)((n1)j=0m2(n2j)(km2j)+(n+k1m1))\displaystyle{=}\frac{n}{{\binom{n+k}{m}}}\left((n-1)\sum_{j=0}^{m-2}\binom{n-2}{j}\binom{k}{m-2-j}+\binom{n+k-1}{m-1}\right) (198)
=n(n+km)((n1)(n+k2m2)+(n+k1m1))\displaystyle{=}\frac{n}{{\binom{n+k}{m}}}\left((n-1)\binom{n+k-2}{m-2}+\binom{n+k-1}{m-1}\right) (199)
=n(n1)(n+k1m1)(n+k2m2)(n+km)(n+k1m1)+n(n+k1m1)(n+km)\displaystyle=n(n-1)\frac{\binom{n+k-1}{m-1}\binom{n+k-2}{m-2}}{\binom{n+k}{m}\binom{n+k-1}{m-1}}+n\frac{\binom{n+k-1}{m-1}}{\binom{n+k}{m}} (200)
=n(n1)m(m1)(n+k)(n+k1)+nmn+k.\displaystyle=n(n-1)\frac{m(m-1)}{(n+k)(n+k-1)}+n\frac{m}{n+k}\,. (201)

Combining everything yields

Varp[X]=𝔼p[X2](𝔼p[X])2=Equations 193 and 201kmn(n+km)(n+k1)(n+k)2.\displaystyle\mathrm{Var}_{p}[X]=\mathbb{E}_{p}[X^{2}]-(\mathbb{E}_{p}[X])^{2}\overset{\textnormal{{\lx@cref{creftypepluralcap~refnum}{eq_expecation_value} and\lx@nobreakspace\lx@cref{refnum}{eq_expecation_value2}}}}{=}\frac{kmn(n+k-m)}{(n+k-1)(n+k)^{2}}\,. (202)

Let QX1n(q)𝒱(𝒳n,qnr)Q^{(q)}_{X_{1}^{n}}\in\mathcal{V}(\mathcal{X}^{n},q^{n-r}) and consider a binary random string X1nQX1n(q)X_{1}^{n}\sim Q^{(q)}_{X_{1}^{n}}. Then without loss of generality assume that the rr defects are at the end of the random string, and hence

Var[i=1nXi]=Var[i=1nrXi]+Var[i=nr+1nXi]0(nr)Varq[X],\displaystyle\mathrm{Var}\Big[\sum_{i=1}^{n}X_{i}\Big]=\mathrm{Var}\Big[\sum_{i=1}^{n-r}X_{i}\Big]+\underbrace{\mathrm{Var}\Big[\sum_{i=n-r+1}^{n}X_{i}\Big]}_{\geq 0}\geq(n-r)\mathrm{Var}_{q}[X]\,, (203)

where the first equality uses that the defects are independent of iid parts.

Example D.2.

Let α(0,1)\alpha\in(0,1), m=n+k2m=\frac{n+k}{2}, k=nαk=n^{\alpha} and r=o(n)r=o(n). Then Fact D.1 gives Varp[X]=n1+α4(n+nα1)=Θ(nα)\mathrm{Var}_{p}[X]=\frac{n^{1+\alpha}}{4(n+n^{\alpha}-1)}=\Theta(n^{\alpha}). Furthermore, Equation 203 yields Var[i=1nXi]Θ(n)\mathrm{Var}\Big[\sum_{i=1}^{n}X_{i}\Big]\geq\Theta(n).

Fact D.3.

Let t[0,1]t\in[0,1] and p,qp,q be two probability distributions. Then,

Vartp+(1t)q[X]tVarp[X]+(1t)Varq[X].\displaystyle\mathrm{Var}_{tp+(1-t)q}[X]\geq t\mathrm{Var}_{p}[X]+(1-t)\mathrm{Var}_{q}[X]\,. (204)
Proof.

By definition of the variance, we have

Vartp+(1t)q[X]\displaystyle\mathrm{Var}_{tp+(1-t)q}[X] =𝔼tp+(1t)q[X2](𝔼tp+(1t)q[X])2\displaystyle=\mathbb{E}_{tp+(1-t)q}[X^{2}]-(\mathbb{E}_{tp+(1-t)q}[X])^{2} (205)
=t𝔼p[X2]+(1t)𝔼q[X2](t𝔼p[X]+(1t)𝔼q[X])2\displaystyle=t\mathbb{E}_{p}[X^{2}]+(1-t)\mathbb{E}_{q}[X^{2}]-(t\mathbb{E}_{p}[X]+(1-t)\mathbb{E}_{q}[X])^{2} (206)
t𝔼p[X2]+(1t)𝔼q[X2](t𝔼p[X]2+(1t)𝔼q[X]2)\displaystyle\geq t\mathbb{E}_{p}[X^{2}]+(1-t)\mathbb{E}_{q}[X^{2}]-(t\mathbb{E}_{p}[X]^{2}+(1-t)\mathbb{E}_{q}[X]^{2}) (207)
=tVarp[X]+(1t)Varq[X].\displaystyle=t\mathrm{Var}_{p}[X]+(1-t)\mathrm{Var}_{q}[X]\,. (208)

Putting everything together, we obtain for α(0,1)\alpha\in(0,1)

Varp[X]=Example D.2Θ(nα)andVarQX(q)ν(𝑑q)[X]Fact D.3ν(𝑑q)VarQX(q)[X]=Example D.2Θ(n).\displaystyle\mathrm{Var}_{p}[X]\overset{\textnormal{{\lx@cref{creftypecap~refnum}{ex_Variance_1}}}}{=}\Theta(n^{\alpha})\quad\textnormal{and}\quad\mathrm{Var}_{\int Q_{X}^{(q)}\nu(\mathrm{d}q)}[X]\overset{\textnormal{{\lx@cref{creftypecap~refnum}{fact_concavity_variance}}}}{\geq}\int\nu(\mathrm{d}q)\mathrm{Var}_{Q_{X}^{(q)}}[X]\overset{\textnormal{{\lx@cref{creftypecap~refnum}{ex_Variance_1}}}}{=}\Theta(n)\,. (209)

This proves the assertion of Proposition 3.1. Note that for better readability, the above proof has been done for the specific choice k=nα=o(n)k=n^{\alpha}=o(n) for α(0,1)\alpha\in(0,1), but the same argument remains valid for an arbitrary k=o(n)k=o(n). ∎

Appendix E Alternative proof of Theorem 4.1

In this section, we present an alternative proof for Theorem 4.1 which uses an entirely different proof technique which may be of independent interest. However, we note that the scaling of the defects rr is slightly worse than in Theorem 4.1. Note that it suffices to prove the following result.

Theorem E.1.

Let σAS(A)\sigma_{A}\in\mathrm{S}(\mathcal{H}_{A}) and ρA1nSn(A,σAnr)\rho_{A_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{A},\sigma_{A}^{\otimes n-r}) for r=o(n)r=o(\sqrt{n}). Then

1nH(A1n)ρ=H(A)σ+o(n)n.\displaystyle\frac{1}{n}H(A_{1}^{n})_{\rho}=H(A)_{\sigma}+\frac{o(n)}{n}\,. (210)

Theorem E.1 implies that the conditional entropy of almost-iid states coincides asymptotically with the conditional entropy of iid states.

Corollary E.2.

Let σABS(AB)\sigma_{AB}\in\mathrm{S}(\mathcal{H}_{AB}) and ρA1nB1nSn(AB,σABnr)\rho_{A_{1}^{n}B_{1}^{n}}\in\mathrm{S}^{n}(\mathcal{H}_{AB},\sigma_{AB}^{\otimes n-r}) for r=o(n)r=o(\sqrt{n}). Then

1nH(A1n|B1n)ρ=H(A|B)σ+o(n)n.\displaystyle\frac{1}{n}H(A_{1}^{n}|B_{1}^{n})_{\rho}=H(A|B)_{\sigma}+\frac{o(n)}{n}\,. (211)
Proof.

By definition of the conditional entropy, we have

1nH(A1n|B1n)ρ=1nH(A1nB1n)ρ1nH(B1n)ρ=Theorem E.1H(AB)σH(B)σ+o(n)n=H(A|B)σ+o(n)n.\displaystyle\frac{1}{n}H(A_{1}^{n}|B_{1}^{n})_{\rho}=\frac{1}{n}H(A_{1}^{n}B_{1}^{n})_{\rho}-\frac{1}{n}H(B_{1}^{n})_{\rho}\overset{\textnormal{{\lx@cref{creftypecap~refnum}{thm_entropy_of_almost_iid_alternative}}}}{=}H(AB)_{\sigma}-H(B)_{\sigma}+\frac{o(n)}{n}=H(A|B)_{\sigma}+\frac{o(n)}{n}\,.

E.1 Proof of Theorem E.1

One direction of Equation 210 is simple. To see this, let εn=2rsn\varepsilon_{n}=2\sqrt{\frac{rs}{n}} for s=ns=\lfloor\sqrt{n}\rfloor and consider

1nH(A1n)ρsubadditivityH(A1)ρProposition 2.6H(A)σ+εnlogd+h(εn)=H(A)σ+o(n)n,\displaystyle\frac{1}{n}H(A_{1}^{n})_{\rho}\overset{\textnormal{subadditivity}}{\leq}H(A_{1})_{\rho}\overset{\textnormal{{{\lx@cref{creftypecap~refnum}{prop_distance}}}}}{\leq}H(A)_{\sigma}+\varepsilon_{n}\log d+h(\varepsilon_{n})=H(A)_{\sigma}+\frac{o(n)}{n}\,, (212)

where the penultimate step uses the continuity of entropy [2, 32, 42].

The other direction is more complicated. Recall that strong subadditivity of quantum entropy (SSA) [27, 28] ensures I(A:C|B)ρ0I(A:C|B)_{\rho}\geq 0. Furthermore, the conditional mutual information satisfies a chain rule I(A:BC)ρ=I(A:B)ρ+I(A:C|B)ρI(A:BC)_{\rho}=I(A:B)_{\rho}+I(A:C|B)_{\rho}. Let |θAE|\theta\rangle_{AE} be a purification of σA\sigma_{A} and let ρA1nE1n\rho_{A_{1}^{n}E_{1}^{n}} be an extension of ρA1n\rho_{A_{1}^{n}} that satisfies the two conditions of Definition 2.1. Since ρA1nE1n\rho_{A_{1}^{n}E_{1}^{n}} is permutation-invariant, for entropy and mutual information terms the indices of the considered subsystems can be changed. Hence, we find for any kn2=:n0k\leq\lfloor\frac{n}{2}\rfloor=:n_{0}, kn0+1k^{\prime}\leq n_{0}+1 and for any =k,,n0\ell=k,\ldots,n_{0}, =k,,n0+1\ell^{\prime}=k^{\prime},\ldots,n_{0}+1

I(A1n0E1n0:An0+1nEn0+1n)ρ\displaystyle I(A_{1}^{n_{0}}E_{1}^{n_{0}}:A_{n_{0}+1}^{n}E_{n_{0}+1}^{n})_{\rho} I(A1n0:An0+1n)ρ\displaystyle{\geq}I(A_{1}^{n_{0}}:A_{n_{0}+1}^{n})_{\rho} (213)
=I(A1:An0+1n|A+1n0)ρ+I(A+1n0:An0+1n)ρ\displaystyle{=}I(A_{1}^{\ell}:A_{n_{0}+1}^{n}|A_{\ell+1}^{n_{0}})_{\rho}+I(A_{\ell+1}^{n_{0}}:A_{n_{0}+1}^{n})_{\rho} (214)
I(A1k:An0+1n|A+1n0)ρ\displaystyle{\geq}I(A_{1}^{k}:A_{n_{0}+1}^{n}|A_{\ell+1}^{n_{0}})_{\rho} (215)
=I(A1k:An0+1n0+|A+1n0An0++1n)ρ+I(A1k:An0++1n|A+1n0)ρ\displaystyle{=}I(A_{1}^{k}:A_{n_{0}+1}^{n_{0}+{\ell^{\prime}}}|A_{\ell+1}^{n_{0}}A_{n_{0}+\ell^{\prime}+1}^{n})_{\rho}+I(A_{1}^{k}:A_{n_{0}+\ell^{\prime}+1}^{n}|A_{\ell+1}^{n_{0}})_{\rho} (216)
I(A1k:An0+1n0+k|A+1n0An0++1n)ρ\displaystyle{\geq}I(A_{1}^{k}:A_{n_{0}+1}^{n_{0}+k^{\prime}}|A_{\ell+1}^{n_{0}}A_{n_{0}+\ell^{\prime}+1}^{n})_{\rho} (217)
=I(A1k:Ak+1k+k|Ak+k+1m)ρ,\displaystyle{=}I(A_{1}^{k}:A_{k+1}^{k+k^{\prime}}|A_{k+k^{\prime}+1}^{m})_{\rho}\,, (218)

where in the last step, we permute all the remaining systems appearing in the conditioning (there are nn-\ell-\ell^{\prime} many) into neighboring systems labeled by indices from (k+k+1)(k+k^{\prime}+1) to m:=(k+k+n)m:=(k+k^{\prime}+n-\ell-\ell^{\prime}).

Thus, for any mk+km\geq k+k^{\prime} we find

I(A1k:Ak+1k+k|Ak+k+1m)ρ\displaystyle I(A_{1}^{k}:A_{k+1}^{k+k^{\prime}}|A_{k+k^{\prime}+1}^{m})_{\rho} I(A1n0B1n0:An0+1nEn0+1n)ρ\displaystyle{\leq}I(A_{1}^{n_{0}}B_{1}^{n_{0}}:A_{n_{0}+1}^{n}E_{n_{0}+1}^{n})_{\rho} (219)
=H(A1n0E1n0)ρ+H(An0+1nEn0+1n)ρH(A1nE1n)ρ\displaystyle=H(A_{1}^{n_{0}}E_{1}^{n_{0}})_{\rho}+H(A_{n_{0}+1}^{n}E_{n_{0}+1}^{n})_{\rho}-H(A_{1}^{n}E_{1}^{n})_{\rho} (220)
H(A1n0E1n0)ρ+H(An0+1nEn0+1n)ρ.\displaystyle\leq H(A_{1}^{n_{0}}E_{1}^{n_{0}})_{\rho}+H(A_{n_{0}+1}^{n}E_{n_{0}+1}^{n})_{\rho}\,. (221)

In the proof of Lemma B.3 it is shown that ρA1n0E1n0span𝒱(AEn0,|θAEn0r)\rho_{A_{1}^{n_{0}}E_{1}^{n_{0}}}\in\mathrm{span}\mathcal{V}(\mathcal{H}_{AE}^{\otimes n_{0}},|\theta\rangle_{AE}^{\otimes n_{0}-r}). This implies

H(A1n0E1n0)ρEquation 9log(2n0h(r/n0)dAEr)=n0h(rn0)+rlogdAE.\displaystyle H(A_{1}^{n_{0}}E_{1}^{n_{0}})_{\rho}\overset{\textnormal{{\lx@cref{creftypecap~refnum}{eq_sizeT}}}}{\leq}\log\left(2^{n_{0}h(r/n_{0})}d_{AE}^{r}\right)=n_{0}h\Big(\frac{r}{n_{0}}\Big)+r\log d_{AE}\,. (222)

Similarly, we obtain

H(An0+1nEn0+1n)ρ(nn0)h(rnn0)+rlogdAE.\displaystyle H(A_{n_{0}+1}^{n}E_{n_{0}+1}^{n})_{\rho}\leq(n-n_{0})h\Big(\frac{r}{n-n_{0}}\Big)+r\log d_{AE}\,. (223)

Hence we find

I(A1k:Ak+1k+k|Ak+k+1m)ρ\displaystyle I(A_{1}^{k}:A_{k+1}^{k+k^{\prime}}|A_{k+k^{\prime}+1}^{m})_{\rho} H(A1n0B1n0)ρ+H(An0+1nBn0+1n)ρ\displaystyle{\leq}H(A_{1}^{n_{0}}B_{1}^{n_{0}})_{\rho}+H(A_{n_{0}+1}^{n}B_{n_{0}+1}^{n})_{\rho} (224)
nh(2rn)+2rlogdAE.\displaystyle{\leq}nh\Big(\frac{2r}{n}\Big)+2r\log d_{AE}\,. (225)

For any <s=(logn)32\ell<s=\lfloor(\log n)^{\frac{3}{2}}\rfloor, we thus have

ρA1+1σA(+1)1ρA1sσAs1Proposition 2.6εn,\displaystyle\left\lVert\rho_{A_{1}^{\ell+1}}-\sigma_{A}^{\otimes(\ell+1)}\right\rVert_{1}\leq\left\lVert\rho_{A_{1}^{s}}-\sigma_{A}^{\otimes s}\right\rVert_{1}\overset{\textnormal{{\lx@cref{creftypecap~refnum}{prop_distance}}}}{\leq}\varepsilon_{n}\,, (226)

where the first step follows since the trace distance is contractive under trace-preserving completely positive maps [43, Theorem 8.16]. The continuity of conditional entropy [1] then implies that

|H(A1|A2+1)ρH(A)σ|=|H(A1|A2+1)ρH(A1|A2+1)σ(+1)|4εnlogdA+2h(εn)=:δn,\displaystyle|H(A_{1}|A_{2}^{\ell+1})_{\rho}-H(A)_{\sigma}|=|H(A_{1}|A_{2}^{\ell+1})_{\rho}-H(A_{1}|A_{2}^{\ell+1})_{\sigma^{\otimes(\ell+1)}}|\leq 4\varepsilon_{n}\log d_{A}+2h(\varepsilon_{n})=:\delta_{n}\,, (227)

where dA=dim(A)d_{A}=\dim(A). The chain rule together with permutation invariance allows us to write

H(A1n)ρ\displaystyle H(A_{1}^{n})_{\rho} =H(A1n|An+1n)ρ+H(An+1n)ρ\displaystyle{=}H(A_{1}^{n-\ell}|A_{n-\ell+1}^{n})_{\rho}+H(A_{n-\ell+1}^{n})_{\rho} (228)
H(A1n|An+1n)ρ\displaystyle\geq H(A_{1}^{n-\ell}|A_{n-\ell+1}^{n})_{\rho} (229)
=i=1nH(Ai|An+1n)ρi=1nI(A1i1:Ai|An+1n)ρ\displaystyle{=}\sum_{i=1}^{n-\ell}H(A_{i}|A_{n-\ell+1}^{n})_{\rho}-\sum_{i=1}^{n-\ell}I(A^{i-1}_{1}:A_{i}|A_{n-\ell+1}^{n})_{\rho} (230)
=(n)H(A1|A2+1)ρi=1nI(A1i1:Ai|An+1n)ρ\displaystyle{=}(n-\ell)H(A_{1}|A_{2}^{\ell+1})_{\rho}-\sum_{i=1}^{n-\ell}I(A^{i-1}_{1}:A_{i}|A_{n-\ell+1}^{n})_{\rho} (231)
(n)H(A)σi=1nI(A1i1:Ai|An+1n)ρnδn\displaystyle{\geq}(n-\ell)H(A)_{\sigma}-\sum_{i=1}^{n-\ell}I(A^{i-1}_{1}:A_{i}|A_{n-\ell+1}^{n})_{\rho}-n\delta_{n} (232)
=(n)H(A)σi=1nsI(A1i1:Ai|An+1n)ρi=ns+1nI(A1i1:Ai|An+1n)ρnδn\displaystyle=(n\!-\!\ell)H(A)_{\sigma}\!-\!\sum_{i=1}^{n-s}I(A^{i-1}_{1}:A_{i}|A_{n-\ell+1}^{n})_{\rho}\!-\!\sum_{i=n-s+1}^{n-\ell}I(A^{i-1}_{1}:A_{i}|A_{n-\ell+1}^{n})_{\rho}\!-\!n\delta_{n} (233)
(n)H(A)σi=1nsI(A1i1:Ai|An+1n)ρ2(s)logdAnδn\displaystyle\geq(n-\ell)H(A)_{\sigma}-\sum_{i=1}^{n-s}I(A^{i-1}_{1}:A_{i}|A_{n-\ell+1}^{n})_{\rho}-2(s-\ell)\log d_{A}-n\delta_{n} (234)
=(n)H(A)σi=1nsI(A1i1:Ai|Ans+1ns+)ρ2(s)logdAnδn.\displaystyle{=}(n-\ell)H(A)_{\sigma}-\sum_{i=1}^{n-s}I(A^{i-1}_{1}:A_{i}|A_{n-s+1}^{n-s+\ell})_{\rho}-2(s-\ell)\log d_{A}-n\delta_{n}\,. (235)
Claim E.3.

Let nn\in\mathbb{N}, exists s=n\ell\leq s=\lfloor\sqrt{n}\rfloor such that for i=1nsI(A1i1:Ai|Ans+1ns+)ρ=:ξn\sum_{i=1}^{n-s}I(A^{i-1}_{1}:A_{i}|A_{n-s+1}^{n-s+\ell})_{\rho}=:\xi_{n} we have limnξnn=0\lim_{n\to\infty}\frac{\xi_{n}}{n}=0.

Proof.

The proof follows the idea from [8, Equation (5)], which proves a variant of Claim E.3 for a purely classical scenario. By the permutation-invariance we have for any 1i<kmn1\leq i<k\leq m\leq n

I(A1i1:Ai|Ak+1m)ρ=I(A1i1:Am|Akm1)ρ.\displaystyle I(A_{1}^{i-1}:A_{i}|A_{k+1}^{m})_{\rho}=I(A_{1}^{i-1}:A_{m}|A_{k}^{m-1})_{\rho}\,. (236)

This allows us to write

m=knI(A1i1:Ai|Ak+1m)ρ=(236)m=knI(A1i1:Am|Akm1)ρ=chain ruleI(A1i1:Akn)ρ.\displaystyle\sum_{m=k}^{n}I(A_{1}^{i-1}:A_{i}|A_{k+1}^{m})_{\rho}\overset{\eqref{eq_step_conc1}}{=}\sum_{m=k}^{n}I(A_{1}^{i-1}:A_{m}|A_{k}^{m-1})_{\rho}\overset{\textnormal{chain rule}}{=}I(A_{1}^{i-1}:A_{k}^{n})_{\rho}\,. (237)

Summing Equation 237 over all 1i<k1\leq i<k and dividing by nk+1n-k+1 gives

1nk+1m=kni=1k1I(A1i1:Ai|Ak+1m)ρ=1nk+1i=1k1I(A1i1:Akn)ρ.\displaystyle\frac{1}{n-k+1}\sum_{m=k}^{n}\sum_{i=1}^{k-1}I(A_{1}^{i-1}:A_{i}|A_{k+1}^{m})_{\rho}=\frac{1}{n-k+1}\sum_{i=1}^{k-1}I(A_{1}^{i-1}:A_{k}^{n})_{\rho}\,. (238)

Hence, there exists m{k,k+1,,n}m^{*}\in\{k,k+1,\ldots,n\} such that

i=1k1I(A1i1:Ai|Ak+1m)ρ1nk+1i=1k1I(A1i1:Akn)ρ.\displaystyle\sum_{i=1}^{k-1}I(A_{1}^{i-1}:A_{i}|A_{k+1}^{m^{*}})_{\rho}\leq\frac{1}{n-k+1}\sum_{i=1}^{k-1}I(A_{1}^{i-1}:A_{k}^{n})_{\rho}\,. (239)

Choosing k=nsk=n-s, Equation 239 can be rewritten (as there exists 0s0\leq\ell\leq s such that m=k+m^{*}=k+\ell) as

i=1ns1I(A1i1:Ai|Ans+1ns+)ρ1s+1i=1nsI(A1i1:Ansn)ρ\displaystyle\sum_{i=1}^{n-s-1}I(A^{i-1}_{1}:A_{i}|A_{n-s+1}^{n-s+\ell})_{\rho}\leq\frac{1}{s+1}\sum_{i=1}^{n-s}I(A_{1}^{i-1}:A_{n-s}^{n})_{\rho} (240)

We next show that Equation 225 implies

I(A1i1:Ansn)ρ2nh(2rn)+4rlogdAEi=1,,ns.\displaystyle I(A_{1}^{i-1}:A_{n-s}^{n})_{\rho}\leq 2nh\Big(\frac{2r}{n}\Big)+4r\log d_{AE}\quad\forall i=1,\ldots,n-s\,. (241)

To see this, we can assume w.l.o.g.88 8 It can be shown that s+1n0+2s+1\leq n_{0}+2 for all nn\in\mathbb{N}. Hence, if s+1=n0+2s+1=n_{0}+2, due to the chain rule we can write I(A1i1:Ansn)ρ=I(A1i1:Ans+1n)ρ+I(A1i1:Ansns|Ans+1n)ρI(A1i1:Ans+1n)ρ+2logdAI(A_{1}^{i-1}:A_{n-s}^{n})_{\rho}=I(A_{1}^{i-1}:A_{n-s+1}^{n})_{\rho}+I(A_{1}^{i-1}:A_{n-s}^{n-s}|A_{n-s+1}^{n})_{\rho}\leq I(A_{1}^{i-1}:A_{n-s+1}^{n})_{\rho}+2\log d_{A}, and the argument above still works. that nn is large enough such that

k:=s+1=(logn)32+1n2+1=n0+1.\displaystyle k^{\prime}:=s+1=\left\lfloor(\log n)^{\frac{3}{2}}\right\rfloor+1\leq\left\lfloor\frac{n}{2}\right\rfloor+1=n_{0}+1\,. (242)

Let us now consider two cases: In case i1n0i-1\leq n_{0}, the assertion follows from Equation 225 by choosing m=k+km=k+k^{\prime}, since

I(A1i1:Ansn)ρ\displaystyle I(A_{1}^{i-1}:A_{n-s}^{n})_{\rho} =I(A1i1:Aii+s)ρnh(2rn)+2rlogdAE.\displaystyle{=}I(A_{1}^{i-1}:A_{i}^{i+s})_{\rho}\leq nh\Big(\frac{2r}{n}\Big)+2r\log d_{AE}\,. (243)

In case i1>n0i-1>n_{0}, we can use the chain rule to write

I(A1i1:Ansn)ρ\displaystyle I(A_{1}^{i-1}:A_{n-s}^{n})_{\rho} =I(A1n0:Ansn)ρ+I(An0+1i1:Ansn|A1n0)ρ\displaystyle=I(A_{1}^{n_{0}}:A_{n-s}^{n})_{\rho}+I(A_{n_{0}+1}^{i-1}:A_{n-s}^{n}|A_{1}^{n_{0}})_{\rho} (244)
=I(A1n0:An0+1n0+1+s)ρ+I(A1i1n0:Ain0in0+s|Ain0+s+1i+s)ρ,\displaystyle{=}I(A_{1}^{n_{0}}:A_{{n_{0}}+1}^{{n_{0}}+1+s})_{\rho}+I(A_{1}^{i-1-{n_{0}}}:A_{i-{n_{0}}}^{i-{n_{0}}+s}|A_{i-n_{0}+s+1}^{i+s})_{\rho}\,, (245)

where both terms on the right-hand side are bounded from above by nh(2rn)+2rlogdAEnh(\frac{2r}{n})+2r\log d_{AE} via Equation 225.

Putting things together yields

i=1nsI(A1i1:Ai|Ans+1ns+)ρ\displaystyle\sum_{i=1}^{n-s}I(A^{i-1}_{1}:A_{i}|A_{n-s+1}^{n-s+\ell})_{\rho} =i=1ns1I(A1i1:Ai|Ans+1ns+)ρ+I(A1ns1:Ans|Ans+1ns+)ρ\displaystyle=\sum_{i=1}^{n-s-1}I(A^{i-1}_{1}:A_{i}|A_{n-s+1}^{n-s+\ell})_{\rho}+I(A_{1}^{n-s-1}:A_{n-s}|A_{n-s+1}^{n-s+\ell})_{\rho} (246)
1s+1i=1nsI(A1i1:Ansn)ρ+2logdA\displaystyle{\leq}\frac{1}{s+1}\sum_{i=1}^{n-s}I(A_{1}^{i-1}:A_{n-s}^{n})_{\rho}+2\log d_{A} (247)
nss+1(2nh(2rn)+4rlogdAE)+2logdA\displaystyle{\leq}\frac{n-s}{s+1}\left(2nh\Big(\frac{2r}{n}\Big)+4r\log d_{AE}\right)+2\log d_{A} (248)
=o(n),\displaystyle=o(n)\,, (249)

where the final step uses that for s=ns=\lfloor\sqrt{n}\rfloor and r=o(n)r=o(n) we have

limnnsh(2rn)=0.\displaystyle\lim_{n\to\infty}\frac{n}{s}h\Big(\frac{2r}{n}\Big)=0\,. (250)

Since limnδn=0\lim_{n\to\infty}\delta_{n}=0, combining Claim E.3 with Equation 235 yields for some s=n\ell\leq s=\lfloor\sqrt{n}\rfloor

H(A1n)ρ(n)H(A)σo(n)=nH(A)σo(n).\displaystyle H(A_{1}^{n})_{\rho}\geq(n-\ell)H(A)_{\sigma}-o(n)=nH(A)_{\sigma}-o(n)\,. (251)

References

  • [1] R. Alicki and M. Fannes. Continuity of quantum conditional information. Journal of Physics A: Mathematical and General, 37(5):55–57, 2004. DOI: doi:10.1088/0305-4470/37/5/L01.
  • [2] K. M. R. Audenaert. A sharp continuity estimate for the von Neumann entropy. Journal of Physics A: Mathematical and Theoretical, 40(28):8127, 2007. Available online: http://stacks.iop.org/1751-8121/40/i=28/a=S18.
  • [3] C. H. Bennett, H. J. Bernstein, S. Popescu, and B. Schumacher. Concentrating partial entanglement by local operations. Phys. Rev. A, 53:2046–2052, 1996. DOI: 10.1103/PhysRevA.53.2046.
  • [4] C. H. Bennett, G. Brassard, S. Popescu, B. Schumacher, J. A. Smolin, and W. K. Wootters. Purification of noisy entanglement and faithful teleportation via noisy channels. Phys. Rev. Lett., 76:722–725, 1996. DOI: 10.1103/PhysRevLett.76.722.
  • [5] C. H. Bennett, D. P. DiVincenzo, J. A. Smolin, and W. K. Wootters. Mixed-state entanglement and quantum error correction. Physical Review A, 54(5):3824–3851, 1996. DOI: 10.1103/PhysRevA.54.3824.
  • [6] M. Berta, F. G. S. L. Brandão, G. Gour, L. Lami, M. B. Plenio, B. Regula, and M. Tomamichel. On a gap in the proof of the generalised quantum Stein’s lemma and its consequences for the reversibility of quantum resources. Quantum, 7:1103, 2023. DOI: 10.22331/q-2023-09-07-1103.
  • [7] M. Berta, F. G. S. L. Brandão, G. Gour, L. Lami, M. B. Plenio, B. Regula, and M. Tomamichel. The tangled state of quantum hypothesis testing. Nature Physics, 20(2):172–175, 2024. DOI: 10.1038/s41567-023-02289-9.
  • [8] M. Berta, L. Gavalakis, and I. Kontoyiannis. A third information-theoretic approach to finite de Finetti theorems, 2023. DOI: 10.48550/arXiv.2304.05360.
  • [9] F. G. S. L. Brandão and M. B. Plenio. A generalization of quantum Stein’s lemma. Communications in Mathematical Physics, 295(3):791–828, 2010. DOI: 10.1007/s00220-010-1005-z.
  • [10] F. Buscemi, D. Sutter, and M. Tomamichel. An information-theoretic treatment of quantum dichotomies. Quantum, 3:209, 2019. DOI: 10.22331/q-2019-12-09-209.
  • [11] E. Carlen. Trace Inequalities and Quantum Entropy: An Introductory Course. Contemporary Mathematics, 2009. DOI: 10.1090/conm/529.
  • [12] C. M. Caves, C. A. Fuchs, and R. Schack. Unknown quantum states: The quantum de Finetti representation. Journal of Mathematical Physics, 43(9):4537–4559, 2002. DOI: 10.1063/1.1494475.
  • [13] M. Christandl, R. König, G. Mitchison, and R. Renner. One-and-a-half quantum de Finetti theorems. Communications in Mathematical Physics, 273(2):473–498, 2007. DOI: 10.1007/s00220-007-0189-3.
  • [14] M. Christandl and A. Winter. “Squashed entanglement”: An additive entanglement measure. Journal of Mathematical Physics, 45(3):829–840, 2004. DOI: 10.1063/1.1643788.
  • [15] T. M. Cover and J. A. Thomas. Elements of Information Theory. Wiley Interscience, 2006. DOI: 10.1002/047174882X.
  • [16] N. Datta. Min- and max-relative entropies and a new entanglement monotone. IEEE Transactions on Information Theory, 55(6):2816–2826, 2009. DOI: 10.1109/TIT.2009.2018325.
  • [17] N. Datta and R. Renner. Smooth entropies and the quantum information spectrum. IEEE Transactions on Information Theory, 55(6):2807–2815, 2009. DOI: 10.1109/TIT.2009.2018340.
  • [18] B. De Finetti. La prévision: ses lois logiques, ses sources subjectives. In Annales de l’institut Henri Poincaré, volume 7, pages 1–68, 1937.
  • [19] B. de Finetti. Logical foundations and measurement of subjective probability. Acta Psychologica, 34:129–145, 1970. DOI: https://doi.org/10.1016/0001-6918(70)90012-0.
  • [20] P. Diaconis and D. Freedman. Finite Exchangeable Sequences. The Annals of Probability, 8(4):745 – 764, 1980. DOI: 10.1214/aop/1176994663.
  • [21] O. Fawzi and R. Renner. Quantum conditional mutual information and approximate Markov chains. Communications in Mathematical Physics, 340(2):575–611, 2015. DOI: 10.1007/s00220-015-2466-x.
  • [22] C. Fuchs and J. van de Graaf. Cryptographic distinguishability measures for quantum-mechanical states. IEEE Transactions on Information Theory, 45(4):1216 –1227, 1999. DOI: 10.1109/18.761271.
  • [23] M. Hayashi and H. Yamasaki. The generalized quantum Stein’s lemma and the second law of quantum resource theories. Nature Physics, 21(12):1988–1993, 2025. DOI: 10.1038/s41567-025-03047-9.
  • [24] P. M. Hayden, M. Horodecki, and B. M. Terhal. The asymptotic entanglement cost of preparing a quantum state. Journal of Physics A: Mathematical and General, 34(35):6891, 2001. DOI: 10.1088/0305-4470/34/35/314.
  • [25] R. König and R. Renner. A de Finetti representation for finite symmetric quantum states. Journal of Mathematical Physics, 46(12):122108, 2005. DOI: 10.1063/1.2146188.
  • [26] L. Lami. A solution of the generalized quantum Stein’s lemma. IEEE Transactions on Information Theory, 71(6):4454–4484, 2025. DOI: 10.1109/TIT.2025.3543610.
  • [27] E. H. Lieb and M. B. Ruskai. A fundamental property of quantum-mechanical entropy. Physical Review Letters, 30:434–436, 1973. DOI: 10.1103/PhysRevLett.30.434.
  • [28] E. H. Lieb and M. B. Ruskai. Proof of the strong subadditivity of quantum-mechanical entropy. Journal of Mathematical Physics, 14(12):1938–1941, 1973. DOI: 10.1063/1.1666274.
  • [29] G. Mazzola and R. Renner. Asymptotic transformation rates with almost iid resources, 2026. in preparation.
  • [30] P. Monari and D. Cocchi. Introduction to Bruno de Finetti’s “probabiliá e induzione”. Cooperativa Libraria Universitaria Editrice, Bologna, 1993.
  • [31] M. Müller-Lennert, F. Dupuis, O. Szehr, S. Fehr, and M. Tomamichel. On quantum Rényi entropies: A new generalization and some properties. Journal of Mathematical Physics, 54(12), 2013. DOI: http://dx.doi.org/10.1063/1.4838856.
  • [32] D. Petz. Quantum Information Theory and Quantum Statistics. Springer, 2008. DOI: 10.1007/978-3-540-74636-2.
  • [33] R. Renner. Security of quantum key distribution. PhD thesis, ETH Zurich, 2005. available at arXiv:quant-ph/0512258.
  • [34] R. Renner. Symmetry of large physical systems implies independence of subsystems. Nature Physics, 3(9):pp. 645–649, 2007. Available online: http://www.nature.com/nphys/journal/v3/n9/suppinfo/nphys684_S1.html.
  • [35] D. Sutter. Approximate Quantum Markov Chains. Springer International Publishing, 2018. DOI: 10.1007/978-3-319-78732-9_5.
  • [36] M. Tomamichel. Quantum Information Processing with Finite Resources, volume 5 of SpringerBriefs in Mathematical Physics. Springer, 2015. DOI: 10.1007/978-3-319-21891-5.
  • [37] M. Tomamichel, R. Colbeck, and R. Renner. A fully quantum asymptotic equipartition property. IEEE Transactions on Information Theory, 55(12):5840–5847, 2009. DOI: 10.1109/TIT.2009.2032797.
  • [38] M. Tomamichel and M. Hayashi. A hierarchy of information quantities for finite block length analysis of quantum tasks. IEEE Transactions on Information Theory, 59(11):7693–7710, 2013. DOI: 10.1109/TIT.2013.2276628.
  • [39] V. Vedral, M. B. Plenio, M. A. Rippin, and P. L. Knight. Quantifying entanglement. Phys. Rev. Lett., 78:2275–2279, 1997. DOI: 10.1103/PhysRevLett.78.2275.
  • [40] K. G. H. Vollbrecht and R. F. Werner. Entanglement measures under symmetry. Phys. Rev. A, 64:062307, 2001. DOI: 10.1103/PhysRevA.64.062307.
  • [41] M. M. Wilde, A. Winter, and D. Yang. Strong converse for the classical capacity of entanglement-breaking and Hadamard channels via a sandwiched Rényi relative entropy. Communications in Mathematical Physics, 331(2):593–622, 2014. DOI: 10.1007/s00220-014-2122-x.
  • [42] A. Winter. Tight uniform continuity bounds for quantum entropies: Conditional entropy, relative entropy distance and energy constraints. Communications in Mathematical Physics, 347(1):291–313, 2016. DOI: 10.1007/s00220-016-2609-8.
  • [43] M. M. Wolf. Quantum channels & operations: Guided tour, 2012. Lecture notes available at https://www-m5.ma.tum.de/foswiki/pub/M5/Allgemeines/MichaelWolf/QChannelLecture.pdf.
  • [44] Y.-D. Wu and G. Chiribella. Detecting quantum capacities of continuous-variable quantum channels. Phys. Rev. Res., 4:043149, 2022. DOI: 10.1103/PhysRevResearch.4.043149.