Almost-iid information theory
Abstract
Information-theoretic techniques are based on the assumption that resources are well characterized by independent and identically distributed (iid) states. This assumption cannot be justified operationally, since, for example, correlations between subsequent systems emitted by a source cannot be detected by any practical tomographic protocol. Operationally motivated symmetry assumptions still imply, via de Finetti theorems, that the resources are described by almost-iid states. This raises the question: Are almost-iid resources as effective as perfect iid resources for information-processing tasks? Here we address this question and prove that the conditional entropy of almost-iid states asymptotically coincides with that of iid states. As an application, this implies that squashed entanglement is robust for almost-iid states, asymptotically matching its value on iid states.
1 Introduction
A common assumption in physics is that the same experiment can be repeated many times independently. More precisely, one often assumes that the outcomes of running the experiment times are described by independent and identically distributed (iid) random variables. In technical terms, this means that their joint distribution is factorized, i.e., . The Italian mathematician Bruno de Finetti cautioned that the “most common and misleading error” in probability theory is treating the iid assumption as fundamental [19]. He provided a mathematically convincing argument to operationally justify the iid assumption, relating it to the permutation invariance of the outcomes [18, 30]. For many realistic systems, permutation invariance follows from natural assumptions such as the indistinguishability of subsystems (see [34] for a more detailed discussion).11 1 Alternatively, permutation invariance can also be enforced by randomly permuting the subsystems.
Suppose that we have a source that generates random variables , where we assume that the underlying joint distribution is permutation-invariant and that each random variable takes values in the set . When selecting of these random variables, i.e. ignoring random variables, de Finetti’s theorem [20] shows that there exists a probability measure on the set of distributions on such that
| (1) |
where denotes the -norm and . This result justifies that when ignoring random variables the remaining ones are approximately a convex combination of iid random variables.
De Finetti theorems have been generalized to the quantum case. In [12] a quantum de Finetti theorem has been obtained as an implication of the classical result in the case where . This yields a result that is applicable under the assumption of having an infinite number of samples. Later in [25, 13] quantum de Finetti theorems have been derived that work for a finite number of samples — analogous to Equation 1. For any state on that is symmetric (i.e. invariant under permutations of the subsystems), there exists a probability measure on the unit sphere such that
| (2) |
where is the trace norm, denotes the partial trace of subsystems, and . It is worth mentioning that the results stated in Equations 1 and 2 are essentially tight [20, 13].
The main obstacle with these de Finetti results (classical and quantum) is that they require to be large. Specifically, if we want the error term to vanish in the limit , we need to be superlinear in , i.e. .22 2 Note that, by definition, . Hence, , or in other words, . Therefore, we need to ignore a large fraction of our data. This may be justified in certain scenarios where the number of data is by default huge33 3 For example, random samples of a coin toss, which in principle can be repeated arbitrarily many times., however, it can be prohibitive in other settings.44 4 For example, in the setting of quantum key distribution. A second, but often less drastic, disadvantage is that, by choosing , the error term is vanishing at a slow rate of order .
The insight of [33, 34] was that we can overcome these limitations when relaxing the iid assumption. More precisely, the exponential de Finetti theorem [33, 34] shows that there exist a probability measure on the unit sphere and a family of almost-iid states in with a defect of size such that
| (3) |
The precise statement is given in Theorem 3.1. Almost-iid states in are superpositions of states that are equal to on subsystems and arbitrary on the remaining subsystems. Note that in each element of the superposition the positions of the defects may be different. Usually, the number of defects is sublinear in , i.e. . The precise definition of almost-iid states is given in Section 2. Equation 3 has a drastically better parameter scaling compared to Equation 2. It allows us to choose sublinear in , i.e. , and still have an error term that vanishes in the limit — even at an exponential rate. To see this, note that, for example, by choosing and , we obtain .
While the exponential de Finetti theorem justifies the importance of almost-iid states, it is natural to ask if these states are also relevant in a purely classical scenario. One may define almost-iid distributions as a convex combination of an iid distribution on subsystems and an arbitrary distribution on the remaining subsystems, where the position of the defects can be different in each element of the convex combination. For such distributions, we show that no classical equivalent to Equation 3 exists. In other words, it is not possible to approximate a classical permutation-invariant distribution by a convex combination of almost-iid distributions for sublinear in . This is made precise in Section 3.
The lack of a classical exponential de Finetti theorem shows that almost-iid states are considerably more powerful than classical almost-iid distributions. This is due to the allowed superpositions which can contain long-range entanglement. For that reason, almost-iid states are more complicated to analyze. In particular, it is unclear if certain (robust) applications could behave differently for almost-iid and iid states. However, it is expected that they should not behave differently for practical purposes. Quantum tomography, which is used to infer the characteristics of a system, cannot distinguish between almost-iid and iid states for practically feasible scenarios. We refer to Proposition 2.7 for a mathematically rigorous statement.
Given the operational relevance of almost-iid states (ensured by the exponential de Finetti theorem), it is crucial to understand which applications behave asymptotically equally under almost-iid and iid states. In this work, we introduce the definition of mixed almost-iid states and prove that for conditional entropies, there is no asymptotic difference between almost-iid and iid states. For and a mixed almost-iid state in with a defect of size , we show that
| (4) |
where is the conditional entropy and for is the von Neumann entropy. The exact result is given in Theorem 4.1. To prove our results, we use information-theoretic tools that have been developed to go beyond the traditionally assumed iid structure [33, 17, 36].
We conclude in Section 5 with a discussion of whether popular entanglement measures asymptotically coincide for almost-iid and iid states. We prove that this holds for the squashed entanglement (see Corollary 5.1); however, it remains an open question for other entanglement measures such as the distillable entanglement, the entanglement cost, and the relative entropy of entanglement.
2 Almost-iid states
In this section, we formally define almost-iid states and discuss some of their properties. Historically [33], almost-iid states were introduced as pure and symmetric states with additional structure, as explained in the following. This family of states appears, for example, in exponential de Finetti theorems such as those presented in Equation 3. However, we will relax the definition from [33] to capture more general families of states, including mixed ones, that have a similar almost-iid structure.
Let be the set of permutations on and let denote the set of density matrices on a Hilbert space whose dimension is denoted by . The symmetric subspace is given by . For a fixed , let and
| (5) |
Let such that and . In [33], an -almost-iid state in was defined as a pure state .
However, for pure states outside the symmetric subspace or mixed states, this definition needs to be relaxed in order to capture states that intuitively have an almost-iid structure such as those mentioned in Example 2.2 below. We therefore next introduce a novel definition for mixed almost-iid states.
Definition 2.1 (Almost-iid states).
Let be a Hilbert space, , and such that . Then, is called a -almost-iid state in if there exists a purification of and an extension of such that
- (i)
is permutation-invariant with respect to ;
- (ii)
.
We then write .
Here, denotes the support of a linear operator and is defined as . We next discuss a few pedagogical examples of mixed almost-iid states.
Example 2.2.
A few simple instances of almost-iid states include:
- (a)
Tensor power states . These are almost-iid states with defect size , i.e. .
- (b)
Convex combinations of tensor power states with a small number of defects, i.e., states of the form , where denotes an arbitrary density operator of size representing the defects. We have . To see this, let and be purifications of and , respectively. Then the extension
(6) of clearly satisfies property (i) since it is permutation-invariant. It also satisfies the property (ii) as it can be written as for and for all .
- (c)
The antisymmetric Bell state with is a -almost-iid state in .
We want to emphasize that the class of almost-iid states contains many more families than the ones presented in Example 2.2. In particular, it allows for superpositions of product states with a small number of defects. These superpositions are crucial for an exponential de Finetti theorem as in Equation 3 to hold. On the other hand, classical almost-iid distributions have a simpler structure as they do not have superpositions and hence are just convex combinations of almost-product distributions. This is discussed in Section 3.
Remark 2.3.
A more restrictive definition of mixed almost-iid states was introduced in [9] by calling an -almost-iid state in if there exists a purification such that for some purification of . Such states clearly satisfy the assumptions of Definition 2.1. However, Definition 2.1 is strictly broader as it includes states that do not meet the definition given in [9]. One such example is given by Example 2.2 (c) as explained in Appendix A.
We call a -generalized-almost-iid-state in if it satisfies (ii) but not necessarily (i). We then write . Trivially, we have . With the following trace-preserving completely positive map
| (7) |
which symmetrizes the input, it is possible to convert a generalized-almost-iid-state into an almost-iid state in the sense that
| (8) |
Remark 2.4 (Properties of almost-iid states).
Let be a mixed almost-iid state according to Definition 2.1. Then it satisfies the following properties:
- (a)
- (b)
is permutation-invariant.
- (c)
- (d)
implies for any . See Lemma B.3.
- (e)
implies . This follows directly from the definition of mixed almost-iid states.
- (f)
For any we have for any that . See Lemma B.4.
- (g)
The set of almost-iid states is convex. This follows directly from the definition of mixed almost-iid states.
In the analysis of almost-iid states, the representation in Equation 10 is particularly useful because it exploits the product structure of the space . However, it is sometimes preferable to go to a purified description, i.e., to work with a purification of Equation 10. The following lemma establishes that the underlying structure of persists even in this purified formulation.
Lemma 2.5 (Purification in a preferred basis).
Let for some composite Hilbert space , and let be any subset of vectors such that admits an orthonormal basis with vectors . Then, the following are equivalent:
- (i)
- (ii)
for coefficients and normalized vectors for all .
In addition, if is normalized and satisfies , we have .
Proof.
: By the Schmidt decomposition, there exist orthonormal sets of vectors and , and scalars such that . Taking the partial trace, we obtain
| (11) |
By assumption (i), it follows that for all such that . (If, by contradiction, there were for some with , then we would have . ) Therefore, we may write for some coefficients depending on . With that, we find
| (12) | ||||
| (13) |
where we have introduced coefficients to normalize the vector (with the convention that and are normalized but arbitrary if ).
: For any orthonormal basis of , we simply evaluate
| (14) |
which implies since . In addition, if is normalized, we find since the vectors are orthonormal and the vectors are normalized for each . ∎
An important property of almost-iid states is that when tracing out many subsystems, we obtain a state that is close to an iid state.
Proposition 2.6.
Let such that , , and . Then
| (15) |
Proof.
Let and define . Let be a purification of . By assumption which according to Remark 2.4 implies . Hence, due to the definition of this vector space, there exists a family of orthonormal vectors from such that
| (16) |
with , where is an extension of . Now, we may interpret as a state on the blocks. Since , we know that is at least in blocks of the form and in the remaining blocks arbitrary. For any , let denote the set of block indices on which is not of the form and note that
| (17) |
Then, for any block specified by , the total weight of the vectors that deviate from for that block is given by
| (18) |
Summing over all blocks yields
| (19) |
where the final step uses . Equation 19 implies that
| (20) |
For a fixed , let be the projector onto the subspace spanned by , i.e.
| (21) |
Note that and . The operator is subnormalized, i.e.,
| (22) |
which can be seen, for example, via Hölder’s inequality [35, Proposition 2.5]. The fidelity between two density operators is defined as . Hence, for any we have
| (23) |
Using the permutation invariance of we obtain
| (24) |
where the final step uses the data-processing inequality for the fidelity [21, Lemma B.4]. By definition of the fidelity we have
| (25) |
By definition of we have
| (26) |
The Fuchs-van der Graaf inequality [22] then gives
| (27) | ||||
| (28) | ||||
| (29) |
Recalling that the trace distance is contractive under the partial trace [43, Theorem 8.16] implies
| (30) |
To conclude the proof of the assertion, note that for we have
| (31) |
Hence
| (32) |
where the final step uses . Putting everything together yields
| (33) |
∎
Another property concerns the statistics of almost-iid and iid states when performing an iid measurement (as is used in a tomography procedure, for example) [33, Theorem 4.5.2]. Let and assume that independent measurements with respect to a POVM are performed on iid copies of , giving the outcomes with . The outcomes can be characterized by a frequency distribution (or type) , which is defined as the following probability distribution on ,
| (34) |
for all . Intuitively, corresponds to the relative number of occurrences of in the sequence of outcomes . For large values of , it follows by the law of large numbers that the frequency distribution is close to the probability distribution defined by . Proposition 2.7 shows that a similar statement holds when performing independent measurements on a -almost-iid state in for small enough values of . This proves that almost-iid and iid states cannot be distinguished by any practically feasible measurement. This is discussed in more detail in [29].
Proposition 2.7 (Statistics of almost-iid states).
Let such that , , and . Let be a POVM on , and for all . Then, for any , we have
| (35) |
where the probability is taken over the outcomes of the product measurement applied to , and
| (36) |
where is the binary entropy and . Furthermore, if , then .
The proof is given in Appendix C.
3 Quantum vs. classical exponential de Finetti theorem
In this section, we formally state the quantum exponential de Finetti theorem [34] which justifies the definition of almost-iid states. We then show that for classical almost-iid distributions, defined as convex combinations of product distributions with a small number of defects, no classical exponential de Finetti theorem can hold.
Theorem 3.1 (Exponential de Finetti [34, Theorem 1]).
Let and let be a -dimensional Hilbert space. For any there exists a probability measure on the unit sphere and a family of states such that and
| (37) |
We next define almost-iid distributions which is the classical counterpart to almost-iid states. For a finite alphabet , let denote the set of permutation-invariant distributions on . Furthermore, for a distribution on let
| (38) |
and
| (39) |
Note that the set is considerably simpler compared to since the former set consists of convex combinations of product distributions, whereas the latter set contains superpositions of product states.
The following proposition states that no classical version of the quantum de Finetti theorem can exist where almost-iid states are replaced with almost-iid distributions.
Proposition 3.1 (No classical exponential de Finetti result).
Let , , , be a finite alphabet of size . There exists such that it is not possible to approximate with a probabilistic mixture of distributions with in the -norm up to such that .
The proof of Proposition 3.1 is given in Appendix D. As mentioned in Equation 1, if we are willing to choose large, more precisely , then a classical de Finetti theorem holds. However, Proposition 3.1 states that this is no longer the case for , even if the error term would decrease at a non-exponential rate. Further note that Proposition 3.1 does not prohibit the existence of a classical exponential de Finetti theorem if the definition of almost-iid distributions, Equation 39, is relaxed. For example, one could define the set in Equation 38 by only requiring the marginals of the distributions to yield an iid state instead of imposing a product structure between the iid part and the defects. However, such a structure would no longer be covered by the quantum version in Definition 2.1 and therefore, may potentially be difficult to work with.
One may still wonder what Theorem 3.1 produces if it is applied to a classical symmetric distribution. Can the resulting family of almost-iid states that approximates the classical distribution become classical as well? To analyze this, let be a symmetric classical distribution. We can apply Theorem 3.1 to by embedding it into a permutation-invariant density matrix,
| (40) |
Let be a symmetric purification of . Consider the measurement channel on the -system
| (41) |
Because is classical, we have
| (42) | ||||
| (43) |
where , and the penultimate step uses that only acts on the -system and hence commutes with the partial trace over the -system. Theorem 3.1 shows that there exist a probability measure on the unit sphere and, for each , a family of states such that , and such that for and denoting the induced measure obtained by taking the partial trace, we have
| (44) | ||||
| (45) |
where the second step uses the fact that the trace distance is contractive under trace-preserving completely positive maps [43, Theorem 8.16]. Proposition 3.1 now implies that the density operators cannot be given almost-iid distributions in the sense of Equation 39. Intuitively, this happens since the states appearing in Equation 44 are not necessarily classical and therefore, the almost-iid structure is not preserved after the measurement.
4 Conditional entropy of almost-iid states
In this section, we prove that the conditional entropy of almost-iid states asymptotically coincides with the conditional entropy of iid states. In the proof, we develop technical tools that may be of independent interest.
For the conditional entropy is defined as , where is the von Neumann entropy. It is straightforward to see that the conditional entropy is additive for iid states, i.e.,
| (46) |
We next show that this property is preserved for almost-iid states in the limit .
Theorem 4.1.
Let and for . Then
| (47) |
The main difficulty in proving the assertion of Theorem 4.1 is the fact that almost-iid states are more general than just convex mixtures of product states where each element in the convex sum has a certain number of defects (see Example 2.2). A look at Definition 2.1 reveals that almost-iid states may contain superpositions which store long range correlations and entanglement. Dealing with these superpositions is the main technical challenge in the proof. We do this utilizing tools from one-shot information theory such as Rényi and smooth entropies. We also want to emphasize that the superpositions in the definition of almost-iid states are crucial for making the exponential de Finetti theorem (Theorem 3.1) possible. As shown in Proposition 3.1, no exponential de Finetti theorem can exist without superpositions.
Before presenting the proof of Theorem 4.1, which is given in Section 4.4, we need to define some entropic quantities. We note that an alternative proof using different techniques that may therefore be of independent interest is given in Appendix E.
4.1 Entropic quantities
For the sandwiched Rényi divergence [31, 41] is given by
| (48) |
For we have . In the limits and the sandwiched Rényi divergence converges to the relative entropy and the max-relative entropy [33, 16]
| (49) |
respectively. For , the Rényi divergence can be used to define a conditional Rényi entropy [36]
| (50) |
which converges to the conditional entropy for and to the conditional min-entropy for . For we obtain the conditional max-entropy . The trace distance between two states is given by and the purified distance [36] is defined as . The Fuchs-van der Graaf inequality [22] implies . For and define the -ball around by . We then define a smooth variant of the min- and max-entropy by
| (51) |
4.2 Asymptotic equipartition property for almost-iid states
Another statement which can be proven using similar techniques and may be of independent interest is a strong asymptotic equipartition property (AEP) for almost-iid states. To understand this, let us recall the AEP for iid states [37, 36]. This fundamental result ensures that for any density operator and any we have
| (52) |
Note that Equation 52 is called a strong AEP as the error term vanishes in the limit for any fixed . Furthermore, it is understood how fast the error term vanishes for finite values of [38]. We show that Equation 52 remains valid when replacing iid states with almost-iid states.
Proposition 4.2 (Strong AEP for almost-iid states).
Let , , and for . Then
| (53) |
In [33, Theorem 4.4.1] it was shown that the conditional smooth min-entropy for almost-iid states, which are classical on one subsystem, asymptotically coincides with the conditional entropy of iid states. This was crucial to prove security of quantum key distribution via a de Finetti argument. Using the duality of conditional entropy, the result can be lifted to the smooth max-entropy of almost-iid states [44]. In addition, for the case of pure almost-iid states, a similar result has been proven based on [33] in [44, Lemma 11]. Here, the presented proof of Proposition 4.2 is more general as it applies for mixed almost-iid states and is also more modular allowing one to distill other results such as Theorem 4.1. Beyond these results, to the best of our knowledge, little is known about how entropic functions behave for almost-iid states.
4.3 Proof of Proposition 4.2
Let . By definition, there exist a purification of and an extension of that can be written as
| (54) |
for a family of orthonormal vectors from with satisfying and
| (55) |
Let
| (56) |
Lemma 4.3.
For the setting defined above, we have
| (57) |
Proof.
The proof idea is similar to [33, Proof of Lemma 3.1.13] but is based on pinching maps and therefore works for a more general setup. Consider the pinching map
| (58) |
Using the fact that is orthonormal, we have
| (59) |
Hence, we find
| (60) |
The interested reader can find more information on pinching maps, including a proof of the pinching inequality in [35, Lemma 3.5]. Since the partial trace is a completely positive map Equation 60 implies
| (61) |
∎
Lemma 4.4.
Let such that , , , , and . Then
| (62) |
and
| (63) |
Proof.
We start by proving Equation 62. For and , defined above, we have
| (64) | ||||
| (65) | ||||
| (66) |
where the inequality step used that the function is monotone [11, Theorem 2.10]. Furthermore, we have
| (67) | ||||
| (68) | ||||
| (69) |
where the final step uses that the logarithm is a quasi-linear function and that with . Recalling that for and using the additivity of the Rényi entropies under tensor products allows us to write for any
| (70) |
Putting everything together yields
| (71) | ||||
| (72) |
where in the final step, we used . This proves Equation 62.
The statement from Equation 63 follows similarly. For we have
| (73) | ||||
| (74) | ||||
| (75) |
where the inequality step used that the function is monotone [11, Theorem 2.10]. In addition, we have
| (76) | ||||
| (77) | ||||
| (78) |
where the final step uses that the logarithm is a quasi-linear function and that with . Since for , the additivity of the Rényi entropies under tensor products implies for any
| (79) |
Combining Equations 75, 78 and 79 yields
| (80) | ||||
| (81) |
where in the final step we used that . ∎
We are now equipped with all the tools we need to prove the assertion of Proposition 4.2. This will be done in four steps, by proving two inequalities (direct and converse part) for both the smooth min- and max-entropy.
- (i)
Direct part for smooth min-entropy: For consider the error term
(82) We can write
(83) (84) (85) for a constant . For a choice and recalling that we see that
(86) where we used that and . Hence, we obtain
(87) - (ii)
Direct part for smooth max-entropy: For consider the error term
(88) We can write
(89) (90) (91) Similarly as above, choosing yields and hence
(92) - (iii)
Converse part for smooth min-entropy: For a fixed consider an arbitrary . Then
(93) - (iv)
Converse part for smooth max-entropy: For a fixed consider an arbitrary . Then
(94)
Combining Equations 87, 92, 93 and 94 completes the proof.∎
4.4 Proof of Theorem 4.1
The monotonicity of the Rényi divergence in [31] implies that for we have
| (95) | ||||
| (96) | ||||
| (97) |
for a constant . Choosing and recalling that yields55 5 Note that and .
| (98) |
To see the other direction, note that for with
| (99) | ||||
| (100) | ||||
| (101) | ||||
| (102) | ||||
| (103) |
where the penultimate step uses the continuity of entropy [42]. Combining Equations 98 and 103 completes the proof.
5 Robustness of information measures for almost-iid states
In this work, we justified the importance of almost-iid states. This prompts the question if almost-iid states are as effective as perfect iid states for information-processing tasks. To answer this, it is crucial to understand if certain functionals (that characterize specific information-processing tasks) behave equally or differently for almost-iid and perfect iid states.
In Section 4 we have seen that the conditional entropy is robust for almost-iid states in the sense that it asymptotically coincides with the entropy of iid states. This implies that also the mutual information is robust for almost-iid states. To make this precise, recall that for a bipartite density matrix the mutual information is defined as
| (104) |
The mutual information is a popular correlation measure in the sense that it satisfies (i) , (ii) iff and (iii) .66 6 Properties (i) and (ii) follow by noting that . The third property follows from strong subadditivity together with the chain rule as . Let and for . Then
| (105) | ||||
| (106) | ||||
| (107) |
It is natural to ask if popular entanglement measures are also robust for almost-iid states. In abstract terms, let be an arbitrary entanglement measure. Let and for . Is it true that
| (108) |
In the following, we discuss the robustness of (a) squashed entanglement, (b) entanglement distillation, (c) entanglement cost, and (d) relative entropy of entanglement.
5.1 Robustness of squashed entanglement
Above we have seen that the mutual information is robust under almost-iid states. The same argument can be extended to see that the conditional mutual information also coincides for almost-iid and iid states. The squashed entanglement [14] is an entanglement measure that is based on the conditional mutual information. Given a biparitite density matrix , the squashed entanglement is defined as
| (109) |
where there is no bound on the dimension of . It features many desirable properties such as being additive on tensor products and superadditive in general. We next show that the squashed entanglement for almost-iid and iid states coincide.
Corollary 5.1.
Let and for . Then
| (110) |
Proof.
Let and . We can employ the permutation invariance of and the superadditivity of the squashed entanglement [14, Proposition 4] to write for
| (111) | ||||
| (112) | ||||
| (113) | ||||
| (114) |
where the continuity of squashed entanglement follows from the continuity of the conditional entropy [1] as explained in [14, Section IV].
It thus remains to prove the other direction. For any there exists an extension or with such that
| (115) |
To see this, note that by the definition of the squashed entanglement there exists an extension of (with possibly unbounded -system) such that
| (116) |
Choose a finite-dimensional projector on the -system such that for some . Let
| (117) |
where is a state orthogonal to the support of . By the continuity of the conditional entropy [1] we can choose such that
| (118) |
where we used that the continuity of the conditional entropy does not depend on the dimension of the conditioning system. The triangle inequality implies
| (119) |
which thus justifies Equation 115.
Due to Lemma B.5, there exists an extension of which is an -almost-iid state in . Hence,
| (120) | ||||
| (121) | ||||
| (122) | ||||
| (123) | ||||
| (124) |
Since this holds for any we can consider , which concludes the proof. ∎
5.2 Robustness of entanglement distillation and entanglement cost
Let denote an entangled Bell state. Given a bipartite density matrix , recall the definitions of entanglement distillation [3, 4, 5]
| (125) |
and entanglement cost [24]
| (126) |
Question 5.2.
Let and for . Is it true that
| (127) |
As discussed in [29], there are strong indications that one direction of Equation 127 holds, namely that
| (128) |
Whether the other direction holds also remains an open question.
The equivalent question for the entanglement cost asks:
Question 5.3.
Let and for . Is it true that
| (129) |
However, none of the two directions of Equation 129 are known to hold.
At this point, we emphasize that, unlike squashed entanglement, already the definitions of both entanglement cost and entanglement distillation rely on a tensor power (iid) structure. This may suggest that, rather than asking about the robustness of Equations 125 and 126, one should incorporate robustness directly into the definition itself. One may therefore wonder how these notions would change if an almost-iid structure were built into the definition from the outset. In [29], the authors investigate this question by introducing new asymptotic state transformation rates that avoid the standard iid assumption. We refer the interested reader to that paper for further details.
5.3 Robustness of relative entropy of entanglement
A popular measure to quantify the amount of entanglement is the relative entropy of entanglement [39] defined as
| (130) |
where . It is known [40] that the relative entropy of entanglement is not additive under the tensor product, which justifies the definition of a regularized version . The limit in the regularization exists due to Fekete’s subadditivity lemma.
Question 5.4.
Let and for . Is it true that
| (131) |
If Equation 131 were true, this would save the original proof of the generalized quantum Stein’s lemma [23, 26] by Brandão and Plenio [9] (see also [6, 7]). One direction of Equation 131 follows from the results developed in this paper. To see this, choose a monotonically increasing sequence of integers such that and . Let . Note that . Hence, using the monotonicity of the entanglement of formation under partial trace,
| (132) | ||||
| (133) |
where the final step uses and . In the following steps, we omit the subscripts for better readability. Let . Then for we have
| (134) | ||||
| (135) | ||||
| (136) | ||||
| (137) | ||||
| (138) | ||||
| (139) | ||||
| (140) | ||||
| (141) | ||||
| (142) |
where the continuity of the von Neumann entropy and the relative entropy of entanglement can be found in [2, 32, 42]. In the last step, we use that and that the limit in the RHS of Equation 131 exists due to Fekete’s subadditivity lemma.
The other direction of Equation 131 appears more complicated and remains an open question.
Acknowledgements
We thank Fernando Brandão for his talk on almost-iid states at the SwissMAP Research Station (SRS) conference in Les Diablerets 2024, which motivated us to write this paper. We further thank Frédéric Dupuis and Ludovico Lami for insightful discussions on this topic at the same conference. GM and RR acknowledge support from the NCCR SwissMAP, the ETH Zurich Quantum Center, the SNSF project No. 20QU-1 225171, and the CHIST-ERA project MoDIC.
Appendix
Appendix A Justification of Remark 2.3
To justify the assertion of Remark 2.3, we need to show that for all purifications of and for all purifications of it follows that
To see this note that, because is pure, . Similarly, because is pure, any purification on must be of the form . To ensure that is an almost-iid state according to the definition from [9], we need . Thus, the swap operation must satisfy
| (143) |
This yields , i.e., is anti-symmetric. Expressing this in an orthonormal basis of with gives with for all . Hence,
| (144) |
which cannot be inside . Thus, the state is not an almost-iid state according to the definition from [9].
Appendix B Properties of almost-iid states
Lemma B.1.
Let with purification of according to Definition 2.1. Then, any other purification of would also satisfy Definition 2.1.
Proof.
Without loss of generality, assume . This can be done since the rank of the reduced state of any purification on the purifying system is bounded by the rank of . Since all purifications are then equal up to unitaries on the purifying system, we have for some unitary on . For being the extension of according to Definition 2.1 we define
| (145) |
Clearly is an extension of , i.e., and . Thus, it remains to show that this extension satisfies the two properties in Definition 2.1.
First, we note that is permutation-invariant. To see this, let be a permutation that swaps . Then
| (146) | ||||
| (147) | ||||
| (148) |
Second, we have
| (149) | ||||
| (150) |
Note that . ∎
Lemma B.2.
There exists an orthonormal basis of with vectors for all , and with
| (151) |
where , and for is the binary entropy function.
Proof.
Let denote an orthonormal basis of such that for some . Any vector can be expanded as
| (152) |
for coefficients and vectors . By definition, for some permutation and vector . Since is a basis of , we can write
| (153) |
for all , where is an -tuple of elements from , , and . Then
| (154) |
with , , and . Intuitively, is labeling the positions of the defects, and labels the state at these defected positions. Note that the vectors are normalized for all , and, for different values of , they are either pairwise orthogonal, or they are equal (e.g. if ). Thus, restricting to a maximal set of pairwise orthogonal vectors , Equation 154 shows that this set forms an orthonormal basis of .
To determine the size 77 7 More precisely, one finds . Since for all , it follows that ., note that for any vector in there are possible combinations for the positions of the defects, and at each such position, the dimension of the Hilbert space is given by . Hence, we get the upper bound stated in Equation 151, where the first bound would correspond to the case where the defects are different from . The second inequality then follows from the relation , e.g. see [15, Example 11.1.3]. ∎
Lemma B.3.
implies for any .
Proof.
Since , there exists an extension that is permutation invariant. We first show that is then also permutation-invariant. To see this, let denote an arbitrary permutation. Then
| (155) | ||||
| (156) | ||||
| (157) | ||||
| (158) |
Thus, it remains to show that satisfies Property (ii) in Definition 2.1. By definition is such that for a purification of . We need to show that . To see this, let be the “standard” tensor product basis of , i.e., for a basis of we have for some permutation . We can observe that
| (159) |
and
| (160) |
where denotes the standard basis of . Note that these basis elements in Equation 160 are independent of and . Hence, we find
| (161) | ||||
| (162) | ||||
| (163) |
where the final step uses . This shows that and hence completes the proof. ∎
Lemma B.4.
Let . Then for any we have .
Proof.
Let denote a purification of . Since there exists an extension that can be written as
| (164) |
for . Furthermore, by Equation 164 we have
| (165) |
where . This shows that , which completes the proof. ∎
Lemma B.5.
Let , , , and . Then, there exists an extension of that is -almost-iid in , i.e. .
Proof.
This is essentially a consequence of Lemma B.1. Formally, there exist a purification of and an extension of such that , by definition. Now fix a purification of such that . Since is also a purification of , there exists an isometry such that . Then we define
| (166) |
and
| (167) |
Clearly, is an extension of . Furthermore, this extension satisfies the two properties in Definition 2.1, namely, it is permutation invariant and , which can be shown by following the same arguments as in the proof of Lemma B.1. On the other hand, is also an extension of , which means that is -almost-iid in . Since is an extension of , the claim follows. ∎
Appendix C Proof of Proposition 2.7
We generalize the statement for pure almost-iid states from [33, Theorem 4.5.2] to the more general setting of mixed almost-iid states. The proof works similarly to the one from [33, Theorem 4.5.2] with a few modifications.
Let be a purification of , where denotes the purifying system of dimension . Since is an almost-iid state, there exists an extension which can be written as
| (168) |
where is an orthonormal basis of with vectors . For any fixed we can assume without loss of generality that , where represents the defects. Let be the outcomes of the measurement applied to . Furthermore, let and . Clearly is distributed according to the product distribution . Hence, for any we have
| (169) |
Using this can be simplified to
| (170) |
Using yields
| (171) |
Hence,
| (172) | ||||
| (173) |
This can be rewritten as
| (174) |
for
| (175) |
The notation indicates that is distributed according to the outcomes of the measurement applied to .
For we find
| (176) | ||||
| (177) | ||||
| (178) | ||||
| (179) | ||||
| (180) |
Choosing yields
| (181) |
which completes the proof. ∎
Appendix D Proof of Proposition 3.1
Consider a -bit string with ones and zeros. If we throw away bits, then the probability of having ones in the remaining -bit string is
| (182) |
Note that follows from a known identity due to Vandermonde which states that for all we have
| (183) |
To see Equation 182, let denote the number of ones that have been thrown away. Hence,
| (184) |
for some normalization constant
| (185) |
Fact D.1.
For the setting above, we have
| (186) |
Proof.
Recall that
| (187) |
With this we can write
| (188) | ||||
| (189) | ||||
| (190) | ||||
| (191) | ||||
| (192) | ||||
| (193) |
Similarly, we find
| (194) | ||||
| (195) | ||||
| (196) | ||||
| (197) | ||||
| (198) | ||||
| (199) | ||||
| (200) | ||||
| (201) |
Combining everything yields
| (202) |
∎
Let and consider a binary random string . Then without loss of generality assume that the defects are at the end of the random string, and hence
| (203) |
where the first equality uses that the defects are independent of iid parts.
Example D.2.
Let , , and . Then Fact D.1 gives . Furthermore, Equation 203 yields .
Fact D.3.
Let and be two probability distributions. Then,
| (204) |
Proof.
By definition of the variance, we have
| (205) | ||||
| (206) | ||||
| (207) | ||||
| (208) |
∎
Putting everything together, we obtain for
| (209) |
This proves the assertion of Proposition 3.1. Note that for better readability, the above proof has been done for the specific choice for , but the same argument remains valid for an arbitrary . ∎
Appendix E Alternative proof of Theorem 4.1
In this section, we present an alternative proof for Theorem 4.1 which uses an entirely different proof technique which may be of independent interest. However, we note that the scaling of the defects is slightly worse than in Theorem 4.1. Note that it suffices to prove the following result.
Theorem E.1.
Let and for . Then
| (210) |
Theorem E.1 implies that the conditional entropy of almost-iid states coincides asymptotically with the conditional entropy of iid states.
Corollary E.2.
Let and for . Then
| (211) |
Proof.
By definition of the conditional entropy, we have
∎
E.1 Proof of Theorem E.1
One direction of Equation 210 is simple. To see this, let for and consider
| (212) |
where the penultimate step uses the continuity of entropy [2, 32, 42].
The other direction is more complicated. Recall that strong subadditivity of quantum entropy (SSA) [27, 28] ensures . Furthermore, the conditional mutual information satisfies a chain rule . Let be a purification of and let be an extension of that satisfies the two conditions of Definition 2.1. Since is permutation-invariant, for entropy and mutual information terms the indices of the considered subsystems can be changed. Hence, we find for any , and for any ,
| (213) | ||||
| (214) | ||||
| (215) | ||||
| (216) | ||||
| (217) | ||||
| (218) |
where in the last step, we permute all the remaining systems appearing in the conditioning (there are many) into neighboring systems labeled by indices from to .
Thus, for any we find
| (219) | ||||
| (220) | ||||
| (221) |
In the proof of Lemma B.3 it is shown that . This implies
| (222) |
Similarly, we obtain
| (223) |
Hence we find
| (224) | ||||
| (225) |
For any , we thus have
| (226) |
where the first step follows since the trace distance is contractive under trace-preserving completely positive maps [43, Theorem 8.16]. The continuity of conditional entropy [1] then implies that
| (227) |
where . The chain rule together with permutation invariance allows us to write
| (228) | ||||
| (229) | ||||
| (230) | ||||
| (231) | ||||
| (232) | ||||
| (233) | ||||
| (234) | ||||
| (235) |
Claim E.3.
Let , exists such that for we have .
Proof.
The proof follows the idea from [8, Equation (5)], which proves a variant of Claim E.3 for a purely classical scenario. By the permutation-invariance we have for any
| (236) |
This allows us to write
| (237) |
Summing Equation 237 over all and dividing by gives
| (238) |
Hence, there exists such that
| (239) |
Choosing , Equation 239 can be rewritten (as there exists such that ) as
| (240) |
We next show that Equation 225 implies
| (241) |
To see this, we can assume w.l.o.g.88 8 It can be shown that for all . Hence, if , due to the chain rule we can write , and the argument above still works. that is large enough such that
| (242) |
Let us now consider two cases: In case , the assertion follows from Equation 225 by choosing , since
| (243) |
In case , we can use the chain rule to write
| (244) | ||||
| (245) |
where both terms on the right-hand side are bounded from above by via Equation 225.
Putting things together yields
| (246) | ||||
| (247) | ||||
| (248) | ||||
| (249) |
where the final step uses that for and we have
| (250) |
∎
References
- [1] R. Alicki and M. Fannes. Continuity of quantum conditional information. Journal of Physics A: Mathematical and General, 37(5):55–57, 2004. DOI: doi:10.1088/0305-4470/37/5/L01.
- [2] K. M. R. Audenaert. A sharp continuity estimate for the von Neumann entropy. Journal of Physics A: Mathematical and Theoretical, 40(28):8127, 2007. Available online: http://stacks.iop.org/1751-8121/40/i=28/a=S18.
- [3] C. H. Bennett, H. J. Bernstein, S. Popescu, and B. Schumacher. Concentrating partial entanglement by local operations. Phys. Rev. A, 53:2046–2052, 1996. DOI: 10.1103/PhysRevA.53.2046.
- [4] C. H. Bennett, G. Brassard, S. Popescu, B. Schumacher, J. A. Smolin, and W. K. Wootters. Purification of noisy entanglement and faithful teleportation via noisy channels. Phys. Rev. Lett., 76:722–725, 1996. DOI: 10.1103/PhysRevLett.76.722.
- [5] C. H. Bennett, D. P. DiVincenzo, J. A. Smolin, and W. K. Wootters. Mixed-state entanglement and quantum error correction. Physical Review A, 54(5):3824–3851, 1996. DOI: 10.1103/PhysRevA.54.3824.
- [6] M. Berta, F. G. S. L. Brandão, G. Gour, L. Lami, M. B. Plenio, B. Regula, and M. Tomamichel. On a gap in the proof of the generalised quantum Stein’s lemma and its consequences for the reversibility of quantum resources. Quantum, 7:1103, 2023. DOI: 10.22331/q-2023-09-07-1103.
- [7] M. Berta, F. G. S. L. Brandão, G. Gour, L. Lami, M. B. Plenio, B. Regula, and M. Tomamichel. The tangled state of quantum hypothesis testing. Nature Physics, 20(2):172–175, 2024. DOI: 10.1038/s41567-023-02289-9.
- [8] M. Berta, L. Gavalakis, and I. Kontoyiannis. A third information-theoretic approach to finite de Finetti theorems, 2023. DOI: 10.48550/arXiv.2304.05360.
- [9] F. G. S. L. Brandão and M. B. Plenio. A generalization of quantum Stein’s lemma. Communications in Mathematical Physics, 295(3):791–828, 2010. DOI: 10.1007/s00220-010-1005-z.
- [10] F. Buscemi, D. Sutter, and M. Tomamichel. An information-theoretic treatment of quantum dichotomies. Quantum, 3:209, 2019. DOI: 10.22331/q-2019-12-09-209.
- [11] E. Carlen. Trace Inequalities and Quantum Entropy: An Introductory Course. Contemporary Mathematics, 2009. DOI: 10.1090/conm/529.
- [12] C. M. Caves, C. A. Fuchs, and R. Schack. Unknown quantum states: The quantum de Finetti representation. Journal of Mathematical Physics, 43(9):4537–4559, 2002. DOI: 10.1063/1.1494475.
- [13] M. Christandl, R. König, G. Mitchison, and R. Renner. One-and-a-half quantum de Finetti theorems. Communications in Mathematical Physics, 273(2):473–498, 2007. DOI: 10.1007/s00220-007-0189-3.
- [14] M. Christandl and A. Winter. “Squashed entanglement”: An additive entanglement measure. Journal of Mathematical Physics, 45(3):829–840, 2004. DOI: 10.1063/1.1643788.
- [15] T. M. Cover and J. A. Thomas. Elements of Information Theory. Wiley Interscience, 2006. DOI: 10.1002/047174882X.
- [16] N. Datta. Min- and max-relative entropies and a new entanglement monotone. IEEE Transactions on Information Theory, 55(6):2816–2826, 2009. DOI: 10.1109/TIT.2009.2018325.
- [17] N. Datta and R. Renner. Smooth entropies and the quantum information spectrum. IEEE Transactions on Information Theory, 55(6):2807–2815, 2009. DOI: 10.1109/TIT.2009.2018340.
- [18] B. De Finetti. La prévision: ses lois logiques, ses sources subjectives. In Annales de l’institut Henri Poincaré, volume 7, pages 1–68, 1937.
- [19] B. de Finetti. Logical foundations and measurement of subjective probability. Acta Psychologica, 34:129–145, 1970. DOI: https://doi.org/10.1016/0001-6918(70)90012-0.
- [20] P. Diaconis and D. Freedman. Finite Exchangeable Sequences. The Annals of Probability, 8(4):745 – 764, 1980. DOI: 10.1214/aop/1176994663.
- [21] O. Fawzi and R. Renner. Quantum conditional mutual information and approximate Markov chains. Communications in Mathematical Physics, 340(2):575–611, 2015. DOI: 10.1007/s00220-015-2466-x.
- [22] C. Fuchs and J. van de Graaf. Cryptographic distinguishability measures for quantum-mechanical states. IEEE Transactions on Information Theory, 45(4):1216 –1227, 1999. DOI: 10.1109/18.761271.
- [23] M. Hayashi and H. Yamasaki. The generalized quantum Stein’s lemma and the second law of quantum resource theories. Nature Physics, 21(12):1988–1993, 2025. DOI: 10.1038/s41567-025-03047-9.
- [24] P. M. Hayden, M. Horodecki, and B. M. Terhal. The asymptotic entanglement cost of preparing a quantum state. Journal of Physics A: Mathematical and General, 34(35):6891, 2001. DOI: 10.1088/0305-4470/34/35/314.
- [25] R. König and R. Renner. A de Finetti representation for finite symmetric quantum states. Journal of Mathematical Physics, 46(12):122108, 2005. DOI: 10.1063/1.2146188.
- [26] L. Lami. A solution of the generalized quantum Stein’s lemma. IEEE Transactions on Information Theory, 71(6):4454–4484, 2025. DOI: 10.1109/TIT.2025.3543610.
- [27] E. H. Lieb and M. B. Ruskai. A fundamental property of quantum-mechanical entropy. Physical Review Letters, 30:434–436, 1973. DOI: 10.1103/PhysRevLett.30.434.
- [28] E. H. Lieb and M. B. Ruskai. Proof of the strong subadditivity of quantum-mechanical entropy. Journal of Mathematical Physics, 14(12):1938–1941, 1973. DOI: 10.1063/1.1666274.
- [29] G. Mazzola and R. Renner. Asymptotic transformation rates with almost iid resources, 2026. in preparation.
- [30] P. Monari and D. Cocchi. Introduction to Bruno de Finetti’s “probabiliá e induzione”. Cooperativa Libraria Universitaria Editrice, Bologna, 1993.
- [31] M. Müller-Lennert, F. Dupuis, O. Szehr, S. Fehr, and M. Tomamichel. On quantum Rényi entropies: A new generalization and some properties. Journal of Mathematical Physics, 54(12), 2013. DOI: http://dx.doi.org/10.1063/1.4838856.
- [32] D. Petz. Quantum Information Theory and Quantum Statistics. Springer, 2008. DOI: 10.1007/978-3-540-74636-2.
- [33] R. Renner. Security of quantum key distribution. PhD thesis, ETH Zurich, 2005. available at arXiv:quant-ph/0512258.
- [34] R. Renner. Symmetry of large physical systems implies independence of subsystems. Nature Physics, 3(9):pp. 645–649, 2007. Available online: http://www.nature.com/nphys/journal/v3/n9/suppinfo/nphys684_S1.html.
- [35] D. Sutter. Approximate Quantum Markov Chains. Springer International Publishing, 2018. DOI: 10.1007/978-3-319-78732-9_5.
- [36] M. Tomamichel. Quantum Information Processing with Finite Resources, volume 5 of SpringerBriefs in Mathematical Physics. Springer, 2015. DOI: 10.1007/978-3-319-21891-5.
- [37] M. Tomamichel, R. Colbeck, and R. Renner. A fully quantum asymptotic equipartition property. IEEE Transactions on Information Theory, 55(12):5840–5847, 2009. DOI: 10.1109/TIT.2009.2032797.
- [38] M. Tomamichel and M. Hayashi. A hierarchy of information quantities for finite block length analysis of quantum tasks. IEEE Transactions on Information Theory, 59(11):7693–7710, 2013. DOI: 10.1109/TIT.2013.2276628.
- [39] V. Vedral, M. B. Plenio, M. A. Rippin, and P. L. Knight. Quantifying entanglement. Phys. Rev. Lett., 78:2275–2279, 1997. DOI: 10.1103/PhysRevLett.78.2275.
- [40] K. G. H. Vollbrecht and R. F. Werner. Entanglement measures under symmetry. Phys. Rev. A, 64:062307, 2001. DOI: 10.1103/PhysRevA.64.062307.
- [41] M. M. Wilde, A. Winter, and D. Yang. Strong converse for the classical capacity of entanglement-breaking and Hadamard channels via a sandwiched Rényi relative entropy. Communications in Mathematical Physics, 331(2):593–622, 2014. DOI: 10.1007/s00220-014-2122-x.
- [42] A. Winter. Tight uniform continuity bounds for quantum entropies: Conditional entropy, relative entropy distance and energy constraints. Communications in Mathematical Physics, 347(1):291–313, 2016. DOI: 10.1007/s00220-016-2609-8.
- [43] M. M. Wolf. Quantum channels & operations: Guided tour, 2012. Lecture notes available at https://www-m5.ma.tum.de/foswiki/pub/M5/Allgemeines/MichaelWolf/QChannelLecture.pdf.
- [44] Y.-D. Wu and G. Chiribella. Detecting quantum capacities of continuous-variable quantum channels. Phys. Rev. Res., 4:043149, 2022. DOI: 10.1103/PhysRevResearch.4.043149.