Shapes and Norms of Random Pairs
Abstract
The shape function of a pair of finite-valued random variables was introduced in [MR26], where it was used to derive a spectral bound on the entanglement of the pair, a quantity measuring the extent to which their mutual information can be extracted. In this article, we further develop the theory of shape functions for pairs of random variables. We prove that, when is uniformly supported on the edges of a biregular bipartite graph, the value of the shape function equals the logarithm of the operator norm of the graph’s incidence matrix with respect to Lebesgue exponents determined by .This identification, in particular, enables the numerical approximation of the shape function and, by duality, of the extension profile, also known as the tension region, of the pair. We also establish a collection of relations and inequalities satisfied by shape functions, including convexity and monotonicity properties, composition inequalities, and relations describing their behavior under conditioning and the adjoining of variables.
1 Shape function and its relation to the norms of incidence operators
The shape function of a pair of jointly distributed random variables , introduced in [MR26], is a function on defined by
where supremum is taken over all extensions . Function is convex on the square and it is affine on the upper-right triangle
The affine part is determined by the entropy profile of the pair, that is entropies of the two variables and their joint. In general, the restriction of to the lower-left triangle is not determined by the entropy profile and contains additional information about the pair. In particular, the entanglement of the pair – the quantity that measures to what extend mutual information is extractable and defined by
can be recovered from , see [MR26, MR26a] and Equation (MMRV) on page MMRV of the present article.
If the pair is uniform on its support, that is, each of the random variables , and their joint are uniform on their respective supports, then it is completely determined by the supporting graph , where
Such a bipartite graph, which is necessarily biregular, can be represented by its bipartite incidence matrix of size .
In this article we evaluate the shape function of a pair uniform on its support in terms of the incidence matrix of the supporting graph. We view as a matrix, with respect to standard bases, of the bilinear form
where and where we use infix notation for evaluation of the form. We define its -norm by
The main result of this article is the following theorem.
Theorem A.
Let be a pair of random variables uniformly supported on a homogeneous bipartite graph with the incidence form . Then for all holds
where and .
By Tropical Asymptotic Equipartition Property, [MP18], a pair obtained by taking independent copies of a pair , can be arbitrarily well approximated on the normalized scale by a pair of random variables uniformly supported on a homogeneous graph. Since the shape is stable, that is,
and continuous with respect to such approximations, Theorem A allows one, in principle, to evaluate shape function for an arbitrary, not necessarily uniformly supported, pair.
Theorem A also gives a practical way to approximate the shape function. Indeed, the norm of the bilinear form is the same as the operator norm , where is the Hölder conjugate of . Lebesgue operator norms can be approximated numerically by robust methods. Thus, for pairs uniformly supported on biregular bipartite graphs, one can numerically approximate both the shape function and, by duality, the extension profile.
In the next section we introduce our notation and conventions. Section 3 we recall some facts about norms on tensor product of normed vector spaces and prove multiplicativity of Lebesgue operator norms in the hypercontractive regime, Proposition 3.3.A. Further we prove our main technical tool, Theorem 3.4.C, that asserts that for high tensor powers of bilinear forms there are uniform vectors which are almost in resonance. In Section 4 we apply the results of the previous sections to bilinear forms associated to bipartite graphs. Theorem A is proven in Section 5. We also derive inequalities satisfied by the shape function in Section 6. Especially interesting is the inequality (COMP) on page COMP. We do not know whether this inequality can be derived in the purely entropic context, that is without using Theorem A.
2 Notation, conventions, recollections
2.1 Notation
For a natural number we denote . We write or for the cardinality of a finite set . Denote the -cardinality of a set by . For a finite set , we write for the vector space of real-valued functions on . For a subset and a point we denote by the indicator function of and by the point mass.
For an extended real we denote by its Hölder conjugate, that is the two extended reals must satisfy .
We use logarithms with the natural base throughout the article and denotes Euler’s number.
2.2 Random variables
All random variables in this article have finite alphabets. As a notational convention, use capitals for random variables and the corresponding san-serif letters for their alphabets. For a random variable and we use . We call an atom of if .The support is the collection of all atoms.
We tacitly assume that the alphabets of different random variables are disjoint and we write for the conditional random variable in lieu of .
Tuples of random variables are written as comma-separated lists, such as, for example, or , while joints of random variables, regarded as a single random variable, where marginalization structures are ignored, are denoted by concatenation of the corresponding letters, such as or , or by using the subset for the subscript, as in for the joint of , . For example, for a triple of random variables, , the notation stands for a pair of variables consisting of variable and the joint variable . This pair is different from the pair . Given a pair , a third random variable jointly distributed with forms an extension of the pair and is called an extending variable.
Supports of the variables in the pair form a bipartite graph , where
We say that is supported on and write
A tuple of random variables is called uniform on its support if all partial joints , are uniform on their respective supports. If the pair is uniform on its support, then the supporting graph is biregular and its combinatorial structure completely determines the pair. In that case we say that uniformly supported on .
2.3 Graphs
All graphs considered in this article have no isolated vertices. A bipartite graph is called biregular if the degrees of vertices are constant within each part. We denote by the left and right degrees of , respectively. The automorphism group of is the group of symmetries of preserving each part. Graph is called homogeneous if acts transitively on the edge set . Homogeneous graphs are biregular. By a subgraph we always mean a nonempty subgraph without isolated vertices. We denote by , and the left, right parts and edge-set of , respectively. We also set
For a pair uniformly supported on a graph the following identities hold:
2.4 Extension profile and shape function
The entropy profile of a pair is a vector
For an extension the conditional entropy profile is
The extension profile of a pair is the set of all conditional entropy profiles for all extensions of the pair.
It is closely related to the tension region; see [PP14, LE17, Csi23]. One may think of the extending variable as a probe testing finer properties of the relation between variables and — those that are not already reflected in the entropy profile of the pair.
The extension profile of any pair is a convex compact subset of . In [MR26] a dual object, the so called shape function, or simply shape is introduced and studied. It is the function on the square in the -plane defined by
where the later supremum is over all extending variables .
For the basic properties of , we refer the reader to [MR26]. There, the shape function is studied in detail, including its upper and lower bounds and, for pairs uniformly supported on graphs, its relation to spectral properties of the supporting graph. In this article we establish some additional properties of the shape.
3 Norms on tensor products
Here we recall several standard facts about tensor products of normed vector spaces. We restrict attention to finite-dimensional spaces, although many of the constructions discussed below extend to general Banach spaces with the usual additional analytic care. Standard references for tensor products of Banach spaces include [DF93, Rya02].
The main results of this section are Theorem 3.4.C and Corollary 3.4.D. The main technical tool, Proposition 3.3.A, is, most likely, known to the specialists, but for the lack of a suitable reference we provide a proof.
3.1 Tensor products
We write for an algebraic isomorphism of vector spaces, and for an isomorphism of Banach spaces.
We use several natural algebraic isomorphisms between tensor products of finite-dimensional vector spaces. Since these isomorphisms are canonical, we shall use them implicitly and suppress them from the notation:
| (commutativity) | ||||
| (associativity) | ||||
| () | ||||
| () |
For example, for our purposes a linear map , its adjoint , the bilinear form
and the tensor in are regarded as different realizations of the same object, denoted by .
We shall also use the following algebraic identifications for spaces of functions on finite sets:
3.2 Cross norms
3.2.A Algebraic definition of cross norms
Let and be Banach spaces. In this article, we call a norm on the algebraic tensor product a cross norm11 1 Sometimes in the literature a cross norm is defined as a norm on the tensor product satisfying only the first condition in (3.2.B), while a norm satisfying both conditions is called a reasonable cross norm. if, for all , , , and , one has
| (3.2.B) | ||||
where stands for the dual norm on , etc.
3.2.C Geometric definition of cross norms
For the geometrically oriented reader, cross norms admit the following equivalent description. For a Banach space , denote by the -unit sphere. The projectivized product
is naturally contained in the tensor product .
A norm on is a cross norm if and only if
3.2.D Injective and projective cross norms
Here we consider two notable examples of cross norms: the least/injective and the greatest/projective cross norms.22 2 These two norms are often denoted by and , respectively. We use different notation to emphasize the dependence of the injective and projective tensor norms on and .
| (injective) | ||||
| (projective) |
where .
The operations and are dual to each other in the sense that
The projective norm can be geometrically defined by the following. We declare the -unit ball to be the smallest convex set containing the projectivized product of the unit balls in and , that is
where stands for the closed unit ball in a Banach space , etc. By duality there is a similar description of the injective norm .
Note that for an operator (which can be considered as an element of ) the operator norm is
It is known and follows directly from the geometric description of the projective and injective norms that a norm on is a cross norm if and only if
3.2.E Cross norms for Lebesgue spaces
For a Banach space , finite set and denote by
the Banach space of -indexed families of vectors in with the norm
with the usual provision for the case . We will write .
An important example for our purposes is
with the usual modifications in the cases where or is infinite.
It is straightforward to establish, that for finite
where are the Hölder conjugates of , respectively. There are canonical algebraic isomorphisms independent of
To avoid ambiguity, we denote the norm on by
Another pair of isomorphisms, that are used in what follows, are
We note that, under the algebraic identification above the Lebesgue -norm on is a cross norm with respect to the tensor product decomposition . This follows directly from the definitions and duality .
3.3 Multiplicative family of norms for bilinear forms
Let be finite sets. For a bilinear form define
where we use infix notation for the value of the form on vectors , . If we view as an element from or as the operator , then
The main technical result, Proposition 3.3.A, asserts that in hypercontractive regime, , the norms form a multiplicative family for bilinear forms on , where range over finite sets.
3.3.A.
Let satisfy and let be finite sets. For two bilinear forms
consider their tensor product
Then
To prove the proposition we shall use the following lemma, a version of the so called Minkowski’s integral inequality, [HLP52], that says that under canonical algebraic isomorphism norms and are comparable.
3.3.B.
Let be finite sets. If , then
We give a proof of the lemma in Section 3.3.E.
3.3.C Proof of Proposition 3.3.A
By renormalization, it is enough to prove that
First we prove the lower bound. Indeed, since the Lebesgue norms are cross norms,
Hence, by testing against decomposable vectors, we have
Now we prove the upper bound. Let
be such that and .
For each and , set
Since , the associated operator
has norm . Hence
and
Since
we can apply Lemma 3.3.B:
| (3.3.D) |
For each , define
Then (3.3.D) says that
Therefore, using , we get (with infix notation)
Here the last inequality uses Hölder’s inequality along -index and the fact that
Thus , and the proof is complete. ∎
3.3.E Proof of Lemma 3.3.B
The lemma is immediate when . Indeed,
By duality, this also implies the case and any , that is for any holds
For the general case, we shall use the elementary identity
valid for every . In particular, when , we may take . Applying the case to and to the exponent , we obtain
This proves the lemma in the general case. ∎
3.4 Resonance vectors for tensor powers of bilinear forms
3.4.A Resonance vectors for bilinear form. Uniform vectors
Consider a bilinear form . We say that a pair of non-zero vectors and is a resonance pair or equivalently a maximizing pair for the form if (with infix notation for the value of bilinear forms)
Essentially, a resonance pair consists of the two optimizers in the definition of the norm of the bilinear form . Clearly any positive multiples of the vectors in a resonance pair also form a resonance pair. In finite dimensions resonance vectors always exist. If coefficients of with respect to natural bases are non-negative then there exist pair of resonance vectors, which also have non-negative coefficients.
We call a nonzero vector uniform if it is a positive scalar multiple of the indicator function of a nonempty subset of .
We show below that for large tensor powers of a coordinate-wise non-negative bilinear form , there exists a pair of uniform “almost” resonance vectors. Here adverb “almost” should be understood on the normalized -scale, see Theorem 3.4.C below.
3.4.B Uniform almost resonance pairs for tensor powers
Throughout this section we fix satisfying , two finite sets and and a bilinear form of norm one and with non-negative coefficients with respect to the standard bases in and . We only provide proofs for finite . The results of this section also hold when one of the exponents is infinite. While proofs in the endpoint cases are different, they are much simpler and are left to the reader.
3.4.C.
Let satisfying and let
be a coordinate-wise non-negative bilinear form of norm one. Then there exist sequences of uniform vectors and , such that
For an arbitrary coordinate-wise non-negative bilinear form, the theorem gives the following corollary by rescaling.
3.4.D.
Let satisfy , and let be a non-zero bilinear form with non-negative coefficients with respect to the standard bases in , , where . Then there exist sequences of uniform vectors and , such that
By renormalization we can choose and for suitable subsets and . Then
3.4.E.
Consider the following extremal problem similar to the definition of the norm of a bilinear form
Clearly we have . In effect, Corollary 3.4.D states that on the normalized -scale, the usual norm and the restricted subnorm defined above are asymptotically the same.
If necessary we may assume that uniform vectors and provided by Corollary 3.4.D are optimizers in the definition of the restricted subnorm .
3.4.F Proof of Theorem 3.4.C
Some recollections.
We start by recalling several notions needed in the proof of Theorem 3.4.C. For a finite set denote by the set of probability distributions on . Define the empirical map
For the value of the distribution at point is
Let be a distribution on . Define the divergence ball of radius around as
where stands for the Kullback–Leibler divergence. Denote by the complement of the divergence ball. Note that the support of every distributions in is contained in the support of .
The proof.
We treat the case , leaving the cases of or infinite to the reader, as they are much simpler.
Uniform asymptotically resonance vectors and are constructed in three steps. We start with the “true” resonance pair of unit vectors and for . By Proposition 3.3.A their tensor powers and form a resonance pair for . In the first step we restrict and on suitable typical subsets in and , respectively. The resulting vectors and are almost unit and quasi-uniform. In the second step we replace and by unit uniform vectors and with the same respective supports. In the third step we combine the bounds in the previous steps to derive the proposition.
Step 1.
Let and be unit, coordinate-wise non-negative, resonance pair for . That is
We use to construct of unit norm and with reduced support which forms any asymptotically resonance pair with . The construction of vector goes along the similar lines.
Since
we regard as a probability distribution on . Consider the sequence of divergence balls
of radius around , and denote their complements.
Let be the preimage of the divergence ball under the empirical map and let be its complement. Define
Here is the typical part and is the atypical remainder. The supports of and are complimentary and we have
By Equation (3.4.G) we have
for some and all sufficiently large . Thus
Define .
In a similar fashion, write
We then similarly decompose into typical part and atypical remainder. Then
for sufficiently large and some . Define .
Clearly and are unit vectors in and , respectively, and for sufficiently large we have
| (3.4.I) | ||||
Step 2.
We will now estimate the quasi-uniformity constant of and show that it grows subexponentially in . For any Equation (3.4.H) gives
First we note that for sufficiently large , the divergence ball around of radius lies in an open face of the simplex — the one that corresponds to the support of . From now on, we assume that is sufficiently large for this assertion to hold. Consequently, there exists a compact subset of this open face that contains all divergence balls for large . The entropy function is smooth on . Let be an upper bound for the norm of differential of restricted on , where the norm is evaluated with respect to the total variation distance on .
Step 3.
Let and be uniform unit vectors in and , whose supports coincide with that of and , respectively. Then for all and holds
Since coefficients of are non-negative, the above pointwise estimates imply
Taking normalized logarithm and combining with inequality in (3.4.I) we get
This finishes the proof of the theorem. ∎
4 Bilinear form associated to a bipartite graph
In this section we consider the bilinear form associated with a biregular bipartite graph and explore the relations between uniform vectors, subgraphs, and extensions of random pairs uniformly supported on the graph.
Let be a biregular bipartite graph. Denote by the bilinear form associated with defined by
We refer to as the incidence bilinear form, or simply the incidence form, for the graph .
4.1 Norms of incidence bilinear forms
Let stand for the power of . The incidence bilinear form of is the tensor power of the incidence form of . By Proposition 3.3.A we have the equality
for any , satisfying .
In the complimentary range , the norms of the incidence form are completely determined by the sizes of the graph.
4.1.A.
Let be the incidence form of a biregular bipartite graph . Let , . Then
Note that the right-hand side in the equality above is multiplicative with respect to product of bipartite graphs. Thus for incidence bilinear forms the multiplicativity of the norms, Proposition 3.3.A, holds without any restriction on the exponents . This property is specific to incidence forms and does not hold for general bilinear forms.
4.1.B Proof of Proposition 4.1.A
We start by evaluating norms for the extremal values of exponents
For any and we have the inequality
with equality for , . Thus
Also
with equality for for some , and . Similar inequality holds with the roles of and switched. Thus
4.2 Uniform almost resonance vectors for incidence forms
Recall that Corollary 3.4.D and Remark 3.4.E allow us to find almost resonance pairs of uniform vectors for large tensor powers of bilinear form , such that in addition they optimize the restricted supremum in the definition of . We now apply these results to the incidence form of a biregular bipartite graph. In that case, the supports of almost resonance uniform pair of vectors have an additional regularity property, that we discuss below.
Consider the incidence form of a biregular bipartite graph . Let and be two uniform vectors attaining the maximum in the definition of . Without loss of generality we may assume that and are indicator functions of some subsets of and , respectively. Denote by the subgraph of induced by the supports and of and , respectively. We choose the supports minimal among optimizers; then the induced subgraph has no isolated vertices.33 3 The fact that has no isolated vertices is automatic for since removing isolated vertices would strictly improve the quotient in the definition of restricted subnorm. We only need to minimize the support in the case when at least one of the exponents is infinite.
Then we have the following relations: For any and
The quotients
are equal to the average left and right degree of , respectively.
For we say that is -quasi-biregular if every vertex in the left part has degree at least -fraction of the average left degree and every vertex in the right part has degree at least -fraction of the average right degree. In other words, for all , holds
In the next proposition we show that for the uniform -optimizing pair of vectors and the graph induced by the supports of the vectors is -quasi-biregular.
4.2.A.
Let and let , and be as above. Then for all and holds
Thus the degree within of any vertex can only deviate down from the average degree by a fixed factor and is -quasi-biregular.
Proposition 4.2.A is not needed for the proof of the main result. Nevertheless, we include it here because it gives useful structural information about uniform optimizers.
4.2.B Proof of Proposition 4.2.A
Let be an arbitrary vertex in the left part of and . Removing the vertex can not improve the optimization quotient, hence
Therefore
which gives
Similar argument works for the right part of the subgraph . ∎
4.3 Subgraphs and extensions
4.3.A Construction of an extension from a subgraph
Suppose is a homogeneous bipartite graph and is uniformly supported on . Denote by the automorphism group of (acting on on the left). Let be a subgraph without isolated vertices. Applying symmetries we obtain a family of subgraphs
We now construct an extension with alphabet . The joint distribution of the triple is defined by
All the pairs for different choices of are isomorphic and have distribution
Thus the supporting graph of is . We also have
Therefore we have the following bounds
| (4.3.B) | ||||||
The next proposition establishes the connection between uniform optimizers and extensions.
4.3.C.
Let be a pair uniformly supported on a homogeneous bipartite graph with the incidence form . Let and , be a pair of indicator vectors which are optimizers in the definition of . Then there exists an extension such that
4.3.D Proof of Proposition 4.3.C
5 Proof of the main theorem
5.1.A.
Let be a pair of random variables uniformly supported on a homogeneous bipartite graph with the incidence form . Let . Then
where and .
The proof of the theorem splits into two cases, the hypercontractive regime, , and Hölder regime, . The proof strategy in the two cases seems rather different, and we do not know whether there exists a unified proof covering both. Most of the considerations above can be generalized to longer tuples of random variables and their shapes. For -tuples the hypercontractive region, , and Hölder region, , are no longer complimentary, if . We do not know, whether generalized theorem holds in the intermediate region of values of the exponents not included in these two regions.
5.1.B Proof of Theorem 5.1.A
Let , and be as in the theorem. Fix and set , . We consider two cases.
Case ; equivalently, .
Case ; equivalently, .
We prove two inequalities, which together imply the theorem,
| (5.1.E) | ||||
| (5.1.F) |
Let be the pair obtained by taking independent copies of . On the one hand by [MR26, Proposition 3.6.A] the shape function is additive under independent products, hence
On the other hand, Proposition 3.3.A implies that the -norm of is also stable in the sense that
Proof of inequality (5.1.E).
Suppose that is an extension, such that
Using stability of the shape function we also have
By Tropical Asymptotic Equipartition Property, [MP18], we can replace the extending variable by another extending variable such that the joint distribution of is uniform on its support and the conditional entropy profiles and are close on the normalized scale. More precisely, for every there exists and an extension , uniform on its support, such that
Consequently,
Choose an arbitrary atom from the alphabet of and let be the subgraph supporting . Write , where
are the left part, the right part and the edge set of . Since is uniform, we have the following identities
Let and be the indicator functions of and , respectively. Then
Therefore
This inequality holds for every , so inequality (5.1.E) follows.
Proof of inequality (5.1.F).
6 Properties of the shape
In this section we establish inequalities and relations satisfied by shapes of pairs of random variables. Most of these relations are elementary and can be derived directly in the context of the definition of the shape by using properties of entropy and Shannon inequalities. However, inequality (COMP) and its implication, inequality (REF), seem to be different. We do not know, whether a proof not referring to Theorem 5.1.A exists.
6.1 Lower and upper bounds for the shape function
We need to recall some definitions from [MR26]: For a pair of random variables and define
The apex point where three affine pieces of meet is
6.2 Properties of the shape
For every tuple of random variables and every
holds:
6.2.A Symmetry
| (SYM) |
6.2.B Convexity
If
then
| (CONV) |
6.2.C Coordinate-wise monotonicity and reverse monotonicity
If and then
| (MO) | ||||
| (RMO) |
Written together
6.2.D Additivity
If is the independent sum, then
| (ADD) |
6.2.E Chain rule
| (CHAIN) |
6.2.F Test inequality
| (TEST) |
6.2.G Upper bound
| (UP) |
The inequality is equality on the upper-right triangle , where the second max-summand dominates. It is also equality on the boundary of the square where is equal to the first max-summand.
The upper bound is attained if and only if mutual information of is extractable, that is there is an extension with
6.2.H Lower bound
| (LO) |
On the upper-right triangle and the boundary of the square we have equality.
Lower bound is attained for rigid pairs. It is attained up to for pairs supported on balanced expanders, see [MR26].
Define the rigidity of the nondegenerate pair by
Thus we have , with if and only if the pair is rigid and if and only if the pair has extractable mutual information.
6.2.I Composition inequality
| (COMP) |
In the Hölder regime (), take to be any value in the interval . The difference between the right-hand side and the left-hand side is then .
6.2.J Refinement and coarsening
For a pair of random variables write if is a refinement of , equivalently .
If and then
| (REF) |
| (CRS) |
We can use symmetry to sequentially coarsen the first and the second variable, but the resulting bound depend on the order of application. However the following symmetric but weaker inequality holds
| (CRS2) |
6.2.K Link and reversed link inequalities
| (LI) | ||||
| (RLI) |
6.2.L Flattening/links and reverse
| (FLI) | ||||
| (RFLI) |
6.2.M MMRV-inequality
| (MMRV) |
6.3 Proofs
Symmetry (SYM) follows directly from the definition of the shape. Convexity of the shape function, (CONV), is proven in [MR26, Section 2.2]. The relations (ADD) and (CHAIN) are proven in [MR26, Section 3.6]. The upper and lower bounds for the shape function, (UP) and (LO), are proven in [MR26, Sections 3.4 and 3.5]. Inequality (TEST) follows directly from the definition.
Proof of (MO) and (RMO).
For and and all ’s extending holds
Substituting in the definition of the shape we obtain the required inequalities. ∎
Proof of (COMP).
By additivity and continuity of the shape function, [MR26, Proposition 3.6.A] and by Tropical Asymptotic Equipartition Property, [MP18, Theorem 6.1] we can assume that the triple is uniform on the support. Let and
be the incidence operators of the graphs supporting pairs , and , respectively. Denote by the composition of the operators. Then
We note that all operators have non-negative coefficients, therefore coordinate-wise domination implies the corresponding inequality for the norms. Thus, taking the operator norms, we obtain inequality
Switching to norms of bilinear forms we obtain
We now take logarithm of the last inequality, use Theorem 5.1.A and the substitutions , and , to obtain inequality (COMP). ∎
Proof of (REF).
Proof of (CRS).
Let be an optimizer in the definition of . Then
∎
Proof of (LI).
Proof of (RLI).
Proof of (FLI).
Suppose is an optimizer in the definition of . Then
∎
Proof of (RFLI).
Let be the optimizer in the definition of . Then
∎
Proof of (MMRV).
In [Mak+02] the following non-Shannon inequality for five random variables is proven
The right-hand side can be rewritten as
Thus
Since , then taking the supremum over ’s gives
∎
References
- [BL76] Jöran Bergh and Jörgen Löfström “Interpolation Spaces: An Introduction” Springer, 1976
- [CK11] Imre Csiszár and János Körner “Information Theory: Coding Theorems for Discrete Memoryless Systems” Cambridge University Press, 2011
- [Csi23] László Csirmaz “A short proof of the Gács–Körner theorem” In arXiv preprint arXiv:2306.14718, 2023
- [Csi98] Imre Csiszár “The Method of Types” In IEEE Transactions on Information Theory 44.6, 1998, pp. 2505–2523
- [DF93] Andreas Defant and Klaus Floret “Tensor Norms and Operator Ideals” North-Holland, 1993
- [HLP52] G.. Hardy, J.. Littlewood and G. Pólya “Inequalities” Cambridge: Cambridge University Press, 1952
- [LE17] Cheuk Li and Abbas El “Extended Gray–Wyner system with complementary causal side information” In IEEE Transactions on Information Theory 64.8 IEEE, 2017, pp. 5862–5878
- [Mak+02] Konstantin Makarychev, Yury Makarychev, Andrei Romashchenko and Nikolai Vereshchagin “A new class of non-Shannon-type inequalities for entropies” In Communications in Information and Systems 2.2 International Press of Boston, 2002, pp. 147–166
- [Mat05] František Matúš “Inequalities for Shannon entropies and adhesivity of polymatroids” In 9th Canadian Workshop on Information Theory, McGill University, Montréal, Québec, Canada 540, 2005
- [Mat07] Frantisek Matúš “Infinitely many information inequalities” In 2007 IEEE International Symposium on Information Theory, 2007, pp. 41–44 IEEE
- [MP18] Rostislav Matveev and Jacobus Portegies “Asymptotic dependency structure of multiple signals: Asymptotic equipartition property for diagrams of probability spaces” In Information Geometry 1.2 Springer, 2018, pp. 237–285
- [MR26] Rostislav Matveev and Andrei Romashchenko “Beyond Mutual Information: Extension Profiles and Shape Functions of Random Variable Pairs” In arXiv preprint arXiv:2606.23849, 2026
- [MR26a] Rostislav Matveev and Andrei Romashchenko “Spectral Conditions for the Ingleton Inequality” In IEEE Transactions on Information Theory, 2026
- [Pin64] Mark Pinsker “Information and Information Stability of Random Variables and Processes” Translated and edited by Amiel Feinstein San Francisco: Holden-Day, 1964
- [PP14] Vinod Prabhakaran and Manoj Prabhakaran “Assisted common information with an application to secure two-party sampling” In IEEE Transactions on Information Theory 60.6 IEEE, 2014, pp. 3413–3434
- [Rie27] Marcel Riesz “Sur les maxima des formes bilinéaires et sur les fonctionnelles linéaires” In Acta Mathematica 49, 1927, pp. 465–497
- [Rya02] Raymond. Ryan “Introduction to Tensor Products of Banach Spaces” Springer, 2002
- [San57] I.. Sanov “On the Probability of Large Deviations of Random Magnitudes” In Matematicheskii Sbornik 42(84).1, 1957, pp. 11–44
- [Tho39] G. Thorin “Convexity theorems generalizing those of M. Riesz and Hadamard with some applications” In Meddelanden från Lunds Universitets Matematiska Seminarium 9, 1939, pp. 1–58